128GB VRAM

DGX Spark 128G — what LLMs can it run?

84 of 90 indexed models fit comfortably in 128GB at 4K context, each at the highest-quality quant that still leaves headroom.

The short answer for DGX Spark 128G

128 GB, 273 GB/s Unified LPDDR5X (One pool shared by the GB10's CPU and GPU and by DGX OS itself. NVIDIA publishes no fixed GPU share, so fits are judged against the whole pool — and nothing larger than it can load, because there is no separate system RAM to spill into.). 84 of 90 models in this index fit comfortably at 4K context; 85 load at all.

Biggest that fits
DBRX Instruct — 84.3 GB at Q4_K_M — 43.7 GB spare at 4K, so longer context comes out of a thin margin.
Room to grow
GPT-OSS 120B — 66.4 GB at MXFP4 — under 60% of the card, which leaves 61.6 GB for a long context window or a second process.
Speed ceiling
Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 273 GB/s puts a hard ceiling near 59 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
The wall you will hit
Bandwidth, and one pool. 128 GB holds models no single consumer card can, but every byte of it — CPU, GPU and DGX OS alike — comes from the same LPDDR5X at 273 GB/s, a fraction of a discrete flagship's. There is no system RAM behind it to spill into, so this site treats anything larger than the pool as not loading at all. The largest dense model here, Llama 3.2 90B Vision Instruct at Q4_K_M, has a ceiling of about 5 tok/s on it; mixture-of-experts models read only their active experts per token and run far faster than their size suggests.

70B+ · 10

DGX Spark 128G — what LLMs can it run? — 70B+
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
DBRX Instruct132B MoEQ4_K_M84.31 GB+43.7 GB—
GPT-OSS 120B117B MoEMXFP466.4 GB+61.6 GB—
Llama 4 Scout 17B (16E)109B MoEQ4_K_M69.88 GB+58.1 GB—
GLM-4.5-Air106B MoEQ4_K_M67.94 GB+60.1 GB—
Llama 3.2 90B Vision Instruct90BQ4_K_M57.5 GB+70.5 GB—
Qwen2.5 72B Instruct72BQ5_K_M55.31 GB+72.7 GB—
Llama 3.1 70B Instruct70BQ5_K_M53.75 GB+74.3 GB—
Llama 3.3 70B Instruct70BQ5_K_M53.75 GB+74.3 GB—
DeepSeek-R1-Distill-Llama-70B70BQ4_K_M46.1 GB+81.9 GB—
Jamba 1.5 Mini52B-A12BQ4_K_M33.01 GB+95 GB—

32B · 24

DGX Spark 128G — what LLMs can it run? — 32B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
Kimi Linear 48B-A3B Instruct48B-A3BQ8_053.33 GB+74.7 GB—
Mixtral 8x7B Instruct47B MoEQ4_K_M30.13 GB+97.9 GB—
Seed-OSS 36B Instruct36BQ5_K_M27.81 GB+100.2 GB—
Command R 35B35BQ4_K_M22.86 GB+105.1 GB42Estimated
Yi 1.5 34B Chat34BQ4_K_M22.82 GB+105.2 GB40Estimated
Qwen3 32B Instruct32BQ4_K_M21.85 GB+106.2 GB42
Qwen2.5 32B Instruct32BQ4_K_M21.69 GB+106.3 GB44
Qwen2.5-Coder 32B Instruct32BQ4_K_M21.69 GB+106.3 GB44Estimated
DeepSeek-R1-Distill-Qwen-32B32BQ4_K_M21.69 GB+106.3 GB42Estimated
Gemma 4 31B IT31BQ8_035.72 GB+92.3 GB—
Qwen3 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+105 GB82
Qwen3-Coder 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+105 GB80
Qwen3-VL 30B-A3B Instruct30B-A3BQ8_034.28 GB+93.7 GB—
GLM-4.7-Flash30B-A3BQ8_033.47 GB+94.5 GB—
Qwen3.8 27B27BQ4_K_M17.47 GB+110.5 GB—
Gemma 3 27B IT27BQ4_K_M18.37 GB+109.6 GB48Community
Gemma 2 27B Instruct27BQ5_K_M21.76 GB+106.2 GB42Estimated
Gemma 4 26B-A4B IT26B-A4BQ8_028.42 GB+99.6 GB—
Mistral Small 24B Instruct24BEXL2 4.65bpw15.26 GB+112.7 GB—
Devstral Small 1.1 24B24BQ6_K20.91 GB+107.1 GB48Estimated
Magistral Small 1.2 24B24BQ6_K20.91 GB+107.1 GB47Estimated
Codestral 22B22BQ4_K_M15.03 GB+113 GB58Estimated
ERNIE 4.5 21B-A3B21B-A3BQ8_024.44 GB+103.6 GB—
GPT-OSS 20B21B MoEMXFP412.76 GB+115.2 GB195Community

14B · 16

DGX Spark 128G — what LLMs can it run? — 14B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
InternLM2 20B Chat20BQ5_K_M15.59 GB+112.4 GB68Estimated
DeepSeek-Coder-V2-Lite Instruct16B-A2.4BQ8_017.56 GB+110.4 GB118Estimated
DeepSeek-V2-Lite Chat16B-A2.4BQ4_K_M10.08 GB+117.9 GB142Estimated
StarCoder2 15B15BQ4_K_M10.19 GB+117.8 GB92Estimated
Qwen3 14B Instruct14BQ5_K_M11.65 GB+116.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+116.3 GB86Estimated
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+118.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+116.2 GB78Estimated
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+115.1 GB88Estimated
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+116.9 GB95Estimated
Gemma 3 12B IT12BQ5_K_M9.84 GB+118.2 GB92Estimated
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+119.7 GB108Estimated
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+115.1 GB72Estimated
Solar 10.7B Instruct11BQ4_K_M7.6 GB+120.4 GB125Estimated
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+120.7 GB118Estimated
Gemma 2 9B Instruct9BQ8_011.7 GB+116.3 GB108Estimated

7B · 23

DGX Spark 128G — what LLMs can it run? — 7B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
GLM-4-9B-Chat9BQ8_010.16 GB+117.8 GB105Estimated
Qwen3-VL 8B Instruct8BQ8_010.39 GB+117.6 GB108Estimated
Ministral 3 8B Instruct8BQ8_010.02 GB+118 GB—
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+122.5 GB72Estimated
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+122.1 GB142Estimated
Qwen3 8B Instruct8BQ6_K7.64 GB+120.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+118.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+122.6 GB232Estimated
Aya 23 8B8BQ4_K_M5.64 GB+122.4 GB145Estimated
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+122.6 GB228Estimated
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+121.5 GB128Estimated
InternLM2 7B Chat7BQ4_K_M5.45 GB+122.6 GB148Estimated
Qwen2.5 7B Instruct7BQ6_K6.77 GB+121.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+123.1 GB248Estimated
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+123.1 GB210Estimated
Gemma 4 E4B ITE4BQ8_08.46 GB+119.5 GB—
OLMo 2 7B Instruct7BQ8_010.3 GB+117.7 GB125Estimated
OLMo 3 7B Instruct7BQ8_010.3 GB+117.7 GB—
WizardLM-2 7B7BQ4_K_M5.14 GB+122.9 GB152Estimated
Mistral 7B Instruct v0.37BQ6_K6.75 GB+121.3 GB135Estimated
Zephyr 7B Beta7BQ6_K6.75 GB+121.3 GB132Estimated
Gemma 4 E2B ITE2BQ8_05.2 GB+122.8 GB—
Gemma 3 4B IT4BQ8_05.05 GB+123 GB145Estimated

≤3B · 11

DGX Spark 128G — what LLMs can it run? — ≤3B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
Qwen3 4B Instruct4BQ6_K4.05 GB+124 GB145
Phi-4 Mini Instruct3.8BQ8_04.81 GB+123.2 GB262Estimated
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+122.1 GB255Estimated
Llama 3.2 3B Instruct3BQ8_04.05 GB+124 GB285Estimated
Qwen2.5 3B Instruct3BQ8_03.59 GB+124.4 GB290Estimated
SmolLM3 3B3BQ8_03.73 GB+124.3 GB—
Gemma 2 2B Instruct2BQ8_03.34 GB+124.7 GB320Estimated
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+125.6 GB240Estimated
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+126.2 GB410Estimated
Llama 3.2 1B Instruct1BQ8_01.51 GB+126.5 GB450Estimated
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+127.4 GB540Estimated

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 1 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a DGX Spark 128G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.

Cards with the same budget

The same 128GB in system memory is a different machine — everything on this page fits there, at a speed set by your DIMMs rather than a graphics bus: 128 GB RAM (CPU)

    Setup guides for this card

    Step-by-step guides written for the DGX Spark 128G's runtime or memory size, with the commands and the checks that prove the model is running on the card.

    Common questions

    What is the best local LLM for a DGX Spark 128G?

    For everyday use, GPT-OSS 120B at MXFP4 — about 66.4 GB of the card's 128 GB at 4K context, leaving 61.6 GB for a longer window. If you want the largest thing that will load, that is DBRX Instruct at Q4_K_M (84.3 GB, 43.7 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.

    Can a DGX Spark 128G run Qwen3 235B-A22B Instruct?

    Not reliably. Its smallest build here, Q3_K_M, needs about 119.6 GB — 93% of the machine's 128 GB, and DGX OS runs from that same pool, so whether it loads depends on what the system itself is holding, with nothing left for a longer context. The largest model this machine does clear is DBRX Instruct.

    How many tokens per second does a DGX Spark 128G do on an 8B model at Q4?

    This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 273 GB/s — a ceiling near 59 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.

    DGX Spark 128G or Mac M5 Max 128G for local LLMs?

    They hold the same models: 128 GB against 128 GB fits 84 and 84 of 90 respectively at 4K. The difference is throughput — 273 GB/s against 614 GB/s, a 2.2× gap in how fast the weights can be read, which is what token generation is bound by. The Mac M5 Max 128G generates faster on any model both can hold.

    84 of 90 indexed models fit comfortably in 128GB at 4K context, each at the highest-quality quant that still leaves headroom.

    This page's figures change when the model or the runtime does.Last updated 2026-10-05 RSS → /feed.xml