DGX Spark 64G — what LLMs can it run?
82 of 90 indexed models fit comfortably in 64GB at 4K context, each at the highest-quality quant that still leaves headroom.
The short answer for DGX Spark 64G
64 GB, 273 GB/s Unified LPDDR5X (Same 273 GB/s as the 128 GB model. One pool shared by the GB10's CPU and GPU and by DGX OS itself. NVIDIA publishes no fixed GPU share, so fits are judged against the whole pool — and nothing larger than it can load, because there is no separate system RAM to spill into.). 82 of 90 models in this index fit comfortably at 4K context; 83 load at all.
- Biggest that fits
- Llama 4 Scout 17B (16E) — 55.9 GB at Q3_K_M — 8.1 GB spare at 4K, so longer context comes out of a thin margin.
- Room to grow
- Jamba 1.5 Mini — 33.0 GB at Q4_K_M — under 60% of the card, which leaves 31.0 GB for a long context window or a second process.
- Speed ceiling
- Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 273 GB/s puts a hard ceiling near 59 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
- The wall you will hit
- Bandwidth, and one pool. 64 GB holds models no single consumer card can, but every byte of it — CPU, GPU and DGX OS alike — comes from the same LPDDR5X at 273 GB/s, a fraction of a discrete flagship's. There is no system RAM behind it to spill into, so this site treats anything larger than the pool as not loading at all. The largest dense model here, Llama 3.2 90B Vision Instruct at Q3_K_M, has a ceiling of about 7 tok/s on it; mixture-of-experts models read only their active experts per token and run far faster than their size suggests.
70B+ · 8
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Llama 4 Scout 17B (16E)109B MoE | Q3_K_M | 55.92 GB | +8.1 GB | — |
| GLM-4.5-Air106B MoE | Q3_K_M | 54.37 GB | +9.6 GB | — |
| Llama 3.2 90B Vision Instruct90B | Q3_K_M | 46.16 GB | +17.8 GB | — |
| Qwen2.5 72B Instruct72B | Q5_K_M | 55.31 GB | +8.7 GB | — |
| Llama 3.1 70B Instruct70B | Q5_K_M | 53.75 GB | +10.3 GB | — |
| Llama 3.3 70B Instruct70B | Q5_K_M | 53.75 GB | +10.3 GB | — |
| DeepSeek-R1-Distill-Llama-70B70B | Q4_K_M | 46.1 GB | +17.9 GB | — |
| Jamba 1.5 Mini52B-A12B | Q4_K_M | 33.01 GB | +31 GB | — |
32B · 24
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Kimi Linear 48B-A3B Instruct48B-A3B | Q8_0 | 53.33 GB | +10.7 GB | — |
| Mixtral 8x7B Instruct47B MoE | Q4_K_M | 30.13 GB | +33.9 GB | — |
| Seed-OSS 36B Instruct36B | Q5_K_M | 27.81 GB | +36.2 GB | — |
| Command R 35B35B | Q4_K_M | 22.86 GB | +41.1 GB | 42Estimated |
| Yi 1.5 34B Chat34B | Q4_K_M | 22.82 GB | +41.2 GB | 40Estimated |
| Qwen3 32B Instruct32B | Q4_K_M | 21.85 GB | +42.2 GB | 42 |
| Qwen2.5 32B Instruct32B | Q4_K_M | 21.69 GB | +42.3 GB | 44 |
| Qwen2.5-Coder 32B Instruct32B | Q4_K_M | 21.69 GB | +42.3 GB | 44Estimated |
| DeepSeek-R1-Distill-Qwen-32B32B | Q4_K_M | 21.69 GB | +42.3 GB | 42Estimated |
| Gemma 4 31B IT31B | Q8_0 | 35.72 GB | +28.3 GB | — |
| Qwen3 30B-A3B Instruct30B-A3B | Q5_K_M | 23.04 GB | +41 GB | 82 |
| Qwen3-Coder 30B-A3B Instruct30B-A3B | Q5_K_M | 23.04 GB | +41 GB | 80 |
| Qwen3-VL 30B-A3B Instruct30B-A3B | Q8_0 | 34.28 GB | +29.7 GB | — |
| GLM-4.7-Flash30B-A3B | Q8_0 | 33.47 GB | +30.5 GB | — |
| Qwen3.8 27B27B | Q4_K_M | 17.47 GB | +46.5 GB | — |
| Gemma 3 27B IT27B | Q4_K_M | 18.37 GB | +45.6 GB | 48Community |
| Gemma 2 27B Instruct27B | Q5_K_M | 21.76 GB | +42.2 GB | 42Estimated |
| Gemma 4 26B-A4B IT26B-A4B | Q8_0 | 28.42 GB | +35.6 GB | — |
| Mistral Small 24B Instruct24B | EXL2 4.65bpw | 15.26 GB | +48.7 GB | — |
| Devstral Small 1.1 24B24B | Q6_K | 20.91 GB | +43.1 GB | 48Estimated |
| Magistral Small 1.2 24B24B | Q6_K | 20.91 GB | +43.1 GB | 47Estimated |
| Codestral 22B22B | Q4_K_M | 15.03 GB | +49 GB | 58Estimated |
| ERNIE 4.5 21B-A3B21B-A3B | Q8_0 | 24.44 GB | +39.6 GB | — |
| GPT-OSS 20B21B MoE | MXFP4 | 12.76 GB | +51.2 GB | 195Community |
14B · 16
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| InternLM2 20B Chat20B | Q5_K_M | 15.59 GB | +48.4 GB | 68Estimated |
| DeepSeek-Coder-V2-Lite Instruct16B-A2.4B | Q8_0 | 17.56 GB | +46.4 GB | 118Estimated |
| DeepSeek-V2-Lite Chat16B-A2.4B | Q4_K_M | 10.08 GB | +53.9 GB | 142Estimated |
| StarCoder2 15B15B | Q4_K_M | 10.19 GB | +53.8 GB | 92Estimated |
| Qwen3 14B Instruct14B | Q5_K_M | 11.65 GB | +52.4 GB | 78 |
| Qwen2.5 14B Instruct14B | Q5_K_M | 11.73 GB | +52.3 GB | 86Estimated |
| DeepSeek-R1-Distill-Qwen-14B14B | EXL2 4.65bpw | 9.75 GB | +54.3 GB | 128 |
| Phi-4 14B14B | Q5_K_M | 11.77 GB | +52.2 GB | 78Estimated |
| Phi-3 Medium 14B Instruct14B | Q6_K | 12.86 GB | +51.1 GB | 88Estimated |
| Mistral Nemo 12B Instruct12B | Q6_K | 11.14 GB | +52.9 GB | 95Estimated |
| Gemma 3 12B IT12B | Q5_K_M | 9.84 GB | +54.2 GB | 92Estimated |
| Stable LM 2 12B Chat12B | Q4_K_M | 8.35 GB | +55.7 GB | 108Estimated |
| Llama 3.2 11B Vision Instruct11B | Q8_0 | 12.9 GB | +51.1 GB | 72Estimated |
| Solar 10.7B Instruct11B | Q4_K_M | 7.6 GB | +56.4 GB | 125Estimated |
| Falcon 3 10B Instruct10B | Q4_K_M | 7.28 GB | +56.7 GB | 118Estimated |
| Gemma 2 9B Instruct9B | Q8_0 | 11.7 GB | +52.3 GB | 108Estimated |
7B · 23
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| GLM-4-9B-Chat9B | Q8_0 | 10.16 GB | +53.8 GB | 105Estimated |
| Qwen3-VL 8B Instruct8B | Q8_0 | 10.39 GB | +53.6 GB | 108Estimated |
| Ministral 3 8B Instruct8B | Q8_0 | 10.02 GB | +54 GB | — |
| Qwen2-VL 7B Instruct7B | Q4_K_M | 5.49 GB | +58.5 GB | 72Estimated |
| Granite 3.1 8B Instruct8B | Q4_K_M | 5.88 GB | +58.1 GB | 142Estimated |
| Qwen3 8B Instruct8B | Q6_K | 7.64 GB | +56.4 GB | 122 |
| Llama 3.1 8B Instruct8B | Q8_0 | 9.47 GB | +54.5 GB | 118 |
| Nous Hermes 3 Llama 3.1 8B8B | EXL2 4.65bpw | 5.43 GB | +58.6 GB | 232Estimated |
| Aya 23 8B8B | Q4_K_M | 5.64 GB | +58.4 GB | 145Estimated |
| OpenChat 3.6 8B8B | EXL2 4.65bpw | 5.43 GB | +58.6 GB | 228Estimated |
| DeepSeek-R1-Distill-Llama-8B8B | Q5_K_M | 6.51 GB | +57.5 GB | 128Estimated |
| InternLM2 7B Chat7B | Q4_K_M | 5.45 GB | +58.6 GB | 148Estimated |
| Qwen2.5 7B Instruct7B | Q6_K | 6.77 GB | +57.2 GB | 132 |
| Qwen2.5-Coder 7B Instruct7B | EXL2 4.65bpw | 4.87 GB | +59.1 GB | 248Estimated |
| DeepSeek-R1-Distill-Qwen-7B7B | EXL2 4.65bpw | 4.87 GB | +59.1 GB | 210Estimated |
| Gemma 4 E4B ITE4B | Q8_0 | 8.46 GB | +55.5 GB | — |
| OLMo 2 7B Instruct7B | Q8_0 | 10.3 GB | +53.7 GB | 125Estimated |
| OLMo 3 7B Instruct7B | Q8_0 | 10.3 GB | +53.7 GB | — |
| WizardLM-2 7B7B | Q4_K_M | 5.14 GB | +58.9 GB | 152Estimated |
| Mistral 7B Instruct v0.37B | Q6_K | 6.75 GB | +57.3 GB | 135Estimated |
| Zephyr 7B Beta7B | Q6_K | 6.75 GB | +57.3 GB | 132Estimated |
| Gemma 4 E2B ITE2B | Q8_0 | 5.2 GB | +58.8 GB | — |
| Gemma 3 4B IT4B | Q8_0 | 5.05 GB | +59 GB | 145Estimated |
≤3B · 11
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Qwen3 4B Instruct4B | Q6_K | 4.05 GB | +60 GB | 145 |
| Phi-4 Mini Instruct3.8B | Q8_0 | 4.81 GB | +59.2 GB | 262Estimated |
| Phi-3.5 Mini Instruct3.8B | Q8_0 | 5.89 GB | +58.1 GB | 255Estimated |
| Llama 3.2 3B Instruct3B | Q8_0 | 4.05 GB | +60 GB | 285Estimated |
| Qwen2.5 3B Instruct3B | Q8_0 | 3.59 GB | +60.4 GB | 290Estimated |
| SmolLM3 3B3B | Q8_0 | 3.73 GB | +60.3 GB | — |
| Gemma 2 2B Instruct2B | Q8_0 | 3.34 GB | +60.7 GB | 320Estimated |
| Qwen3 1.7B Instruct1.7B | Q8_0 | 2.39 GB | +61.6 GB | 240Estimated |
| Qwen2.5 1.5B Instruct1.5B | Q8_0 | 1.83 GB | +62.2 GB | 410Estimated |
| Llama 3.2 1B Instruct1B | Q8_0 | 1.51 GB | +62.5 GB | 450Estimated |
| Qwen2.5 0.5B Instruct0.5B | Q8_0 | 0.6 GB | +63.4 GB | 540Estimated |
How this list is built
Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.
A further 1 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.
Measured on this card
No benchmark runs in this index were recorded on a DGX Spark 64G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.
Cards with the same budget
The same 64GB in system memory is a different machine — everything on this page fits there, at a speed set by your DIMMs rather than a graphics bus: 64 GB RAM (CPU)
Setup guides for this card
Step-by-step guides written for the DGX Spark 64G's runtime or memory size, with the commands and the checks that prove the model is running on the card.
Stepping up
A DGX Spark 128G (128GB) fits 2 more of the indexed models than this card. DGX Spark 128G →
Common questions
What is the best local LLM for a DGX Spark 64G?
For everyday use, Jamba 1.5 Mini at Q4_K_M — about 33.0 GB of the card's 64 GB at 4K context, leaving 31.0 GB for a longer window. If you want the largest thing that will load, that is Llama 4 Scout 17B (16E) at Q3_K_M (55.9 GB, 8.1 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.
Can a DGX Spark 64G run GPT-OSS 120B?
Not reliably. Its smallest build here, Q4_K_M, needs about 62.7 GB — 98% of the machine's 64 GB, and DGX OS runs from that same pool, so whether it loads depends on what the system itself is holding, with nothing left for a longer context. The native MXFP4 release is 66.4 GB, larger than the whole pool. The largest model this machine does clear is Llama 4 Scout 17B (16E).
How many tokens per second does a DGX Spark 64G do on an 8B model at Q4?
This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 273 GB/s — a ceiling near 59 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.
DGX Spark 64G or Mac M1 Max 64G for local LLMs?
They do not hold the same set: DGX Spark 64G fits 82 and Mac M1 Max 64G 76 of 90 at 4K — macOS keeps part of the Mac M1 Max 64G's unified memory from the GPU, leaving about 48 GB usable. Speed is the other difference — 273 GB/s against 400 GB/s, a 1.5× gap in how fast the weights can be read, which is what token generation is bound by. The Mac M1 Max 64G generates faster on any model both can hold.
82 of 90 indexed models fit comfortably in 64GB at 4K context, each at the highest-quality quant that still leaves headroom.
This page's figures change when the model or the runtime does.Last updated 2026-10-05 RSS → /feed.xml