The short answer
Both DGX Spark models use the same GB10 chip and the same 273 GB/s memory; the only difference that matters for local models is capacity. NVIDIA announced the 64GB model at $4,999, on sale from 23 October 2026 through Acer, ASUS, Dell, GIGABYTE, HP and MSI. At 4K context, 82 of the 90 models in this index fit the 64GB Spark comfortably and 84 fit the 128GB one — against 74 for a 32GB RTX 5090. The gap between the two Sparks is narrow but it sits exactly on the models people buy a Spark for: the 100B-class mixture-of-experts releases.
GPT-OSS 120B does not fit the 64GB model
GPT-OSS 120B at MXFP4 with an 8K context is about 66.5 GB — more than the whole machine has, and the CPU and DGX OS live in that same pool. The Q4_K_M conversions do not rescue it: GPT-OSS 120B at Q4_K_M with an 8K context is about 62.9 GB, 98% of the pool, which leaves the operating system nothing. On the 128GB model the native build is about half the memory. GLM-4.5-Air and Llama 4 Scout follow the same pattern at Q4_K_M; on 64GB they need their Q3_K_M builds. Note that this site treats anything over the pool as not loading at all: on a discrete card a small overrun can spill into system RAM, but a Spark has no system RAM besides this.
GGUF 8K 32K 128K 64GB 128GB
GPT-OSS 20B MXFP4 12.9 GB 13.5 GB 16.0 GB ok ok
Llama 3.3 70B Q4_K_M 47.5 GB 55.7 GB 88.7 GB ok* ok
GLM-4.5-Air Q3_K_M 55.2 GB 59.9 GB 78.9 GB ok† ok
GLM-4.5-Air Q4_K_M 68.7 GB 73.5 GB 92.5 GB ✗ ok
Llama 4 Scout Q4_K_M 70.7 GB 72.0 GB 77.0 GB ✗ ok
GPT-OSS 120B Q4_K_M 62.9 GB 63.8 GB 67.5 GB ✗ ok
GPT-OSS 120B MXFP4 66.5 GB 67.5 GB 71.2 GB ✗ ok
Qwen3 235B-A22B Q3_K_M 120.4 GB 125.3 GB — ✗ ✗
ok = within 88% of the pool * up to the 32K column † 8K column only
✗ = over 88%; past 100% it cannot load at allWhat 64GB is good for
The 64GB model is the dense-70B machine. Llama 3.3 70B at Q4_K_M with a 32K context is about 55.7 GB, inside the comfortable line, and no single consumer card holds it at all. GPT-OSS 20B at MXFP4 with a 128K context is about 16.0 GB, so it runs its full window with most of the pool to spare. GLM-4.5-Air at Q3_K_M with an 8K context is about 55.2 GB, which is the largest mixture-of-experts option, at a lower quant than you would choose with room to spare. If the plan is GPT-OSS 120B, GLM-4.5-Air at Q4_K_M or long context on 70B, buy the 128GB model.
Speed: 273 GB/s is the ceiling for both
Generation reads the active weights once per token, so bandwidth divided by the bytes read is a hard ceiling. For a dense model that is the whole file: Llama 3.3 70B at Q4_K_M tops out around 7 tok/s on either Spark. That is the trade the machine makes — an RTX 5090 moves 1,792 GB/s, about 6.6× more, on anything it can hold. Mixture-of-experts models read only their active experts, which is why GPT-OSS and GLM-4.5-Air are the natural fit here. This site has no measured run on a Spark and publishes no tok/s figure for one; measure your own with llama-bench.
Setup: Ollama or llama.cpp
The Spark is an Arm machine (arm64), so anything you install must have an arm64 build. Ollama's GPU documentation lists the GB10 (DGX Spark) as compute capability 12.1, and its standard Linux installer handles arm64. For llama.cpp, build with GGML_CUDA=ON: its CMake adds the GB10's native architecture only when the CUDA toolkit is 12.9 or newer. The prebuilt :server-cuda13 container image is published for arm64; the older :server-cuda image is built against CUDA 12.8 and lacks GB10 native kernels. Bind servers to 127.0.0.1 unless you mean to serve your network.
# Ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama run gpt-oss:20b # 64GB and 128GB
ollama run gpt-oss:120b # 128GB only
# llama.cpp, built on the Spark itself
nvcc --version # 12.9 or newer for GB10 native kernels
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j
./build/bin/llama-server -hf ggml-org/gpt-oss-120b-GGUF --jinja \
-ngl all -c 16384 --host 127.0.0.1 --port 8080 # 128GB modelCheck it actually ran on the GPU
Unified memory changes the usual checks. nvidia-smi lists processes but reports no memory-usage figure on this integrated GPU, so it cannot tell you whether a model fits. With Ollama, ollama ps shows the split in its PROCESSOR column — 100% GPU is what you want. With llama.cpp, the startup log should report every layer offloaded. One habit from discrete cards does not carry over: --n-cpu-moe keeps experts in system RAM to save VRAM, but on a Spark that is the same memory, so it saves nothing and only slows generation.
ollama ps
NAME SIZE PROCESSOR CONTEXT
gpt-oss:20b ... 100% GPU ...
llama.cpp log:
load_tensors: offloaded 37/37 layers to GPU
Fails to load on 64GB with a 100B-class model
→ expected; see the table above (use Q3_K_M, or the 128GB model)
"no kernel image is available"
→ llama.cpp built against CUDA older than 12.9; rebuild, or use :server-cuda13Common questions
Can the 64GB DGX Spark run GPT-OSS 120B?
No. This site sizes GPT-OSS 120B at MXFP4 with an 8K context at about 66.5 GB, more than the machine's 64 GB in total, and the Q4_K_M conversion at about 62.9 GB still leaves DGX OS — which shares the same memory — nothing to run in. GPT-OSS 20B fits easily, and the 128GB model runs the 120B with about half its memory to spare.
Should I buy the 64GB or the 128GB DGX Spark?
Decide by the largest model you want. At 4K context the 64GB model fits 82 of the 90 models in this index comfortably and the 128GB fits 84, with the same chip and the same 273 GB/s, so neither is faster. The ones only the 128GB holds are the 100B-class mixture-of-experts models at 4-bit — GPT-OSS 120B, GLM-4.5-Air and Llama 4 Scout at Q4_K_M — plus long context on 70B.
Is a DGX Spark faster than an RTX 5090 for local LLMs?
Not on any model both can hold. Token generation is bound by memory bandwidth, and the 5090 moves 1,792 GB/s against the Spark's 273, about 6.6× more. The Spark's advantage is capacity: 70B dense models and 100B-class mixture-of-experts models that no 32GB card holds at all.
What this guide uses
- Hardware
- DGX Spark 64G64GB
Next steps
See everything that fits 🟩 DGX Spark 64GReverse lookup — 64GB at 8192 context, ranked by qualityModels covered in this guide
Did this actually run?
Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.
Related guides
Run GPT-OSS 20B (and 120B) locally without re-quantizing
GPT-OSS ships natively in MXFP4, so the usual "download the Q4_K_M" habit makes it bigger and worse. Sizing, the right flags, and how MoE expert-offload puts the 120B on a 24GB card.
Intel Arc: llama.cpp (SYCL) or Ollama (Vulkan)
Run GGUF models on an Arc B580, B570, A770 or A750 — the SYCL build, the Ollama route, how to prove the GPU is doing the work, and what fits in 8–16 GB.
RTX 5090: What 32GB Actually Buys for Local LLMs
What fits on an RTX 5090 that does not fit on a 24GB card: not many more models, but much longer context and higher speed ceilings. Sizes from the calculator, setup for Blackwell.