IntermediateEdge / Local 5 min readPublished

DGX Spark 64GB vs 128GB: What Each One Runs

Which models fit the 64GB and 128GB DGX Spark, from the calculator: GPT-OSS 120B needs the 128GB model, 70B dense fits both, and 273 GB/s sets the speed. Setup for Ollama and llama.cpp.

Written against

DGX OS (Ubuntu-based, arm64) · current Ollama, or llama.cpp built with GGML_CUDA=ON against CUDA 12.9+ · GGUF

Not re-run since it was last edited — treat the commands as a starting point, not a tested recipe.

DGX SparkGB10unified memory64GB128GBGPT-OSSOllamallama.cpp

The short answer

Both DGX Spark models use the same GB10 chip and the same 273 GB/s memory; the only difference that matters for local models is capacity. NVIDIA announced the 64GB model at $4,999, on sale from 23 October 2026 through Acer, ASUS, Dell, GIGABYTE, HP and MSI. At 4K context, 82 of the 90 models in this index fit the 64GB Spark comfortably and 84 fit the 128GB one — against 74 for a 32GB RTX 5090. The gap between the two Sparks is narrow but it sits exactly on the models people buy a Spark for: the 100B-class mixture-of-experts releases.

GPT-OSS 120B does not fit the 64GB model

GPT-OSS 120B at MXFP4 with an 8K context is about 66.5 GB — more than the whole machine has, and the CPU and DGX OS live in that same pool. The Q4_K_M conversions do not rescue it: GPT-OSS 120B at Q4_K_M with an 8K context is about 62.9 GB, 98% of the pool, which leaves the operating system nothing. On the 128GB model the native build is about half the memory. GLM-4.5-Air and Llama 4 Scout follow the same pattern at Q4_K_M; on 64GB they need their Q3_K_M builds. Note that this site treats anything over the pool as not loading at all: on a discrete card a small overrun can spill into system RAM, but a Spark has no system RAM besides this.

text
GGUF                         8K        32K       128K     64GB   128GB
GPT-OSS 20B      MXFP4      12.9 GB   13.5 GB   16.0 GB   ok     ok
Llama 3.3 70B    Q4_K_M     47.5 GB   55.7 GB   88.7 GB   ok*    ok
GLM-4.5-Air      Q3_K_M     55.2 GB   59.9 GB   78.9 GB   ok†    ok
GLM-4.5-Air      Q4_K_M     68.7 GB   73.5 GB   92.5 GB   ✗      ok
Llama 4 Scout    Q4_K_M     70.7 GB   72.0 GB   77.0 GB   ✗      ok
GPT-OSS 120B     Q4_K_M     62.9 GB   63.8 GB   67.5 GB   ✗      ok
GPT-OSS 120B     MXFP4      66.5 GB   67.5 GB   71.2 GB   ✗      ok
Qwen3 235B-A22B  Q3_K_M    120.4 GB  125.3 GB     —       ✗      ✗

ok = within 88% of the pool   * up to the 32K column   † 8K column only
✗ = over 88%; past 100% it cannot load at all

What 64GB is good for

The 64GB model is the dense-70B machine. Llama 3.3 70B at Q4_K_M with a 32K context is about 55.7 GB, inside the comfortable line, and no single consumer card holds it at all. GPT-OSS 20B at MXFP4 with a 128K context is about 16.0 GB, so it runs its full window with most of the pool to spare. GLM-4.5-Air at Q3_K_M with an 8K context is about 55.2 GB, which is the largest mixture-of-experts option, at a lower quant than you would choose with room to spare. If the plan is GPT-OSS 120B, GLM-4.5-Air at Q4_K_M or long context on 70B, buy the 128GB model.

Speed: 273 GB/s is the ceiling for both

Generation reads the active weights once per token, so bandwidth divided by the bytes read is a hard ceiling. For a dense model that is the whole file: Llama 3.3 70B at Q4_K_M tops out around 7 tok/s on either Spark. That is the trade the machine makes — an RTX 5090 moves 1,792 GB/s, about 6.6× more, on anything it can hold. Mixture-of-experts models read only their active experts, which is why GPT-OSS and GLM-4.5-Air are the natural fit here. This site has no measured run on a Spark and publishes no tok/s figure for one; measure your own with llama-bench.

Setup: Ollama or llama.cpp

The Spark is an Arm machine (arm64), so anything you install must have an arm64 build. Ollama's GPU documentation lists the GB10 (DGX Spark) as compute capability 12.1, and its standard Linux installer handles arm64. For llama.cpp, build with GGML_CUDA=ON: its CMake adds the GB10's native architecture only when the CUDA toolkit is 12.9 or newer. The prebuilt :server-cuda13 container image is published for arm64; the older :server-cuda image is built against CUDA 12.8 and lacks GB10 native kernels. Bind servers to 127.0.0.1 unless you mean to serve your network.

bash
# Ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama run gpt-oss:20b          # 64GB and 128GB
ollama run gpt-oss:120b         # 128GB only

# llama.cpp, built on the Spark itself
nvcc --version                  # 12.9 or newer for GB10 native kernels
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

./build/bin/llama-server -hf ggml-org/gpt-oss-120b-GGUF --jinja \
  -ngl all -c 16384 --host 127.0.0.1 --port 8080   # 128GB model

Check it actually ran on the GPU

Unified memory changes the usual checks. nvidia-smi lists processes but reports no memory-usage figure on this integrated GPU, so it cannot tell you whether a model fits. With Ollama, ollama ps shows the split in its PROCESSOR column — 100% GPU is what you want. With llama.cpp, the startup log should report every layer offloaded. One habit from discrete cards does not carry over: --n-cpu-moe keeps experts in system RAM to save VRAM, but on a Spark that is the same memory, so it saves nothing and only slows generation.

text
ollama ps
  NAME            SIZE     PROCESSOR    CONTEXT
  gpt-oss:20b     ...      100% GPU     ...

llama.cpp log:
  load_tensors: offloaded 37/37 layers to GPU

Fails to load on 64GB with a 100B-class model
  → expected; see the table above (use Q3_K_M, or the 128GB model)

"no kernel image is available"
  → llama.cpp built against CUDA older than 12.9; rebuild, or use :server-cuda13

Common questions

Can the 64GB DGX Spark run GPT-OSS 120B?

No. This site sizes GPT-OSS 120B at MXFP4 with an 8K context at about 66.5 GB, more than the machine's 64 GB in total, and the Q4_K_M conversion at about 62.9 GB still leaves DGX OS — which shares the same memory — nothing to run in. GPT-OSS 20B fits easily, and the 128GB model runs the 120B with about half its memory to spare.

Should I buy the 64GB or the 128GB DGX Spark?

Decide by the largest model you want. At 4K context the 64GB model fits 82 of the 90 models in this index comfortably and the 128GB fits 84, with the same chip and the same 273 GB/s, so neither is faster. The ones only the 128GB holds are the 100B-class mixture-of-experts models at 4-bit — GPT-OSS 120B, GLM-4.5-Air and Llama 4 Scout at Q4_K_M — plus long context on 70B.

Is a DGX Spark faster than an RTX 5090 for local LLMs?

Not on any model both can hold. Token generation is bound by memory bandwidth, and the 5090 moves 1,792 GB/s against the Spark's 273, about 6.6× more. The Spark's advantage is capacity: 70B dense models and 100B-class mixture-of-experts models that no 32GB card holds at all.

What this guide uses

Format
GGUFAWQ

Next steps

See everything that fits 🟩 DGX Spark 64GReverse lookup — 64GB at 8192 context, ranked by quality

Models covered in this guide

Did this actually run?

Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.

Related guides

Deployment guides are educational. Each model is subject to its own license — read the official Hugging Face model card before downloading or deploying.