IntermediateEdge / Local 4 min readPublished

RTX 5090: What 32GB Actually Buys for Local LLMs

What fits on an RTX 5090 that does not fit on a 24GB card: not many more models, but much longer context and higher speed ceilings. Sizes from the calculator, setup for Blackwell.

Written against

Windows or Linux · NVIDIA driver with CUDA 12.8+ · current Ollama, or llama.cpp built with GGML_CUDA=ON against CUDA 12.8+ · GGUF

Not re-run since it was last edited — treat the commands as a starting point, not a tested recipe.

RTX 509032GBBlackwellllama.cppOllamaGGUFlong context

The honest answer: not many more models

At 4K context, 73 of the 89 models in this index fit an RTX 5090 comfortably, against 71 for a 24GB RTX 4090. The models that fit on neither are the 70B-and-up class: Llama 3.3 70B at Q4_K_M with an 8K context is about 47.5 GB, so a single 5090 does not change that — it takes two cards, or offloading layers to system RAM. What the extra 8GB does buy is room for context on the 27B–36B models that a 24GB card can only run with a short one, and the bandwidth — 1,792 GB/s against 1,008 — raises the speed ceiling on everything by about 1.8×.

Where the extra 8GB goes: context

The table below is where 24GB and 32GB actually differ. Qwen3 32B at Q4_K_M with an 8K context is about 22.9 GB, which is comfortable here and tight on a 24GB card. Gemma 4 31B at Q4_K_M with a 32K context is about 23.5 GB: comfortable on the 5090, 98% of a 4090. Qwen3.8 27B at Q4_K_M with a 128K context is about 26.0 GB, which a 24GB card cannot hold at all — it stays this small because only a quarter of its layers keep a growing cache. Dense 32B models still run out at long context: Qwen3 32B at Q4_K_M with a 32K context is about 29.6 GB, tight even here.

text
GGUF Q4_K_M               8K        32K       64K       128K
Qwen3.8 27B (hybrid)      17.8 GB   19.4 GB   21.6 GB   26.0 GB
Gemma 3 27B               18.7 GB   20.8 GB   23.5 GB   29.0 GB ~
Gemma 4 31B               21.4 GB   23.5 GB   26.2 GB   31.7 GB ~
Qwen3-Coder 30B-A3B       20.1 GB   22.6 GB   25.9 GB   32.5 GB ~
Qwen3 32B (40K max)       22.9 GB   29.6 GB ~
Seed-OSS 36B              25.0 GB   31.6 GB ~ 40.4 GB ✗

~ = tight (88–105% of 32 GB)   ✗ = does not fit

Blackwell needs a current stack

The RTX 50 series is a new architecture (compute capability 12.0), and NVIDIA’s CUDA 12.8 is the first toolkit that supports it. In practice that means: a current driver, a current Ollama, and if you build llama.cpp yourself, a CUDA 12.8 or newer toolkit. An older build either fails to load the CUDA backend or falls back to the CPU — which is why the GPU check below matters more on a brand-new card than on an old one.

bash
nvidia-smi                 # driver version, and the CUDA version it supports
nvcc --version             # if building llama.cpp: needs 12.8 or newer

git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j

Run a 32B model with a real context

Pass -ngl all and an explicit -c to llama.cpp. Recent builds fit the model to the card automatically and keep about 1 GiB free, which on a model near the limit quietly moves layers to the CPU; asking for every layer makes an oversized setting fail at load instead. Bind to 127.0.0.1 unless you intend to serve your network. With Ollama, pull the exact quant from the GGUF repo and raise num_ctx inside the session.

bash
hf download Qwen/Qwen3-32B-GGUF --include "Qwen3-32B-Q4_K_M.gguf" --local-dir ./models

./build/bin/llama-server -m ./models/Qwen3-32B-Q4_K_M.gguf \
  -ngl all -c 16384 --host 127.0.0.1 --port 8080
#   load_tensors: offloaded 65/65 layers to GPU

# or with Ollama:
ollama run hf.co/Qwen/Qwen3-32B-GGUF:Q4_K_M
/set parameter num_ctx 16384

What the speed should look like

Generation reads the whole weight file once per token, so bandwidth divided by file size is a hard ceiling. At 1,792 GB/s that puts Qwen3 32B at Q4_K_M under about 95 tok/s and Gemma 3 27B under about 114. These are ceilings, not measurements — this site has no runs on a 5090, and real throughput lands below them. Mixture-of-experts models such as Qwen3-Coder 30B-A3B read only their active experts per token, so they run well above what their total size suggests. Measure your own with llama-bench.

bash
./build/bin/llama-bench -m ./models/Qwen3-32B-Q4_K_M.gguf -ngl 99   # pp512 / tg128

When it does not work

On a new card most failures are the software stack, not the model.

text
CUDA backend fails to load, or "no kernel image is available"
  → toolkit / build older than CUDA 12.8; rebuild, or update Ollama

Runs, but at CPU speed
  → check ollama ps (100% GPU) or the load_tensors line
  → Windows: System Memory Fallback is hiding an out-of-memory

32B fine at 8K, fails at 32K
  → expected for dense 32B (29.6 GB); use 16K, or a hybrid model like Qwen3.8 27B

Want 70B
  → does not fit one 32 GB card at Q4_K_M; see the dual-GPU guide

Common questions

Is the RTX 5090 worth it over the RTX 4090 for local LLMs?

Only if you need the context or the speed. At 4K context the 5090 fits 73 of the 89 models in this index comfortably and the 4090 fits 71, so the list of models barely changes. What changes is long context on 27B–32B models — Gemma 4 31B at 32K is comfortable on the 5090 and 98% of a 4090 — and the speed ceiling, which rises with bandwidth from 1,008 to 1,792 GB/s.

Can an RTX 5090 run a 70B model?

Not on the card alone at a usable quality. This site sizes Llama 3.3 70B at Q4_K_M at about 47.5 GB with an 8K context, half again more than the card holds. Two cards, or offloading part of the model to system RAM at a large speed cost, are the realistic routes — the dual-GPU guide covers the first.

What do I need to run llama.cpp on an RTX 5090?

A driver and toolkit that know Blackwell: NVIDIA's CUDA 12.8 is the first toolkit to support the RTX 50 series, so build llama.cpp with GGML_CUDA=ON against 12.8 or newer, or use a current Ollama release. Then confirm the startup log reports every layer offloaded to the GPU — on a new card, a stack that is too old can fall back to the CPU without an obvious error.

What this guide uses

Hardware
RTX 509032GB

Next steps

See everything that fits 🟢 RTX 5090Reverse lookup — 32GB at 32768 context, ranked by quality

Models covered in this guide

Did this actually run?

Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.

Related guides

Deployment guides are educational. Each model is subject to its own license — read the official Hugging Face model card before downloading or deploying.