The honest answer: not many more models
At 4K context, 73 of the 89 models in this index fit an RTX 5090 comfortably, against 71 for a 24GB RTX 4090. The models that fit on neither are the 70B-and-up class: Llama 3.3 70B at Q4_K_M with an 8K context is about 47.5 GB, so a single 5090 does not change that — it takes two cards, or offloading layers to system RAM. What the extra 8GB does buy is room for context on the 27B–36B models that a 24GB card can only run with a short one, and the bandwidth — 1,792 GB/s against 1,008 — raises the speed ceiling on everything by about 1.8×.
Where the extra 8GB goes: context
The table below is where 24GB and 32GB actually differ. Qwen3 32B at Q4_K_M with an 8K context is about 22.9 GB, which is comfortable here and tight on a 24GB card. Gemma 4 31B at Q4_K_M with a 32K context is about 23.5 GB: comfortable on the 5090, 98% of a 4090. Qwen3.8 27B at Q4_K_M with a 128K context is about 26.0 GB, which a 24GB card cannot hold at all — it stays this small because only a quarter of its layers keep a growing cache. Dense 32B models still run out at long context: Qwen3 32B at Q4_K_M with a 32K context is about 29.6 GB, tight even here.
GGUF Q4_K_M 8K 32K 64K 128K
Qwen3.8 27B (hybrid) 17.8 GB 19.4 GB 21.6 GB 26.0 GB
Gemma 3 27B 18.7 GB 20.8 GB 23.5 GB 29.0 GB ~
Gemma 4 31B 21.4 GB 23.5 GB 26.2 GB 31.7 GB ~
Qwen3-Coder 30B-A3B 20.1 GB 22.6 GB 25.9 GB 32.5 GB ~
Qwen3 32B (40K max) 22.9 GB 29.6 GB ~
Seed-OSS 36B 25.0 GB 31.6 GB ~ 40.4 GB ✗
~ = tight (88–105% of 32 GB) ✗ = does not fitBlackwell needs a current stack
The RTX 50 series is a new architecture (compute capability 12.0), and NVIDIA’s CUDA 12.8 is the first toolkit that supports it. In practice that means: a current driver, a current Ollama, and if you build llama.cpp yourself, a CUDA 12.8 or newer toolkit. An older build either fails to load the CUDA backend or falls back to the CPU — which is why the GPU check below matters more on a brand-new card than on an old one.
nvidia-smi # driver version, and the CUDA version it supports
nvcc --version # if building llama.cpp: needs 12.8 or newer
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -jRun a 32B model with a real context
Pass -ngl all and an explicit -c to llama.cpp. Recent builds fit the model to the card automatically and keep about 1 GiB free, which on a model near the limit quietly moves layers to the CPU; asking for every layer makes an oversized setting fail at load instead. Bind to 127.0.0.1 unless you intend to serve your network. With Ollama, pull the exact quant from the GGUF repo and raise num_ctx inside the session.
hf download Qwen/Qwen3-32B-GGUF --include "Qwen3-32B-Q4_K_M.gguf" --local-dir ./models
./build/bin/llama-server -m ./models/Qwen3-32B-Q4_K_M.gguf \
-ngl all -c 16384 --host 127.0.0.1 --port 8080
# load_tensors: offloaded 65/65 layers to GPU
# or with Ollama:
ollama run hf.co/Qwen/Qwen3-32B-GGUF:Q4_K_M
/set parameter num_ctx 16384What the speed should look like
Generation reads the whole weight file once per token, so bandwidth divided by file size is a hard ceiling. At 1,792 GB/s that puts Qwen3 32B at Q4_K_M under about 95 tok/s and Gemma 3 27B under about 114. These are ceilings, not measurements — this site has no runs on a 5090, and real throughput lands below them. Mixture-of-experts models such as Qwen3-Coder 30B-A3B read only their active experts per token, so they run well above what their total size suggests. Measure your own with llama-bench.
./build/bin/llama-bench -m ./models/Qwen3-32B-Q4_K_M.gguf -ngl 99 # pp512 / tg128When it does not work
On a new card most failures are the software stack, not the model.
CUDA backend fails to load, or "no kernel image is available"
→ toolkit / build older than CUDA 12.8; rebuild, or update Ollama
Runs, but at CPU speed
→ check ollama ps (100% GPU) or the load_tensors line
→ Windows: System Memory Fallback is hiding an out-of-memory
32B fine at 8K, fails at 32K
→ expected for dense 32B (29.6 GB); use 16K, or a hybrid model like Qwen3.8 27B
Want 70B
→ does not fit one 32 GB card at Q4_K_M; see the dual-GPU guideCommon questions
Is the RTX 5090 worth it over the RTX 4090 for local LLMs?
Only if you need the context or the speed. At 4K context the 5090 fits 73 of the 89 models in this index comfortably and the 4090 fits 71, so the list of models barely changes. What changes is long context on 27B–32B models — Gemma 4 31B at 32K is comfortable on the 5090 and 98% of a 4090 — and the speed ceiling, which rises with bandwidth from 1,008 to 1,792 GB/s.
Can an RTX 5090 run a 70B model?
Not on the card alone at a usable quality. This site sizes Llama 3.3 70B at Q4_K_M at about 47.5 GB with an 8K context, half again more than the card holds. Two cards, or offloading part of the model to system RAM at a large speed cost, are the realistic routes — the dual-GPU guide covers the first.
What do I need to run llama.cpp on an RTX 5090?
A driver and toolkit that know Blackwell: NVIDIA's CUDA 12.8 is the first toolkit to support the RTX 50 series, so build llama.cpp with GGML_CUDA=ON against 12.8 or newer, or use a current Ollama release. Then confirm the startup log reports every layer offloaded to the GPU — on a new card, a stack that is too old can fall back to the CPU without an obvious error.
What this guide uses
- Hardware
- RTX 509032GB
- Picks for it
- Best local LLM for 32GB
Next steps
See everything that fits 🟢 RTX 5090Reverse lookup — 32GB at 32768 context, ranked by qualityModels covered in this guide
Did this actually run?
Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.
Related guides
Run GPT-OSS 20B (and 120B) locally without re-quantizing
GPT-OSS ships natively in MXFP4, so the usual "download the Q4_K_M" habit makes it bigger and worse. Sizing, the right flags, and how MoE expert-offload puts the 120B on a 24GB card.
Intel Arc: llama.cpp (SYCL) or Ollama (Vulkan)
Run GGUF models on an Arc B580, B570, A770 or A750 — the SYCL build, the Ollama route, how to prove the GPU is doing the work, and what fits in 8–16 GB.
What to Run on a 12GB GPU (RTX 3060 12G, 4070, 5070)
Which models fit a 12GB NVIDIA card at 8K–32K context, where 14B stops fitting, how to run them with Ollama or llama.cpp, and how to prove they are on the GPU.