RTX 5050 — what LLMs can it run?
36 of 87 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.
The short answer for RTX 5050
8 GB, bandwidth not listed. 36 of 87 models in this index fit comfortably at 4K context; 43 load at all.
- Biggest that fits
- Stable LM 2 12B Chat — 7.0 GB at AWQ INT4 — 1.0 GB spare at 4K, so longer context comes out of a thin margin.
- Room to grow
- Qwen3 4B Instruct — 4.0 GB at Q6_K — under 60% of the card, which leaves 4.0 GB for a long context window or a second process.
- The wall you will hit
- Capacity. 36 of 87 models clear 8 GB comfortably at 4K context, and the ones that do not are short by whole gigabytes rather than by a rounding error. Moving up a size class is what changes the list, not a different quant.
14B · 4
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Stable LM 2 12B Chat12B | AWQ INT4 | 7.01 GB | +1 GB | 142Estimated |
| Solar 10.7B Instruct11B | AWQ INT4 | 6.42 GB | +1.6 GB | 168Estimated |
| Falcon 3 10B Instruct10B | GPTQ INT4 | 6.12 GB | +1.9 GB | 155Estimated |
| Gemma 2 9B Instruct9B | AWQ INT4 | 6.27 GB | +1.7 GB | 188Estimated |
7B · 22
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| GLM-4-9B-Chat9B | Q4_K_M | 5.87 GB | +2.1 GB | 135Estimated |
| Qwen3-VL 8B Instruct8B | Q4_K_M | 6.19 GB | +1.8 GB | 140Community |
| Ministral 3 8B Instruct8B | Q5_K_M | 6.92 GB | +1.1 GB | — |
| Qwen2-VL 7B Instruct7B | Q4_K_M | 5.49 GB | +2.5 GB | 72Estimated |
| Granite 3.1 8B Instruct8B | Q4_K_M | 5.88 GB | +2.1 GB | 142Estimated |
| Qwen3 8B Instruct8B | EXL2 4.65bpw | 5.59 GB | +2.4 GB | 228 |
| Llama 3.1 8B Instruct8B | EXL2 4.65bpw | 5.43 GB | +2.6 GB | 235 |
| Nous Hermes 3 Llama 3.1 8B8B | EXL2 4.65bpw | 5.43 GB | +2.6 GB | 232Estimated |
| Aya 23 8B8B | Q4_K_M | 5.64 GB | +2.4 GB | 145Estimated |
| OpenChat 3.6 8B8B | EXL2 4.65bpw | 5.43 GB | +2.6 GB | 228Estimated |
| DeepSeek-R1-Distill-Llama-8B8B | Q5_K_M | 6.51 GB | +1.5 GB | 128Estimated |
| InternLM2 7B Chat7B | Q4_K_M | 5.45 GB | +2.6 GB | 148Estimated |
| Qwen2.5 7B Instruct7B | Q6_K | 6.77 GB | +1.2 GB | 132 |
| Qwen2.5-Coder 7B Instruct7B | EXL2 4.65bpw | 4.87 GB | +3.1 GB | 248Estimated |
| DeepSeek-R1-Distill-Qwen-7B7B | EXL2 4.65bpw | 4.87 GB | +3.1 GB | 210Estimated |
| Gemma 4 E4B ITE4B | Q4_K_M | 4.88 GB | +3.1 GB | — |
| OLMo 2 7B Instruct7B | Q4_K_M | 6.82 GB | +1.2 GB | 150Estimated |
| WizardLM-2 7B7B | Q4_K_M | 5.14 GB | +2.9 GB | 152Estimated |
| Mistral 7B Instruct v0.37B | Q6_K | 6.75 GB | +1.3 GB | 135Estimated |
| Zephyr 7B Beta7B | Q6_K | 6.75 GB | +1.3 GB | 132Estimated |
| Gemma 4 E2B ITE2B | Q8_0 | 5.2 GB | +2.8 GB | — |
| Gemma 3 4B IT4B | Q8_0 | 5.05 GB | +3 GB | 145Estimated |
≤3B · 10
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Qwen3 4B Instruct4B | Q6_K | 4.05 GB | +4 GB | 145 |
| Phi-4 Mini Instruct3.8B | Q8_0 | 4.81 GB | +3.2 GB | 262Estimated |
| Phi-3.5 Mini Instruct3.8B | Q8_0 | 5.89 GB | +2.1 GB | 255Estimated |
| Llama 3.2 3B Instruct3B | Q8_0 | 4.05 GB | +4 GB | 285Estimated |
| Qwen2.5 3B Instruct3B | Q8_0 | 3.59 GB | +4.4 GB | 290Estimated |
| Gemma 2 2B Instruct2B | Q8_0 | 3.34 GB | +4.7 GB | 320Estimated |
| Qwen3 1.7B Instruct1.7B | Q8_0 | 2.39 GB | +5.6 GB | 240Estimated |
| Qwen2.5 1.5B Instruct1.5B | Q8_0 | 1.83 GB | +6.2 GB | 410Estimated |
| Llama 3.2 1B Instruct1B | Q8_0 | 1.51 GB | +6.5 GB | 450Estimated |
| Qwen2.5 0.5B Instruct0.5B | Q8_0 | 0.6 GB | +7.4 GB | 540Estimated |
How this list is built
Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.
A further 7 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.
Measured on this card
No benchmark runs in this index were recorded on a RTX 5050. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.
Cards with the same budget
What fits is decided by memory, so every 8GB card of this type returns the same list. These pages are not different answers — they differ in throughput, which this index does not measure per card.
Stepping up
A RTX 3080 10G (10GB) fits 11 more of the indexed models than this card. RTX 3080 10G →
Common questions
What is the best local LLM for an RTX 5050?
For everyday use, Qwen3 4B Instruct at Q6_K — about 4.0 GB of the card's 8 GB at 4K context, leaving 4.0 GB for a longer window. If you want the largest thing that will load, that is Stable LM 2 12B Chat at AWQ INT4 (7.0 GB, 1.0 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.
Can an RTX 5050 run Mistral Nemo 12B Instruct?
Only just. Its smallest build here, AWQ INT4, needs about 7.1 GB against 8.0 GB usable — that loads with nothing else running, with no margin for a longer context window. It is not on the list above, which requires a model to stay inside 88% of usable capacity.
RTX 5050 or RTX 5060 Ti 8G for local LLMs?
They hold the same models: 8 GB against 8 GB fits 36 and 36 of 87 respectively at 4K. Bandwidth is not recorded for both cards here, so this page will not rank them on speed.
36 of 87 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.
This page's figures change when the model or the runtime does.Last updated 2026-10-02 RSS → /feed.xml