16GB VRAM

RTX 4080 Super — what LLMs can it run?

51 of 79 indexed models fit comfortably in 16GB at 4K context, each at the highest-quality quant that still leaves headroom.

32B · 5

RTX 4080 Super — what LLMs can it run?32B
ModelQuantEst. VRAMHeadroomtok/s
Mistral Small 24B Instruct24BAWQ INT413.23 GB+2.8 GB78
Devstral Small 1.1 24B24BAWQ INT413.02 GB+3 GB78
Magistral Small 1.2 24B24BAWQ INT413.02 GB+3 GB76
Codestral 22B22BAWQ INT412.56 GB+3.4 GB72
GPT-OSS 20B21B MoEMXFP411.81 GB+4.2 GB195

14B · 17

RTX 4080 Super — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
InternLM2 20B Chat20BQ4_K_M13.43 GB+2.6 GB78
DeepSeek-Coder-V2-Lite Instruct16BQ4_K_M10.87 GB+5.1 GB145
DeepSeek-V2-Lite Chat16BQ4_K_M10.87 GB+5.1 GB142
StarCoder2 15B15BQ4_K_M10.19 GB+5.8 GB92
Qwen3 14B Instruct14BQ5_K_M11.65 GB+4.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+4.3 GB86
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+6.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+4.2 GB78
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+3.1 GB88
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+4.9 GB95
Gemma 3 12B IT12BQ5_K_M10.6 GB+5.4 GB92
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+7.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+7.9 GB95
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+3.1 GB72
Solar 10.7B Instruct11BQ4_K_M7.6 GB+8.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+8.7 GB118
Gemma 2 9B Instruct9BQ8_011.7 GB+4.3 GB108

7B · 19

RTX 4080 Super — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+5.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+5.6 GB108
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+10.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+10.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+8.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+6.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+10.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+10.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+10.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+9.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+10.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+9.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+11.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+10.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+11.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+5.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+9.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+9.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+10.6 GB145

≤3B · 10

RTX 4080 Super — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+12 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+11.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+10.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+12 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+12.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+12.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+13.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+14.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+14.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+15.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

Stepping up

A RTX 4090 (24GB) fits 12 more of the indexed models than this card. RTX 4090

51 of 79 indexed models fit comfortably in 16GB at 4K context, each at the highest-quality quant that still leaves headroom.