12GB VRAM

RTX 3060 12G — what LLMs can it run?

47 of 87 indexed models fit comfortably in 12GB at 4K context, each at the highest-quality quant that still leaves headroom.

The short answer for RTX 3060 12G

12 GB, 360 GB/s GDDR6. 47 of 87 models in this index fit comfortably at 4K context; 49 load at all.

Biggest that fits
DeepSeek-Coder-V2-Lite Instruct — 10.1 GB at Q4_K_M — 1.9 GB spare at 4K, so longer context comes out of a thin margin.
Room to grow
Qwen2-VL 7B Instruct — 5.5 GB at Q4_K_M — under 60% of the card, which leaves 6.5 GB for a long context window or a second process.
Speed ceiling
Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 360 GB/s puts a hard ceiling near 78 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
The wall you will hit
Bandwidth, not capacity. The models on this page fit, but at 360 GB/s this card reads the whole weight set once per generated token — a 912 GB/s RTX 3080 Ti holds exactly the same 12 GB and moves those bytes 2.5× faster. Expect the same quant to generate proportionally slower here.
Best local LLM for 12GB

14B · 15

RTX 3060 12G — what LLMs can it run? — 14B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
DeepSeek-Coder-V2-Lite Instruct16B-A2.4BQ4_K_M10.08 GB+1.9 GB145Estimated
DeepSeek-V2-Lite Chat16B-A2.4BQ4_K_M10.08 GB+1.9 GB142Estimated
StarCoder2 15B15BQ4_K_M10.19 GB+1.8 GB92Estimated
Qwen3 14B Instruct14BEXL2 4.65bpw9.66 GB+2.3 GB125
Qwen2.5 14B Instruct14BEXL2 4.65bpw9.75 GB+2.3 GB—
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+2.3 GB128
Phi-4 14B14BQ4_K_M10.17 GB+1.8 GB88Community
Phi-3 Medium 14B Instruct14BQ4_K_M9.73 GB+2.3 GB102Estimated
Mistral Nemo 12B Instruct12BQ4_K_M8.42 GB+3.6 GB112Estimated
Gemma 3 12B IT12BQ5_K_M9.84 GB+2.2 GB92Estimated
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+3.7 GB108Estimated
Llama 3.2 11B Vision Instruct11BQ4_K_M7.66 GB+4.3 GB88Estimated
Solar 10.7B Instruct11BQ4_K_M7.6 GB+4.4 GB125Estimated
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+4.7 GB118Estimated
Gemma 2 9B Instruct9BQ4_K_M7.3 GB+4.7 GB132Estimated

7B · 22

RTX 3060 12G — what LLMs can it run? — 7B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
GLM-4-9B-Chat9BQ8_010.16 GB+1.8 GB105Estimated
Qwen3-VL 8B Instruct8BQ8_010.39 GB+1.6 GB108Estimated
Ministral 3 8B Instruct8BQ8_010.02 GB+2 GB—
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+6.5 GB72Estimated
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+6.1 GB142Estimated
Qwen3 8B Instruct8BQ6_K7.64 GB+4.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+2.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+6.6 GB232Estimated
Aya 23 8B8BQ4_K_M5.64 GB+6.4 GB145Estimated
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+6.6 GB228Estimated
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+5.5 GB128Estimated
InternLM2 7B Chat7BQ4_K_M5.45 GB+6.6 GB148Estimated
Qwen2.5 7B Instruct7BQ6_K6.77 GB+5.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+7.1 GB248Estimated
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+7.1 GB210Estimated
Gemma 4 E4B ITE4BQ8_08.46 GB+3.5 GB—
OLMo 2 7B Instruct7BQ8_010.3 GB+1.7 GB125Estimated
WizardLM-2 7B7BQ4_K_M5.14 GB+6.9 GB152Estimated
Mistral 7B Instruct v0.37BQ6_K6.75 GB+5.3 GB135Estimated
Zephyr 7B Beta7BQ6_K6.75 GB+5.3 GB132Estimated
Gemma 4 E2B ITE2BQ8_05.2 GB+6.8 GB—
Gemma 3 4B IT4BQ8_05.05 GB+7 GB145Estimated

≤3B · 10

RTX 3060 12G — what LLMs can it run? — ≤3B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
Qwen3 4B Instruct4BQ6_K4.05 GB+8 GB145
Phi-4 Mini Instruct3.8BQ8_04.81 GB+7.2 GB262Estimated
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+6.1 GB255Estimated
Llama 3.2 3B Instruct3BQ8_04.05 GB+8 GB285Estimated
Qwen2.5 3B Instruct3BQ8_03.59 GB+8.4 GB290Estimated
Gemma 2 2B Instruct2BQ8_03.34 GB+8.7 GB320Estimated
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+9.6 GB240Estimated
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+10.2 GB410Estimated
Llama 3.2 1B Instruct1BQ8_01.51 GB+10.5 GB450Estimated
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+11.4 GB540Estimated

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 2 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a RTX 3060 12G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.

Cards with the same budget

What fits is decided by memory, so every 12GB card of this type returns the same list. These pages are not different answers — they differ in throughput, which this index does not measure per card.

Stepping up

A RTX 5080 (16GB) fits 6 more of the indexed models than this card. RTX 5080 →

Common questions

What is the best local LLM for an RTX 3060 12G?

For everyday use, Qwen2-VL 7B Instruct at Q4_K_M — about 5.5 GB of the card's 12 GB at 4K context, leaving 6.5 GB for a longer window. If you want the largest thing that will load, that is DeepSeek-Coder-V2-Lite Instruct at Q4_K_M (10.1 GB, 1.9 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.

Can an RTX 3060 12G run InternLM2 20B Chat?

No. Its smallest build here, Q4_K_M, needs about 13.4 GB and this card has 12.0 GB usable — short by 1.4 GB before any context beyond 4K. The largest model this card does clear is DeepSeek-Coder-V2-Lite Instruct.

How many tokens per second does an RTX 3060 12G do on an 8B model at Q4?

This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 360 GB/s — a ceiling near 78 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.

RTX 3060 12G or RTX 3080 Ti for local LLMs?

They hold the same models: 12 GB against 12 GB fits 47 and 47 of 87 respectively at 4K. The difference is throughput — 360 GB/s against 912 GB/s, a 2.5× gap in how fast the weights can be read, which is what token generation is bound by. The RTX 3080 Ti generates faster on any model both can hold.

47 of 87 indexed models fit comfortably in 12GB at 4K context, each at the highest-quality quant that still leaves headroom.

This page's figures change when the model or the runtime does.Last updated 2026-10-02 RSS → /feed.xml