Which cards this covers
Seven NVIDIA cards in this index have 12GB: the RTX 3060 12G, RTX 4070, 4070 Super and 4070 Ti, RTX 5070, and RTX 3080 12G and 3080 Ti. They fit exactly the same models — only capacity decides what fits — but they do not run them at the same speed, because token generation is limited by memory bandwidth, and that ranges from 360 GB/s on the 3060 to 912 GB/s on the 3080s. A 12GB Radeon or an Arc B580 fits the same list too; the AMD and Intel Arc guides cover their runtimes.
Card Bandwidth ceiling for Qwen3 8B Q4_K_M
RTX 3060 12G 360 GB/s 76 tok/s
RTX 4070 / Super / Ti 504 GB/s 107 tok/s
RTX 5070 672 GB/s 142 tok/s
RTX 3080 12G / Ti 912 GB/s 193 tok/s
Ceilings = bandwidth ÷ weight file size. Not measured here; real speed is lower.What fits, and where 14B stops fitting
48 of the 89 models in this index fit a 12GB card comfortably at 4K context, and 50 load at all. 7B–12B models are the comfortable zone, and the context length is what decides the rest. Gemma 3 12B at Q4_K_M with a 32K context is about 10.4 GB, still comfortable, because most of its layers use a short sliding window. Mistral Nemo 12B at Q4_K_M with a 32K context is about 13.2 GB, which does not fit. 14B is the edge: Qwen3 14B at Q4_K_M with a 4K context is about 10.0 GB, comfortable, but Qwen3 14B at Q4_K_M with an 8K context is about 10.7 GB — 89% of the card, past this site’s comfortable line — and with 16K it no longer fits.
GGUF Q4_K_M 8K 16K 32K
Llama 3.1 8B 6.2 GB 7.3 GB 9.5 GB
Qwen3 8B 6.4 GB 7.7 GB 10.1 GB
Gemma 3 12B 8.8 GB 9.3 GB 10.4 GB
Mistral Nemo 12B 9.1 GB 10.5 GB 13.2 GB ✗
Qwen3 14B 10.7 GB ~ 12.1 GB ✗ 14.9 GB ✗
Phi-4 14B 11.0 GB ~ 12.8 GB ✗ (16K max)
~ = tight (88–105% of the card) ✗ = does not fitRun it with Ollama
Ollama is the shortest path on Windows or Linux. Pulling straight from a Hugging Face GGUF repo pins the exact quant, which matters here: the difference between Q4_K_M and Q8_0 on a 12B model is the difference between fitting and not. Ollama sets its own context length; raise it inside the session when you need more, and check the table above before you do.
ollama run hf.co/Qwen/Qwen3-8B-GGUF:Q4_K_M
# inside the session, for a longer context:
/set parameter num_ctx 16384
ollama ps # PROCESSOR should read 100% GPUOr run it with llama.cpp
With llama.cpp, pass -ngl all and an explicit -c. Recent builds fit the model to the card automatically and keep about 1 GiB free, so a model near the limit silently loses layers to the CPU; asking for every layer makes a model that does not fit fail loudly instead of running slowly. Bind to 127.0.0.1 unless you mean to serve your network.
pip install -U huggingface_hub
hf download Qwen/Qwen3-8B-GGUF --include "Qwen3-8B-Q4_K_M.gguf" --local-dir ./models
./build/bin/llama-server -m ./models/Qwen3-8B-Q4_K_M.gguf \
-ngl all -c 16384 --host 127.0.0.1 --port 8080
# startup log: load_tensors: offloaded 37/37 layers to GPUCheck it is really on the GPU
Three checks, any one of which is enough: ollama ps shows 100% GPU; llama.cpp’s startup log reports every layer offloaded; nvidia-smi shows the memory held while the model is loaded. On Windows, watch for the slow-but-working case: the NVIDIA control panel’s System Memory Fallback is on by default and spills VRAM into system RAM instead of failing, so a model that does not fit turns into a model that crawls. Set CUDA Sysmem Fallback Policy to prefer no fallback while you are sizing things.
nvidia-smi --query-gpu=name,memory.used,memory.total --format=csv
# measure speed on your own card:
./build/bin/llama-bench -m ./models/Qwen3-8B-Q4_K_M.gguf -ngl 99 # pp512 / tg128When it does not work
Most problems on a 12GB card are the context, not the model.
Out of memory after a while, not at load
→ the KV cache grows with the conversation; lower -c / num_ctx
Loads, but generates at a few tok/s
→ layers on the CPU: check ollama ps / the load_tensors line
→ Windows: System Memory Fallback is hiding an out-of-memory
14B fits at 4K, fails at 16K
→ expected (see the table); drop to 8K, or use a 12B model
A 20B MoE model "almost" fits
→ keep some experts on the CPU (--n-cpu-moe), see the GPT-OSS guideCommon questions
Can a 12GB GPU run a 14B model?
Yes, at Q4_K_M with a short context. This site sizes Qwen3 14B at Q4_K_M at about 10.0 GB at 4K context, which is comfortable, and about 10.7 GB at 8K, which is tight. At 16K it no longer fits. If you need a long context on a 12GB card, Gemma 3 12B is the larger model that stays comfortable at 32K, because most of its layers use a sliding window.
RTX 3060 12G or RTX 4060 Ti 16G for local LLMs?
The 16GB card fits more: 55 of the 89 models in this index comfortably at 4K, against 48 for any 12GB card, and it holds 14B models at long contexts that a 12GB card cannot. On speed the 3060 is not behind — its 360 GB/s of memory bandwidth is higher than the 4060 Ti's 288 GB/s, and bandwidth is what limits generation once a model fits. If the models you want already fit in 12GB, the extra 4GB buys nothing.
Can I run GPT-OSS 20B on a 12GB card?
Not entirely on the card: this site sizes GPT-OSS 20B at its native MXFP4 at about 12.8 GB at 4K context, just over the limit. Because it is a mixture-of-experts model, it runs well with some experts kept in system RAM using llama.cpp's --n-cpu-moe option — the GPT-OSS guide covers the settings.
What this guide uses
- Picks for it
- Best local LLM for 12GB
Next steps
See everything that fits 🟢 RTX 3060 12GReverse lookup — 12GB at 8192 context, ranked by qualityModels covered in this guide
Did this actually run?
Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.
Related guides
8GB GPU Starter Guide: 3060 / 4060 / 3070
The most common local LLM hardware tier — which models, quants, and context lengths actually fit in 8GB VRAM.
Run GPT-OSS 20B (and 120B) locally without re-quantizing
GPT-OSS ships natively in MXFP4, so the usual "download the Q4_K_M" habit makes it bigger and worse. Sizing, the right flags, and how MoE expert-offload puts the 120B on a 24GB card.
Intel Arc: llama.cpp (SYCL) or Ollama (Vulkan)
Run GGUF models on an Arc B580, B570, A770 or A750 — the SYCL build, the Ollama route, how to prove the GPU is doing the work, and what fits in 8–16 GB.