BeginnerEdge / Local 5 min readPublished

What to Run on a 12GB GPU (RTX 3060 12G, 4070, 5070)

Which models fit a 12GB NVIDIA card at 8K–32K context, where 14B stops fitting, how to run them with Ollama or llama.cpp, and how to prove they are on the GPU.

Written against

Windows or Linux · NVIDIA driver with CUDA · Ollama, or llama.cpp built with GGML_CUDA=ON · GGUF Q4_K_M

Not re-run since it was last edited — treat the commands as a starting point, not a tested recipe.

12GBRTX 3060RTX 4070RTX 5070Ollamallama.cppGGUF

Which cards this covers

Seven NVIDIA cards in this index have 12GB: the RTX 3060 12G, RTX 4070, 4070 Super and 4070 Ti, RTX 5070, and RTX 3080 12G and 3080 Ti. They fit exactly the same models — only capacity decides what fits — but they do not run them at the same speed, because token generation is limited by memory bandwidth, and that ranges from 360 GB/s on the 3060 to 912 GB/s on the 3080s. A 12GB Radeon or an Arc B580 fits the same list too; the AMD and Intel Arc guides cover their runtimes.

text
Card            Bandwidth   ceiling for Qwen3 8B Q4_K_M
RTX 3060 12G    360 GB/s    76 tok/s
RTX 4070 / Super / Ti   504 GB/s    107 tok/s
RTX 5070        672 GB/s    142 tok/s
RTX 3080 12G / Ti       912 GB/s    193 tok/s

Ceilings = bandwidth ÷ weight file size. Not measured here; real speed is lower.

What fits, and where 14B stops fitting

48 of the 89 models in this index fit a 12GB card comfortably at 4K context, and 50 load at all. 7B–12B models are the comfortable zone, and the context length is what decides the rest. Gemma 3 12B at Q4_K_M with a 32K context is about 10.4 GB, still comfortable, because most of its layers use a short sliding window. Mistral Nemo 12B at Q4_K_M with a 32K context is about 13.2 GB, which does not fit. 14B is the edge: Qwen3 14B at Q4_K_M with a 4K context is about 10.0 GB, comfortable, but Qwen3 14B at Q4_K_M with an 8K context is about 10.7 GB — 89% of the card, past this site’s comfortable line — and with 16K it no longer fits.

text
GGUF Q4_K_M            8K        16K       32K
Llama 3.1 8B           6.2 GB    7.3 GB    9.5 GB
Qwen3 8B               6.4 GB    7.7 GB    10.1 GB
Gemma 3 12B            8.8 GB    9.3 GB    10.4 GB
Mistral Nemo 12B       9.1 GB    10.5 GB   13.2 GB  ✗
Qwen3 14B              10.7 GB ~ 12.1 GB ✗ 14.9 GB ✗
Phi-4 14B              11.0 GB ~ 12.8 GB ✗ (16K max)

~ = tight (88–105% of the card)   ✗ = does not fit

Run it with Ollama

Ollama is the shortest path on Windows or Linux. Pulling straight from a Hugging Face GGUF repo pins the exact quant, which matters here: the difference between Q4_K_M and Q8_0 on a 12B model is the difference between fitting and not. Ollama sets its own context length; raise it inside the session when you need more, and check the table above before you do.

bash
ollama run hf.co/Qwen/Qwen3-8B-GGUF:Q4_K_M

# inside the session, for a longer context:
/set parameter num_ctx 16384

ollama ps          # PROCESSOR should read 100% GPU

Or run it with llama.cpp

With llama.cpp, pass -ngl all and an explicit -c. Recent builds fit the model to the card automatically and keep about 1 GiB free, so a model near the limit silently loses layers to the CPU; asking for every layer makes a model that does not fit fail loudly instead of running slowly. Bind to 127.0.0.1 unless you mean to serve your network.

bash
pip install -U huggingface_hub
hf download Qwen/Qwen3-8B-GGUF --include "Qwen3-8B-Q4_K_M.gguf" --local-dir ./models

./build/bin/llama-server -m ./models/Qwen3-8B-Q4_K_M.gguf \
  -ngl all -c 16384 --host 127.0.0.1 --port 8080

# startup log:  load_tensors: offloaded 37/37 layers to GPU

Check it is really on the GPU

Three checks, any one of which is enough: ollama ps shows 100% GPU; llama.cpp’s startup log reports every layer offloaded; nvidia-smi shows the memory held while the model is loaded. On Windows, watch for the slow-but-working case: the NVIDIA control panel’s System Memory Fallback is on by default and spills VRAM into system RAM instead of failing, so a model that does not fit turns into a model that crawls. Set CUDA Sysmem Fallback Policy to prefer no fallback while you are sizing things.

bash
nvidia-smi --query-gpu=name,memory.used,memory.total --format=csv

# measure speed on your own card:
./build/bin/llama-bench -m ./models/Qwen3-8B-Q4_K_M.gguf -ngl 99   # pp512 / tg128

When it does not work

Most problems on a 12GB card are the context, not the model.

text
Out of memory after a while, not at load
  → the KV cache grows with the conversation; lower -c / num_ctx

Loads, but generates at a few tok/s
  → layers on the CPU: check ollama ps / the load_tensors line
  → Windows: System Memory Fallback is hiding an out-of-memory

14B fits at 4K, fails at 16K
  → expected (see the table); drop to 8K, or use a 12B model

A 20B MoE model "almost" fits
  → keep some experts on the CPU (--n-cpu-moe), see the GPT-OSS guide

Common questions

Can a 12GB GPU run a 14B model?

Yes, at Q4_K_M with a short context. This site sizes Qwen3 14B at Q4_K_M at about 10.0 GB at 4K context, which is comfortable, and about 10.7 GB at 8K, which is tight. At 16K it no longer fits. If you need a long context on a 12GB card, Gemma 3 12B is the larger model that stays comfortable at 32K, because most of its layers use a sliding window.

RTX 3060 12G or RTX 4060 Ti 16G for local LLMs?

The 16GB card fits more: 55 of the 89 models in this index comfortably at 4K, against 48 for any 12GB card, and it holds 14B models at long contexts that a 12GB card cannot. On speed the 3060 is not behind — its 360 GB/s of memory bandwidth is higher than the 4060 Ti's 288 GB/s, and bandwidth is what limits generation once a model fits. If the models you want already fit in 12GB, the extra 4GB buys nothing.

Can I run GPT-OSS 20B on a 12GB card?

Not entirely on the card: this site sizes GPT-OSS 20B at its native MXFP4 at about 12.8 GB at 4K context, just over the limit. Because it is a mixture-of-experts model, it runs well with some experts kept in system RAM using llama.cpp's --n-cpu-moe option — the GPT-OSS guide covers the settings.

What this guide uses

Next steps

See everything that fits 🟢 RTX 3060 12GReverse lookup — 12GB at 8192 context, ranked by quality

Models covered in this guide

Did this actually run?

Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.

Related guides

Deployment guides are educational. Each model is subject to its own license — read the official Hugging Face model card before downloading or deploying.