Tesla P100 16G — what LLMs can it run?
52 of 81 indexed models fit comfortably in 16GB at 4K context, each at the highest-quality quant that still leaves headroom.
The short answer for Tesla P100 16G
16 GB, 732 GB/s HBM2. 52 of 81 models in this index fit comfortably at 4K context; 61 load at all.
- Biggest that fits
- Mistral Small 24B Instruct — 13.2 GB at AWQ INT4 — 2.8 GB spare at 4K, so longer context comes out of a thin margin.
- Room to grow
- Stable LM 2 12B Chat — 8.3 GB at Q4_K_M — under 60% of the card, which leaves 7.7 GB for a long context window or a second process.
- Speed ceiling
- Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 732 GB/s puts a hard ceiling near 158 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
- The wall you will hit
- Bandwidth, not capacity. The models on this page fit, but at 732 GB/s this card reads the whole weight set once per generated token — a 960 GB/s RTX 5080 holds exactly the same 16 GB and moves those bytes 1.3× faster. Expect the same quant to generate proportionally slower here.
32B · 5
| Model | Quant | Est. VRAM | Headroom | tok/s |
|---|---|---|---|---|
| Mistral Small 24B Instruct24B | AWQ INT4 | 13.23 GB | +2.8 GB | 78 |
| Devstral Small 1.1 24B24B | AWQ INT4 | 13.02 GB | +3 GB | 78 |
| Magistral Small 1.2 24B24B | AWQ INT4 | 13.02 GB | +3 GB | 76 |
| Codestral 22B22B | AWQ INT4 | 12.56 GB | +3.4 GB | 72 |
| GPT-OSS 20B21B MoE | MXFP4 | 11.81 GB | +4.2 GB | 195 |
14B · 17
| Model | Quant | Est. VRAM | Headroom | tok/s |
|---|---|---|---|---|
| InternLM2 20B Chat20B | Q4_K_M | 13.43 GB | +2.6 GB | 78 |
| DeepSeek-Coder-V2-Lite Instruct16B | Q4_K_M | 10.87 GB | +5.1 GB | 145 |
| DeepSeek-V2-Lite Chat16B | Q4_K_M | 10.87 GB | +5.1 GB | 142 |
| StarCoder2 15B15B | Q4_K_M | 10.19 GB | +5.8 GB | 92 |
| Qwen3 14B Instruct14B | Q5_K_M | 11.65 GB | +4.4 GB | 78 |
| Qwen2.5 14B Instruct14B | Q5_K_M | 11.73 GB | +4.3 GB | 86 |
| DeepSeek-R1-Distill-Qwen-14B14B | EXL2 4.65bpw | 9.75 GB | +6.3 GB | 128 |
| Phi-4 14B14B | Q5_K_M | 11.77 GB | +4.2 GB | 78 |
| Phi-3 Medium 14B Instruct14B | Q6_K | 12.86 GB | +3.1 GB | 88 |
| Mistral Nemo 12B Instruct12B | Q6_K | 11.14 GB | +4.9 GB | 95 |
| Gemma 3 12B IT12B | Q5_K_M | 10.6 GB | +5.4 GB | 92 |
| Stable LM 2 12B Chat12B | Q4_K_M | 8.35 GB | +7.7 GB | 108 |
| Jamba 1.5 Mini12B | Q4_K_M | 8.15 GB | +7.9 GB | 95 |
| Llama 3.2 11B Vision Instruct11B | Q8_0 | 12.9 GB | +3.1 GB | 72 |
| Solar 10.7B Instruct11B | Q4_K_M | 7.6 GB | +8.4 GB | 125 |
| Falcon 3 10B Instruct10B | Q4_K_M | 7.28 GB | +8.7 GB | 118 |
| Gemma 2 9B Instruct9B | Q8_0 | 11.7 GB | +4.3 GB | 108 |
7B · 20
| Model | Quant | Est. VRAM | Headroom | tok/s |
|---|---|---|---|---|
| GLM-4-9B-Chat9B | Q8_0 | 10.16 GB | +5.8 GB | 105 |
| Qwen3-VL 8B Instruct8B | Q8_0 | 10.39 GB | +5.6 GB | 108 |
| Ministral 3 8B Instruct8B | Q8_0 | 10.02 GB | +6 GB | — |
| Qwen2-VL 7B Instruct7B | Q4_K_M | 5.49 GB | +10.5 GB | 72 |
| Granite 3.1 8B Instruct8B | Q4_K_M | 5.88 GB | +10.1 GB | 142 |
| Qwen3 8B Instruct8B | Q6_K | 7.64 GB | +8.4 GB | 122 |
| Llama 3.1 8B Instruct8B | Q8_0 | 9.47 GB | +6.5 GB | 118 |
| Nous Hermes 3 Llama 3.1 8B8B | EXL2 4.65bpw | 5.43 GB | +10.6 GB | 232 |
| Aya 23 8B8B | Q4_K_M | 5.64 GB | +10.4 GB | 145 |
| OpenChat 3.6 8B8B | EXL2 4.65bpw | 5.43 GB | +10.6 GB | 228 |
| DeepSeek-R1-Distill-Llama-8B8B | Q5_K_M | 6.51 GB | +9.5 GB | 128 |
| InternLM2 7B Chat7B | Q4_K_M | 5.45 GB | +10.6 GB | 148 |
| Qwen2.5 7B Instruct7B | Q6_K | 6.77 GB | +9.2 GB | 132 |
| Qwen2.5-Coder 7B Instruct7B | EXL2 4.65bpw | 4.87 GB | +11.1 GB | 248 |
| WizardLM-2 7B7B | Q4_K_M | 5.07 GB | +10.9 GB | 152 |
| DeepSeek-R1-Distill-Qwen-7B7B | EXL2 4.65bpw | 4.87 GB | +11.1 GB | 210 |
| OLMo 2 7B Instruct7B | Q8_0 | 10.3 GB | +5.7 GB | 125 |
| Mistral 7B Instruct v0.37B | Q6_K | 6.75 GB | +9.3 GB | 135 |
| Zephyr 7B Beta7B | Q6_K | 6.75 GB | +9.3 GB | 132 |
| Gemma 3 4B IT4B | Q8_0 | 5.36 GB | +10.6 GB | 145 |
≤3B · 10
| Model | Quant | Est. VRAM | Headroom | tok/s |
|---|---|---|---|---|
| Qwen3 4B Instruct4B | Q6_K | 4.05 GB | +12 GB | 145 |
| Phi-4 Mini Instruct3.8B | Q8_0 | 4.68 GB | +11.3 GB | 262 |
| Phi-3.5 Mini Instruct3.8B | Q8_0 | 5.89 GB | +10.1 GB | 255 |
| Llama 3.2 3B Instruct3B | Q8_0 | 4.05 GB | +12 GB | 285 |
| Qwen2.5 3B Instruct3B | Q8_0 | 3.67 GB | +12.3 GB | 290 |
| Gemma 2 2B Instruct2B | Q8_0 | 3.34 GB | +12.7 GB | 320 |
| Qwen3 1.7B Instruct1.7B | Q8_0 | 2.39 GB | +13.6 GB | 240 |
| Qwen2.5 1.5B Instruct1.5B | Q8_0 | 1.83 GB | +14.2 GB | 410 |
| Llama 3.2 1B Instruct1B | Q8_0 | 1.51 GB | +14.5 GB | 450 |
| Qwen2.5 0.5B Instruct0.5B | Q8_0 | 0.6 GB | +15.4 GB | 540 |
How this list is built
Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.
A further 9 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.
Measured on this card
No benchmark runs in this index were recorded on a Tesla P100 16G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements.
Cards with the same budget
The same 16GB in system memory is a different machine — everything on this page fits there, at a speed set by your DIMMs rather than a graphics bus: 16 GB RAM (CPU)
Stepping up
A Tesla P40 24G (24GB) fits 13 more of the indexed models than this card. Tesla P40 24G →
Common questions
What is the best local LLM for a Tesla P100 16G?
For everyday use, Stable LM 2 12B Chat at Q4_K_M — about 8.3 GB of the card's 16 GB at 4K context, leaving 7.7 GB for a longer window. If you want the largest thing that will load, that is Mistral Small 24B Instruct at AWQ INT4 (13.2 GB, 2.8 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.
Can a Tesla P100 16G run Gemma 2 27B Instruct?
Only just. Its smallest build here, AWQ INT4, needs about 15.8 GB against 16 GB — that loads on a card with nothing else on it, with no margin for a longer context window. It is not on the list above, which requires a model to stay inside 88% of the card.
How many tokens per second does a Tesla P100 16G do on an 8B model at Q4?
This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 732 GB/s — a ceiling near 158 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.
Tesla P100 16G or RTX 5080 for local LLMs?
They hold the same models: 16 GB against 16 GB fits 52 and 52 of 81 respectively at 4K. The difference is throughput — 732 GB/s against 960 GB/s, a 1.3× gap in how fast the weights can be read, which is what token generation is bound by. The RTX 5080 generates faster on any model both can hold.
52 of 81 indexed models fit comfortably in 16GB at 4K context, each at the highest-quality quant that still leaves headroom.
This page's figures change when the model or the runtime does.Last updated 2026-09-14 RSS → /feed.xml