24GB VRAM

Tesla P40 24G — what LLMs can it run?

65 of 81 indexed models fit comfortably in 24GB at 4K context, each at the highest-quality quant that still leaves headroom.

The short answer for Tesla P40 24G

24 GB, 346 GB/s GDDR5. 65 of 81 models in this index fit comfortably at 4K context; 66 load at all.

Biggest that fits
Seed-OSS 36B Instruct19.9 GB at AWQ INT4 — 4.1 GB spare at 4K, so longer context comes out of a thin margin.
Room to grow
GPT-OSS 20B11.8 GB at MXFP4 — under 60% of the card, which leaves 12.2 GB for a long context window or a second process.
Speed ceiling
Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 346 GB/s puts a hard ceiling near 75 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
The wall you will hit
Bandwidth, not capacity. The models on this page fit, but at 346 GB/s this card reads the whole weight set once per generated token — a 1,008 GB/s RTX 4090 holds exactly the same 24 GB and moves those bytes 2.9× faster. Expect the same quant to generate proportionally slower here.
Best local LLM for 24GB

32B · 18

Tesla P40 24G — what LLMs can it run?32B
ModelQuantEst. VRAMHeadroomtok/s
Seed-OSS 36B Instruct36BAWQ INT419.91 GB+4.1 GB55
Command R 35B35BGPTQ INT418.97 GB+5 GB55
Yi 1.5 34B Chat34BAWQ INT419 GB+5 GB52
Qwen3 32B Instruct32BAWQ INT418.22 GB+5.8 GB55
Qwen2.5 32B Instruct32BEXL2 3.5bpw15.96 GB+8 GB68
Qwen2.5-Coder 32B Instruct32BAWQ INT418.08 GB+5.9 GB52
DeepSeek-R1-Distill-Qwen-32B32BEXL2 3.5bpw15.96 GB+8 GB65
Qwen3 30B-A3B Instruct30B-A3BQ4_K_M19.73 GB+4.3 GB95
Qwen3-Coder 30B-A3B Instruct30B-A3BQ4_K_M19.73 GB+4.3 GB92
Qwen3-VL 30B-A3B Instruct30B-A3BQ4_K_M19.73 GB+4.3 GB95
Qwen3.8 27B27BQ4_K_M17.47 GB+6.5 GB
Gemma 3 27B IT27BQ4_K_M19.49 GB+4.5 GB48
Gemma 2 27B Instruct27BQ4_K_M18.81 GB+5.2 GB48
Mistral Small 24B Instruct24BEXL2 4.65bpw15.26 GB+8.7 GB88
Devstral Small 1.1 24B24BQ6_K20.91 GB+3.1 GB48
Magistral Small 1.2 24B24BQ6_K20.91 GB+3.1 GB47
Codestral 22B22BQ4_K_M15.03 GB+9 GB58
GPT-OSS 20B21B MoEMXFP411.81 GB+12.2 GB195

14B · 17

Tesla P40 24G — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
InternLM2 20B Chat20BQ5_K_M15.59 GB+8.4 GB68
DeepSeek-Coder-V2-Lite Instruct16BQ8_018.36 GB+5.6 GB118
DeepSeek-V2-Lite Chat16BQ4_K_M10.87 GB+13.1 GB142
StarCoder2 15B15BQ4_K_M10.19 GB+13.8 GB92
Qwen3 14B Instruct14BQ5_K_M11.65 GB+12.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+12.3 GB86
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+14.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+12.2 GB78
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+11.1 GB88
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+12.9 GB95
Gemma 3 12B IT12BQ5_K_M10.6 GB+13.4 GB92
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+15.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+15.9 GB95
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+11.1 GB72
Solar 10.7B Instruct11BQ4_K_M7.6 GB+16.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+16.7 GB118
Gemma 2 9B Instruct9BQ8_011.7 GB+12.3 GB108

7B · 20

Tesla P40 24G — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+13.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+13.6 GB108
Ministral 3 8B Instruct8BQ8_010.02 GB+14 GB
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+18.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+18.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+16.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+14.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+18.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+18.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+18.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+17.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+18.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+17.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+19.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+18.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+19.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+13.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+17.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+17.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+18.6 GB145

≤3B · 10

Tesla P40 24G — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+20 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+19.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+18.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+20 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+20.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+20.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+21.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+22.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+22.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+23.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 1 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a Tesla P40 24G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements.

Stepping up

A A100 40G (40GB) fits 3 more of the indexed models than this card. A100 40G

Common questions

What is the best local LLM for a Tesla P40 24G?

For everyday use, GPT-OSS 20B at MXFP4 — about 11.8 GB of the card's 24 GB at 4K context, leaving 12.2 GB for a longer window. If you want the largest thing that will load, that is Seed-OSS 36B Instruct at AWQ INT4 (19.9 GB, 4.1 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.

Can a Tesla P40 24G run Mixtral 8x7B Instruct?

Only just. Its smallest build here, AWQ INT4, needs about 24.9 GB against 24 GB — that loads on a card with nothing else on it, with no margin for a longer context window. It is not on the list above, which requires a model to stay inside 88% of the card.

How many tokens per second does a Tesla P40 24G do on an 8B model at Q4?

This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 346 GB/s — a ceiling near 75 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.

Tesla P40 24G or RTX 4090 for local LLMs?

They hold the same models: 24 GB against 24 GB fits 65 and 65 of 81 respectively at 4K. The difference is throughput — 346 GB/s against 1,008 GB/s, a 2.9× gap in how fast the weights can be read, which is what token generation is bound by. The RTX 4090 generates faster on any model both can hold.

65 of 81 indexed models fit comfortably in 24GB at 4K context, each at the highest-quality quant that still leaves headroom.

This page's figures change when the model or the runtime does.Last updated 2026-09-14 RSS → /feed.xml