Llama 3.1 405B Instruct

405B

Meta Llama 3.1

Meta frontier dense 405B. Q4 needs ~230GB+ VRAM; dual H100 80G or 8× consumer GPU.

Pro GPU

131K

Max Context

3

Quant Variants

GGUF Q4_K_M

Best Quality

97.7%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.85238.0 GB2.3%6 tok/sEstimated
CalcHF
GGUFQ3_K_M3.87192.0 GB5.0%8 tok/sEstimated
CalcHF
AWQINT44210.0 GB3.5%10 tok/sEstimated
CalcHF

Running Llama 3.1 405B Instruct locally

At Q4_K_M and 4K of context, Llama 3.1 405B Instruct needs about 258.7 GB — 233.2 GB of weights, 1.97 GB of KV cache and a 23.5 GB activation buffer. The smallest card in this index that clears that comfortably is the Mac M5 Ultra 512G at 512 GB, and 1 of the 63 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.

What longer context costs

Going from 4K to 32K adds about 13.78 GB, taking the total to 273.9 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 128K; holding all of it at Q4_K_M would need about 326 GB.

Which build to download

This index tracks 2 formats for it — GGUF, AWQ — across 3 levels. Q4_K_M carries the lowest published perplexity loss at 2.3%. The fastest level measured here is AWQ INT4 at 10 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.

Common questions

How much VRAM does Llama 3.1 405B Instruct need?
About 258.7 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 273.9 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
Will Llama 3.1 405B Instruct run on a 512GB GPU?
Yes — at Q4_K_M and 4K context it needs about 258.7 GB, which leaves 253.3 GB spare on a Mac M5 Ultra 512G. That is the smallest card in this index that clears it comfortably; 1 of 63 do.
Which quantization of Llama 3.1 405B Instruct should I use?
Q4_K_M has the lowest published quality loss (2.3%), and Q4_K_M is the level most people run. All 3 levels in the index are Q4_K_M, Q3_K_M, AWQ INT4.

Where this model fits

Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.

Ships in
GGUFAWQ
Compared here
GGUF vs AWQ

This page's figures change when the model or the runtime does.Last updated 2026-09-23 RSS → /feed.xml