Mistral Small 24B Instruct
24BMistral AI
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
Consumer GPUPro GPU
33K
Max Context
3
Quant Variants
EXL2 4.65bpw
Best Quality
97.8%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Codestral 22B22B
Codestral 22B
Mistral AI
Consumer GPUPro GPU
13.2 GBmin VRAM·97.0%accuracy
Mistral's dedicated code model. 80+ language support, Fill-in-the-Middle capable.
12B
Mistral Nemo 12B Instruct
Mistral AI
Consumer GPUMac / Apple Silicon
7.8 GBmin VRAM·99.1%accuracy
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
47B MoE
Mixtral 8x7B Instruct
Mistral AI
Pro GPU
25.2 GBmin VRAM·97.2%accuracy
Classic MoE model. ~13B active params per token; needs 32GB+ VRAM for Q4.
7B
Mistral 7B Instruct v0.3
Mistral AI
Consumer GPUMac / Apple Silicon
4.6 GBmin VRAM·99.1%accuracy
Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.