Mistral Nemo 12B Instruct
12BMistral AI
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
Consumer GPUMac / Apple Silicon
131K
Max Context
3
Quant Variants
GGUF Q6_K
Best Quality
99.1%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Mistral 7B7B
Mistral 7B Instruct v0.3
Mistral AI
Consumer GPUMac / Apple Silicon
4.6 GBmin VRAM·99.1%accuracy
Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.
24B
Mistral Small 24B Instruct
Mistral AI
Consumer GPUPro GPU
13.5 GBmin VRAM·97.8%accuracy
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
47B MoE
Mixtral 8x7B Instruct
Mistral AI
Pro GPU
25.2 GBmin VRAM·97.2%accuracy
Classic MoE model. ~13B active params per token; needs 32GB+ VRAM for Q4.
675B MoE
Mistral Large 3 675B Instruct
Mistral AI
Pro GPU
312.0 GBmin VRAM·97.9%accuracy
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.