Mixtral 8x7B Instruct
47B MoEMistral AI
Classic MoE model. ~13B active params per token; needs 32GB+ VRAM for Q4.
Pro GPU
33K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
97.2%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Mistral Large675B MoE
Mistral Large 3 675B Instruct
Mistral AI
Pro GPU
312.0 GBmin VRAM·97.9%accuracy
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.
24B
Mistral Small 24B Instruct
Mistral AI
Consumer GPUPro GPU
13.5 GBmin VRAM·97.8%accuracy
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
12B
Mistral Nemo 12B Instruct
Mistral AI
Consumer GPUMac / Apple Silicon
7.8 GBmin VRAM·99.1%accuracy
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
7B
Mistral 7B Instruct v0.3
Mistral AI
Consumer GPUMac / Apple Silicon
4.6 GBmin VRAM·99.1%accuracy
Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.