Mistral Large 3 675B Instruct

675B MoE

Mistral AI

Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.

1.1K HF downloads241 likesmistralai/Mistral-Large-3-675B-Instruct-2512· stats from 8/8/2026
Pro GPU

262K

Max Context

2

Quant Variants

GGUF Q4_K_M

Best Quality

97.9%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.85388.0 GB2.1%4 tok/sEstimated
CalcHF
GGUFQ3_K_M3.87312.0 GB4.5%5 tok/sEstimated
CalcHF