Devstral Small 1.1 24B

24B

Mistral Devstral

Mistral + All Hands agentic coding model on a Mistral Small 3.1 base. Built for repo-scale tool use rather than single-file completion. Q4 ~14GB fits a 16GB card.

109.7K HF downloads368 likesmistralai/Devstral-Small-2507· stats from 8/8/2026
Consumer GPUMac / Apple Silicon

131K

Max Context

5

Quant Variants

GGUF Q6_K

Best Quality

99.4%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.8514.3 GB2.9%62 tok/sCommunity
CalcHF
GGUFQ5_K_M5.6816.8 GB1.4%55 tok/sEstimated
CalcHF
GGUFQ6_K6.5619.4 GB0.6%48 tok/sEstimated
CalcHF
AWQINT4413.0 GB4.0%78 tok/sEstimated
CalcHF
EXL24.65bpw4.6513.9 GB2.5%88 tok/sEstimated
CalcHF