Devstral Small 1.1 24B
24BMistral Devstral
Mistral + All Hands agentic coding model on a Mistral Small 3.1 base. Built for repo-scale tool use rather than single-file completion. Q4 ~14GB fits a 16GB card.
131K
Max Context
5
Quant Variants
GGUF Q6_K
Best Quality
99.4%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 30B-A3BQwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Qwen3-Coder 30B-A3B Instruct
Alibaba Qwen3
Agentic coding MoE with 3.3B active params and 256K native context. Top open coder for 16–24GB cards.
GPT-OSS 20B
OpenAI GPT-OSS
OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.
Qwen2.5 32B Instruct
Alibaba Qwen2.5
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.