DBRX Instruct
132BDatabricks
MoE flagship (~36B active). Needs multi-GPU; strong code and reasoning at scale.
33K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
97.5%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 235B-A22BQwen3 235B-A22B Instruct
Alibaba Qwen3
Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.
DeepSeek-V3
DeepSeek
DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.
DeepSeek-R1
DeepSeek
DeepSeek-R1 reasoning model built on V3 MoE. Chain-of-thought at frontier scale — use distill variants for local GPUs.
Mistral Large 3 675B Instruct
Mistral AI
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.