DBRX Instruct
132BDatabricks
⚠ Superseded · Prefer GLM-4.5-Air
This index would suggest GLM-4.5-Air instead today — longer context (32K → 128K tokens). The numbers below are still accurate; they are just for a model you probably should not start with. Compare the two →
MoE flagship (~36B active). Needs multi-GPU; strong code and reasoning at scale.
33K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
97.5%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 235B-A22BQwen3 235B-A22B Instruct
Alibaba Qwen3
Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.
DeepSeek-V3
DeepSeek
DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.
DeepSeek-R1
DeepSeek
DeepSeek-R1 reasoning model built on V3 MoE. Chain-of-thought at frontier scale — use distill variants for local GPUs.
Mistral Large 3 675B Instruct
Mistral AI
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.
Running DBRX Instruct locally
At Q4_K_M and 4K of context, DBRX Instruct needs about 84.3 GB — 76.0 GB of weights, 0.63 GB of KV cache and a 7.7 GB activation buffer. The smallest card in this index that clears that comfortably is the Mac M5 Max 128G at 128 GB, and 8 of the 63 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.
What longer context costs
Going from 4K to 32K adds about 4.37 GB, taking the total to 89.1 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 32K; holding all of it at Q4_K_M would need about 89 GB.
Which build to download
This index only tracks DBRX Instruct in GGUF, across 2 levels (Q4_K_M, Q3_K_M). Q4_K_M carries the lowest published perplexity loss at 2.5%. The fastest level measured here is Q3_K_M at 18 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.
Common questions
- How much VRAM does DBRX Instruct need?
- About 84.3 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 89.1 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
- Will DBRX Instruct run on a 128GB GPU?
- Yes — at Q4_K_M and 4K context it needs about 84.3 GB, which leaves 43.7 GB spare on a Mac M5 Max 128G. That is the smallest card in this index that clears it comfortably; 8 of 63 do.
- Which quantization of DBRX Instruct should I use?
- Q4_K_M has the lowest published quality loss (2.5%), and Q4_K_M is the level most people run. All 2 levels in the index are Q4_K_M, Q3_K_M.
Where this model fits
Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.
How to actually run this
Deployment guides for this model and this class of hardware.
This page's figures change when the model or the runtime does.Last updated 2026-09-21 RSS → /feed.xml