GPT-OSS 120B
117B MoEOpenAI GPT-OSS
The big GPT-OSS (117B total / 5.1B active). Native MXFP4 checkpoint is ~61GB — fits one 80GB card or a 128GB unified-memory Mac. Partial offload on 24GB consumer cards is slow but works.
131K
Max Context
2
Quant Variants
GGUF MXFP4
Best Quality
100.0%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with GPT-OSS 20BGPT-OSS 20B
OpenAI GPT-OSS
OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.
GLM-4.5-Air
Zhipu GLM-4.5
Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.
DBRX Instruct
Databricks
MoE flagship (~36B active). Needs multi-GPU; strong code and reasoning at scale.
Qwen3 235B-A22B Instruct
Alibaba Qwen3
Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.