GPT-OSS 20B

21B MoE

OpenAI GPT-OSS

OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.

8.1M HF downloads4891 likesopenai/gpt-oss-20b· stats from 8/8/2026
Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

3

Quant Variants

GGUF MXFP4

Best Quality

100.0%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFMXFP44.2512.8 GB0.0%195 tok/sCommunity
CalcHF
GGUFQ8_05.113.8 GB0.0%178 tok/sEstimated
CalcHF
GGUFQ4_K_M4.111.9 GB1.4%205 tok/sEstimated
CalcHF