GLM-4.5-Air
106B MoEZhipu GLM-4.5
Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.
131K
Max Context
3
Quant Variants
GGUF Q4_K_M
Best Quality
97.6%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with GPT-OSS 120BGPT-OSS 120B
OpenAI GPT-OSS
The big GPT-OSS (117B total / 5.1B active). Native MXFP4 checkpoint is ~61GB — fits one 80GB card or a 128GB unified-memory Mac. Partial offload on 24GB consumer cards is slow but works.
DBRX Instruct
Databricks
MoE flagship (~36B active). Needs multi-GPU; strong code and reasoning at scale.
Qwen3 235B-A22B Instruct
Alibaba Qwen3
Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.
DeepSeek-V3
DeepSeek
DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.