GLM-4.5-Air

106B MoE

Zhipu GLM-4.5

Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.

371.0K HF downloads629 likeszai-org/GLM-4.5-Air· stats from 8/8/2026
Pro GPUMac / Apple Silicon

131K

Max Context

3

Quant Variants

GGUF Q4_K_M

Best Quality

97.6%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.8564.3 GB2.4%18 tok/sCommunity
CalcHF
GGUFQ3_K_M3.8751.3 GB5.1%21 tok/sEstimated
CalcHF
GGUFQ2_K2.6334.8 GB12.0%26 tok/sEstimated
CalcHF