Phi-3.5 Mini Instruct

3.8B

Microsoft Phi

Microsoft's tiny powerhouse. Best 4B model for on-device deployment.

134.5K HF downloads85 likesbartowski/Phi-3.5-mini-instruct-GGUF· stats from 8/8/2026
Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

3

Quant Variants

GGUF Q8_0

Best Quality

99.8%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.852.8 GB3.8%298 tok/sEstimated
CalcHF
GGUFQ8_08.54.2 GB0.2%255 tok/sEstimated
CalcHF
AWQINT442.5 GB5.1%385 tok/sEstimated
CalcHF