Benchmarks & Insights

Hard numbers, no hype — measured on real hardware

Hardware × Format Matrix

Multi-model speed & VRAM on real hardware — same quant levels where comparable

ModelHardwareFrameworkQuantSpeed (tok/s)VRAM UsedNotes
Llama 3.1 8BRTX 4090 24GExLlamaV2EXL2 4.65bpw2355.4 GBPeak consumer performance
Llama 3.1 8BRTX 4090 24GvLLMAWQ INT42184.9 GBBest for batch API
Llama 3.1 8BRTX 4090 24Gllama.cppGGUF Q4_K_M1485.7 GBEasiest setup
Qwen2.5 7BRTX 4090 24Gllama.cppGGUF Q4_K_M1555.4 GBStrong coding; similar VRAM to 8B
DeepSeek-R1 14BRTX 4090 24GExLlamaV2EXL2 4.65bpw1289.8 GBReasoning distill; hot in 2026
Qwen2.5 32BRTX 4090 24Gllama.cppGGUF Q4_K_M4422 GBTight fit at 4K ctx; use Q3 for headroom
Llama 3.1 8BRTX 4060 Ti 16GExLlamaV2EXL2 4.65bpw985.4 GBGreat budget option
Llama 3.1 8BRTX 4060 Ti 16Gllama.cppGGUF Q4_K_M785.7 GBBudget-friendly
Qwen2.5 7BRTX 4060 Ti 16Gllama.cppGGUF Q4_K_M825.4 GBSweet spot on 16GB cards
Qwen3 8BRTX 4090 24Gllama.cppGGUF Q4_K_M1425.8 GBQwen3 thinking mode; ~2026 flagship 8B
Qwen3 14BRTX 4090 24GExLlamaV2EXL2 4.65bpw11810 GBStrong reasoning; 16GB+ sweet spot
Qwen3 32BRTX 4090 24Gllama.cppGGUF Q4_K_M4222.5 GBDense 32B successor to Qwen2.5-32B
Qwen3 30B-A3BRTX 4090 24Gllama.cppGGUF Q4_K_M9519 GBMoE 3B active — fast on 16GB cards
R1-Distill-Llama-8BRTX 4090 24Gllama.cppGGUF Q4_K_M1455.6 GBR1 reasoning on 8B footprint
Phi-4 14BRTX 4090 24Gllama.cppGGUF Q4_K_M889.1 GBStrong dense 14B; 12GB with short ctx
Qwen3-Coder 30B-A3BRTX 4090 24Gllama.cppGGUF Q4_K_M9219.2 GBAgentic coding MoE; 256K native ctx
Llama 3.1 8BRTX 3090 24GExLlamaV2EXL2 4.65bpw1755.4 GBOlder but capable
Llama 3.1 8BM3 Max 48GOllamaGGUF Q4_K_M685.7 GBUnified memory advantage
Llama 3.1 8BM2 Ultra 192Gllama.cppGGUF Q4_K_M905.7 GBCan run 70B models solo