Gemma 3 27B IT
27BGoogle Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
Consumer GPUPro GPU
131K
Max Context
3
Quant Variants
GGUF Q4_K_M
Best Quality
97.2%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 34B
Gemma 3 4B IT
Google Gemma 3
Consumer GPUMac / Apple Silicon
3.0 GBmin VRAM·99.8%accuracy
Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.
12B
Gemma 3 12B IT
Google Gemma 3
Consumer GPUMac / Apple Silicon
8.0 GBmin VRAM·98.7%accuracy
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
32B
Qwen2.5 32B Instruct
Alibaba Qwen2.5
Consumer GPUPro GPU
16.4 GBmin VRAM·97.3%accuracy
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.
24B
Mistral Small 24B Instruct
Mistral AI
Consumer GPUPro GPU
13.5 GBmin VRAM·97.8%accuracy
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.