Gemma 2 9B Instruct
9BGoogle Gemma 2
Google's compact Gemma 2 with sliding window attention. Punches above 9B.
Consumer GPUMac / Apple Silicon
8K
Max Context
3
Quant Variants
GGUF Q8_0
Best Quality
99.8%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 22B
Gemma 2 2B Instruct
Google Gemma 2
Consumer GPUMac / Apple Silicon
1.8 GBmin VRAM·99.7%accuracy
Ultra-compact Gemma 2. Runs on 4GB VRAM; great for edge prototyping.
27B
Gemma 2 27B Instruct
Google Gemma 2
Consumer GPUPro GPU
16.2 GBmin VRAM·98.7%accuracy
Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.
14B
Qwen2.5 14B Instruct
Alibaba Qwen2.5
Consumer GPUMac / Apple Silicon
9.2 GBmin VRAM·98.6%accuracy
The sweet spot between performance and resource usage. 16GB VRAM with Q4.
12B
Mistral Nemo 12B Instruct
Mistral AI
Consumer GPUMac / Apple Silicon
7.8 GBmin VRAM·99.1%accuracy
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.