Gemma 2 2B Instruct
2BGoogle Gemma 2
Ultra-compact Gemma 2. Runs on 4GB VRAM; great for edge prototyping.
Consumer GPUMac / Apple SiliconCPU / VPS
8K
Max Context
3
Quant Variants
GGUF Q8_0
Best Quality
99.7%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Gemma 29B
Gemma 2 9B Instruct
Google Gemma 2
Consumer GPUMac / Apple Silicon
5.8 GBmin VRAM·99.8%accuracy
Google's compact Gemma 2 with sliding window attention. Punches above 9B.
27B
Gemma 2 27B Instruct
Google Gemma 2
Consumer GPUPro GPU
16.2 GBmin VRAM·98.7%accuracy
Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.
3B
Llama 3.2 3B Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
2.0 GBmin VRAM·99.8%accuracy
Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.
3B
Qwen2.5 3B Instruct
Alibaba Qwen2.5
Consumer GPUMac / Apple Silicon
2.1 GBmin VRAM·99.7%accuracy
Tiny Qwen2.5 for edge devices. Runs on 4GB VRAM or Raspberry Pi class hardware.