Stable LM 2 12B Chat
12BStability AI
Stability AI's 12B chat model. Solid general-purpose option for 16GB GPUs.
Consumer GPUMac / Apple Silicon
4K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
96.8%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen2.5 14B14B
Qwen2.5 14B Instruct
Alibaba Qwen2.5
Consumer GPUMac / Apple Silicon
9.2 GBmin VRAM·98.6%accuracy
The sweet spot between performance and resource usage. 16GB VRAM with Q4.
12B
Mistral Nemo 12B Instruct
Mistral AI
Consumer GPUMac / Apple Silicon
7.8 GBmin VRAM·99.1%accuracy
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
9B
Gemma 2 9B Instruct
Google Gemma 2
Consumer GPUMac / Apple Silicon
5.8 GBmin VRAM·99.8%accuracy
Google's compact Gemma 2 with sliding window attention. Punches above 9B.
11B
Solar 10.7B Instruct
Upstage
Consumer GPUMac / Apple Silicon
6.5 GBmin VRAM·97.0%accuracy
Depth-upscaled 10.7B punching above weight. Strong on reasoning benchmarks.