Command R 35B
35BCohere
Cohere's RAG-optimised model. Excellent retrieval-augmented generation.
Pro GPU
131K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
97.0%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen2.5 32B32B
Qwen2.5 32B Instruct
Alibaba Qwen2.5
Consumer GPUPro GPU
16.4 GBmin VRAM·97.3%accuracy
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.
24B
Mistral Small 24B Instruct
Mistral AI
Consumer GPUPro GPU
13.5 GBmin VRAM·97.8%accuracy
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
34B
Yi 1.5 34B Chat
01.AI Yi
Consumer GPUPro GPU
19.8 GBmin VRAM·97.2%accuracy
01.AI's strong bilingual (EN/ZH) model. Competitive with Qwen 32B.
27B
Gemma 2 27B Instruct
Google Gemma 2
Consumer GPUPro GPU
16.2 GBmin VRAM·98.7%accuracy
Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.