InternLM2 20B Chat
20BShanghai AI Lab
Mid-size InternLM2 with excellent Chinese comprehension. Fits 24GB at Q4.
Consumer GPUPro GPU
33K
Max Context
2
Quant Variants
GGUF Q5_K_M
Best Quality
98.6%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with InternLM2 7B7B
InternLM2 7B Chat
Shanghai AI Lab
Consumer GPUMac / Apple Silicon
4.9 GBmin VRAM·96.9%accuracy
Strong bilingual (EN/ZH) 7B from Shanghai AI Lab. Competitive with Qwen 7B.
32B
Qwen2.5 32B Instruct
Alibaba Qwen2.5
Consumer GPUPro GPU
16.4 GBmin VRAM·97.3%accuracy
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.
24B
Mistral Small 24B Instruct
Mistral AI
Consumer GPUPro GPU
13.5 GBmin VRAM·97.8%accuracy
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
34B
Yi 1.5 34B Chat
01.AI Yi
Consumer GPUPro GPU
19.8 GBmin VRAM·97.2%accuracy
01.AI's strong bilingual (EN/ZH) model. Competitive with Qwen 32B.