Seed-OSS 36B Instruct
36BByteDance Seed
Dense 36B with a native 512K context. Q4 weights are ~22GB, which lands on a 24GB card or a 2×16GB split — but the context is the real cost: 512K of KV cache is roughly 128GB on its own, so budget context first and weights second.
524K
Max Context
4
Quant Variants
GGUF Q5_K_M
Best Quality
98.7%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen2.5 32BQwen2.5 32B Instruct
Alibaba Qwen2.5
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.
DeepSeek-R1-Distill-Qwen-32B
DeepSeek
R1 distilled to 32B. Near-frontier reasoning on a single 24GB card (Q3/Q4).
Qwen3 32B Instruct
Alibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
How to actually run this
Deployment guides for this model and this class of hardware.