Qwen3 8B Instruct
8BAlibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
41K
Max Context
4
Quant Variants
GGUF Q6_K
Best Quality
99.4%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 4BQwen3 4B Instruct
Alibaba Qwen3
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
Qwen3 14B Instruct
Alibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Qwen3 32B Instruct
Alibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
Running Qwen3 8B Instruct locally
At Q4_K_M and 4K of context, Qwen3 8B Instruct needs about 5.8 GB — 4.7 GB of weights, 0.56 GB of KV cache and a 0.5 GB activation buffer. The smallest card in this index that clears that comfortably is the RTX 5060 Ti 8G at 8 GB, and 62 of the 63 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.
What longer context costs
Going from 4K to 32K adds about 3.94 GB, taking the total to 10.1 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 40K; holding all of it at Q4_K_M would need about 11 GB.
Which build to download
This index tracks 3 formats for it — GGUF, AWQ, EXL2 — across 4 levels. Q6_K carries the lowest published perplexity loss at 0.6%. The fastest level measured here is EXL2 4.65bpw at 228 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.
Common questions
- How much VRAM does Qwen3 8B Instruct need?
- About 5.8 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 10.1 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
- Will Qwen3 8B Instruct run on a 8GB GPU?
- Yes — at Q4_K_M and 4K context it needs about 5.8 GB, which leaves 2.2 GB spare on a RTX 5060 Ti 8G. That is the smallest card in this index that clears it comfortably; 62 of 63 do.
- Which quantization of Qwen3 8B Instruct should I use?
- Q6_K has the lowest published quality loss (0.6%), and Q4_K_M is the level most people run. All 4 levels in the index are Q4_K_M, Q6_K, AWQ INT4, EXL2 4.65bpw.
Where this model fits
Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.
- Runs comfortably on
- RTX 5060 Ti 8G8GB · +2.2RTX 50608GB · +2.2RTX 4060 Ti 8G8GB · +2.2RTX 40608GB · +2.2RTX 3070 Ti8GB · +2.2
- Tight but possible
- Mac M3 8G8GB
- Compared here
- GGUF vs AWQGGUF vs EXL2AWQ vs EXL2
- Try it yourself
- Size it in the calculatorGenerate the run command
How to actually run this
Deployment guides for this model and this class of hardware.
- Docker Compose LLM Stack: Ollama + Open WebUI 5 min read·Beginner·covers this model
- Ollama on Windows (Native, No WSL) 3 min read·Beginner·covers this model
- Mac M3 Max: The Ultimate Local LLM Setup 3 min read·Beginner·covers this model
- WSL2 + Ollama GPU Passthrough on Windows 4 min read·Intermediate·covers this model
This page's figures change when the model or the runtime does.Last updated 2026-09-23 RSS → /feed.xml