Qwen3 235B-A22B Instruct

235B-A22B

Alibaba Qwen3

Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.

18.7K HF downloads10 likesQwen/Qwen3-235B-A22B-GGUF· stats from 9/23/2026
Pro GPU

41K

Max Context

3

Quant Variants

GGUF Q4_K_M

Best Quality

97.8%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.85142.0 GB2.2%10 tok/sEstimated
CalcHF
GGUFQ3_K_M3.87115.0 GB4.8%12 tok/sEstimated
CalcHF
AWQINT44125.0 GB3.2%14 tok/sEstimated
CalcHF

Running Qwen3 235B-A22B Instruct locally

At Q4_K_M and 4K of context, Qwen3 235B-A22B Instruct needs about 149.7 GB — 135.3 GB of weights, 0.73 GB of KV cache and a 13.6 GB activation buffer. The smallest card in this index that clears that comfortably is the Mac M5 Ultra 256G at 256 GB, and 2 of the 63 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.

What longer context costs

Going from 4K to 32K adds about 5.15 GB, taking the total to 155.3 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 40K; holding all of it at Q4_K_M would need about 157 GB.

Which build to download

This index tracks 2 formats for it — GGUF, AWQ — across 3 levels. Q4_K_M carries the lowest published perplexity loss at 2.2%. The fastest level measured here is AWQ INT4 at 14 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.

Common questions

How much VRAM does Qwen3 235B-A22B Instruct need?
About 149.7 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 155.3 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
Will Qwen3 235B-A22B Instruct run on a 256GB GPU?
Yes — at Q4_K_M and 4K context it needs about 149.7 GB, which leaves 106.3 GB spare on a Mac M5 Ultra 256G. That is the smallest card in this index that clears it comfortably; 2 of 63 do.
Which quantization of Qwen3 235B-A22B Instruct should I use?
Q4_K_M has the lowest published quality loss (2.2%), and Q4_K_M is the level most people run. All 3 levels in the index are Q4_K_M, Q3_K_M, AWQ INT4. Note this is a mixture-of-experts model — all parameters must be resident even though only a fraction are active per token, so the memory cost follows the total, not the active count.

Where this model fits

Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.

Ships in
GGUFAWQ
Compared here
GGUF vs AWQ

This page's figures change when the model or the runtime does.Last updated 2026-09-23 RSS → /feed.xml