AWQ vs EXL2: which quantization format should you use?
Compared on hardware support, runtime, quality and what the 79-model index actually ships.
At a glance
| AWQ | EXL2 | |
|---|---|---|
| Runs on | NVIDIA GPU (CUDA 11.8+) | NVIDIA GPU (Ampere+ recommended) |
| Runtime | vLLM · AutoAWQ · TGI | ExLlamaV2 · TabbyAPI |
| Best for | High-throughput API server | Max single-GPU performance |
| Adoption estimate | 45% | 32% |
| Models in this index | 53 / 79 | 21 / 79 |
AWQ
Activation-Aware Weight Quantization. High-accuracy INT4 for NVIDIA. Pairs perfectly with vLLM for server deployment.
Strengths
- Best accuracy at 4-bit
- Blazing fast with vLLM
- Excellent batch throughput
Trade-offs
- NVIDIA only
- More setup than GGUF
EXL2
ExLlamaV2 format. Mixed-precision per-layer quantization — best accuracy-per-bit ratio. Fastest single-GPU inference available.
Strengths
- Fastest GPU inference
- Best accuracy per bit
- Ultra-low 2bpw option
Trade-offs
- NVIDIA only
- Steeper learning curve
Models that ship both
These 16 models publish weights in both formats, so the two rows describe the same model rather than two different ones — the only place this comparison is measurable rather than editorial.
| Model | AWQ Level | Quality loss | EXL2 Level | Quality loss |
|---|---|---|---|---|
| Qwen2.5 72B Instruct | AWQ INT4 | 3.5% | EXL2 3.5bpw | 4.8% |
| Llama 3.1 70B Instruct | AWQ INT4 | 3.9% | EXL2 3.5bpw | 5.2% |
| Qwen3 32B Instruct | AWQ INT4 | 3.6% | EXL2 3.5bpw | 4.5% |
| Qwen2.5-Coder 32B Instruct | AWQ INT4 | 3.5% | EXL2 3.5bpw | 4.5% |
| Mistral Small 24B Instruct | AWQ INT4 | 3.8% | EXL2 4.65bpw | 2.2% |
| Devstral Small 1.1 24B | AWQ INT4 | 4% | EXL2 4.65bpw | 2.5% |
| Magistral Small 1.2 24B | AWQ INT4 | 4% | EXL2 4.65bpw | 2.5% |
| Qwen3 14B Instruct | AWQ INT4 | 3.5% | EXL2 4.65bpw | 1.8% |
| Qwen2.5 14B Instruct | AWQ INT4 | 3.8% | EXL2 4.65bpw | 2.1% |
| DeepSeek-R1-Distill-Qwen-14B | AWQ INT4 | 3.6% | EXL2 4.65bpw | 2% |
| Qwen3 8B Instruct | AWQ INT4 | 3.8% | EXL2 4.65bpw | 2% |
| Llama 3.1 8B Instruct | AWQ INT4 | 4.5% | EXL2 4.65bpw | 2.5% |
| DeepSeek-R1-Distill-Llama-8B | AWQ INT4 | 3.5% | EXL2 4.65bpw | 1.9% |
| Qwen2.5 7B Instruct | AWQ INT4 | 4.2% | EXL2 4.65bpw | 2.2% |
| Qwen2.5-Coder 7B Instruct | AWQ INT4 | 4% | EXL2 4.65bpw | 2% |
| Qwen3 4B Instruct | AWQ INT4 | 3.8% | EXL2 4.65bpw | 2.1% |
Quality loss is perplexity increase against the unquantized weights — lower is better, and only comparable within one model.