Quant Hub
Structured index of open-source quantized models — no file hosting, just precise metadata · 81 models in index
Data updated 2026-09-21
Quick GPU filter (4K context) · incl. tight fits
Parameters
Category
Hardware
Format
Recency
66 / 81 models
Showing all 66 models
Llama 3.1 8B Instruct
8BMeta Llama 3.1
Meta's flagship 8B model with 128K context. Best-in-class for local deployment.
Q4_K_M · RTX 4090 · batch 1
5.7 GB
VRAM
131K
ctx
148
tok/s
Formats
Llama 3.1 70B Instruct
70BMeta Llama 3.1
Meta's frontier 70B model. Requires 40GB+ VRAM; dual 3090 or M2 Ultra.
Q4_K_M · RTX 4090 · batch 1
43.5 GB
VRAM
131K
ctx
38
tok/s
Formats
Llama 3.2 3B Instruct
3BMeta Llama 3.2
Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.
Q4_K_M · RTX 4090 · batch 1
2.2 GB
VRAM
131K
ctx
320
tok/s
Formats
Qwen2.5 7B Instruct
7BAlibaba Qwen2.5
Alibaba's highly optimized 7B. Punches well above its weight, especially in coding.
Q4_K_M · RTX 4090 · batch 1
5.4 GB
VRAM
131K
ctx
155
tok/s
Formats
Qwen2.5 14B Instruct
14BAlibaba Qwen2.5
The sweet spot between performance and resource usage. 16GB VRAM with Q4.
Q4_K_M · RTX 4090 · batch 1
10.2 GB
VRAM
131K
ctx
98
tok/s
Formats
Qwen2.5 32B Instruct
32BAlibaba Qwen2.5
Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.
Q4_K_M · RTX 4090 · batch 1
22.0 GB
VRAM
131K
ctx
44
tok/s
Formats
DeepSeek-Coder-V2-Lite Instruct
16BDeepSeek
MoE architecture coding model. Active params ~2.4B, total ~16B. Exceptional code quality.
Q4_K_M · RTX 4090 · batch 1
11.1 GB
VRAM
164K
ctx
145
tok/s
Formats
Phi-3.5 Mini Instruct
3.8BMicrosoft Phi
Microsoft's tiny powerhouse. Best 4B model for on-device deployment.
Q4_K_M · RTX 4090 · batch 1
2.8 GB
VRAM
131K
ctx
298
tok/s
Formats
Mistral Nemo 12B Instruct
12BMistral AI
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
Q4_K_M · RTX 4090 · batch 1
8.5 GB
VRAM
131K
ctx
112
tok/s
Formats
Gemma 2 9B Instruct
9BGoogle Gemma 2
Google's compact Gemma 2 with sliding window attention. Punches above 9B.
Q4_K_M · RTX 4090 · batch 1
6.5 GB
VRAM
8K
ctx
132
tok/s
Formats
Qwen2.5 72B Instruct
72BAlibaba Qwen2.5
Flagship Qwen2.5. Requires dual 4090 or A100 80G. Exceptional reasoning at scale.
Q4_K_M · RTX 4090 · batch 1
43.6 GB
VRAM
131K
ctx
28
tok/s
Formats
DeepSeek-R1-Distill-Qwen-14B
14BDeepSeek
R1 reasoning distilled into 14B. Huge community interest; excellent chain-of-thought.
Q4_K_M · RTX 4090 · batch 1
10.2 GB
VRAM
131K
ctx
95
tok/s
Formats
Llama 3.3 70B Instruct
70BMeta Llama 3.3
Latest Meta 70B with improved multilingual. Drop-in upgrade from Llama 3.1 70B.
Q4_K_M · RTX 4090 · batch 1
43.5 GB
VRAM
131K
ctx
38
tok/s
Formats
Mistral Small 24B Instruct
24BMistral AI
Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.
Q4_K_M · RTX 4090 · batch 1
14.2 GB
VRAM
33K
ctx
62
tok/s
Formats
Qwen2.5-Coder 32B Instruct
32BAlibaba Qwen2.5
Top-tier open coding model. HumanEval competitive with GPT-4o on 32B scale.
Q4_K_M · RTX 4090 · batch 1
22.0 GB
VRAM
131K
ctx
44
tok/s
Formats
Qwen2.5-Coder 7B Instruct
7BAlibaba Qwen2.5
Best 7B coding model. Ideal for local dev assistants on 8–16GB VRAM.
Q4_K_M · RTX 4090 · batch 1
5.4 GB
VRAM
131K
ctx
158
tok/s
Formats
Qwen2.5 3B Instruct
3BAlibaba Qwen2.5
Tiny Qwen2.5 for edge devices. Runs on 4GB VRAM or Raspberry Pi class hardware.
Q4_K_M · RTX 4090 · batch 1
2.1 GB
VRAM
33K
ctx
340
tok/s
Formats
Llama 3.2 1B Instruct
1BMeta Llama 3.2
Ultra-light Llama for mobile and embedded. Sub-2GB VRAM with Q4.
Q4_K_M · RTX 4090 · batch 1
1.0 GB
VRAM
131K
ctx
520
tok/s
Formats
DeepSeek-R1-Distill-Llama-70B
70BDeepSeek
R1 reasoning in Llama 70B architecture. Top open reasoning model for dual-GPU setups.
Q4_K_M · RTX 4090 · batch 1
43.5 GB
VRAM
131K
ctx
36
tok/s
Formats
Codestral 22B
22BMistral AI
Mistral's dedicated code model. 80+ language support, Fill-in-the-Middle capable.
Q4_K_M · RTX 4090 · batch 1
14.8 GB
VRAM
33K
ctx
58
tok/s
Formats
Llama 3.2 11B Vision Instruct
11BMeta Llama 3.2
Multimodal Llama with image understanding. Vision encoder adds ~2GB VRAM overhead.
Q4_K_M · RTX 4090 · batch 1
9.5 GB
VRAM
131K
ctx
88
tok/s
Formats
Nous Hermes 3 Llama 3.1 8B
8BNousResearch
Fine-tuned Llama 3.1 8B with improved roleplay and instruction following.
Q4_K_M · RTX 4090 · batch 1
5.7 GB
VRAM
131K
ctx
148
tok/s
Formats
Granite 3.1 8B Instruct
8BIBM Granite
IBM's enterprise-grade 8B. Strong RAG and tool-use; permissive Apache 2.0 license.
Q4_K_M · RTX 4090 · batch 1
5.8 GB
VRAM
131K
ctx
142
tok/s
Formats
Gemma 2 2B Instruct
2BGoogle Gemma 2
Ultra-compact Gemma 2. Runs on 4GB VRAM; great for edge prototyping.
Q4_K_M · RTX 4090 · batch 1
2.0 GB
VRAM
8K
ctx
380
tok/s
Formats
Gemma 2 27B Instruct
27BGoogle Gemma 2
Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.
Q4_K_M · RTX 4090 · batch 1
18.5 GB
VRAM
8K
ctx
48
tok/s
Formats
Qwen2.5 0.5B Instruct
0.5BAlibaba Qwen2.5
Smallest Qwen2.5. Ideal for Raspberry Pi, phones, and ultra-low-latency demos.
Q4_K_M · RTX 4090 · batch 1
0.6 GB
VRAM
33K
ctx
620
tok/s
Formats
Qwen2.5 1.5B Instruct
1.5BAlibaba Qwen2.5
Tiny Qwen with 128K context. Surprisingly capable for summarisation and chat.
Q4_K_M · RTX 4090 · batch 1
1.4 GB
VRAM
131K
ctx
480
tok/s
Formats
Phi-3 Medium 14B Instruct
14BMicrosoft Phi
Microsoft's mid-size Phi-3. Excellent quality-per-GB on 16GB cards.
Q4_K_M · RTX 4090 · batch 1
9.8 GB
VRAM
131K
ctx
102
tok/s
Formats
Phi-4 Mini Instruct
3.8BMicrosoft Phi
Latest Phi mini with improved math and code. Strong 4B-class performer.
Q4_K_M · RTX 4090 · batch 1
2.8 GB
VRAM
131K
ctx
305
tok/s
Formats
Mistral 7B Instruct v0.3
7BMistral AI
Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.
Q4_K_M · RTX 4090 · batch 1
5.2 GB
VRAM
33K
ctx
158
tok/s
Formats
DeepSeek-V2-Lite Chat
16BDeepSeek
MoE general model (~2.4B active). Long context and strong multilingual chat.
Q4_K_M · RTX 4090 · batch 1
11.0 GB
VRAM
164K
ctx
142
tok/s
Formats
DeepSeek-R1-Distill-Qwen-7B
7BDeepSeek
R1 reasoning in a 7B footprint. Best value for 8–12GB VRAM CoT experiments.
Q4_K_M · RTX 4090 · batch 1
5.4 GB
VRAM
131K
ctx
152
tok/s
Formats
DeepSeek-R1-Distill-Qwen-32B
32BDeepSeek
R1 distilled to 32B. Near-frontier reasoning on a single 24GB card (Q3/Q4).
Q4_K_M · RTX 4090 · batch 1
22.2 GB
VRAM
131K
ctx
42
tok/s
Formats
Llama 3.2 90B Vision Instruct
90BMeta Llama 3.2
Flagship multimodal Llama. Requires dual 4090 or A100; vision adds ~3GB overhead.
Q4_K_M · RTX 4090 · batch 1
54.8 GB
VRAM
131K
ctx
22
tok/s
Formats
OLMo 2 7B Instruct
7BAllen AI OLMo
Fully open training pipeline from Allen AI. Great for reproducibility research.
Q4_K_M · RTX 4090 · batch 1
5.3 GB
VRAM
4K
ctx
150
tok/s
Formats
Falcon 3 10B Instruct
10BTII UAE
Technology Innovation Institute's latest Falcon. Good multilingual and code mix.
Q4_K_M · RTX 4090 · batch 1
7.0 GB
VRAM
33K
ctx
118
tok/s
Formats
Qwen3 8B Instruct
8BAlibaba Qwen3
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
Q4_K_M · RTX 4090 · batch 1
5.8 GB
VRAM
41K
ctx
142
tok/s
Formats
Qwen3 14B Instruct
14BAlibaba Qwen3
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
Q4_K_M · RTX 4090 · batch 1
10.5 GB
VRAM
41K
ctx
88
tok/s
Formats
Gemma 3 4B IT
4BGoogle Gemma 3
Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.
Q4_K_M · RTX 4090 · batch 1
3.4 GB
VRAM
131K
ctx
175
tok/s
Formats
Gemma 3 12B IT
12BGoogle Gemma 3
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
Q4_K_M · RTX 4090 · batch 1
8.8 GB
VRAM
131K
ctx
105
tok/s
Formats
Llama 4 Scout 17B (16E)
109B MoEMeta Llama 4
Meta Llama 4 Scout MoE (17B active / 109B total). Multimodal; needs ~68GB VRAM at Q4_K_M.
Q4_K_M · RTX 4090 · batch 1
68.0 GB
VRAM
10486K
ctx
22
tok/s
Formats
Llama 4 Maverick 17B (128E)
400B MoEMeta Llama 4
Llama 4 Maverick flagship MoE (17B active / 400B total). Multi-GPU or H100 cluster territory.
Q4_K_M · RTX 4090 · batch 1
245.0 GB
VRAM
1049K
ctx
8
tok/s
Formats
Llama 3.1 405B Instruct
405BMeta Llama 3.1
Meta frontier dense 405B. Q4 needs ~230GB+ VRAM; dual H100 80G or 8× consumer GPU.
Q4_K_M · RTX 4090 · batch 1
238.0 GB
VRAM
131K
ctx
6
tok/s
Formats
Qwen3 32B Instruct
32BAlibaba Qwen3
Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.
Q4_K_M · RTX 4090 · batch 1
22.5 GB
VRAM
41K
ctx
42
tok/s
Formats
Qwen3 30B-A3B Instruct
30B-A3BAlibaba Qwen3
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.
Q4_K_M · RTX 4090 · batch 1
19.0 GB
VRAM
41K
ctx
95
tok/s
Formats
Qwen3 235B-A22B Instruct
235B-A22BAlibaba Qwen3
Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.
Q4_K_M · RTX 4090 · batch 1
142.0 GB
VRAM
41K
ctx
10
tok/s
Formats
DeepSeek-V3
671B MoEDeepSeek
DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.
Q4_K_M · RTX 4090 · batch 1
385.0 GB
VRAM
164K
ctx
4
tok/s
Formats
DeepSeek-R1
671B MoEDeepSeek
DeepSeek-R1 reasoning model built on V3 MoE. Chain-of-thought at frontier scale — use distill variants for local GPUs.
Q4_K_M · RTX 4090 · batch 1
385.0 GB
VRAM
164K
ctx
4
tok/s
Formats
Qwen3 4B Instruct
4BAlibaba Qwen3
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
Q4_K_M · RTX 4090 · batch 1
3.2 GB
VRAM
33K
ctx
168
tok/s
Formats
Qwen3-Coder 30B-A3B Instruct
30B-A3BAlibaba Qwen3
Agentic coding MoE with 3.3B active params and 256K native context. Top open coder for 16–24GB cards.
Q4_K_M · RTX 4090 · batch 1
19.2 GB
VRAM
262K
ctx
92
tok/s
Formats
Mistral Large 3 675B Instruct
675B MoEMistral AI
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.
Q4_K_M · RTX 4090 · batch 1
388.0 GB
VRAM
262K
ctx
4
tok/s
Formats
GLM-4-9B-Chat
9BZhipu GLM-4
Zhipu GLM-4 open 9B with 128K context, tool calling, and strong bilingual (EN/ZH) performance.
Q4_K_M · RTX 4090 · batch 1
6.2 GB
VRAM
131K
ctx
135
tok/s
Formats
Gemma 3 27B IT
27BGoogle Gemma 3
Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.
Q4_K_M · RTX 4090 · batch 1
16.2 GB
VRAM
131K
ctx
48
tok/s
Formats
DeepSeek-R1-Distill-Llama-8B
8BDeepSeek
R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.
Q4_K_M · RTX 4090 · batch 1
5.6 GB
VRAM
131K
ctx
145
tok/s
Formats
Phi-4 14B
14BMicrosoft Phi
Microsoft Phi-4 dense 14B — strong reasoning for size. Q4 ~9GB fits 12GB cards with moderate context.
Q4_K_M · RTX 4090 · batch 1
9.1 GB
VRAM
16K
ctx
88
tok/s
Formats
Qwen3 1.7B Instruct
1.7BAlibaba Qwen3
Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.
Q4_K_M · RTX 4090 · batch 1
1.4 GB
VRAM
33K
ctx
280
tok/s
Formats
GPT-OSS 20B
21B MoEOpenAI GPT-OSS
OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.
Q4_K_M · RTX 4090 · batch 1
11.9 GB
VRAM
131K
ctx
205
tok/s
Formats
GPT-OSS 120B
117B MoEOpenAI GPT-OSS
The big GPT-OSS (117B total / 5.1B active). Native MXFP4 checkpoint is ~61GB — fits one 80GB card or a 128GB unified-memory Mac. Partial offload on 24GB consumer cards is slow but works.
Q4_K_M · RTX 4090 · batch 1
58.5 GB
VRAM
131K
ctx
24
tok/s
Formats
GLM-4.5-Air
106B MoEZhipu GLM-4.5
Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.
Q4_K_M · RTX 4090 · batch 1
64.3 GB
VRAM
131K
ctx
18
tok/s
Formats
Devstral Small 1.1 24B
24BMistral Devstral
Mistral + All Hands agentic coding model on a Mistral Small 3.1 base. Built for repo-scale tool use rather than single-file completion. Q4 ~14GB fits a 16GB card.
Q4_K_M · RTX 4090 · batch 1
14.3 GB
VRAM
131K
ctx
62
tok/s
Formats
Qwen3-VL 8B Instruct
8BAlibaba Qwen3-VL
Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.
Q4_K_M · RTX 4090 · batch 1
5.9 GB
VRAM
41K
ctx
140
tok/s
Formats
Qwen3-VL 30B-A3B Instruct
30B-A3BAlibaba Qwen3-VL
Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.
Q4_K_M · RTX 4090 · batch 1
19.0 GB
VRAM
41K
ctx
95
tok/s
Formats
Magistral Small 1.2 24B
24BMistral Magistral
Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.
Q4_K_M · RTX 4090 · batch 1
14.3 GB
VRAM
131K
ctx
60
tok/s
Formats
Seed-OSS 36B Instruct
36BByteDance Seed
Dense 36B with a native 512K context. Q4 weights are ~22GB, which lands on a 24GB card or a 2×16GB split — but the context is the real cost: 512K of KV cache is roughly 128GB on its own, so budget context first and weights second.
Q4_K_M · RTX 4090 · batch 1
21.8 GB
VRAM
524K
ctx
42
tok/s
Formats
Ministral 3 8B Instruct
8BMistral Ministral 3
Mistral ships the GGUF itself, in 3B / 8B / 14B and Instruct or Reasoning variants — the 8B is the one that fits an 8GB card at Q4 with room for real context. Plain full attention on all 34 layers, so the memory estimates below behave the way the calculator assumes. Quality loss per level has not been published.
Q4_K_M · RTX 4090 · batch 1
5.2 GB
VRAM
262K
ctx
—
tok/s
Formats
Qwen3.8 27B
27BAlibaba Qwen3.8
Hybrid attention: only 16 of its 64 layers keep a KV cache, which is why a 27B model holds a 262K context in about 16GB of cache rather than 64GB. Sizing it as a conventional stack overstates the cache fourfold — the estimates here count the 16 and match published measurements at 8K, 32K and 262K. Quality loss per level has not been published.
Q4_K_M · RTX 4090 · batch 1
15.3 GB
VRAM
262K
ctx
—
tok/s
Formats