Quant Hub

Structured index of open-source quantized models — no file hosting, just precise metadata · 81 models in index

Data updated 2026-09-21

81 models37 families4 quant formats≤3B · 107B · 2014B · 1732B · 1970B+ · 15

Quick GPU filter (4K context) · incl. tight fits

Parameters

Category

Hardware

Format

Recency

66 / 81 models

Showing all 66 models

Llama 3.1 8B Instruct

8B

Meta Llama 3.1

Meta's flagship 8B model with 128K context. Best-in-class for local deployment.

Q4_K_M · RTX 4090 · batch 1

5.7 GB

VRAM

131K

ctx

148

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Llama 3.1 70B Instruct

70B

Meta Llama 3.1

Meta's frontier 70B model. Requires 40GB+ VRAM; dual 3090 or M2 Ultra.

Q4_K_M · RTX 4090 · batch 1

43.5 GB

VRAM

131K

ctx

38

tok/s

Formats

GGUFAWQEXL2
Pro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Llama 3.2 3B Instruct

3B

Meta Llama 3.2

Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.

Q4_K_M · RTX 4090 · batch 1

2.2 GB

VRAM

131K

ctx

320

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen2.5 7B Instruct

7B

Alibaba Qwen2.5

Alibaba's highly optimized 7B. Punches well above its weight, especially in coding.

Q4_K_M · RTX 4090 · batch 1

5.4 GB

VRAM

131K

ctx

155

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen2.5 14B Instruct

14B

Alibaba Qwen2.5

The sweet spot between performance and resource usage. 16GB VRAM with Q4.

Q4_K_M · RTX 4090 · batch 1

10.2 GB

VRAM

131K

ctx

98

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen2.5 32B Instruct

32B

Alibaba Qwen2.5

Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.

Q4_K_M · RTX 4090 · batch 1

22.0 GB

VRAM

131K

ctx

44

tok/s

Formats

GGUFEXL2
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-Coder-V2-Lite Instruct

16B

DeepSeek

MoE architecture coding model. Active params ~2.4B, total ~16B. Exceptional code quality.

Q4_K_M · RTX 4090 · batch 1

11.1 GB

VRAM

164K

ctx

145

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Phi-3.5 Mini Instruct

3.8B

Microsoft Phi

Microsoft's tiny powerhouse. Best 4B model for on-device deployment.

Q4_K_M · RTX 4090 · batch 1

2.8 GB

VRAM

131K

ctx

298

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Mistral Nemo 12B Instruct

12B

Mistral AI

Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.

Q4_K_M · RTX 4090 · batch 1

8.5 GB

VRAM

131K

ctx

112

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Gemma 2 9B Instruct

9B

Google Gemma 2

Google's compact Gemma 2 with sliding window attention. Punches above 9B.

Q4_K_M · RTX 4090 · batch 1

6.5 GB

VRAM

8K

ctx

132

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen2.5 72B Instruct

72B

Alibaba Qwen2.5

Flagship Qwen2.5. Requires dual 4090 or A100 80G. Exceptional reasoning at scale.

Q4_K_M · RTX 4090 · batch 1

43.6 GB

VRAM

131K

ctx

28

tok/s

Formats

GGUFAWQEXL2
Pro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Qwen-14B

14B

DeepSeek

R1 reasoning distilled into 14B. Huge community interest; excellent chain-of-thought.

Q4_K_M · RTX 4090 · batch 1

10.2 GB

VRAM

131K

ctx

95

tok/s

Formats

GGUFEXL2AWQ
Consumer GPU
GGUF Q4_K_M
Hugging Face

Llama 3.3 70B Instruct

70B

Meta Llama 3.3

Latest Meta 70B with improved multilingual. Drop-in upgrade from Llama 3.1 70B.

Q4_K_M · RTX 4090 · batch 1

43.5 GB

VRAM

131K

ctx

38

tok/s

Formats

GGUFAWQ
Pro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Mistral Small 24B Instruct

24B

Mistral AI

Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.

Q4_K_M · RTX 4090 · batch 1

14.2 GB

VRAM

33K

ctx

62

tok/s

Formats

GGUFAWQEXL2
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen2.5-Coder 32B Instruct

32B

Alibaba Qwen2.5

Top-tier open coding model. HumanEval competitive with GPT-4o on 32B scale.

Q4_K_M · RTX 4090 · batch 1

22.0 GB

VRAM

131K

ctx

44

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen2.5-Coder 7B Instruct

7B

Alibaba Qwen2.5

Best 7B coding model. Ideal for local dev assistants on 8–16GB VRAM.

Q4_K_M · RTX 4090 · batch 1

5.4 GB

VRAM

131K

ctx

158

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen2.5 3B Instruct

3B

Alibaba Qwen2.5

Tiny Qwen2.5 for edge devices. Runs on 4GB VRAM or Raspberry Pi class hardware.

Q4_K_M · RTX 4090 · batch 1

2.1 GB

VRAM

33K

ctx

340

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Llama 3.2 1B Instruct

1B

Meta Llama 3.2

Ultra-light Llama for mobile and embedded. Sub-2GB VRAM with Q4.

Q4_K_M · RTX 4090 · batch 1

1.0 GB

VRAM

131K

ctx

520

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Llama-70B

70B

DeepSeek

R1 reasoning in Llama 70B architecture. Top open reasoning model for dual-GPU setups.

Q4_K_M · RTX 4090 · batch 1

43.5 GB

VRAM

131K

ctx

36

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Codestral 22B

22B

Mistral AI

Mistral's dedicated code model. 80+ language support, Fill-in-the-Middle capable.

Q4_K_M · RTX 4090 · batch 1

14.8 GB

VRAM

33K

ctx

58

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Llama 3.2 11B Vision Instruct

11B

Meta Llama 3.2

Multimodal Llama with image understanding. Vision encoder adds ~2GB VRAM overhead.

Q4_K_M · RTX 4090 · batch 1

9.5 GB

VRAM

131K

ctx

88

tok/s

Formats

GGUF
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Nous Hermes 3 Llama 3.1 8B

8B

NousResearch

Fine-tuned Llama 3.1 8B with improved roleplay and instruction following.

Q4_K_M · RTX 4090 · batch 1

5.7 GB

VRAM

131K

ctx

148

tok/s

Formats

GGUFEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Granite 3.1 8B Instruct

8B

IBM Granite

IBM's enterprise-grade 8B. Strong RAG and tool-use; permissive Apache 2.0 license.

Q4_K_M · RTX 4090 · batch 1

5.8 GB

VRAM

131K

ctx

142

tok/s

Formats

GGUFGPTQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Gemma 2 2B Instruct

2B

Google Gemma 2

Ultra-compact Gemma 2. Runs on 4GB VRAM; great for edge prototyping.

Q4_K_M · RTX 4090 · batch 1

2.0 GB

VRAM

8K

ctx

380

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Gemma 2 27B Instruct

27B

Google Gemma 2

Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.

Q4_K_M · RTX 4090 · batch 1

18.5 GB

VRAM

8K

ctx

48

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen2.5 0.5B Instruct

0.5B

Alibaba Qwen2.5

Smallest Qwen2.5. Ideal for Raspberry Pi, phones, and ultra-low-latency demos.

Q4_K_M · RTX 4090 · batch 1

0.6 GB

VRAM

33K

ctx

620

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen2.5 1.5B Instruct

1.5B

Alibaba Qwen2.5

Tiny Qwen with 128K context. Surprisingly capable for summarisation and chat.

Q4_K_M · RTX 4090 · batch 1

1.4 GB

VRAM

131K

ctx

480

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Phi-3 Medium 14B Instruct

14B

Microsoft Phi

Microsoft's mid-size Phi-3. Excellent quality-per-GB on 16GB cards.

Q4_K_M · RTX 4090 · batch 1

9.8 GB

VRAM

131K

ctx

102

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Phi-4 Mini Instruct

3.8B

Microsoft Phi

Latest Phi mini with improved math and code. Strong 4B-class performer.

Q4_K_M · RTX 4090 · batch 1

2.8 GB

VRAM

131K

ctx

305

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Mistral 7B Instruct v0.3

7B

Mistral AI

Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.

Q4_K_M · RTX 4090 · batch 1

5.2 GB

VRAM

33K

ctx

158

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

DeepSeek-V2-Lite Chat

16B

DeepSeek

MoE general model (~2.4B active). Long context and strong multilingual chat.

Q4_K_M · RTX 4090 · batch 1

11.0 GB

VRAM

164K

ctx

142

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Qwen-7B

7B

DeepSeek

R1 reasoning in a 7B footprint. Best value for 8–12GB VRAM CoT experiments.

Q4_K_M · RTX 4090 · batch 1

5.4 GB

VRAM

131K

ctx

152

tok/s

Formats

GGUFEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Qwen-32B

32B

DeepSeek

R1 distilled to 32B. Near-frontier reasoning on a single 24GB card (Q3/Q4).

Q4_K_M · RTX 4090 · batch 1

22.2 GB

VRAM

131K

ctx

42

tok/s

Formats

GGUFEXL2
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Llama 3.2 90B Vision Instruct

90B

Meta Llama 3.2

Flagship multimodal Llama. Requires dual 4090 or A100; vision adds ~3GB overhead.

Q4_K_M · RTX 4090 · batch 1

54.8 GB

VRAM

131K

ctx

22

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

OLMo 2 7B Instruct

7B

Allen AI OLMo

Fully open training pipeline from Allen AI. Great for reproducibility research.

Q4_K_M · RTX 4090 · batch 1

5.3 GB

VRAM

4K

ctx

150

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Falcon 3 10B Instruct

10B

TII UAE

Technology Innovation Institute's latest Falcon. Good multilingual and code mix.

Q4_K_M · RTX 4090 · batch 1

7.0 GB

VRAM

33K

ctx

118

tok/s

Formats

GGUFGPTQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen3 8B Instruct

8B

Alibaba Qwen3

Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.

Q4_K_M · RTX 4090 · batch 1

5.8 GB

VRAM

41K

ctx

142

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen3 14B Instruct

14B

Alibaba Qwen3

Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.

Q4_K_M · RTX 4090 · batch 1

10.5 GB

VRAM

41K

ctx

88

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Gemma 3 4B IT

4B

Google Gemma 3

Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.

Q4_K_M · RTX 4090 · batch 1

3.4 GB

VRAM

131K

ctx

175

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Gemma 3 12B IT

12B

Google Gemma 3

Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.

Q4_K_M · RTX 4090 · batch 1

8.8 GB

VRAM

131K

ctx

105

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Llama 4 Scout 17B (16E)

109B MoE

Meta Llama 4

Meta Llama 4 Scout MoE (17B active / 109B total). Multimodal; needs ~68GB VRAM at Q4_K_M.

Q4_K_M · RTX 4090 · batch 1

68.0 GB

VRAM

10486K

ctx

22

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Llama 4 Maverick 17B (128E)

400B MoE

Meta Llama 4

Llama 4 Maverick flagship MoE (17B active / 400B total). Multi-GPU or H100 cluster territory.

Q4_K_M · RTX 4090 · batch 1

245.0 GB

VRAM

1049K

ctx

8

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

Llama 3.1 405B Instruct

405B

Meta Llama 3.1

Meta frontier dense 405B. Q4 needs ~230GB+ VRAM; dual H100 80G or 8× consumer GPU.

Q4_K_M · RTX 4090 · batch 1

238.0 GB

VRAM

131K

ctx

6

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 32B Instruct

32B

Alibaba Qwen3

Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.

Q4_K_M · RTX 4090 · batch 1

22.5 GB

VRAM

41K

ctx

42

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 30B-A3B Instruct

30B-A3B

Alibaba Qwen3

Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.

Q4_K_M · RTX 4090 · batch 1

19.0 GB

VRAM

41K

ctx

95

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen3 235B-A22B Instruct

235B-A22B

Alibaba Qwen3

Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.

Q4_K_M · RTX 4090 · batch 1

142.0 GB

VRAM

41K

ctx

10

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-V3

671B MoE

DeepSeek

DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.

Q4_K_M · RTX 4090 · batch 1

385.0 GB

VRAM

164K

ctx

4

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-R1

671B MoE

DeepSeek

DeepSeek-R1 reasoning model built on V3 MoE. Chain-of-thought at frontier scale — use distill variants for local GPUs.

Q4_K_M · RTX 4090 · batch 1

385.0 GB

VRAM

164K

ctx

4

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 4B Instruct

4B

Alibaba Qwen3

Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.

Q4_K_M · RTX 4090 · batch 1

3.2 GB

VRAM

33K

ctx

168

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen3-Coder 30B-A3B Instruct

30B-A3B

Alibaba Qwen3

Agentic coding MoE with 3.3B active params and 256K native context. Top open coder for 16–24GB cards.

Q4_K_M · RTX 4090 · batch 1

19.2 GB

VRAM

262K

ctx

92

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Mistral Large 3 675B Instruct

675B MoE

Mistral AI

Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.

Q4_K_M · RTX 4090 · batch 1

388.0 GB

VRAM

262K

ctx

4

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

GLM-4-9B-Chat

9B

Zhipu GLM-4

Zhipu GLM-4 open 9B with 128K context, tool calling, and strong bilingual (EN/ZH) performance.

Q4_K_M · RTX 4090 · batch 1

6.2 GB

VRAM

131K

ctx

135

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Gemma 3 27B IT

27B

Google Gemma 3

Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.

Q4_K_M · RTX 4090 · batch 1

16.2 GB

VRAM

131K

ctx

48

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Llama-8B

8B

DeepSeek

R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.

Q4_K_M · RTX 4090 · batch 1

5.6 GB

VRAM

131K

ctx

145

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Phi-4 14B

14B

Microsoft Phi

Microsoft Phi-4 dense 14B — strong reasoning for size. Q4 ~9GB fits 12GB cards with moderate context.

Q4_K_M · RTX 4090 · batch 1

9.1 GB

VRAM

16K

ctx

88

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen3 1.7B Instruct

1.7B

Alibaba Qwen3

Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.

Q4_K_M · RTX 4090 · batch 1

1.4 GB

VRAM

33K

ctx

280

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

GPT-OSS 20B

21B MoE

OpenAI GPT-OSS

OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.

Q4_K_M · RTX 4090 · batch 1

11.9 GB

VRAM

131K

ctx

205

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

GPT-OSS 120B

117B MoE

OpenAI GPT-OSS

The big GPT-OSS (117B total / 5.1B active). Native MXFP4 checkpoint is ~61GB — fits one 80GB card or a 128GB unified-memory Mac. Partial offload on 24GB consumer cards is slow but works.

Q4_K_M · RTX 4090 · batch 1

58.5 GB

VRAM

131K

ctx

24

tok/s

Formats

GGUF
Pro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

GLM-4.5-Air

106B MoE

Zhipu GLM-4.5

Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.

Q4_K_M · RTX 4090 · batch 1

64.3 GB

VRAM

131K

ctx

18

tok/s

Formats

GGUF
Pro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Devstral Small 1.1 24B

24B

Mistral Devstral

Mistral + All Hands agentic coding model on a Mistral Small 3.1 base. Built for repo-scale tool use rather than single-file completion. Q4 ~14GB fits a 16GB card.

Q4_K_M · RTX 4090 · batch 1

14.3 GB

VRAM

131K

ctx

62

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Qwen3-VL 8B Instruct

8B
New

Alibaba Qwen3-VL

Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.

Q4_K_M · RTX 4090 · batch 1

5.9 GB

VRAM

41K

ctx

140

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen3-VL 30B-A3B Instruct

30B-A3B
New

Alibaba Qwen3-VL

Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.

Q4_K_M · RTX 4090 · batch 1

19.0 GB

VRAM

41K

ctx

95

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Magistral Small 1.2 24B

24B
New

Mistral Magistral

Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.

Q4_K_M · RTX 4090 · batch 1

14.3 GB

VRAM

131K

ctx

60

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Seed-OSS 36B Instruct

36B
New

ByteDance Seed

Dense 36B with a native 512K context. Q4 weights are ~22GB, which lands on a 24GB card or a 2×16GB split — but the context is the real cost: 512K of KV cache is roughly 128GB on its own, so budget context first and weights second.

Q4_K_M · RTX 4090 · batch 1

21.8 GB

VRAM

524K

ctx

42

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Ministral 3 8B Instruct

8B
New

Mistral Ministral 3

Mistral ships the GGUF itself, in 3B / 8B / 14B and Instruct or Reasoning variants — the 8B is the one that fits an 8GB card at Q4 with room for real context. Plain full attention on all 34 layers, so the memory estimates below behave the way the calculator assumes. Quality loss per level has not been published.

Q4_K_M · RTX 4090 · batch 1

5.2 GB

VRAM

262K

ctx

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Qwen3.8 27B

27B
New

Alibaba Qwen3.8

Hybrid attention: only 16 of its 64 layers keep a KV cache, which is why a 27B model holds a 262K context in about 16GB of cache rather than 64GB. Sizing it as a conventional stack overstates the cache fourfold — the estimates here count the 16 and match published measurements at 8K, 32K and 262K. Quality loss per level has not been published.

Q4_K_M · RTX 4090 · batch 1

15.3 GB

VRAM

262K

ctx

tok/s

Formats

GGUF
Consumer GPUPro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face