Quant Hub

Structured index of open-source quantized models — no file hosting, just precise metadata · 79 models in index

Data updated 2026-09-01

79 models35 families4 quant formats≤3B · 107B · 1914B · 1732B · 1870B+ · 15

Quick GPU filter (4K context)

Parameters

Category

Hardware

Format

Recency

79 / 79 models

Showing all 79 models

Llama 3.1 8B Instruct

8B

Meta Llama 3.1

Meta's flagship 8B model with 128K context. Best-in-class for local deployment.

3.2 GB

min VRAM

131K

ctx

235

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Llama 3.1 70B Instruct

70B

Meta Llama 3.1

Meta's frontier 70B model. Requires 40GB+ VRAM; dual 3090 or M2 Ultra.

33.4 GB

min VRAM

131K

ctx

62

tok/s

Formats

GGUFAWQEXL2
Pro GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Llama 3.2 3B Instruct

3B

Meta Llama 3.2

Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.

2.0 GB

min VRAM

131K

ctx

420

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Qwen2.5 7B Instruct

7B

Alibaba Qwen2.5

Alibaba's highly optimized 7B. Punches well above its weight, especially in coding.

4.8 GB

min VRAM

131K

ctx

245

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q6_K
Hugging Face

Qwen2.5 14B Instruct

14B

Alibaba Qwen2.5

The sweet spot between performance and resource usage. 16GB VRAM with Q4.

9.2 GB

min VRAM

131K

ctx

138

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Qwen2.5 32B Instruct

32B

Alibaba Qwen2.5

Near-GPT-4 reasoning on a 24GB VRAM card (Q4_K_S). Groundbreaking value.

16.4 GB

min VRAM

131K

ctx

68

tok/s

Formats

GGUFEXL2
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-Coder-V2-Lite Instruct

16B

DeepSeek

MoE architecture coding model. Active params ~2.4B, total ~16B. Exceptional code quality.

9.8 GB

min VRAM

164K

ctx

192

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q8_0
Hugging Face

Phi-3.5 Mini Instruct

3.8B

Microsoft Phi

Microsoft's tiny powerhouse. Best 4B model for on-device deployment.

2.5 GB

min VRAM

131K

ctx

385

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Mistral Nemo 12B Instruct

12B

Mistral AI

Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.

7.8 GB

min VRAM

131K

ctx

148

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q6_K
Hugging Face

Gemma 2 9B Instruct

9B

Google Gemma 2

Google's compact Gemma 2 with sliding window attention. Punches above 9B.

5.8 GB

min VRAM

8K

ctx

188

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q8_0
Hugging Face

Qwen2.5 72B Instruct

72B

Alibaba Qwen2.5

Flagship Qwen2.5. Requires dual 4090 or A100 80G. Exceptional reasoning at scale.

33.8 GB

min VRAM

131K

ctx

48

tok/s

Formats

GGUFAWQEXL2
Pro GPU
GGUF Q5_K_M
Hugging Face

DeepSeek-R1-Distill-Qwen-14B

14B

DeepSeek

R1 reasoning distilled into 14B. Huge community interest; excellent chain-of-thought.

9.2 GB

min VRAM

131K

ctx

128

tok/s

Formats

GGUFEXL2AWQ
Consumer GPU
EXL2 4.65bpw
Hugging Face

Llama 3.3 70B Instruct

70B

Meta Llama 3.3

Latest Meta 70B with improved multilingual. Drop-in upgrade from Llama 3.1 70B.

38.2 GB

min VRAM

131K

ctx

54

tok/s

Formats

GGUFAWQ
Pro GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Mistral Small 24B Instruct

24B

Mistral AI

Mistral's efficient 24B. Strong multilingual; fits on 24GB with Q4.

13.5 GB

min VRAM

33K

ctx

88

tok/s

Formats

GGUFAWQEXL2
Consumer GPUPro GPU
EXL2 4.65bpw
Hugging Face

Qwen2.5-Coder 32B Instruct

32B

Alibaba Qwen2.5

Top-tier open coding model. HumanEval competitive with GPT-4o on 32B scale.

16.4 GB

min VRAM

131K

ctx

65

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen2.5-Coder 7B Instruct

7B

Alibaba Qwen2.5

Best 7B coding model. Ideal for local dev assistants on 8–16GB VRAM.

4.8 GB

min VRAM

131K

ctx

248

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
EXL2 4.65bpw
Hugging Face

Qwen2.5 3B Instruct

3B

Alibaba Qwen2.5

Tiny Qwen2.5 for edge devices. Runs on 4GB VRAM or Raspberry Pi class hardware.

2.1 GB

min VRAM

33K

ctx

340

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Llama 3.2 1B Instruct

1B

Meta Llama 3.2

Ultra-light Llama for mobile and embedded. Sub-2GB VRAM with Q4.

1.0 GB

min VRAM

131K

ctx

520

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

DeepSeek-R1-Distill-Llama-70B

70B

DeepSeek

R1 reasoning in Llama 70B architecture. Top open reasoning model for dual-GPU setups.

38.2 GB

min VRAM

131K

ctx

52

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Codestral 22B

22B

Mistral AI

Mistral's dedicated code model. 80+ language support, Fill-in-the-Middle capable.

13.2 GB

min VRAM

33K

ctx

72

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Mixtral 8x7B Instruct

47B MoE

Mistral AI

Classic MoE model. ~13B active params per token; needs 32GB+ VRAM for Q4.

25.2 GB

min VRAM

33K

ctx

62

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Command R 35B

35B

Cohere

Cohere's RAG-optimised model. Excellent retrieval-augmented generation.

20.5 GB

min VRAM

131K

ctx

55

tok/s

Formats

GGUFGPTQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Yi 1.5 34B Chat

34B
Superseded

01.AI Yi

Prefer Qwen3 32B Instruct

01.AI's strong bilingual (EN/ZH) model. Competitive with Qwen 32B.

19.8 GB

min VRAM

4K

ctx

52

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Solar 10.7B Instruct

11B
Superseded

Upstage

Prefer Qwen3 14B Instruct

Depth-upscaled 10.7B punching above weight. Strong on reasoning benchmarks.

6.5 GB

min VRAM

4K

ctx

168

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

StarCoder2 15B

15B

BigCode

BigCode's open code model trained on 600+ languages. Great for polyglot dev.

9.2 GB

min VRAM

16K

ctx

115

tok/s

Formats

GGUFGPTQ
Consumer GPU
GGUF Q4_K_M
Hugging Face

Llama 3.2 11B Vision Instruct

11B

Meta Llama 3.2

Multimodal Llama with image understanding. Vision encoder adds ~2GB VRAM overhead.

9.5 GB

min VRAM

131K

ctx

88

tok/s

Formats

GGUF
Consumer GPUMac / Apple Silicon
GGUF Q8_0
Hugging Face

Qwen2-VL 7B Instruct

7B
Superseded

Alibaba Qwen2

Prefer Qwen3-VL 8B Instruct

Vision-language model with video understanding. Strong OCR and chart reading.

6.0 GB

min VRAM

33K

ctx

95

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Nous Hermes 3 Llama 3.1 8B

8B

NousResearch

Fine-tuned Llama 3.1 8B with improved roleplay and instruction following.

5.4 GB

min VRAM

131K

ctx

232

tok/s

Formats

GGUFEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
EXL2 4.65bpw
Hugging Face

WizardLM-2 7B

7B
Superseded

Microsoft / WizardLM

Prefer Qwen3 8B Instruct

Evol-Instruct fine-tuned Mistral-based 7B. Strong complex instruction handling.

4.8 GB

min VRAM

33K

ctx

218

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Granite 3.1 8B Instruct

8B

IBM Granite

IBM's enterprise-grade 8B. Strong RAG and tool-use; permissive Apache 2.0 license.

5.0 GB

min VRAM

131K

ctx

195

tok/s

Formats

GGUFGPTQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

Gemma 2 2B Instruct

2B

Google Gemma 2

Ultra-compact Gemma 2. Runs on 4GB VRAM; great for edge prototyping.

1.8 GB

min VRAM

8K

ctx

450

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Gemma 2 27B Instruct

27B

Google Gemma 2

Largest open Gemma 2. Strong reasoning; needs 24GB+ VRAM at Q4.

16.2 GB

min VRAM

8K

ctx

58

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q5_K_M
Hugging Face

Qwen2.5 0.5B Instruct

0.5B

Alibaba Qwen2.5

Smallest Qwen2.5. Ideal for Raspberry Pi, phones, and ultra-low-latency demos.

0.6 GB

min VRAM

33K

ctx

620

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Qwen2.5 1.5B Instruct

1.5B

Alibaba Qwen2.5

Tiny Qwen with 128K context. Surprisingly capable for summarisation and chat.

1.4 GB

min VRAM

131K

ctx

480

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Phi-3 Medium 14B Instruct

14B

Microsoft Phi

Microsoft's mid-size Phi-3. Excellent quality-per-GB on 16GB cards.

8.8 GB

min VRAM

131K

ctx

135

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q6_K
Hugging Face

Phi-4 Mini Instruct

3.8B

Microsoft Phi

Latest Phi mini with improved math and code. Strong 4B-class performer.

2.5 GB

min VRAM

131K

ctx

395

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Mistral 7B Instruct v0.3

7B

Mistral AI

Classic Mistral 7B v0.3. Still a reliable baseline for local chat APIs.

4.6 GB

min VRAM

33K

ctx

225

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q6_K
Hugging Face

DeepSeek-V2-Lite Chat

16B

DeepSeek

MoE general model (~2.4B active). Long context and strong multilingual chat.

9.6 GB

min VRAM

164K

ctx

188

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Qwen-7B

7B

DeepSeek

R1 reasoning in a 7B footprint. Best value for 8–12GB VRAM CoT experiments.

5.2 GB

min VRAM

131K

ctx

210

tok/s

Formats

GGUFEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
EXL2 4.65bpw
Hugging Face

DeepSeek-R1-Distill-Qwen-32B

32B

DeepSeek

R1 distilled to 32B. Near-frontier reasoning on a single 24GB card (Q3/Q4).

16.8 GB

min VRAM

131K

ctx

65

tok/s

Formats

GGUFEXL2
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Llama 3.2 90B Vision Instruct

90B

Meta Llama 3.2

Flagship multimodal Llama. Requires dual 4090 or A100; vision adds ~3GB overhead.

44.2 GB

min VRAM

131K

ctx

28

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

OLMo 2 7B Instruct

7B

Allen AI OLMo

Fully open training pipeline from Allen AI. Great for reproducibility research.

5.3 GB

min VRAM

4K

ctx

150

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

InternLM2 7B Chat

7B

Shanghai AI Lab

Strong bilingual (EN/ZH) 7B from Shanghai AI Lab. Competitive with Qwen 7B.

4.9 GB

min VRAM

33K

ctx

215

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

InternLM2 20B Chat

20B

Shanghai AI Lab

Mid-size InternLM2 with excellent Chinese comprehension. Fits 24GB at Q4.

13.8 GB

min VRAM

33K

ctx

78

tok/s

Formats

GGUF
Consumer GPUPro GPU
GGUF Q5_K_M
Hugging Face

Aya 23 8B

8B

Cohere For AI

Multilingual specialist covering 23 languages. Strong for non-English local apps.

5.0 GB

min VRAM

8K

ctx

205

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q4_K_M
Hugging Face

OpenChat 3.6 8B

8B
Superseded

OpenChat

Prefer Qwen3 8B Instruct

C-RLFT fine-tuned Llama 3.1 8B. Known for natural conversational tone.

5.4 GB

min VRAM

8K

ctx

228

tok/s

Formats

GGUFEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
EXL2 4.65bpw
Hugging Face

Zephyr 7B Beta

7B
Superseded

HuggingFaceH4

Prefer Qwen3 8B Instruct

DPO-aligned Mistral 7B. Classic choice for helpful, harmless chat baselines.

5.2 GB

min VRAM

33K

ctx

155

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q6_K
Hugging Face

Stable LM 2 12B Chat

12B

Stability AI

Stability AI's 12B chat model. Solid general-purpose option for 16GB GPUs.

7.2 GB

min VRAM

4K

ctx

142

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Falcon 3 10B Instruct

10B

TII UAE

Technology Innovation Institute's latest Falcon. Good multilingual and code mix.

6.2 GB

min VRAM

33K

ctx

155

tok/s

Formats

GGUFGPTQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Jamba 1.5 Mini

12B

AI21 Labs

Hybrid SSM-Transformer with 256K context. Efficient long-document QA on 16GB.

7.5 GB

min VRAM

262K

ctx

125

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

DBRX Instruct

132B

Databricks

MoE flagship (~36B active). Needs multi-GPU; strong code and reasoning at scale.

63.2 GB

min VRAM

33K

ctx

18

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 8B Instruct

8B

Alibaba Qwen3

Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.

5.1 GB

min VRAM

41K

ctx

228

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q6_K
Hugging Face

Qwen3 14B Instruct

14B

Alibaba Qwen3

Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.

9.5 GB

min VRAM

41K

ctx

125

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Gemma 3 4B IT

4B

Google Gemma 3

Google Gemma 3 multimodal 4B. 128K context; strong vision + text on 8GB cards.

3.0 GB

min VRAM

131K

ctx

210

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Gemma 3 12B IT

12B

Google Gemma 3

Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.

8.0 GB

min VRAM

131K

ctx

128

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Llama 4 Scout 17B (16E)

109B MoE

Meta Llama 4

Meta Llama 4 Scout MoE (17B active / 109B total). Multimodal; needs ~68GB VRAM at Q4_K_M.

55.0 GB

min VRAM

10486K

ctx

28

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Llama 4 Maverick 17B (128E)

400B MoE

Meta Llama 4

Llama 4 Maverick flagship MoE (17B active / 400B total). Multi-GPU or H100 cluster territory.

198.0 GB

min VRAM

1049K

ctx

10

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

Llama 3.1 405B Instruct

405B

Meta Llama 3.1

Meta frontier dense 405B. Q4 needs ~230GB+ VRAM; dual H100 80G or 8× consumer GPU.

192.0 GB

min VRAM

131K

ctx

10

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 32B Instruct

32B

Alibaba Qwen3

Qwen3 dense 32B — successor to Qwen2.5-32B with stronger reasoning and thinking mode.

16.8 GB

min VRAM

41K

ctx

62

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 30B-A3B Instruct

30B-A3B

Alibaba Qwen3

Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.

17.5 GB

min VRAM

41K

ctx

118

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Qwen3 235B-A22B Instruct

235B-A22B

Alibaba Qwen3

Qwen3 flagship MoE (22B active / 235B total). Q4_K_M ~142GB; rivals DeepSeek-R1 class models.

115.0 GB

min VRAM

41K

ctx

14

tok/s

Formats

GGUFAWQ
Pro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-V3

671B MoE

DeepSeek

DeepSeek-V3 frontier MoE (~37B active / 671B total). MLA + FP8; multi-node GPU cluster required at Q4.

310.0 GB

min VRAM

164K

ctx

5

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-R1

671B MoE

DeepSeek

DeepSeek-R1 reasoning model built on V3 MoE. Chain-of-thought at frontier scale — use distill variants for local GPUs.

310.0 GB

min VRAM

164K

ctx

5

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

Qwen3 4B Instruct

4B

Alibaba Qwen3

Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.

2.8 GB

min VRAM

33K

ctx

248

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q6_K
Hugging Face

Qwen3-Coder 30B-A3B Instruct

30B-A3B

Alibaba Qwen3

Agentic coding MoE with 3.3B active params and 256K native context. Top open coder for 16–24GB cards.

17.8 GB

min VRAM

262K

ctx

115

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Mistral Large 3 675B Instruct

675B MoE

Mistral AI

Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.

312.0 GB

min VRAM

262K

ctx

5

tok/s

Formats

GGUF
Pro GPU
GGUF Q4_K_M
Hugging Face

GLM-4-9B-Chat

9B

Zhipu GLM-4

Zhipu GLM-4 open 9B with 128K context, tool calling, and strong bilingual (EN/ZH) performance.

5.5 GB

min VRAM

131K

ctx

178

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q8_0
Hugging Face

Gemma 3 27B IT

27B
New

Google Gemma 3

Gemma 3 large instruct with long context and multimodal support. Q4 ~16GB — dual-GPU or 24GB card with short ctx.

13.0 GB

min VRAM

131K

ctx

62

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPU
GGUF Q4_K_M
Hugging Face

DeepSeek-R1-Distill-Llama-8B

8B
New

DeepSeek

R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.

4.9 GB

min VRAM

131K

ctx

210

tok/s

Formats

GGUFEXL2AWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q5_K_M
Hugging Face

Phi-4 14B

14B
New

Microsoft Phi

Microsoft Phi-4 dense 14B — strong reasoning for size. Q4 ~9GB fits 12GB cards with moderate context.

8.2 GB

min VRAM

16K

ctx

112

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face

Qwen3 1.7B Instruct

1.7B
New

Alibaba Qwen3

Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.

1.2 GB

min VRAM

33K

ctx

320

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

GPT-OSS 20B

21B MoE
New

OpenAI GPT-OSS

OpenAI open-weight MoE (21B total / 3.6B active), shipped natively in MXFP4 — ~12.8GB runs on a 16GB card with no quality tax. Only 3.6B active params means CPU-offload stays usable.

11.9 GB

min VRAM

131K

ctx

205

tok/s

Formats

GGUF
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF MXFP4
Hugging Face

GPT-OSS 120B

117B MoE
New

OpenAI GPT-OSS

The big GPT-OSS (117B total / 5.1B active). Native MXFP4 checkpoint is ~61GB — fits one 80GB card or a 128GB unified-memory Mac. Partial offload on 24GB consumer cards is slow but works.

58.5 GB

min VRAM

131K

ctx

24

tok/s

Formats

GGUF
Pro GPUMac / Apple Silicon
GGUF MXFP4
Hugging Face

GLM-4.5-Air

106B MoE
New

Zhipu GLM-4.5

Zhipu's agentic/reasoning MoE (106B total / 12B active). Q4 ~64GB — the sweet spot is a 96GB+ unified Mac or 2× 48GB cards. Strong tool-calling for its class.

34.8 GB

min VRAM

131K

ctx

26

tok/s

Formats

GGUF
Pro GPUMac / Apple Silicon
GGUF Q4_K_M
Hugging Face

Devstral Small 1.1 24B

24B
New

Mistral Devstral

Mistral + All Hands agentic coding model on a Mistral Small 3.1 base. Built for repo-scale tool use rather than single-file completion. Q4 ~14GB fits a 16GB card.

13.0 GB

min VRAM

131K

ctx

88

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q6_K
Hugging Face

Qwen3-VL 8B Instruct

8B
New

Alibaba Qwen3-VL

Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.

5.3 GB

min VRAM

41K

ctx

165

tok/s

Formats

GGUFAWQ
Consumer GPUMac / Apple SiliconCPU / VPS
GGUF Q8_0
Hugging Face

Qwen3-VL 30B-A3B Instruct

30B-A3B
New

Alibaba Qwen3-VL

Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.

15.2 GB

min VRAM

41K

ctx

112

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPUMac / Apple Silicon
GGUF Q8_0
Hugging Face

Magistral Small 1.2 24B

24B
New

Mistral Magistral

Mistral's reasoning model on a Mistral Small 3.2 base, with reasoning wrapped in [THINK] tags. Q4 ~14GB puts explicit chain-of-thought on a 16GB card — the level below the 70B-class reasoners most guides assume.

13.0 GB

min VRAM

131K

ctx

86

tok/s

Formats

GGUFAWQEXL2
Consumer GPUMac / Apple Silicon
GGUF Q6_K
Hugging Face

Seed-OSS 36B Instruct

36B
New

ByteDance Seed

Dense 36B with a native 512K context. Q4 weights are ~22GB, which lands on a 24GB card or a 2×16GB split — but the context is the real cost: 512K of KV cache is roughly 128GB on its own, so budget context first and weights second.

17.4 GB

min VRAM

524K

ctx

55

tok/s

Formats

GGUFAWQ
Consumer GPUPro GPUMac / Apple Silicon
GGUF Q5_K_M
Hugging Face