8GB VRAM

Mac M3 8G — what LLMs can it run?

34 of 79 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.

14B · 5

Mac M3 8G — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
Stable LM 2 12B Chat12BAWQ INT47.01 GB+1 GB142
Jamba 1.5 Mini12BAWQ INT46.82 GB+1.2 GB125
Solar 10.7B Instruct11BAWQ INT46.42 GB+1.6 GB168
Falcon 3 10B Instruct10BGPTQ INT46.12 GB+1.9 GB155
Gemma 2 9B Instruct9BAWQ INT46.27 GB+1.7 GB188

7B · 19

Mac M3 8G — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ4_K_M5.87 GB+2.1 GB135
Qwen3-VL 8B Instruct8BQ4_K_M6.19 GB+1.8 GB140
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+2.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+2.1 GB142
Qwen3 8B Instruct8BEXL2 4.65bpw5.59 GB+2.4 GB228
Llama 3.1 8B Instruct8BEXL2 4.65bpw5.43 GB+2.6 GB235
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+2.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+2.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+2.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+1.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+2.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+1.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+3.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+2.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+3.1 GB210
OLMo 2 7B Instruct7BQ4_K_M6.82 GB+1.2 GB150
Mistral 7B Instruct v0.37BQ6_K6.75 GB+1.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+1.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+2.6 GB145

≤3B · 10

Mac M3 8G — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+4 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+3.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+2.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+4 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+4.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+4.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+5.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+6.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+6.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+7.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

Stepping up

A Mac M3 16G (16GB) fits 17 more of the indexed models than this card. Mac M3 16G

34 of 79 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.