18GB VRAM

Mac M3 Pro 18G — what LLMs can it run?

53 of 79 indexed models fit comfortably in 18GB at 4K context, each at the highest-quality quant that still leaves headroom.

32B · 7

Mac M3 Pro 18G — what LLMs can it run?32B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3-VL 30B-A3B Instruct30B-A3BQ3_K_M15.83 GB+2.2 GB104
Gemma 2 27B Instruct27BAWQ INT415.79 GB+2.2 GB58
Mistral Small 24B Instruct24BEXL2 4.65bpw15.26 GB+2.7 GB88
Devstral Small 1.1 24B24BEXL2 4.65bpw15.02 GB+3 GB88
Magistral Small 1.2 24B24BEXL2 4.65bpw15.02 GB+3 GB86
Codestral 22B22BQ4_K_M15.03 GB+3 GB58
GPT-OSS 20B21B MoEMXFP411.81 GB+6.2 GB195

14B · 17

Mac M3 Pro 18G — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
InternLM2 20B Chat20BQ5_K_M15.59 GB+2.4 GB68
DeepSeek-Coder-V2-Lite Instruct16BQ4_K_M10.87 GB+7.1 GB145
DeepSeek-V2-Lite Chat16BQ4_K_M10.87 GB+7.1 GB142
StarCoder2 15B15BQ4_K_M10.19 GB+7.8 GB92
Qwen3 14B Instruct14BQ5_K_M11.65 GB+6.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+6.3 GB86
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+8.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+6.2 GB78
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+5.1 GB88
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+6.9 GB95
Gemma 3 12B IT12BQ5_K_M10.6 GB+7.4 GB92
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+9.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+9.9 GB95
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+5.1 GB72
Solar 10.7B Instruct11BQ4_K_M7.6 GB+10.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+10.7 GB118
Gemma 2 9B Instruct9BQ8_011.7 GB+6.3 GB108

7B · 19

Mac M3 Pro 18G — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+7.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+7.6 GB108
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+12.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+12.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+10.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+8.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+12.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+12.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+12.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+11.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+12.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+11.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+13.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+12.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+13.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+7.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+11.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+11.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+12.6 GB145

≤3B · 10

Mac M3 Pro 18G — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+14 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+13.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+12.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+14 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+14.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+14.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+15.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+16.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+16.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+17.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

Stepping up

A Mac M3 Pro 36G (36GB) fits 11 more of the indexed models than this card. Mac M3 Pro 36G

53 of 79 indexed models fit comfortably in 18GB at 4K context, each at the highest-quality quant that still leaves headroom.