128GB VRAM

Mac M4 Max 128G — what LLMs can it run?

75 of 81 indexed models fit comfortably in 128GB at 4K context, each at the highest-quality quant that still leaves headroom.

70B+ · 9

Mac M4 Max 128G — what LLMs can it run?70B+
ModelQuantEst. VRAMHeadroomtok/s
DBRX Instruct132BQ4_K_M84.31 GB+43.7 GB15
GPT-OSS 120B117B MoEMXFP465.15 GB+62.9 GB22
Llama 4 Scout 17B (16E)109B MoEQ4_K_M69.88 GB+58.1 GB22
GLM-4.5-Air106B MoEQ4_K_M67.94 GB+60.1 GB18
Llama 3.2 90B Vision Instruct90BQ4_K_M57.5 GB+70.5 GB22
Qwen2.5 72B Instruct72BQ5_K_M55.31 GB+72.7 GB24
Llama 3.1 70B Instruct70BQ5_K_M53.75 GB+74.3 GB32
Llama 3.3 70B Instruct70BQ5_K_M53.75 GB+74.3 GB32
DeepSeek-R1-Distill-Llama-70B70BQ4_K_M46.1 GB+81.9 GB36

32B · 19

Mac M4 Max 128G — what LLMs can it run?32B
ModelQuantEst. VRAMHeadroomtok/s
Mixtral 8x7B Instruct47B MoEQ4_K_M30.13 GB+97.9 GB48
Seed-OSS 36B Instruct36BQ5_K_M27.81 GB+100.2 GB37
Command R 35B35BQ4_K_M22.86 GB+105.1 GB42
Yi 1.5 34B Chat34BQ4_K_M22.82 GB+105.2 GB40
Qwen3 32B Instruct32BQ4_K_M21.85 GB+106.2 GB42
Qwen2.5 32B Instruct32BQ4_K_M21.69 GB+106.3 GB44
Qwen2.5-Coder 32B Instruct32BQ4_K_M21.69 GB+106.3 GB44
DeepSeek-R1-Distill-Qwen-32B32BQ4_K_M21.69 GB+106.3 GB42
Qwen3 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+105 GB82
Qwen3-Coder 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+105 GB80
Qwen3-VL 30B-A3B Instruct30B-A3BQ8_034.28 GB+93.7 GB78
Qwen3.8 27B27BQ4_K_M17.47 GB+110.5 GB
Gemma 3 27B IT27BQ4_K_M19.49 GB+108.5 GB48
Gemma 2 27B Instruct27BQ5_K_M21.76 GB+106.2 GB42
Mistral Small 24B Instruct24BEXL2 4.65bpw15.26 GB+112.7 GB88
Devstral Small 1.1 24B24BQ6_K20.91 GB+107.1 GB48
Magistral Small 1.2 24B24BQ6_K20.91 GB+107.1 GB47
Codestral 22B22BQ4_K_M15.03 GB+113 GB58
GPT-OSS 20B21B MoEMXFP411.81 GB+116.2 GB195

14B · 17

Mac M4 Max 128G — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
InternLM2 20B Chat20BQ5_K_M15.59 GB+112.4 GB68
DeepSeek-Coder-V2-Lite Instruct16BQ8_018.36 GB+109.6 GB118
DeepSeek-V2-Lite Chat16BQ4_K_M10.87 GB+117.1 GB142
StarCoder2 15B15BQ4_K_M10.19 GB+117.8 GB92
Qwen3 14B Instruct14BQ5_K_M11.65 GB+116.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+116.3 GB86
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+118.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+116.2 GB78
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+115.1 GB88
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+116.9 GB95
Gemma 3 12B IT12BQ5_K_M10.6 GB+117.4 GB92
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+119.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+119.9 GB95
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+115.1 GB72
Solar 10.7B Instruct11BQ4_K_M7.6 GB+120.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+120.7 GB118
Gemma 2 9B Instruct9BQ8_011.7 GB+116.3 GB108

7B · 20

Mac M4 Max 128G — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+117.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+117.6 GB108
Ministral 3 8B Instruct8BQ8_010.02 GB+118 GB
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+122.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+122.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+120.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+118.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+122.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+122.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+122.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+121.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+122.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+121.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+123.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+122.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+123.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+117.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+121.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+121.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+122.6 GB145

≤3B · 10

Mac M4 Max 128G — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+124 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+123.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+122.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+124 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+124.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+124.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+125.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+126.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+126.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+127.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 1 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a Mac M4 Max 128G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements.

Cards with the same budget

What fits is decided by memory, so every 128GB card of this type returns the same list. These pages are not different answers — they differ in throughput, which this index does not measure per card.

Stepping up

A Mac M3 Ultra 192G (192GB) fits 1 more of the indexed models than this card. Mac M3 Ultra 192G

75 of 81 indexed models fit comfortably in 128GB at 4K context, each at the highest-quality quant that still leaves headroom.