192GB VRAM

Mac M3 Ultra 192G — what LLMs can it run?

74 of 79 indexed models fit comfortably in 192GB at 4K context, each at the highest-quality quant that still leaves headroom.

70B+ · 10

Mac M3 Ultra 192G — what LLMs can it run?70B+
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 235B-A22B Instruct235B-A22BQ4_K_M149.68 GB+42.3 GB10
DBRX Instruct132BQ4_K_M84.31 GB+107.7 GB15
GPT-OSS 120B117B MoEMXFP465.15 GB+126.9 GB22
Llama 4 Scout 17B (16E)109B MoEQ4_K_M69.88 GB+122.1 GB22
GLM-4.5-Air106B MoEQ4_K_M67.94 GB+124.1 GB18
Llama 3.2 90B Vision Instruct90BQ4_K_M57.5 GB+134.5 GB22
Qwen2.5 72B Instruct72BQ5_K_M55.31 GB+136.7 GB24
Llama 3.1 70B Instruct70BQ5_K_M53.75 GB+138.3 GB32
Llama 3.3 70B Instruct70BQ5_K_M53.75 GB+138.3 GB32
DeepSeek-R1-Distill-Llama-70B70BQ4_K_M46.1 GB+145.9 GB36

32B · 18

Mac M3 Ultra 192G — what LLMs can it run?32B
ModelQuantEst. VRAMHeadroomtok/s
Mixtral 8x7B Instruct47B MoEQ4_K_M30.13 GB+161.9 GB48
Seed-OSS 36B Instruct36BQ5_K_M27.81 GB+164.2 GB37
Command R 35B35BQ4_K_M22.86 GB+169.1 GB42
Yi 1.5 34B Chat34BQ4_K_M22.82 GB+169.2 GB40
Qwen3 32B Instruct32BQ4_K_M21.85 GB+170.2 GB42
Qwen2.5 32B Instruct32BQ4_K_M21.69 GB+170.3 GB44
Qwen2.5-Coder 32B Instruct32BQ4_K_M21.69 GB+170.3 GB44
DeepSeek-R1-Distill-Qwen-32B32BQ4_K_M21.69 GB+170.3 GB42
Qwen3 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+169 GB82
Qwen3-Coder 30B-A3B Instruct30B-A3BQ5_K_M23.04 GB+169 GB80
Qwen3-VL 30B-A3B Instruct30B-A3BQ8_034.28 GB+157.7 GB78
Gemma 3 27B IT27BQ4_K_M19.49 GB+172.5 GB48
Gemma 2 27B Instruct27BQ5_K_M21.76 GB+170.2 GB42
Mistral Small 24B Instruct24BEXL2 4.65bpw15.26 GB+176.7 GB88
Devstral Small 1.1 24B24BQ6_K20.91 GB+171.1 GB48
Magistral Small 1.2 24B24BQ6_K20.91 GB+171.1 GB47
Codestral 22B22BQ4_K_M15.03 GB+177 GB58
GPT-OSS 20B21B MoEMXFP411.81 GB+180.2 GB195

14B · 17

Mac M3 Ultra 192G — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
InternLM2 20B Chat20BQ5_K_M15.59 GB+176.4 GB68
DeepSeek-Coder-V2-Lite Instruct16BQ8_018.36 GB+173.6 GB118
DeepSeek-V2-Lite Chat16BQ4_K_M10.87 GB+181.1 GB142
StarCoder2 15B15BQ4_K_M10.19 GB+181.8 GB92
Qwen3 14B Instruct14BQ5_K_M11.65 GB+180.4 GB78
Qwen2.5 14B Instruct14BQ5_K_M11.73 GB+180.3 GB86
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+182.3 GB128
Phi-4 14B14BQ5_K_M11.77 GB+180.2 GB78
Phi-3 Medium 14B Instruct14BQ6_K12.86 GB+179.1 GB88
Mistral Nemo 12B Instruct12BQ6_K11.14 GB+180.9 GB95
Gemma 3 12B IT12BQ5_K_M10.6 GB+181.4 GB92
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+183.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+183.9 GB95
Llama 3.2 11B Vision Instruct11BQ8_012.9 GB+179.1 GB72
Solar 10.7B Instruct11BQ4_K_M7.6 GB+184.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+184.7 GB118
Gemma 2 9B Instruct9BQ8_011.7 GB+180.3 GB108

7B · 19

Mac M3 Ultra 192G — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+181.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+181.6 GB108
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+186.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+186.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+184.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+182.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+186.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+186.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+186.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+185.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+186.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+185.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+187.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+186.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+187.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+181.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+185.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+185.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+186.6 GB145

≤3B · 10

Mac M3 Ultra 192G — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+188 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+187.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+186.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+188 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+188.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+188.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+189.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+190.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+190.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+191.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

74 of 79 indexed models fit comfortably in 192GB at 4K context, each at the highest-quality quant that still leaves headroom.