8GB VRAM

Mac M1 8G — what LLMs can it run?

20 of 87 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.

The short answer for Mac M1 8G

8 GB, 68 GB/s (Derived from Apple's statements that the M2 (100 GB/s) is 50% faster and the M1 Max (400 GB/s) nearly 6× faster; Apple does not list the M1 figure directly.). 20 of 87 models in this index fit comfortably at 4K context; 31 load at all.

Biggest that fits
Llama 3.1 8B Instruct — 3.3 GB at Q2_K — 2.7 GB spare at 4K, so longer context comes out of a thin margin.
Speed ceiling
Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 68 GB/s puts a hard ceiling near 15 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
The wall you will hit
The GPU's share of unified memory, not the 8 GB on the box. macOS reserves part of the pool for the system and caps what a single process may wire down, so the practical budget is meaningfully below nameplate — the limit is adjustable (`iogpu.wired_limit_mb`) but it is not absent. Bandwidth is 68 GB/s, which is the number that decides tok/s once a model fits.
Best local LLM for Apple silicon

7B · 10

Mac M1 8G — what LLMs can it run? — 7B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
Llama 3.1 8B Instruct8BQ2_K3.31 GB+2.7 GB210
Qwen2.5 7B Instruct7BQ4_K_M5.07 GB+0.9 GB155
Qwen2.5-Coder 7B Instruct7BQ4_K_M5.07 GB+0.9 GB158Estimated
DeepSeek-R1-Distill-Qwen-7B7BQ4_K_M5.07 GB+0.9 GB152Estimated
Gemma 4 E4B ITE4BQ4_K_M4.88 GB+1.1 GB—
WizardLM-2 7B7BQ4_K_M5.14 GB+0.9 GB152Estimated
Mistral 7B Instruct v0.37BQ4_K_M5.14 GB+0.9 GB158Estimated
Zephyr 7B Beta7BQ4_K_M5.14 GB+0.9 GB155Estimated
Gemma 4 E2B ITE2BQ8_05.2 GB+0.8 GB—
Gemma 3 4B IT4BQ8_05.05 GB+1 GB145Estimated

≤3B · 10

Mac M1 8G — what LLMs can it run? — ≤3B
ModelQuantEst. VRAMHeadroomtok/s on RTX 4090
Qwen3 4B Instruct4BQ6_K4.05 GB+2 GB145
Phi-4 Mini Instruct3.8BQ8_04.81 GB+1.2 GB262Estimated
Phi-3.5 Mini Instruct3.8BQ4_K_M4.07 GB+1.9 GB298Estimated
Llama 3.2 3B Instruct3BQ8_04.05 GB+2 GB285Estimated
Qwen2.5 3B Instruct3BQ8_03.59 GB+2.4 GB290Estimated
Gemma 2 2B Instruct2BQ8_03.34 GB+2.7 GB320Estimated
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+3.6 GB240Estimated
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+4.2 GB410Estimated
Llama 3.2 1B Instruct1BQ8_01.51 GB+4.5 GB450Estimated
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+5.4 GB540Estimated

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 11 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a Mac M1 8G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.

Cards with the same budget

What fits is decided by memory, so every 8GB card of this type returns the same list. These pages are not different answers — they differ in throughput, which this index does not measure per card.

Stepping up

A Mac M3 16G (16GB) fits 27 more of the indexed models than this card. Mac M3 16G →

Common questions

What is the best local LLM for a Mac M1 8G?

For everyday use, Llama 3.1 8B Instruct at Q2_K — about 3.3 GB of the card's 8 GB at 4K context, leaving 2.7 GB for a longer window. If you want the largest thing that will load, that is Llama 3.1 8B Instruct at Q2_K (3.3 GB, 2.7 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.

Can a Mac M1 8G run Qwen3 8B Instruct?

Only just. Its smallest build here, Q4_K_M, needs about 5.8 GB against 6.0 GB usable — that loads with nothing else running, with no margin for a longer context window. It is not on the list above, which requires a model to stay inside 88% of usable capacity.

How many tokens per second does a Mac M1 8G do on an 8B model at Q4?

This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 68 GB/s — a ceiling near 15 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.

Mac M1 8G or RTX 3070 Ti for local LLMs?

They hold the same models: 8 GB against 8 GB fits 20 and 36 of 87 respectively at 4K. The difference is throughput — 68 GB/s against 608 GB/s, a 8.9× gap in how fast the weights can be read, which is what token generation is bound by. The RTX 3070 Ti generates faster on any model both can hold.

20 of 87 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.

This page's figures change when the model or the runtime does.Last updated 2026-10-02 RSS → /feed.xml