Arc A750 8G — what LLMs can it run?
33 of 89 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.
The short answer for Arc A750 8G
8 GB, 512 GB/s GDDR6. 33 of 89 models in this index fit comfortably at 4K context; 38 load at all.
- Biggest that fits
- GLM-4-9B-Chat — 5.9 GB at Q4_K_M — 2.1 GB spare at 4K, so longer context comes out of a thin margin.
- Room to grow
- Qwen3 4B Instruct — 4.0 GB at Q6_K — under 60% of the card, which leaves 4.0 GB for a long context window or a second process.
- Speed ceiling
- Llama 3.1 8B Instruct at Q4_K_M reads 4.6 GB of weights per token, so 512 GB/s puts a hard ceiling near 111 tok/s. That is arithmetic on two published numbers, not a benchmark — real throughput lands below it. No run on this card has been measured here, so there is nothing to compare the ceiling against.
- The wall you will hit
- Backend coverage, not capacity. At 512 GB/s the card has the memory system for its size class; what limits it is which runtimes reach it. llama.cpp's SYCL backend lists Arc A-series and B580 cards as verified, and Ollama runs them through Vulkan — so GGUF is the format this site recommends here. AWQ, GPTQ and EXL2 are not.
7B · 22
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| GLM-4-9B-Chat9B | Q4_K_M | 5.87 GB | +2.1 GB | 135Estimated |
| Qwen3-VL 8B Instruct8B | Q4_K_M | 6.19 GB | +1.8 GB | 140Community |
| Ministral 3 8B Instruct8B | Q5_K_M | 6.92 GB | +1.1 GB | — |
| Qwen2-VL 7B Instruct7B | Q4_K_M | 5.49 GB | +2.5 GB | 72Estimated |
| Granite 3.1 8B Instruct8B | Q4_K_M | 5.88 GB | +2.1 GB | 142Estimated |
| Qwen3 8B Instruct8B | Q4_K_M | 5.81 GB | +2.2 GB | 142 |
| Llama 3.1 8B Instruct8B | Q4_K_M | 5.64 GB | +2.4 GB | 148 |
| Nous Hermes 3 Llama 3.1 8B8B | Q4_K_M | 5.64 GB | +2.4 GB | 148Estimated |
| Aya 23 8B8B | Q4_K_M | 5.64 GB | +2.4 GB | 145Estimated |
| OpenChat 3.6 8B8B | Q4_K_M | 5.64 GB | +2.4 GB | 146Estimated |
| DeepSeek-R1-Distill-Llama-8B8B | Q5_K_M | 6.51 GB | +1.5 GB | 128Estimated |
| InternLM2 7B Chat7B | Q4_K_M | 5.45 GB | +2.6 GB | 148Estimated |
| Qwen2.5 7B Instruct7B | Q6_K | 6.77 GB | +1.2 GB | 132 |
| Qwen2.5-Coder 7B Instruct7B | Q4_K_M | 5.07 GB | +2.9 GB | 158Estimated |
| DeepSeek-R1-Distill-Qwen-7B7B | Q4_K_M | 5.07 GB | +2.9 GB | 152Estimated |
| Gemma 4 E4B ITE4B | Q4_K_M | 4.88 GB | +3.1 GB | — |
| OLMo 2 7B Instruct7B | Q4_K_M | 6.82 GB | +1.2 GB | 150Estimated |
| WizardLM-2 7B7B | Q4_K_M | 5.14 GB | +2.9 GB | 152Estimated |
| Mistral 7B Instruct v0.37B | Q6_K | 6.75 GB | +1.3 GB | 135Estimated |
| Zephyr 7B Beta7B | Q6_K | 6.75 GB | +1.3 GB | 132Estimated |
| Gemma 4 E2B ITE2B | Q8_0 | 5.2 GB | +2.8 GB | — |
| Gemma 3 4B IT4B | Q8_0 | 5.05 GB | +3 GB | 145Estimated |
≤3B · 11
| Model | Quant | Est. VRAM | Headroom | tok/s on RTX 4090 |
|---|---|---|---|---|
| Qwen3 4B Instruct4B | Q6_K | 4.05 GB | +4 GB | 145 |
| Phi-4 Mini Instruct3.8B | Q8_0 | 4.81 GB | +3.2 GB | 262Estimated |
| Phi-3.5 Mini Instruct3.8B | Q8_0 | 5.89 GB | +2.1 GB | 255Estimated |
| Llama 3.2 3B Instruct3B | Q8_0 | 4.05 GB | +4 GB | 285Estimated |
| Qwen2.5 3B Instruct3B | Q8_0 | 3.59 GB | +4.4 GB | 290Estimated |
| SmolLM3 3B3B | Q8_0 | 3.73 GB | +4.3 GB | — |
| Gemma 2 2B Instruct2B | Q8_0 | 3.34 GB | +4.7 GB | 320Estimated |
| Qwen3 1.7B Instruct1.7B | Q8_0 | 2.39 GB | +5.6 GB | 240Estimated |
| Qwen2.5 1.5B Instruct1.5B | Q8_0 | 1.83 GB | +6.2 GB | 410Estimated |
| Llama 3.2 1B Instruct1B | Q8_0 | 1.51 GB | +6.5 GB | 450Estimated |
| Qwen2.5 0.5B Instruct0.5B | Q8_0 | 0.6 GB | +7.4 GB | 540Estimated |
How this list is built
Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.
A further 5 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.
Measured on this card
No benchmark runs in this index were recorded on a Arc A750 8G. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements. The speed column in the tables above is an RTX 4090 figure shown for reference — it is not a speed on this card.
Stepping up
A Arc B570 10G (10GB) fits 7 more of the indexed models than this card. Arc B570 10G →
Common questions
What is the best local LLM for an Arc A750 8G?
For everyday use, Qwen3 4B Instruct at Q6_K — about 4.0 GB of the card's 8 GB at 4K context, leaving 4.0 GB for a longer window. If you want the largest thing that will load, that is GLM-4-9B-Chat at Q4_K_M (5.9 GB, 2.1 GB spare). "Best" here means best fit for the memory budget — this index does not run task benchmarks, so it cannot tell you which model is smarter.
Can an Arc A750 8G run Gemma 2 9B Instruct?
Only just. Its smallest build here, Q4_K_M, needs about 7.3 GB against 8.0 GB usable — that loads with nothing else running, with no margin for a longer context window. It is not on the list above, which requires a model to stay inside 88% of usable capacity.
How many tokens per second does an Arc A750 8G do on an 8B model at Q4?
This index has no measured run on this card, so it does not publish a figure. What can be stated from specifications: generating a token requires reading every weight once, Llama 3.1 8B Instruct at Q4_K_M is 4.6 GB of weights, and this card moves 512 GB/s — a ceiling near 111 tok/s. Batch size, context length, the runtime and how much of the model sits in cache all take you below it.
Arc A750 8G or Arc B570 10G for local LLMs?
They hold the same models: 8 GB against 10 GB fits 33 and 40 of 89 respectively at 4K. The difference is throughput — 512 GB/s against 380 GB/s, a 1.3× gap in how fast the weights can be read, which is what token generation is bound by. The Arc A750 8G generates faster on any model both can hold.
33 of 89 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.
This page's figures change when the model or the runtime does.Last updated 2026-10-03 RSS → /feed.xml