What works on Arc, and what does not
GGUF works, through two routes. llama.cpp has a SYCL backend built on Intel’s oneAPI, and its documentation lists the Arc A770, A750 and B580 as verified cards; the B570 is the same Xe2 generation as the B580 but is not on that list. Ollama has no SYCL build — it reaches Arc through its Vulkan backend, which is on by default. AWQ, GPTQ and EXL2 are not options here: ExLlamaV2 is CUDA-only, and this site has not confirmed which quantized formats vLLM’s Intel backend serves on consumer Arc cards, so the calculator and every recommendation on an Arc page use GGUF only.
llama.cpp SYCL, verified list: Arc A770, A750, B580
Same generation, not listed: Arc B570
Ollama: Vulkan backend, on by default
Formats this site recommends: GGUF onlyPrerequisites on Linux
Install Intel’s GPU driver by following dgpu-docs.intel.com for your distribution, then add your user to the render and video groups and log out and back in. Then install the oneAPI Base Toolkit, keeping the default path /opt/intel/oneapi. Check the card is visible before building anything: sycl-ls should list a level_zero GPU entry. If it does not, the SYCL docs’ own advice is to run sudo sycl-ls and check the group membership — a missing group is the usual cause, not the build.
sudo usermod -aG render $USER
sudo usermod -aG video $USER # then log out and back in
source /opt/intel/oneapi/setvars.sh
sycl-ls # expect a [level_zero:gpu] line for the Arc cardBuild llama.cpp with SYCL
The switch is GGML_SYCL=ON, compiled with Intel’s icx/icpx from the oneAPI environment you just sourced; GGML_SYCL_F16=ON is what the docs recommend for speed. Every new shell needs setvars.sh sourced again before you build or run. The first run is slower than later ones: the build has no ahead-of-time kernels, so they are compiled on first use.
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
source /opt/intel/oneapi/setvars.sh
cmake -B build -DGGML_SYCL=ON -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DGGML_SYCL_F16=ON
cmake --build build --config Release -j
./build/bin/llama-ls-sycl-device # lists the SYCL devices the build can seeStart the server and check it is on the GPU
Bind to 127.0.0.1 — llama-server has no authentication unless you pass --api-key. The confirmation is in the startup log, not in the fact that it answers: look for the SYCL device detection line and load_tensors reporting every layer offloaded. A build that fell back to the CPU still answers, slowly. On a machine with Intel integrated graphics as well, pin the discrete card with ONEAPI_DEVICE_SELECTOR, or list devices with --list-devices and choose one with --device.
ONEAPI_DEVICE_SELECTOR="level_zero:0" ./build/bin/llama-server \
-m ./models/Qwen3-8B-Q4_K_M.gguf \
-ngl 99 -c 8192 --host 127.0.0.1 --port 8080
# In the startup log:
# detect 1 SYCL GPUs: [0] with top Max compute units:...
# load_tensors: offloaded 37/37 layers to GPUThe Ollama route
If you would rather not install oneAPI, Ollama runs Arc through Vulkan with no build step. On Windows the GPU driver already ships Vulkan support; on Linux install Intel’s driver first. Ollama’s docs note that without root or the cap_perfmon capability it cannot read free VRAM from Vulkan and falls back to approximate sizes when scheduling, so grant it. ollama ps is the check: the PROCESSOR column should say GPU, not a CPU/GPU split. In a container, pass /dev/dri only — /dev/kfd is AMD’s device, and Docker will not start a container whose device does not exist.
curl -fsSL https://ollama.com/install.sh | sh
sudo setcap cap_perfmon+ep /usr/local/bin/ollama
ollama run hf.co/Qwen/Qwen3-8B-GGUF:Q4_K_M
ollama ps # PROCESSOR should read 100% GPU
# With an Intel iGPU too, pick the discrete card by its Vulkan index:
# GGML_VK_VISIBLE_DEVICES=1 ollama serveWhat fits, and the speed ceiling
Memory arithmetic does not care which vendor made the card. The figures below come from this site’s calculator at Q4_K_M, which every one of these models ships. Qwen3 8B at Q4_K_M with an 8K context is about 6.4 GB. Gemma 3 12B at Q4_K_M with an 8K context is about 8.8 GB. Qwen3 14B at Q4_K_M with an 8K context is about 10.7 GB, which is 89% of a B580 — past this site’s comfortable line, so keep the context short or move to the 16 GB A770. The speed column is a ceiling, not a measurement: generation reads the whole weight file once per token, so bandwidth divided by file size bounds tok/s. Nobody here has run these on Arc; real throughput will be lower.
Model, Q4_K_M, 8K context size B580 12G A770 16G ceiling B580 / A770
Qwen3 8B 6.4 GB 54% 40% 97 / 119 tok/s
Gemma 3 12B 8.8 GB 73% 55% 65 / 80 tok/s
Qwen3 14B 10.7 GB 89% 67% 54 / 66 tok/s
Measure your own: ./build/bin/llama-bench -m <model.gguf> -ngl 99 (pp512 / tg128)When it does not work
Most first-day failures on Arc are one of a few things, and only the last is the model.
sycl-ls shows no level_zero GPU
→ driver not installed, or user not in render/video (log out and back in)
"setvars.sh" forgotten → icx not found at build, or SYCL libraries missing at run
First run very slow, later runs normal
→ kernels compiled on first use (no ahead-of-time build) — expected
Runs, but on the integrated GPU
→ ONEAPI_DEVICE_SELECTOR="level_zero:0" (llama.cpp) / GGML_VK_VISIBLE_DEVICES (Ollama)
Out of device memory
→ shorten -c, or pick a smaller quant (the SYCL docs' own advice)Common questions
Should I use llama.cpp SYCL or Ollama on an Intel Arc card?
Ollama is the shorter path: no oneAPI install, no build, and it reaches Arc through its Vulkan backend, which is on by default. llama.cpp with SYCL takes more setup but is the route llama.cpp documents for Intel GPUs, with the A770, A750 and B580 on its verified list. This site has not benchmarked either on Arc, so it cannot tell you which is faster on your card — llama-bench will.
Can I use llama.cpp SYCL on Windows without installing oneAPI?
Yes. llama.cpp publishes a Windows SYCL build (the bin-win-sycl-x64 zip on its releases page), and its SYCL docs say the package carries the runtime it needs, so no separate oneAPI install is required. Building from source on Windows does need Visual Studio and the oneAPI Base Toolkit. Windows 11 is the version the docs name.
How many models fit on an Arc B580?
48 of the 89 models in this index fit the B580 comfortably at 4K context, all as GGUF — the same count as any other 12 GB card, because only capacity decides what fits. The A770 16G fits 51, the B570 10G 40, and the A750 8G 33. Speed is a different question: the B580 has 456 GB/s of memory bandwidth, which sets the ceiling on tokens per second once a model fits.
What this guide uses
- Hardware
- Arc B580 12G12GB
- Format
- GGUF
- Picks for it
- Best local LLM for 12GB
Next steps
See everything that fits 🔵 Arc B580 12GReverse lookup — 12GB at 8192 context, ranked by qualityModels covered in this guide
Did this actually run?
Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.
Related guides
Run GPT-OSS 20B (and 120B) locally without re-quantizing
GPT-OSS ships natively in MXFP4, so the usual "download the Q4_K_M" habit makes it bigger and worse. Sizing, the right flags, and how MoE expert-offload puts the 120B on a 24GB card.
What to Run on a 12GB GPU (RTX 3060 12G, 4070, 5070)
Which models fit a 12GB NVIDIA card at 8K–32K context, where 14B stops fitting, how to run them with Ollama or llama.cpp, and how to prove they are on the GPU.
llama.cpp on Windows with CUDA
Build llama.cpp with NVIDIA GPU support on Windows 11 — the path of least resistance for PC gamers.