IntermediateEdge / Local 5 min readPublished

Intel Arc: llama.cpp (SYCL) or Ollama (Vulkan)

Run GGUF models on an Arc B580, B570, A770 or A750 — the SYCL build, the Ollama route, how to prove the GPU is doing the work, and what fits in 8–16 GB.

Written against

Linux (Ubuntu) or Windows 11 · Intel GPU driver · oneAPI Base Toolkit · llama.cpp built with GGML_SYCL=ON · or Ollama with Vulkan · GGUF

Not re-run since it was last edited — treat the commands as a starting point, not a tested recipe.

Intel ArcSYCLllama.cppOllamaVulkanGGUF

What works on Arc, and what does not

GGUF works, through two routes. llama.cpp has a SYCL backend built on Intel’s oneAPI, and its documentation lists the Arc A770, A750 and B580 as verified cards; the B570 is the same Xe2 generation as the B580 but is not on that list. Ollama has no SYCL build — it reaches Arc through its Vulkan backend, which is on by default. AWQ, GPTQ and EXL2 are not options here: ExLlamaV2 is CUDA-only, and this site has not confirmed which quantized formats vLLM’s Intel backend serves on consumer Arc cards, so the calculator and every recommendation on an Arc page use GGUF only.

text
llama.cpp SYCL, verified list:  Arc A770, A750, B580
Same generation, not listed:     Arc B570
Ollama:                          Vulkan backend, on by default
Formats this site recommends:    GGUF only

Prerequisites on Linux

Install Intel’s GPU driver by following dgpu-docs.intel.com for your distribution, then add your user to the render and video groups and log out and back in. Then install the oneAPI Base Toolkit, keeping the default path /opt/intel/oneapi. Check the card is visible before building anything: sycl-ls should list a level_zero GPU entry. If it does not, the SYCL docs’ own advice is to run sudo sycl-ls and check the group membership — a missing group is the usual cause, not the build.

bash
sudo usermod -aG render $USER
sudo usermod -aG video $USER   # then log out and back in

source /opt/intel/oneapi/setvars.sh
sycl-ls                        # expect a [level_zero:gpu] line for the Arc card

Build llama.cpp with SYCL

The switch is GGML_SYCL=ON, compiled with Intel’s icx/icpx from the oneAPI environment you just sourced; GGML_SYCL_F16=ON is what the docs recommend for speed. Every new shell needs setvars.sh sourced again before you build or run. The first run is slower than later ones: the build has no ahead-of-time kernels, so they are compiled on first use.

bash
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
source /opt/intel/oneapi/setvars.sh
cmake -B build -DGGML_SYCL=ON -DCMAKE_C_COMPILER=icx -DCMAKE_CXX_COMPILER=icpx -DGGML_SYCL_F16=ON
cmake --build build --config Release -j

./build/bin/llama-ls-sycl-device   # lists the SYCL devices the build can see

Start the server and check it is on the GPU

Bind to 127.0.0.1 — llama-server has no authentication unless you pass --api-key. The confirmation is in the startup log, not in the fact that it answers: look for the SYCL device detection line and load_tensors reporting every layer offloaded. A build that fell back to the CPU still answers, slowly. On a machine with Intel integrated graphics as well, pin the discrete card with ONEAPI_DEVICE_SELECTOR, or list devices with --list-devices and choose one with --device.

bash
ONEAPI_DEVICE_SELECTOR="level_zero:0" ./build/bin/llama-server \
  -m ./models/Qwen3-8B-Q4_K_M.gguf \
  -ngl 99 -c 8192 --host 127.0.0.1 --port 8080

# In the startup log:
#   detect 1 SYCL GPUs: [0] with top Max compute units:...
#   load_tensors: offloaded 37/37 layers to GPU

The Ollama route

If you would rather not install oneAPI, Ollama runs Arc through Vulkan with no build step. On Windows the GPU driver already ships Vulkan support; on Linux install Intel’s driver first. Ollama’s docs note that without root or the cap_perfmon capability it cannot read free VRAM from Vulkan and falls back to approximate sizes when scheduling, so grant it. ollama ps is the check: the PROCESSOR column should say GPU, not a CPU/GPU split. In a container, pass /dev/dri only — /dev/kfd is AMD’s device, and Docker will not start a container whose device does not exist.

bash
curl -fsSL https://ollama.com/install.sh | sh
sudo setcap cap_perfmon+ep /usr/local/bin/ollama

ollama run hf.co/Qwen/Qwen3-8B-GGUF:Q4_K_M
ollama ps                          # PROCESSOR should read 100% GPU

# With an Intel iGPU too, pick the discrete card by its Vulkan index:
#   GGML_VK_VISIBLE_DEVICES=1 ollama serve

What fits, and the speed ceiling

Memory arithmetic does not care which vendor made the card. The figures below come from this site’s calculator at Q4_K_M, which every one of these models ships. Qwen3 8B at Q4_K_M with an 8K context is about 6.4 GB. Gemma 3 12B at Q4_K_M with an 8K context is about 8.8 GB. Qwen3 14B at Q4_K_M with an 8K context is about 10.7 GB, which is 89% of a B580 — past this site’s comfortable line, so keep the context short or move to the 16 GB A770. The speed column is a ceiling, not a measurement: generation reads the whole weight file once per token, so bandwidth divided by file size bounds tok/s. Nobody here has run these on Arc; real throughput will be lower.

text
Model, Q4_K_M, 8K context     size     B580 12G   A770 16G   ceiling B580 / A770
Qwen3 8B                       6.4 GB   54%        40%        97 / 119 tok/s
Gemma 3 12B                    8.8 GB   73%        55%        65 / 80 tok/s
Qwen3 14B                      10.7 GB  89%        67%        54 / 66 tok/s

Measure your own: ./build/bin/llama-bench -m <model.gguf> -ngl 99   (pp512 / tg128)

When it does not work

Most first-day failures on Arc are one of a few things, and only the last is the model.

text
sycl-ls shows no level_zero GPU
  → driver not installed, or user not in render/video (log out and back in)

"setvars.sh" forgotten → icx not found at build, or SYCL libraries missing at run

First run very slow, later runs normal
  → kernels compiled on first use (no ahead-of-time build) — expected

Runs, but on the integrated GPU
  → ONEAPI_DEVICE_SELECTOR="level_zero:0" (llama.cpp) / GGML_VK_VISIBLE_DEVICES (Ollama)

Out of device memory
  → shorten -c, or pick a smaller quant (the SYCL docs' own advice)

Common questions

Should I use llama.cpp SYCL or Ollama on an Intel Arc card?

Ollama is the shorter path: no oneAPI install, no build, and it reaches Arc through its Vulkan backend, which is on by default. llama.cpp with SYCL takes more setup but is the route llama.cpp documents for Intel GPUs, with the A770, A750 and B580 on its verified list. This site has not benchmarked either on Arc, so it cannot tell you which is faster on your card — llama-bench will.

Can I use llama.cpp SYCL on Windows without installing oneAPI?

Yes. llama.cpp publishes a Windows SYCL build (the bin-win-sycl-x64 zip on its releases page), and its SYCL docs say the package carries the runtime it needs, so no separate oneAPI install is required. Building from source on Windows does need Visual Studio and the oneAPI Base Toolkit. Windows 11 is the version the docs name.

How many models fit on an Arc B580?

48 of the 89 models in this index fit the B580 comfortably at 4K context, all as GGUF — the same count as any other 12 GB card, because only capacity decides what fits. The A770 16G fits 51, the B570 10G 40, and the A750 8G 33. Speed is a different question: the B580 has 456 GB/s of memory bandwidth, which sets the ceiling on tokens per second once a model fits.

What this guide uses

Format
GGUF

Next steps

See everything that fits 🔵 Arc B580 12GReverse lookup — 12GB at 8192 context, ranked by quality

Models covered in this guide

Did this actually run?

Copying a command is not the same as it working, so this is the only place the site asks. Nothing is collected beyond the answer itself.

Related guides

Deployment guides are educational. Each model is subject to its own license — read the official Hugging Face model card before downloading or deploying.