Independent knowledge for Intel Arc Pro

Put your
B70 to work.

Big memory. Open possibilities.
Benchmarks, working guides, and lessons from building on Intel GPUs.

ARC PROB70 / FIELD GUIDEXe32 GB GDDR6ILLUSTRATIVE SCHEMATIC · NOT TO SCALE
32 GB GDDR6 / PER CARDLLAMA.CPP · PYTORCH XPUINTEL SPECIFICATIONS ↗

Start with what you want to do

Your hardware. Your next step.

From the test bench

Useful numbers.
With the fine print.

What fits, how it runs, and exactly what was measured. Start with a workload, then open the configuration.

Explore the dataset
DECODE / TG128TOKENS / SECOND →
Qwen 3.6 35B-A3BUD-Q4_K_M · 1 × B70
54.65
Qwen3-Coder-Next 80B-A3BQ4_K_M · 2 × B70
43.35
DeepSeek R1 Distill 70BQ4_K_M · 2 × B70
11.47

TESTED APR 21, 2026 · SYCL · ec6f7a6a5c (dirty)
Different models and GPU counts. Throughput observations, not quality rankings. Local build diff unavailable. How we measure ↗

Beyond the benchmark

A faster kernel is only half the story.

Q8_0 weight reordering improved the decode path. Then the second prompt exposed a correctness bug. Follow the diagnosis, the fix, and the upstream work.

Inside the Q8_0 work

New in the field guide / September 10

Go a little deeper.

All guides

One card. Two cards.

More memory.
Different possibilities.

A second GPU opens up more than one kind of workload. Explore what actually changes.

The dual-GPU guide

What does a second card change?

Two 32 GB cards provide 64 GB of physical memory across two devices. Each allocation still belongs to a device; software must explicitly distribute a workload. Leave room on each card for cache and compute buffers.

CONCEPTUAL FLOW · NO TRANSFER-RATE OR SPEEDUP CLAIM

The notebook

Recent field notes.

Editorial review: six selected records, two distinct cohorts, and an explicit list of missing metadata. Browse the evidence →

Historical test: June’s Qwen 3.6 decode and prefill measurements answer different questions. See the June cohort →