The measurement notebook / 01

The numbers.
And what’s behind them.

Six reviewed configurations. Two historical test cohorts. Every result comes with its source, settings, and limitations.

6 reviewed records · April and June shown separately

Historical observations, not a model-quality ranking. April uses a dirty build and F16 KV; June uses a different build and Q4_0 KV. Do not interpret the cohorts as a controlled speedup.

DECODE / TG128 · TOKENS PER SECOND · HISTORICAL RESULTS
MODEL / QUANTIZATIONCONFIGURATIONDECODE · TOK/SEVIDENCE
Qwen 3.6 35B-A3BUD-Q4_K_M · MoE1 × B70 · SYCL
2026-04-21 · DIRTY BUILD
54.65
Qwen 3.6 35B-A3B · 2026-04-21 — full configuration

Qwen 3.6 35B-A3B

Historical · tested 2026-04-21
Quantization
UD-Q4_K_M
GPU count
1
Backend
llama.cpp / SYCL
Build commit
ec6f7a6a5c
Dirty build
Yes
Model revision
Unknown
Weight size (GiB)
20.61
Configured context
4096
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
Unknown / Unknown
KV types K / V
f16 / f16
Flash attention
Yes
Threads
1
Batch / microbatch
Unknown / Unknown
Concurrency
Unknown
Warmup
Unknown
Repetitions
Unknown
Decode tok/s ± reported SD
54.65 ± 0.03
Prefill tok/s ± reported SD
615.3 ± 2.82
OS
Ubuntu 26.04 (cohort report)
Kernel
7.0.0-10-generic (cohort report)
Driver
xe / compute-runtime 26.09 (cohort report)
Runtime
oneAPI 2025.3.3 (cohort report)
CPU
Ryzen 5 9600X (cohort report)
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Dirty build: the exact local patch diff has not been recovered.
  • One thread is recorded in this result; methodology prose says six. Per-result metadata is used.
  • Warmup and five repetitions are described by methodology, but not recorded per result.
  • Configured 4K context is not a full-context generation test.
  • Energy and VRAM telemetry are excluded: device inclusion and measurement windows are unresolved.
  • Environment is reported by cohort documentation, not captured in this result. PCIe topology conflicts remain unresolved.

Original benchmark JSON ↗ · Reviewed data ↓

Original record: intel-arc-pro-b70-qwen3-6-35b-a3b-ud-q4-k-m-sycl
Original SHA-256: 2ec17f8960d7032c7d46b515ec7975d62f816e19e777ca1c6e6054fc9e61d15a

Qwen3-Coder-Next 80B-A3BQ4_K_M · MoE2 × B70 · SYCL
2026-04-21 · DIRTY BUILD
43.35
Qwen3-Coder-Next 80B-A3B · 2026-04-21 — full configuration

Qwen3-Coder-Next 80B-A3B

Historical · tested 2026-04-21
Quantization
Q4_K_M
GPU count
2
Backend
llama.cpp / SYCL
Build commit
ec6f7a6a5c
Dirty build
Yes
Model revision
Unknown
Weight size (GiB)
Unknown
Configured context
4096
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
Unknown / Unknown
KV types K / V
f16 / f16
Flash attention
Yes
Threads
1
Batch / microbatch
Unknown / Unknown
Concurrency
Unknown
Warmup
Unknown
Repetitions
Unknown
Decode tok/s ± reported SD
43.35 ± 0.07
Prefill tok/s ± reported SD
304.98 ± 2.22
OS
Ubuntu 26.04 (cohort report)
Kernel
7.0.0-10-generic (cohort report)
Driver
xe / compute-runtime 26.09 (cohort report)
Runtime
oneAPI 2025.3.3 (cohort report)
CPU
Ryzen 5 9600X (cohort report)
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Dirty build: the exact local patch diff has not been recovered.
  • One thread is recorded in this result; methodology prose says six. Per-result metadata is used.
  • Warmup and five repetitions are described by methodology, but not recorded per result.
  • Configured 4K context is not a full-context generation test.
  • Energy and VRAM telemetry are excluded: device inclusion and measurement windows are unresolved.
  • Environment is reported by cohort documentation, not captured in this result. PCIe topology conflicts remain unresolved.
  • Weight size unresolved: this JSON says 14.46 GiB while the report says about 45.1 GiB. Neither is used as an audited weight-size claim.

Original benchmark JSON ↗ · Reviewed data ↓

Original record: intel-arc-pro-b70-qwen3-coder-next-80b-a3b-q4-k-m-sycl-2gpu
Original SHA-256: e99cbe3e1dbb76154c48f16eabd3811f91844b4de1849f1a063470222ba7f902

DeepSeek R1 Distill 70BQ4_K_M · Dense2 × B70 · SYCL
2026-04-21 · DIRTY BUILD
11.47
DeepSeek R1 Distill 70B · 2026-04-21 — full configuration

DeepSeek R1 Distill 70B

Historical · tested 2026-04-21
Quantization
Q4_K_M
GPU count
2
Backend
llama.cpp / SYCL
Build commit
ec6f7a6a5c
Dirty build
Yes
Model revision
Unknown
Weight size (GiB)
39.6
Configured context
4096
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
Unknown / Unknown
KV types K / V
f16 / f16
Flash attention
Yes
Threads
1
Batch / microbatch
Unknown / Unknown
Concurrency
Unknown
Warmup
Unknown
Repetitions
Unknown
Decode tok/s ± reported SD
11.47 ± 0
Prefill tok/s ± reported SD
336.06 ± 3.82
OS
Ubuntu 26.04 (cohort report)
Kernel
7.0.0-10-generic (cohort report)
Driver
xe / compute-runtime 26.09 (cohort report)
Runtime
oneAPI 2025.3.3 (cohort report)
CPU
Ryzen 5 9600X (cohort report)
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Dirty build: the exact local patch diff has not been recovered.
  • One thread is recorded in this result; methodology prose says six. Per-result metadata is used.
  • Warmup and five repetitions are described by methodology, but not recorded per result.
  • Configured 4K context is not a full-context generation test.
  • Energy and VRAM telemetry are excluded: device inclusion and measurement windows are unresolved.
  • Environment is reported by cohort documentation, not captured in this result. PCIe topology conflicts remain unresolved.

Original benchmark JSON ↗ · Reviewed data ↓

Original record: intel-arc-pro-b70-deepseek-r1-distill-llama-70b-q4-k-m-sycl-2gpu
Original SHA-256: 24e822a19618191617fca349f133dba1864bd5fee900fa783a1868d7f537e466

Qwen 3.5 27BQ4_K_M · Dense1 × B70 · SYCL
2026-04-21 · DIRTY BUILD
20.35
Qwen 3.5 27B · 2026-04-21 — full configuration

Qwen 3.5 27B

Historical · tested 2026-04-21
Quantization
Q4_K_M
GPU count
1
Backend
llama.cpp / SYCL
Build commit
ec6f7a6a5c
Dirty build
Yes
Model revision
Unknown
Weight size (GiB)
15.59
Configured context
4096
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
Unknown / Unknown
KV types K / V
f16 / f16
Flash attention
Yes
Threads
1
Batch / microbatch
Unknown / Unknown
Concurrency
Unknown
Warmup
Unknown
Repetitions
Unknown
Decode tok/s ± reported SD
20.35 ± 0.03
Prefill tok/s ± reported SD
718.21 ± 3.16
OS
Ubuntu 26.04 (cohort report)
Kernel
7.0.0-10-generic (cohort report)
Driver
xe / compute-runtime 26.09 (cohort report)
Runtime
oneAPI 2025.3.3 (cohort report)
CPU
Ryzen 5 9600X (cohort report)
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Dirty build: the exact local patch diff has not been recovered.
  • One thread is recorded in this result; methodology prose says six. Per-result metadata is used.
  • Warmup and five repetitions are described by methodology, but not recorded per result.
  • Configured 4K context is not a full-context generation test.
  • Energy and VRAM telemetry are excluded: device inclusion and measurement windows are unresolved.
  • Environment is reported by cohort documentation, not captured in this result. PCIe topology conflicts remain unresolved.

Original benchmark JSON ↗ · Reviewed data ↓

Original record: intel-arc-pro-b70-qwen3-5-27b-q4-k-m-sycl
Original SHA-256: e24e66b6dd5965219902775116f2fe733d81c87ba75cfc4dcedd5a740150a196

Gemma 4 31BQ4_K_M · Dense1 × B70 · SYCL
2026-04-21 · DIRTY BUILD
21.70
Gemma 4 31B · 2026-04-21 — full configuration

Gemma 4 31B

Historical · tested 2026-04-21
Quantization
Q4_K_M
GPU count
1
Backend
llama.cpp / SYCL
Build commit
ec6f7a6a5c
Dirty build
Yes
Model revision
Unknown
Weight size (GiB)
17.07
Configured context
4096
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
Unknown / Unknown
KV types K / V
f16 / f16
Flash attention
Yes
Threads
1
Batch / microbatch
Unknown / Unknown
Concurrency
Unknown
Warmup
Unknown
Repetitions
Unknown
Decode tok/s ± reported SD
21.7 ± 0.03
Prefill tok/s ± reported SD
600.78 ± 0.87
OS
Ubuntu 26.04 (cohort report)
Kernel
7.0.0-10-generic (cohort report)
Driver
xe / compute-runtime 26.09 (cohort report)
Runtime
oneAPI 2025.3.3 (cohort report)
CPU
Ryzen 5 9600X (cohort report)
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Dirty build: the exact local patch diff has not been recovered.
  • One thread is recorded in this result; methodology prose says six. Per-result metadata is used.
  • Warmup and five repetitions are described by methodology, but not recorded per result.
  • Configured 4K context is not a full-context generation test.
  • Energy and VRAM telemetry are excluded: device inclusion and measurement windows are unresolved.
  • Environment is reported by cohort documentation, not captured in this result. PCIe topology conflicts remain unresolved.

Original benchmark JSON ↗ · Reviewed data ↓

Original record: intel-arc-pro-b70-gemma-4-31b-q4-k-m-sycl
Original SHA-256: 8d0b5ebabde9a515fcdd38538c435fd660dec076b3d6a98758aab58ded890470

Qwen 3.6 35B-A3BUD-Q4_K_M · MoE1 × B70 · SYCL
2026-06-13 · BUILD d8a24cc
68.88
Qwen 3.6 35B-A3B · 2026-06-13 — full configuration

Qwen 3.6 35B-A3B

Historical · tested 2026-06-13
Quantization
UD-Q4_K_M
GPU count
1
Backend
llama.cpp / SYCL
Build commit
d8a24cc
Dirty build
Unknown
Model revision
Unknown
Weight size (GiB)
20.604151248931885
Configured context
32768
Exercised prefill tokens
512
Decode tokens
128
Decode prompt / depth
0 / 0
KV types K / V
q4_0 / q4_0
Flash attention
Unknown
Threads
6
Batch / microbatch
2048 / 512
Concurrency
Unknown
Warmup
Unknown
Repetitions
3
Decode tok/s ± reported SD
68.880402 ± 0.411744
Prefill tok/s ± reported SD
1030.696373 ± 6.635591
OS
Unknown
Kernel
Unknown
Driver
Unknown
Runtime
oneAPI 2026.0 (cohort report)
CPU
AMD Ryzen 5 9600X 6-Core Processor
PCIe topology
Unknown
Editorial review
2026-09-10

What this result does not establish

  • Decode is a separate n_prompt=0, n_gen=128, n_depth=0 row. It is not generation after filling 32K context.
  • The 32K prefill row measures 747.127922 tok/s; explorer prefill consistently shows pp512.
  • Build cleanliness, model revision, OS, kernel, driver and warmup are not established by this raw file.
  • Flash attention records -1 (automatic); effective enablement is unknown.
  • This is a separate cohort with different cache and software settings. Do not calculate a controlled speedup against April.

Reviewed June measurement excerpt ↗ · Reviewed data ↓

Original record: 20260613T085749Z_llamacpp_qwen36-moe-35b-q4_single_32768_q4_0
Original SHA-256: 9558dbe958e1f04cd61190a0be5341377a380749dab2c5720b64f17942dafdf4

Energy rankings are withheld. The archived power summaries mix device inclusion and timing windows; zero VRAM readings are missing telemetry. Unknown metadata stays unknown. Read the measurement standards →