Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
18 commits
Select commit Hold shift + click to select a range
38a718a
spec(ORACLE-LLAMACPP-REPIN-STOCK): the llama.cpp floor was our own fo…
mudler Aug 16, 2026
a2ede63
record(ORACLE-LLAMACPP-REPIN-STOCK): repin llama.cpp to stock b10451,…
mudler Aug 16, 2026
3f63066
merge: origin/main into the ORACLE-LLAMACPP-REPIN-STOCK repin
mudler Aug 16, 2026
24e47f2
record(ORACLE-LLAMACPP-REPIN-STOCK): the repin's own claims, re-deriv…
mudler Aug 16, 2026
fa79980
merge: origin/main into the ORACLE-LLAMACPP-REPIN-STOCK re-derivation
mudler Aug 16, 2026
bab8e1f
record(ORACLE-LLAMACPP-REPIN-STOCK): the gate section said two reds, …
mudler Aug 16, 2026
489e93f
record(ORACLE-LLAMACPP-REPIN-STOCK): derive the contaminated set inst…
mudler Aug 16, 2026
e34e3e4
merge: origin/main into the ORACLE-LLAMACPP-REPIN-STOCK sweep, keyed …
mudler Aug 16, 2026
bf62128
record(ORACLE-LLAMACPP-REPIN-STOCK): the gate section is green now, a…
mudler Aug 16, 2026
5de5236
record(ORACLE-LLAMACPP-REPIN-STOCK): the sweep's own path set was a h…
mudler Aug 16, 2026
85a9a7a
record(ORACLE-LLAMACPP-REPIN-STOCK): mark the front page, and the one…
mudler Aug 16, 2026
8e0bf0e
record(ORACLE-LLAMACPP-REPIN-STOCK): the × token was present and dead…
mudler Aug 16, 2026
b60ad20
merge: origin/main into the ORACLE-LLAMACPP-REPIN-STOCK review repairs
mudler Aug 16, 2026
7395b77
record(ORACLE-LLAMACPP-REPIN-STOCK): the gate block headings name the…
mudler Aug 16, 2026
2f0cd62
record(ORACLE-LLAMACPP-REPIN-STOCK): the same head is green at load 1…
mudler Aug 16, 2026
fa94b10
record(ORACLE-LLAMACPP-REPIN-STOCK): the self-test covered 6 of 18 to…
mudler Aug 16, 2026
adcf2bb
record(ORACLE-LLAMACPP-REPIN-STOCK): the self-test proved liveness, n…
mudler Aug 16, 2026
cb07436
merge: origin/main into the #857 landing
mudler Aug 16, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .agents/backend-matrix.md

Large diffs are not rendered by default.

17 changes: 15 additions & 2 deletions .agents/environment.md
Original file line number Diff line number Diff line change
Expand Up @@ -431,8 +431,21 @@ inner 4096, state 128; context 262144.
only `GL_KHR_cooperative_matrix`, and the build still configures and runs. A
source-built shaderc `v2026.4-dev` lives at `/tmp/shaderc/b/glslc/glslc`; pass
`-DVulkan_GLSLC_EXECUTABLE=` to it. **Verify the runtime banner says
`matrix cores: NV_coopmat2` before trusting any number.** llama.cpp at pin
`237ad9b96` is unpacked at `~/lcpp-vk` with `build-vk/bin/llama-bench` built.
`matrix cores: NV_coopmat2` before trusting any number.** llama.cpp is
unpacked at `~/lcpp-vk` with `build-vk/bin/llama-bench` built, **and that tree
is the SUPERSEDED fork `237ad9b96`, not the pin.** Do not reuse it. The
llama.cpp oracle is stock `b10451` since 2026-08-16
([`oracles/llama-cpp.md`](oracles/llama-cpp.md)). `237ad9b96` is a local-only
commit on the developer's `localai-paged` branch, 65 of our own performance
commits past upstream `b9827`, built from a working tree with 27 uncommitted
entries. `~/lcpp-vk` therefore reproduces neither the pin nor any identifiable
object. The `BENCH-VK-LLAMA` decode `4.36 vs 4.35 MET` measured with it is the
**most fragile verdict** in the enumeration, a 0.23% margin inside a 0.69%
spread, and re-taking it is owed under
[#1003](https://github.com/mudler/vllm.cpp/issues/1003). Unpack the pinned SHA
fresh and assert `git status --porcelain` empty before recording any number.
Enumeration and the clean-tree rule:
[`specs/oracle-llamacpp-repin-stock.md`](specs/oracle-llamacpp-repin-stock.md).

- **No Intel GPU exists on any box here**, so `BACKEND-XPU` end-to-end work is
HW-BLOCKED; only policy-port, compile coverage and oneAPI CPU-device unit
Expand Down
4 changes: 2 additions & 2 deletions .agents/feature-matrix.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,7 +173,7 @@ and reference-engine performance.
| ID | Block | State | Grounded summary | Detailed evidence / spike |
|---|---|---|---|---|
| `QUANT-CUDA-GATES` | NVFP4 W4A16, NVFP4 W4A4, gate-specific FP8 W8A8 | `DONE` | support/correctness stays closed; performance remains `ACTIVE` at the current binding `9ecd9d0` **114/124**. FP4 tactics match, and clean `f344dec` closes W1D2/G2 with 27B **235/235** in default+rollback arms plus 35B/GGUF inertness. Component evidence remains open; no quantization speed credit follows | quant matrix §2 + [coverage spike](specs/quantization-coverage.md) |
| `QUANT-GGUF` | llama.cpp encodings and output presets | `PARTIAL` | F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated. **The B4 speed/RSS checkpoint is NO LONGER pending and this cell was two tracks stale:** direct compute-in-quant has been live since CIQ `G4`, and the 20-core Arm/i8mm arm reached llama.cpp parity or better on every axis (`BACKEND-GATE-CPU-LLAMACPP`). Still open: the four-core A76 arm on speed, and the **x86_64** arm, whose quant path is portable-tier only (CIQ `G5`, [#433](https://github.com/mudler/vllm.cpp/issues/433)) | quant matrix §1 |
| `QUANT-GGUF` | llama.cpp encodings and output presets | `PARTIAL` | F32/Q4_0/Q8_0/Q3_K/Q4_K/Q5_K/Q6_K materialize; F16 was corrected to reader-only; CPU threadpool/chunked dispatch is correctness-gated. **The B4 speed/RSS checkpoint is NO LONGER pending and this cell was two tracks stale:** direct compute-in-quant has been live since CIQ `G4`, and the 20-core Arm/i8mm arm reached llama.cpp parity or better on every axis (`BACKEND-GATE-CPU-LLAMACPP`). **That denominator is SUPERSEDED:** it was `237ad9b96`, our own local-only fork, and the oracle repinned to stock `b10451` on 2026-08-16, so every one of those axes is owed a re-take ([#1003](https://github.com/mudler/vllm.cpp/issues/1003), [#857](https://github.com/mudler/vllm.cpp/issues/857)). Still open: the four-core A76 arm on speed, and the **x86_64** arm, whose quant path is portable-tier only (CIQ `G5`, [#433](https://github.com/mudler/vllm.cpp/issues/433)) | quant matrix §1 |
| `QUANT-VLLM-BREADTH` | generic FP8/MX, AWQ/GPTQ, CT integer, vendor methods, KV | `PARTIAL` | gate-specific implementations exist; generic dispatch/modes remain inventoried | quant matrix §§2-3 |
| `QUANT-MLX` | affine Q2-8, MXFP4/MXFP8/NVFP4, QQ, mixed recipes/imports | `INVENTORIED` | required for Apple backend; no MLX runtime yet | quant matrix §4 |

Expand Down Expand Up @@ -283,7 +283,7 @@ evidence.
|---|---|---|---|---|
| `BACKEND-CUDA-SM121` | GB10/sm121a | `PARTIAL` | gate workload built, traced, token/perf gated; full component-family coverage is open | [backend row](backend-matrix.md#cuda-target-rows) |
| `BACKEND-CUDA-OTHER` | vLLM sm70/75/80/86/87/89/90/100/101/103/110/120 targets | `ACTIVE` | 9 CUDA arches build-supported (sm80/86/87/89/90a/100a/103a/110/120a, single-arch portable-kernels-only, `-Werror` clean, SASS emitted); sm70/75/101 not build-supported; no non-121a target is runtime-validated | [backend matrix](backend-matrix.md), [CUDA inventory](specs/cuda-architecture-inventory.md) |
| `BACKEND-CPU` | production CPU | `PARTIAL` | shared CPU path is correctness-gated and the 20-core Arm/i8mm Qwen3.5-2B single-stream llama.cpp floor is closed; server concurrency is open. Raspberry Pi 5 Cortex-A76 R4-R5 is green: QEMU-built artifact, exact output, and an AAPCS64 Q8 leaf that beats compiler SDOT 3.66-5.08% on M1/T1 and M128. Its separate same-file llama.cpp floor is now measured/NOT MET on speed: vllm.cpp is 0.461x prefill and 0.653x decode/E2E, while using 24.2% less RSS; exact-prompt 64-token output matches. M1/T4 is −2.43%; BF16 GEMM, speed closure and concurrency remain open | [backend matrix](backend-matrix.md), [A76 Q8 dot](specs/cpu-a76-q8-dot.md), [assembly evidence](../docs/bench-evidence/rpi5-a76-q8-dot-20260806.md), [llama.cpp evidence](../docs/bench-evidence/rpi5-a76-llamacpp-20260806.md) |
| `BACKEND-CPU` | production CPU | `PARTIAL` | shared CPU path is correctness-gated and the 20-core Arm/i8mm Qwen3.5-2B single-stream llama.cpp floor is closed **against a SUPERSEDED denominator**. It was our own local-only fork `237ad9b96`, the oracle repinned to stock `b10451` on 2026-08-16, and both the closing prefill `1.18x` and the `1.01x`/`0.97x` ties are owed a re-take ([#1003](https://github.com/mudler/vllm.cpp/issues/1003)); server concurrency is open. Raspberry Pi 5 Cortex-A76 R4-R5 is green: QEMU-built artifact, exact output, and an AAPCS64 Q8 leaf that beats compiler SDOT 3.66-5.08% on M1/T1 and M128. Its separate same-file llama.cpp floor is now measured/NOT MET on speed: vllm.cpp is 0.461x prefill and 0.653x decode/E2E, while using 24.2% less RSS; exact-prompt 64-token output matches. That arm ran stock `b9892`, a third revision that is neither pin, so its RSS win is owed a re-take too (#1003). M1/T4 is −2.43%; BF16 GEMM, speed closure and concurrency remain open | [backend matrix](backend-matrix.md), [A76 Q8 dot](specs/cpu-a76-q8-dot.md), [assembly evidence](../docs/bench-evidence/rpi5-a76-q8-dot-20260806.md), [llama.cpp evidence](../docs/bench-evidence/rpi5-a76-llamacpp-20260806.md) |
| `BACKEND-ROCM` | ROCm | `INVENTORIED` | source/dispatch spike required; no "one flag" support claim | [backend matrix](backend-matrix.md) |
| `BACKEND-MLX` | Apple Metal through MLX | `ACTIVE` | Metal/MLX skeleton ACTIVE: two models (OPT-125m, Qwen3-0.6B) run e2e + pass correctness, native-MSL GEMM, batched command buffers 1.50x (compute-bound at 98%+); optional MLX GEMM provider | [backend matrix](backend-matrix.md) |
| `BACKEND-VULKAN` | Vulkan | `ACTIVE` | **22 NATIVE kernels**; **opt-125m RUNS END TO END, STRICT token-exact 6/6 prompts / 96/96 tokens vs the vLLM 0.25.0 oracle, 0 provider declines** of the CPU backend's 87 registered ops (GEMM both orientations, embedding, greedy argmax, block-paged attention + KV write, QKV split, rotary apply, elementwise/norm/fusion, and the GDN/conv1d glue: sigmoid gate, gated RMSNorm, state gather/scatter, decode conv1d update, fused post-conv); the other 65 served by the portable reference tier (CPU kernel, unified memory). `kGdnPrefill`/`kGdnDecode`/`kCausalConv1dFwd` deliberately still host-tier; `kRopeCosSinCache` stays host by design (double-precision table, mirrors vLLM). No speed number owed | [backend matrix](backend-matrix.md), [campaign spec](specs/vulkan-full-support.md) |
Expand Down
Loading
Loading