Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
219 commits
Select commit Hold shift + click to select a range
f445fa8
spec(KERNEL-QUANT-CIQ-GEMM-ROCM): commit the keep-quant provider spec
ghazni101 Aug 21, 2026
6236e9e
feat(KERNEL-QUANT-CIQ-GEMM-ROCM): land the W1 keep-quant providers on…
ghazni101 Aug 21, 2026
2578c9b
spec(ROCM-QUANT-GEMM-BW): commit the keep-quant bandwidth spec
ghazni101 Aug 21, 2026
8e78dfa
perf(ROCM-QUANT-GEMM-BW): split each super-block across the warp's lanes
ghazni101 Aug 22, 2026
0783930
spec(GFX1100-TG150): commit the 150 tok/s campaign spec
ghazni101 Aug 22, 2026
41d060b
perf(ROCM-QUANT-GEMM-BW): branch-free Q6K scalar decode + gated qg4 a…
ghazni101 Aug 22, 2026
eb9e46f
perf(ROCM-QUANT-GEMM-BW): split-K decode arm for the keep-quant GEMM
ghazni101 Aug 22, 2026
c112d88
perf(BACKEND-ROCM): f32-query decode-GQA arm behind VT_ATTN_DECODE_GQ…
ghazni101 Aug 22, 2026
094f603
perf(BACKEND-ROCM): row-permuted keep-quant in_proj + tiny-N f32-out …
ghazni101 Aug 22, 2026
592afd3
spec(GFX1100-TG200): commit the 200 tok/s campaign spec
ghazni101 Aug 22, 2026
767d369
measure(GFX1100-TG200): T1 re-prices the tip -- 40.65 tok/s median, b…
ghazni101 Aug 22, 2026
f758643
measure(GFX1100-TG200): T2a splits the GdnPostConv symbol -- the grid…
ghazni101 Aug 22, 2026
5e57df7
research(GFX1100-TG200): rank vLLM/SGLang mechanisms against our meas…
ghazni101 Aug 22, 2026
2034e17
record(GFX1100-TG200): reject the pointer-keyed quant cache -- alloca…
ghazni101 Aug 23, 2026
bdc8114
perf(GFX1100-TG200): merged keep-quant gate_up -- one quant GEMM per …
ghazni101 Aug 23, 2026
64f38f1
perf(GFX1100-TG200): T2b flips ROCm support_static_graph_mode -- deco…
ghazni101 Aug 23, 2026
68316bd
record(GFX1100-TG200): session-state note appended to t2b evidence (h…
ghazni101 Aug 23, 2026
8a2ee61
perf(GFX1100-TG200): T3a ports the f32-query decode-GQA attention arm…
ghazni101 Aug 23, 2026
239c106
record(GFX1100-TG200): t3a evidence session-state note (hindsight 500s)
ghazni101 Aug 23, 2026
9df8f0a
test(GFX1100-TG200): T4a lands the red-first ROCm quant-dot gate
ghazni101 Aug 23, 2026
205df07
perf(GFX1100-TG200): T4a adds the VT_GEMV_MMVQ=1 K-quant decode GEMV …
ghazni101 Aug 23, 2026
9ce2055
perf(GFX1100-TG200): T4a folds activation quant into the MMVQ GEMV pr…
ghazni101 Aug 23, 2026
554e01d
record(GFX1100-TG200): T4a evidence -- decode GEMV lever closed negat…
ghazni101 Aug 23, 2026
60dac6c
fix(GFX1100-TG200): T4a repairs the MMVQ arm -- m-gates the whole dis…
ghazni101 Aug 23, 2026
baf97fe
record(GFX1100-TG200): T4a repair-round evidence -- lever adopted at …
ghazni101 Aug 23, 2026
0ebe869
test(GFX1100-TG200): T4a repair-2 adds host-side dispatch-route count…
ghazni101 Aug 23, 2026
28cbc6d
perf(GFX1100-TG200): T4a lever-B1 makes the fold crossover tunable be…
ghazni101 Aug 23, 2026
26e14bb
record(GFX1100-TG200): T4a lever-B1 evidence -- fold-crossover re-tun…
ghazni101 Aug 23, 2026
5140bd1
record(GFX1100-TG200): T4a lever-B2 attributes all 21.6 Cijk calls/to…
ghazni101 Aug 23, 2026
6c55cdf
test(GFX1100-TG200): T4a lever-B2 adds the red-first f32-out decode-s…
ghazni101 Aug 23, 2026
8dcac46
perf(GFX1100-TG200): T4a lever-B2 adds the VT_SKINNY_BF16=1 f32-out d…
ghazni101 Aug 23, 2026
6218d9b
test(GFX1100-TG200): T4a lever-B2 restores the routing-witness case p…
ghazni101 Aug 24, 2026
9f5d9da
test(GFX1100-TG200): T4a lever-B2 corrects the routing-witness expect…
ghazni101 Aug 24, 2026
cecceb7
record(GFX1100-TG200): T4a lever-B2 adopts the VT_SKINNY_BF16=1 f32-o…
ghazni101 Aug 24, 2026
a17a209
record(GFX1100-TG200): T4a lever-B2 closes the loop -- ON-arm capture…
ghazni101 Aug 24, 2026
aab871f
test(GFX1100-TG200): T4a lever-B2 makes the TRUE-unset routing window…
ghazni101 Aug 24, 2026
242e8da
record(GFX1100-TG200): T4a lever-B2 re-pins the default-routing witne…
ghazni101 Aug 24, 2026
c574c7c
attribution(GFX1100-TG200): lever-C maps all 97 QuantizeQ8KK decode l…
ghazni101 Aug 24, 2026
cfbba00
test(GFX1100-TG200): lever-C adds red-first witnesses for the fused n…
ghazni101 Aug 24, 2026
25f511f
feat(GFX1100-TG200): lever-C fuses Q8_K activation quant into the Rms…
ghazni101 Aug 24, 2026
79a992d
record(GFX1100-TG200): lever-C adopts VT_NORM_QUANT_FUSED=1 at +7.3% …
ghazni101 Aug 24, 2026
65d8555
perf(GFX1100-TG200): T5a vectorizes the shared Q8_K quant superblock …
ghazni101 Aug 25, 2026
53eb36a
perf(GFX1100-TG200): T5b extends the f32-Q DecodeGqa arm to head_dim 128
ghazni101 Aug 25, 2026
bd6551f
record(GFX1100-TG200): T5b evidence — attention routing hole closed, …
ghazni101 Aug 25, 2026
580c185
record(GFX1100-TG200): T5c closed negative — MMVQ nontemporal weight …
ghazni101 Aug 25, 2026
314dd01
perf(GFX1100-TG200): T6a adds the warp-per-row cooperative GDN scan arm
ghazni101 Aug 25, 2026
b8608a2
record(GFX1100-TG200): T6a evidence — cooperative scan adopted at +4.…
ghazni101 Aug 25, 2026
0a7ac39
perf(GFX1100-TG200): T6b adds the warp-per-item cooperative attn prea…
ghazni101 Aug 25, 2026
3945b7f
record(GFX1100-TG200): T6b evidence — cooperative preamble adopted at…
ghazni101 Aug 25, 2026
1cee023
record(GFX1100-TG200): session-close attribution — 76.6 tok/s median …
ghazni101 Aug 25, 2026
9d028b7
record(GFX1100-TG200): classify the campaign env knobs kernel-internal
ghazni101 Aug 25, 2026
4793e87
Merge upstream main into row/GFX1100-TG200
ghazni101 Aug 25, 2026
bfaff5b
record(GFX1100-TG200): T7 evidence — COALK load topology closed wash,…
ghazni101 Aug 25, 2026
2c30b3e
perf(GFX1100-TG200): T8 adds a cooperative single-row rmsnorm arm
ghazni101 Aug 25, 2026
90ae363
record(GFX1100-TG200): T8 evidence — cooperative rmsnorm adopted at +…
ghazni101 Aug 25, 2026
36a6b87
perf(GFX1100-TG200): T9 gives the gated norm a per-row block
ghazni101 Aug 25, 2026
23f1858
record(GFX1100-TG200): T9 evidence — cooperative gated norm adopted a…
ghazni101 Aug 25, 2026
7c518f6
perf(GFX1100-TG200): T10 and T11 add warp postconv and row-split scan…
ghazni101 Aug 26, 2026
c5827b3
record(GFX1100-TG200): T12 evidence — gated-quant fusion not adopted,…
ghazni101 Aug 26, 2026
cce72ca
record(GFX1100-TG200): dispatch-gap audit names the sampling round trip
ghazni101 Aug 26, 2026
38b08ed
record(GFX1100-TG200): retract T10/T11 engine claims — corrupted outp…
ghazni101 Aug 26, 2026
35fc458
test(GFX1100-TG200): close the gate gap that let the T10 stride bug ship
ghazni101 Aug 26, 2026
b078310
record(GFX1100-TG200): add mechanical decision rules for the T10/T11 …
ghazni101 Aug 26, 2026
920994b
record(GFX1100-TG200): refine dispatch-gap into three measured sub-ta…
ghazni101 Aug 26, 2026
9b75706
perf(GFX1100-TG200): T14 adds a row-split greedy argmax arm
ghazni101 Aug 26, 2026
1033485
record(GFX1100-TG200): verify at ISA level that the dp4a core uses v_…
ghazni101 Aug 26, 2026
317b198
record(GFX1100-TG200): T10/T11 re-measured in a clean window — adopte…
ghazni101 Aug 26, 2026
af50d9d
record(GFX1100-TG200): full-config verification — 92.7-92.9 tok/s can…
ghazni101 Aug 26, 2026
92392bb
record(GFX1100-TG200): async-serving A/B is a wash under HTTP overhea…
ghazni101 Aug 26, 2026
7972238
record(GFX1100-TG200): LDS epilogue closed negative; host-load sensit…
ghazni101 Aug 26, 2026
3be8d2f
record(GFX1100-TG200): position-resolved wvSplitKSml audit corrects t…
ghazni101 Aug 26, 2026
0081ed9
perf(GFX1100-TG200): T16 adds wvSplitK launch-config sweep knobs
ghazni101 Aug 26, 2026
cf03a41
record(GFX1100-TG200): T14 stacked engine A/B closed — adopted at +0.9%
ghazni101 Aug 26, 2026
2ae5496
record(GFX1100-TG200): core pinning does not isolate host-memory cont…
ghazni101 Aug 26, 2026
acfc6ee
record(GFX1100-TG200): T13 implementation plan scoped with file:line …
ghazni101 Aug 26, 2026
39b8b32
adopt(GFX1100-TG200): T16 YTILE=4 default — wins 5/5 paired, bit-iden…
ghazni101 Aug 26, 2026
5d0bc50
record(GFX1100-TG200): idle-window gate 100.4 tok/s, T13 wash, copy s…
ghazni101 Aug 26, 2026
4f3c87e
record(GFX1100-TG200): T17 v_dot2_f32_bf16 closed not-adopted — memor…
ghazni101 Aug 26, 2026
e016838
feat(GFX1100-TG200): T18 v_dot4 instruction selection in KQuantGemvMm…
ghazni101 Aug 26, 2026
2016c3a
record(GFX1100-TG200): T20 full-warp cooperative GEMV closed not-adop…
ghazni101 Aug 26, 2026
08ba905
feat(GFX1100-TG200): T21 keep-quant for V-head row-permuted GDN proje…
ghazni101 Aug 26, 2026
da60bbf
fix(GFX1100-TG200): T22 NormQuant bridge token survives non-matching …
ghazni101 Aug 26, 2026
914cd3c
feat(GFX1100-TG200): T24 LDS-buffered quant epilogue in RmsNormRowCoo…
ghazni101 Aug 26, 2026
dc8a3ba
T25: keep ssm_out as Q5_K with runtime input permutation
ghazni101 Aug 26, 2026
89170de
T27: warp-cooperative QuantizeQ8KK for decode (+2.06%, byte-identical)
ghazni101 Aug 27, 2026
eb0645d
feat(GFX1100-TG150): fuse Q6_K bias correction into single dot product
ghazni101 Aug 27, 2026
7f4b907
spec(GFX1100-TG150): add Outcome, update Now after attempt cap
ghazni101 Aug 27, 2026
0e4fe55
merge(GFX1100-TG200): integrate TG150-SPEC campaign spec
ghazni101 Aug 27, 2026
1a80549
merge(GFX1100-TG200): integrate ROCM-QUANT-GEMM-BW keep-quant work
ghazni101 Aug 27, 2026
001239f
cherry-pick(KV-FP8): W6 ROCm fp8-e4m3 KV cache store and read onto TG200
ghazni101 Aug 27, 2026
c0af8c1
spec(rocm): fp8 KV cache decode attention for PagedAttnDecodeGqaF32Q …
ghazni101 Aug 27, 2026
2cb34e1
feat(rocm): fp8 KV cache support in PagedAttnDecodeGqaF32Q (#7)
ghazni101 Aug 27, 2026
12e1c53
perf(ROCm): parallel random sample with shared primitives
ghazni101 Aug 27, 2026
b417b4a
feat(rocm): advertise fp8 KV cache dtype support and add GGUF chat te…
ghazni101 Aug 27, 2026
37bbb24
merge(GFX1100-TG200): integrate upstream main to clear #1936
ghazni101 Aug 28, 2026
2e8d193
fix(GFX1100-TG200): repair the record and env gates the branch carrie…
ghazni101 Aug 28, 2026
9f5fd75
merge(GFX1100-TG200): integrate upstream main at e551cf8e4
ghazni101 Aug 28, 2026
7884402
fix(GFX1100-TG200): drop fp8 KV decode-attn extras onto row/fp8-kv-de…
ghazni101 Aug 28, 2026
7d9a690
fix(ENG-MM-INPUT-PIPELINE): store the Qwen3-VL, Gemma-4 and GGUF visi…
localai-bot Aug 28, 2026
97464e4
measure(PERF-LAGUNA-FUSED-GATEUP): W3 -- the lever is worth ~4% and i…
localai-bot Aug 28, 2026
4248d0c
spec(BENCH-C8-ADMISSIBILITY): every c=8 number this repository quotes…
localai-bot Aug 28, 2026
0fc10d6
feat(MODEL-MM-QWEN4-EXP): W5b-3 — the PLE dilated depthwise conv is n…
localai-bot Aug 28, 2026
4c47e2f
fix(ENG-ATTN-OPTIN-SWEEP): sweep every remaining `vt::Attention` call…
localai-bot Aug 28, 2026
ecbc8db
fix(LTX25-TEXT-PROJ-DTYPE): resolve a caption projection's storage fo…
localai-bot Aug 28, 2026
d322f6b
feat(LTX25-ORACLE-ABSOLUTE): the blockiness ratios gate against #1864…
localai-bot Aug 28, 2026
5a1d653
spec(SPEC-DFLASH2): the selector's edge kernel reads every successor …
localai-bot Aug 28, 2026
099bea4
record(MODEL-MM-GLM53-FLASH): W0 -- pin transformers 5.16.1 for the g…
localai-bot Aug 28, 2026
a838591
spec(SPEC-DFLASH2): retract the selector-edge mechanism — `sample=` t…
localai-bot Aug 28, 2026
8c87f33
feat(MODEL-MM-QWEN4-EXP): W5b-4 — Qwen Sparse Attention as two `vt::`…
localai-bot Aug 28, 2026
c89b599
feat(MODEL-MM-GLM53-FLASH): W2 — the KDA forget gate is the sigmoid b…
localai-bot Aug 28, 2026
822005f
record(ENG-LTX-RECORD-RECONCILE): the trailer walk has no merge-commi…
localai-bot Aug 28, 2026
c5ef97d
feat(QUANT-EXL3): W1a — EXL3 becomes a scheme on vLLM's LinearMethod …
localai-bot Aug 28, 2026
5813711
fix(ENG-RECURRENT-MULTISTATE): a recurrent layer carries N states, an…
localai-bot Aug 28, 2026
9e91414
feat(MODEL-MM-GLM53-FLASH): W4 -- the mHC wiring, and a head collapse…
localai-bot Aug 28, 2026
0ff5b91
record(MODEL-TEXT-GLM-MOE-DSA): the row's two upstream anchors were e…
localai-bot Aug 28, 2026
0718922
fix(BACKEND-TENSTORRENT-QWEN35): the host-free opt-out leg gates agai…
lu-zero Aug 28, 2026
33b7871
record(BACKEND-TENSTORRENT-QWEN35): reconcile the spec's Now after W4…
lu-zero Aug 28, 2026
1277ea3
record(MODEL-MM-QWEN4-EXP): a speed denominator exists, and the targe…
localai-bot Aug 28, 2026
a50ae93
perf(SPEC-DFLASH2): the batched draft's context gather was the identi…
localai-bot Aug 28, 2026
5d27ddc
spec(SPEC-DFLASH2): the draft forward is 76% of the draft phase, and …
localai-bot Aug 28, 2026
210c882
fix(SPEC-DFLASH2): route the two HOT draft forward bodies through the…
localai-bot Aug 28, 2026
e4fbdf2
record(SPEC-DFLASH2): O3 said the draft's dimensions were unrecorded …
localai-bot Aug 28, 2026
db6215b
perf(SPEC-DFLASH2): tile the draft attention's query axis PER REQUEST…
localai-bot Aug 28, 2026
9fe1924
measure(PERF-LAGUNA-FUSED-GATEUP): W4 partial -- prompt 0 reproduces …
localai-bot Aug 28, 2026
c316e83
record(ORACLE-LLAMA-CPP-GLM5NEXT): pin the llama.cpp that can open th…
localai-bot Aug 28, 2026
52be09a
feat(QUANT-EXL3): W1b — a stock EXL3 checkpoint loads and generates, …
localai-bot Aug 28, 2026
c297db4
feat(SPEC-DFLASH2): split `fwd` into its op groups, because every lev…
localai-bot Aug 28, 2026
b5b9e48
feat(BENCH-C8-ADMISSIBILITY): a leg ledger that survives the host reb…
localai-bot Aug 28, 2026
323e336
feat(MODEL-MM-QWEN4-EXP): W5c-1 — the KV-cache spec is three groups, …
localai-bot Aug 28, 2026
fa9c344
feat(MODEL-MM-GLM53-FLASH): W3 makes the NoPE MLA geometry representa…
localai-bot Aug 29, 2026
7fd4de1
feat(MODEL-MM-dots3-note): W5 — the MoE layer reaches the decode path…
localai-bot Aug 29, 2026
061e563
record(BENCH-C8-ADMISSIBILITY): the spec named the wrong committed ha…
localai-bot Aug 29, 2026
b2dd85b
record(BENCH-C8-ADMISSIBILITY): the right instrument cannot express t…
localai-bot Aug 29, 2026
871e70c
feat(BENCH-C8-ADMISSIBILITY): give the leg ledger its caller, so it i…
localai-bot Aug 29, 2026
82c2dd0
feat(BENCH-C8-ADMISSIBILITY): give the serving driver a speculative a…
localai-bot Aug 29, 2026
bd64366
docs(SPEC-DFLASH2): the batched-lane spec still refused a merge that …
localai-bot Aug 29, 2026
060042d
feat(MODEL-MM-QWEN4-EXP): W5b-5 — the QSA indexer composition moves o…
localai-bot Aug 29, 2026
906a57f
record(BACKEND-TENSTORRENT-QWEN35): index the W3 leftovers issue (#2201)
lu-zero Aug 28, 2026
e53e356
fix(BACKEND-TENSTORRENT-QWEN35): count the two missing GDN d2h paths …
lu-zero Aug 28, 2026
2f7a26c
fix(BACKEND-TENSTORRENT-QWEN35): scope the GDN cache role refusal to …
lu-zero Aug 28, 2026
82c3021
measure(PERF-LAGUNA-FUSED-GATEUP): W4 -- 6 of 6 prompts diverge, and …
localai-bot Aug 29, 2026
10ec490
feat(LOADER-GGUF-IQ): port the IQ2_XS and IQ4_XS dequantizers, the tw…
localai-bot Aug 29, 2026
f218ed8
feat(QUANT-EXL3): W3 — a stock EXL3 checkpoint could not run on a GPU…
localai-bot Aug 29, 2026
bbef971
fix(SPEC-DFLASH2): the draft's paged attention synchronized inside th…
localai-bot Aug 29, 2026
2fd27a2
spec(PERF-LAGUNA-GROUPED-GEMV): measure what bounds the grouped Q4_K/…
localai-bot Aug 29, 2026
f321f54
Port async device-mirror combine/scatter kernels to ROCm
ghazni101 Aug 29, 2026
6933c6e
spec(MODEL-TEXT-GLM-MOE-DSA): GLM-5.3 is 97.49% routed experts, so th…
localai-bot Aug 29, 2026
746b969
Fuse silu-mul with the Q8_K quant epilogue on ROCm (VT_SILU_QUANT_FUSED)
ghazni101 Aug 29, 2026
eebe010
measure(LTX25-ORACLE-ABSOLUTE): #1854's reading is taken, and our ren…
localai-bot Aug 29, 2026
7fe3d3d
record(GFX1100-TG200): three-point branch audit — merges help, 103→91…
ghazni101 Aug 29, 2026
aaec261
fix(MODEL-MM-QWEN4-EXP): one gamma polarity for the whole architectur…
localai-bot Aug 29, 2026
0606559
fix(SPEC-DFLASH2): the capture-safe bound was a per-STEP value baked …
localai-bot Aug 29, 2026
0d9ef2f
record(BACKEND-TENSTORRENT-QWEN35): index the mesh CQ staging wave (#…
lu-zero Aug 29, 2026
a0a0b5a
perf(BACKEND-TENSTORRENT-QWEN35): stage bf16 uploads into a per-slot …
lu-zero Aug 29, 2026
95aa758
record(BACKEND-TENSTORRENT-QWEN35): W5 lands allocation-free staging,…
lu-zero Aug 29, 2026
9458f1c
fix(MODEL-MM-GLM53-FLASH): read the layer schedule out of `attention.…
localai-bot Aug 29, 2026
aed3c9a
fix(MODEL-MM-GLM53-FLASH): read `attention.key_length` the way llama.…
localai-bot Aug 29, 2026
9a2b300
record(BACKEND-TENSTORRENT-QWEN35): index the batched staging wave (#…
lu-zero Aug 29, 2026
e0341ea
record(BACKEND-TENSTORRENT-QWEN35): W6 is not expressible — the trace…
lu-zero Aug 29, 2026
ef66c68
record(BACKEND-TENSTORRENT-QWEN35): repair the W6 evidence citations …
lu-zero Aug 29, 2026
7aab6c9
record(BACKEND-TENSTORRENT-QWEN35): reconcile the W6 bullet's superse…
lu-zero Aug 29, 2026
c5e8032
record(GFX1100-TG200): repair the three staged-preflight gate reds
ghazni101 Aug 29, 2026
e367784
record(GFX1100-TG200): T34 splits the launch-bound residual into in-g…
ghazni101 Aug 29, 2026
dcf6565
fix(MODEL-DSV4-EXL3): the carried tower's FP8 half is held at bf16 (#…
localai-bot Aug 29, 2026
f8ac511
fix(MODEL-MM-GLM53-FLASH): map the published GGUF's `glm4` pre name o…
localai-bot Aug 29, 2026
dc25354
T36: prefill M-tiled K-quant GEMM streams weight rows once (VT_PREFIL…
ghazni101 Aug 29, 2026
3b87aeb
feat(MODEL-MM-QWEN4-EXP): W5d-2 — one interleaved-mRoPE table builder…
localai-bot Aug 29, 2026
bf70c00
record(GFX1100-TG200): T35 b/a GEMV merge closed red on token coheren…
ghazni101 Aug 29, 2026
394667f
feat(QUANT-GGUF-IQ-VECDOT): keep the IQ2_XS and IQ4_XS blocks through…
localai-bot Aug 29, 2026
b6816ad
measure(PERF-LAGUNA-GROUPED-GEMV): W1 -- the grouped GEMV is latency-…
localai-bot Aug 29, 2026
80eedfd
record(GFX1100-TG200): near-tie adjudication harness — teacher-forced…
ghazni101 Aug 29, 2026
068bf4b
record(BACKEND-ROCM): index #9, the never-run quant gate red on row/G…
ghazni101 Aug 29, 2026
0269efc
record(GFX1100-TG200): allowlist VT_PREFILL_TILE, the T36 same-binary…
ghazni101 Aug 29, 2026
4eabe4d
feat(MODEL-MM-GLM53-FLASH): W5c -- the weight tower, and the model LO…
localai-bot Aug 29, 2026
a7973f6
fix(SPEC-DFLASH2): check the paged draft block's bounds on EVERY back…
localai-bot Aug 29, 2026
54d9303
feat(MODEL-MM-QWEN4-EXP): W5d-1 — the ungated grouped RMS norm the PL…
localai-bot Aug 29, 2026
9154071
record(PERF-LAGUNA-GROUPED-GEMV): the W11 lever list is exhausted, an…
localai-bot Aug 29, 2026
a9c8ddf
spec(MODEL-DSV4-DSA-COMPOSE): the DSA composition gets the owning row…
localai-bot Aug 29, 2026
419ddd2
fix(SPEC-DFLASH2): read the DEVICE bounds back, which is the one clas…
localai-bot Aug 29, 2026
c766509
fix(ENGINE-HYBRID-PLACEMENT): refuse an unplaceable MoE arm at the se…
localai-bot Aug 29, 2026
a9affc5
fix(GFX1100-TG200): stop the q6_K bias borrow corrupting the MMVQ arm
ghazni101 Aug 29, 2026
81800df
record(GFX1100-TG200): retire the stale 132,094 gate figure, log the …
ghazni101 Aug 29, 2026
82bfdce
record(GFX1100-TG200): re-mint the campaign reference on the fixed Q6…
ghazni101 Aug 29, 2026
567ad15
feat(MODEL-MM-GLM53-FLASH): W5 lands the 288+1 MoE and the KV-cache s…
localai-bot Aug 29, 2026
f1af567
docs(MODEL-DSV4-EXL3): the Owed entries pointed at an issue W1d close…
localai-bot Aug 29, 2026
dd0df93
T35-r3: GDN b/a same-input GEMV merge, adjudicated bit-identical, clo…
ghazni101 Aug 29, 2026
51513da
T37: small-N GEMV geometry levers closed negative — warps/split-K win…
ghazni101 Aug 29, 2026
8374ffc
feat(MODEL-MM-QWEN4-EXP): W5d-4 — the MoE weight adapter, and the fou…
localai-bot Aug 29, 2026
3f9177f
feat(MODEL-MM-QWEN4-EXP): W5d-4 — the MoE weight adapter, and the fou…
localai-bot Aug 29, 2026
f4bcd59
Merge upstream/main into row/GFX1100-TG200
ghazni101 Aug 29, 2026
d68be86
fix(vt): close the PermuteVHeads scopes the upstream merge unified
ghazni101 Aug 29, 2026
cc33e68
record(GFX1100-TG200): T38 sync gates — upstream merge, reference intact
ghazni101 Aug 29, 2026
28c1689
fix(ENV-DOC): allowlist the three TG200 tuning knobs the checker flagged
ghazni101 Aug 29, 2026
45711c2
record(GFX1100-TG200): pin the rebuilt sync SHAs and the owed anchor rot
ghazni101 Aug 29, 2026
982e40c
record(ENG-MM-INPUT-PIPELINE): the runner drops mm_features, so a Qwe…
localai-bot Aug 29, 2026
cff2576
record(ENG-MM-INPUT-PIPELINE): the runner drops mm_features, so a Qwe…
localai-bot Aug 29, 2026
73d6409
perf(PERF-QWEN35-STAGE-WEIGHTS): stage dense decode weights to the de…
localai-bot Aug 29, 2026
207c129
perf(PERF-QWEN35-STAGE-WEIGHTS): stage dense decode weights to the de…
localai-bot Aug 29, 2026
101a71a
feat(MODEL-MM-GLM53-FLASH): W5b-1 lands the DSA attention block, and …
localai-bot Aug 30, 2026
10b5cab
feat(MODEL-MM-GLM53-FLASH): W5b-1 lands the DSA attention block, and …
localai-bot Aug 30, 2026
91d4fac
record(GFX1100-TG200): T38 cross_device gate + near-tie adjudication
ghazni101 Aug 30, 2026
dd91846
docs(ENV-DOC): document VT_QWEN35_STAGE_MIN_FREE_FRAC so the base gat…
localai-bot Aug 30, 2026
ebc1543
docs(ENV-DOC): document VT_QWEN35_STAGE_MIN_FREE_FRAC so the base gat…
localai-bot Aug 30, 2026
e2466a7
feat(MODEL-MM-QWEN4-EXP): W5d-3 — the QSA consumer can now read the P…
localai-bot Aug 30, 2026
7873736
feat(MODEL-MM-QWEN4-EXP): W5d-3 — the QSA consumer can now read the P…
localai-bot Aug 30, 2026
edb3061
fix(GFX1100-TG200): three pre-existing bugs + record-anchor ratchet
ghazni101 Aug 30, 2026
a63b1fc
feat(MODEL-MM-QWEN4-EXP): W5c-2 gathers EVERY published KV group's bl…
localai-bot Aug 30, 2026
d858fa9
feat(MODEL-MM-QWEN4-EXP): W5c-2 gathers EVERY published KV group's bl…
localai-bot Aug 30, 2026
8f61f7a
fix(runner): guard async dispatch helpers for non-GPU builds
ghazni101 Aug 30, 2026
25915ec
feat(MODEL-MM-GLM53-FLASH): W5b-2a threads the four-stream manifold t…
localai-bot Aug 30, 2026
c0fa299
feat(MODEL-MM-GLM53-FLASH): W5b-2a threads the four-stream manifold t…
localai-bot Aug 30, 2026
44ece0f
fix(PERF-QWEN35-STAGE-WEIGHTS): CUDA never answered the memory budget…
localai-bot Aug 30, 2026
a220309
fix(PERF-QWEN35-STAGE-WEIGHTS): CUDA never answered the memory budget…
localai-bot Aug 30, 2026
b628c9e
fix(MODEL-MM-QWEN4-EXP): the llama.cpp arm's KV guard was failing ope…
localai-bot Aug 30, 2026
bd90b92
fix(MODEL-MM-QWEN4-EXP): the llama.cpp arm's KV guard was failing ope…
localai-bot Aug 30, 2026
4e65261
Merge remote-tracking branch 'upstream/main' into row/GFX1100-TG200
ghazni101 Aug 30, 2026
c653e14
fix(ci): MSVC shadowing, stale anchor, stray file, upstream sync
ghazni101 Aug 30, 2026
349df8e
feat(MODEL-MM-GLM53-FLASH): W5b-2b makes `ModelRegistry::Forward` rea…
localai-bot Aug 30, 2026
6493e2b
merge(GFX1100-TG200): resolve conflicts with upstream/main 349df8e9a
ghazni101 Aug 30, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions .agents/backend-matrix.md

Large diffs are not rendered by default.

5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-GLM53-FLASH-W5B1.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-GLM53-FLASH-W5B1

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-GLM53-FLASH-W5B1` | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` (`ACTIVE`) | Claude Code (opus-5), helper role — a fresh implementer working from the committed spec | local linked worktree `/home/mudler/_git/vllm.cpp-glmw5b1`, base SHA `df024dce466fcfde9b3d2fd40e55d2c25b48e96e`. CPU only: no `rc` lease was taken, no GPU was used, no `ssh` to a fleet device was attempted and no file mutex was needed. No checkpoint download and NO materialising load: `/mnt/nas_share/rc/ckpt/GLM-5.3-Flash-UD-Q2_K_XL/` was not opened at all, because W5c already measured that a materialising load on this box stops at 8.09 GiB RSS in uninterruptible CIFS I/O (O22) and this wave needs no artifact — its substrate is the synthetic `glm5next` GGUF miniature W5c gates its loader against. The `transformers` `v5.16.1` oracle was installed into a throwaway venv under the session scratchpad with `--system-site-packages` so the resident torch was reused rather than re-downloaded, and its `modeling_glm5_next.py` was verified to hash `2092bbb4efa2a8087b74f4a4da37635c503fe1df9ae73f1e6e8342af8b4b8e8b` — the value W3 and W5c both recorded — READ OFF THE INSTALLED MODULE and not off a downloaded copy | `row/MODEL-MM-GLM53-FLASH-W5B1`, issue [#2324](https://github.com/mudler/vllm.cpp/issues/2324) | Owns W5b-1 of [glm5-next-flash.md](../specs/glm5-next-flash.md) `## Work breakdown`: the `Glm5NextTextAttention` block and the `OwnedTensor` -> host f32 bridge. That is the new `src/vllm/model_executor/models/glm5_next_attn.{h,cpp}` and `glm5_next_bridge.{h,cpp}`, `tests/vllm/models/test_glm5_next_attn.cpp` and `test_glm5_next_bridge.cpp`, `tests/vllm/models/fixtures/gen_glm5_next_attn_goldens.py` and its emitted `.inc`, four CMake registrations, this claim, one appended `.agents/issue-index.md` row, and the spec's `### W5b` split, `### W5b-1`, `### W5b-2`, `## Owed` O25 and `## Now`. **EXCLUDES the rest of W5b and says so rather than narrowing silently**: the decoder layer, the mHC stream threading, `Glm5NextTextModel::Forward` and the binding to `MakeGlm5NextKVCache` are W5b-2's, still tracked by [#2241](https://github.com/mudler/vllm.cpp/issues/2241), which this pull request references and does NOT close. The split is not a size decision: W5b-1 answers to `transformers` v5.16.1 and the llama.cpp #27752 container and needs no cache over it, while W5b-2 answers additionally to `MakeGlm5NextKVCache` and the `[T, hc_mult, hidden]` manifold and is where reachability lands. EXCLUDES the KDA arm (W2), the DSA indexer (W3) — `SelectIndexerTopk` is CALLED and not touched — the mHC bricks (W4), the MoE (W5), the weight tower (W5c), the vision tower (W6) and the converter (W7). EXCLUDES `deepseek_v4_dsa.cpp` and `mla_attention.{h,cpp}`, which are the documented wrong reuse and W3's surface respectively, both untouched. EXCLUDES any routing of the experts through `layers::MlpGateUpMethodBase` / `vt::MergedGemmGroup`: O19 / [#2260](https://github.com/mudler/vllm.cpp/issues/2260) records that doing so makes `MoeGateUpSwiGLUGroupedCuda` throw for this 101.14 GiB-resident model, and the bridge has NO overload taking an expert bank. EXCLUDES any parity-pin advance and the `.agents/model-matrix.md` row: the row's lifecycle state does not move | `ACTIVE` | 2026-08-29 — landed both deliverables. **RED captured FIRST on the same tree, in one build**, from the plausible wrong port: with `expand_kv` reading `k_b` UNTRANSPOSED, the all-masked row filled with `-inf` instead of `finfo.min`, `IndexerRoleFor` never reporting `shared`, and the bridge's ceiling not checked, `test_glm5_next_attn` read 9/14 cases and 63/150 assertions failed and `test_glm5_next_bridge` read 3/8 and 3/55. That red also found TWO defects in the tests rather than the product — the refusal golden carried huggingface_hub's wrapper class name, and the bridge's shape case moved a dim `q_b_proj` also depends on, so it threw on the wrong tensor — both repaired before green. Green is 14/14 + 160 assertions and 8/8 + 56 assertions, both exit 0. **Eighteen negative mutations, each sha256-proved applied, built and restored byte-for-byte; SEVENTEEN kill their gate.** The eighteenth is recorded as an EQUIVALENT mutant with its reason in O25, not as a pass: `host_f32_bytes` computed from the dims instead of from the buffers is indistinguishable while `DecodeShaped` refuses a shape disagreement, and the test now pins the sum against the buffers themselves — an earlier version pinned it against the predictor and that mutation passed it. Two more findings came out of the same run and are repaired: the fixture could not distinguish `min(l+1, n-1)` from a wrapping `(l+1) % n`, so a schedule where they disagree was added, and one mutant failed to BUILD under `-Werror` on an unused parameter, which is a passing mutant proving nothing and was rewritten to keep the parameter used and wrong. **NOT REACHED from a production entry point** and O25 carries the disclosure: `grep` over `src/`, `include/` and `examples/` for the five new symbols and the two headers returns nothing outside the four files of this change, so there is no production call site to delete and `.agents/reachability.md`'s mutation is already answered. W5b-2 owns the wiring. **No token, no load and no speed number is claimed**, and none was observed: no oracle for this model runs on any device this project reaches (the reference needs 305.78 GiB FP8 or 598.5 GiB BF16 against ~119.63 GiB), so what is gated is the NUMERICS of one block against a tiny-shape reference and nothing about the MODEL. GPU gate `PENDING`, reason recorded above, no result invented |
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-GLM53-FLASH-W5B2.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-GLM53-FLASH-W5B2

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-GLM53-FLASH-W5B2` | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` (`ACTIVE`) | Claude Code (opus-5), helper role — a fresh implementer working from the committed spec | local linked worktree `/home/mudler/_git/vllm.cpp-glmw5b2`, base SHA `10b5cabb01eac7e24c5aa781538577b5f47c2217`. CPU only: no `rc` lease was taken, no GPU was used, no `ssh` to a fleet device was attempted and no file mutex was needed, because nothing in this wave has a device arm to measure. NO checkpoint download and NO materialising load: `/mnt/nas_share/rc/ckpt/GLM-5.3-Flash-UD-Q2_K_XL/` was opened only through the committed 1412-tensor header manifest and the committed `config.json` fixture, never as bytes — W5c measured that a materialising load on this box stops at 8.09 GiB RSS in uninterruptible CIFS I/O (O22). The `transformers` `v5.16.1` oracle ran from a `git worktree` at that tag plus a throwaway `--system-site-packages` venv holding only `tokenizers>=0.23.1` and `safetensors>=0.8.0`, so the resident torch 2.11.0 was reused rather than re-downloaded; its `modeling_glm5_next.py` was verified to hash `2092bbb4efa2a8087b74f4a4da37635c503fe1df9ae73f1e6e8342af8b4b8e8b` — the value W3, W5b-1 and W5c all recorded — READ OFF THE MODULE THE GENERATOR IMPORTED, and re-asserted a second time by `git cat-file -p v5.16.1:...` in the source clone | `row/MODEL-MM-GLM53-FLASH-W5B2`, issue [#2241](https://github.com/mudler/vllm.cpp/issues/2241) | Owns W5b-2**a** of [glm5-next-flash.md](../specs/glm5-next-flash.md) `## Work breakdown`: the decoder layer, the mHC stream threading, `Glm5NextTextModel::Forward` and the binding of the DSA block to `MakeGlm5NextKVCache`'s groups 0 and 2. That is the new `src/vllm/model_executor/models/glm5_next_layer.{h,cpp}`, an additive `DsaCache*` parameter on `glm5_next::Attention` and a `SelectIndexerTopkFromPacked` lifted out of `SelectIndexerTopk`'s body (both null-default and byte-identical on the uncached path W5b-1 gated, proved by rerunning W5b-1's two suites unchanged), `tests/vllm/models/test_glm5_next_layer.cpp`, `tests/vllm/models/fixtures/gen_glm5_next_layer_goldens.py` and its emitted `.inc`, two CMake registrations, this claim, one appended `.agents/issue-index.md` row, and the spec's `### W5b-2` split into `### W5b-2a` / `### W5b-2b` plus `## Owed` O26 and `## Now`. **EXCLUDES making `ForwardGlm5NextForConditionalGeneration` stop refusing by name, and says so rather than narrowing silently**: that is W5b-2b's, still tracked by [#2241](https://github.com/mudler/vllm.cpp/issues/2241), which this pull request references and does NOT close. The split is not a size decision — it is that reachability needs a weight bridge for the FOUR arms `BridgeDsaLayer` does not cover, and the MoE half of that is a residency design problem (~1,150 GiB of f32 routed experts across the 42 sparse layers against a ~119.63 GiB box, with `kBridgeTensorF32ByteCeiling` correctly refusing a 9.0 GiB bank today), not four more `BridgeDsaLayer`s. EXCLUDES the KDA arm's numerics (W2), the DSA indexer's selection (W3) — `SelectIndexerTopk`'s BODY is lifted into a second entry point and not otherwise touched, and its 1934 assertions are unchanged — the mHC bricks (W4), the MoE block (W5) — `MoeForward` is CALLED and not modified — the weight tower (W5c), the vision tower (W6) and the converter (W7). EXCLUDES any routing of the experts through `layers::MlpGateUpMethodBase` / `vt::MergedGemmGroup`: O19 / [#2260](https://github.com/mudler/vllm.cpp/issues/2260) records that doing so makes `MoeGateUpSwiGLUGroupedCuda` throw for this 101.14 GiB-resident model, and `glm5_next_layer.cpp` reaches only `vt::MoeRouterTopK`, `vt::MoeCombine` and `vt::KdaGatedDeltaRule` through the blocks W2 and W5 already landed. EXCLUDES any parity-pin advance and the `.agents/model-matrix.md` row: the row's lifecycle state does not move | `ACTIVE` | 2026-08-30 — landed all four deliverables. **RED captured FIRST**, on the first run of the new suite: 4 of 10 cases and 7 of 1647 assertions failed, layer 0 (KDA + dense) green and every DSA layer red by 2.7 to 8.1. Bisected against oracle intermediates rather than guessed — the mHC pre plus `input_layernorm` agreed to 4.8e-07 and the attention output did not — and the cause was in the ORACLE CONFIGURATION, not the port: `Glm5NextPreTrainedModel` sets `_supports_sdpa = True`, so a default `Glm5NextTextConfig` resolves `_attn_implementation` to `sdpa`, whose `build_attention_mask_from_topk` returns a BOOLEAN mask (`:1249-1250`) instead of the additive `finfo.min` one the eager arm builds (`:1252-1256`). The two backends DISAGREE on a left-padded query row where every key is masked — torch's SDPA emits 0.0 and eager's uniform softmax emits the mean of the values, measured 0.0 against our 0.509 — so the generator now pins `cfg._attn_implementation = "eager"`, which is the arm W5b-1 gated and the one `:1227-1228` names as the only interface a 3-D per-(query, key) mask can reach. Recorded in O26 as a fact about `glm5_next`, not about this fixture. Green is 10/10 cases and 1656/1656 assertions, exit 0, with the eight sibling glm5 suites unchanged and green (attn 160, dsa 1934, bridge 96, moe 1614, mhc 98, scaffold 2660, gguf_load 8731, kda 342). **Fourteen negative mutations, each sha256-proved applied, built and restored byte-for-byte; ALL FOURTEEN kill their gate — but only after a fifteenth finding, which is the one worth reading.** The mutation that truncates the attention's key range to the current window under a filled cache SURVIVED at 1647 of 1647 on the first pass. Its output is all-NaN; `NaN - want` is NaN and `NaN > x` is FALSE for every x, so the running maximum in the test's own `MaxGap` helper never left its initial zero and an ALL-NaN FORWARD READ AS A PERFECT MATCH on every gap assertion in the file. `MinStreamSeparation` and the cached-tail loop were blind the same way (`std::max(m, NaN)` returns `m`). All three now treat a non-finite value as an INFINITE gap and report the count separately, so a failure distinguishes "wrong number" from "not a number"; the mutation then reds 3 assertions and the suite grew from 1647 to 1656. One earlier mutant also failed to BUILD under `-Werror` on an unused variable, which is a passing mutant proving nothing, and was rewritten to keep the variable used and the arithmetic wrong. **NOT REACHED from a production entry point, and this wave's own scope said it would be** — O26 carries the disclosure in the strong form: `ForwardGlm5NextForConditionalGeneration` still refuses by name, so the only call site of `glm5_next_layer.{h,cpp}` is its own gate's, and O15, O16, O17, O23 and O25 are NOT discharged, because `.agents/reachability.md` is explicit that "an intermediate hop that is itself unreached does not carry". What changed is that the five primitives had five separate dead ends and now have ONE assembly point, gated against the reference. **No token, no load and no speed number is claimed**, and none was observed: no oracle for this model runs on any device this project reaches, so what is gated is the CONTROL FLOW and the NUMERICS of a five-layer stack at tiny shapes and nothing about the 321.32B model. GPU gate `PENDING`, reason recorded above, no result invented |
5 changes: 5 additions & 0 deletions .agents/claims/CLAIM-GLM53-FLASH-W5B2B.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
# CLAIM-GLM53-FLASH-W5B2B

| Claim | Row IDs | Agent | Worktree / remote dir | Branch | Owned scope | State | Last update |
|---|---|---|---|---|---|---|---|
| `CLAIM-GLM53-FLASH-W5B2B` | `MODEL-MM-glm5-next-glm5-next-for-conditional-generation` (`ACTIVE`) | Claude Code (opus-5), helper role — a fresh implementer working from the committed spec | local linked worktree `/home/mudler/_git/vllm.cpp-glmw5b2b`, base SHA `c0fa299b1a03225e88854a605a07a3d726ae2062`. CPU only for every gate: the whole suite is host f32 and the forward REFUSES a non-CPU queue by name, so no `rc` lease was needed to gate it and none was taken for that purpose. A `dgx:gpu0` lease WAS taken, for one thing only — the decisive check, an attempted real generation at the staged artifact — through `rc run -d dgx:gpu0`, never `ssh`. NO checkpoint download and NO materialising load on this box: `/mnt/nas_share/rc/ckpt/GLM-5.3-Flash-UD-Q2_K_XL/` was read only for its shard SIZES, because W5c measured that a materialising load here stops at 8.09 GiB RSS in uninterruptible CIFS I/O. No `transformers` run and no golden regeneration: this wave adds no numerics, so its oracle is the RESIDENT tower assembled from the same loader tensors, and the goldens W5b-2a captured at `v5.16.1` are consumed unchanged by `test_glm5_next_layer` (10 cases, 1656 assertions, unmoved) | `row/MODEL-MM-GLM53-FLASH-W5B2B`, issue [#2337](https://github.com/mudler/vllm.cpp/issues/2337), split out of [#2241](https://github.com/mudler/vllm.cpp/issues/2241) whose append-only index row is spent on W5b-2a | Owns W5b-2**b** of [glm5-next-flash.md](../specs/glm5-next-flash.md) `## Work breakdown`: the weight bridge for the four arms `BridgeDsaLayer` does not cover, and the engine binding that makes `ForwardGlm5NextForConditionalGeneration` stop refusing by name. That is `DecodeOwnedTensorRowsToF32` / `HostF32RowBytes` / `BridgeKdaLayer` / `BridgeMlp` / `BridgeMhcSite` / `BridgeMoeLayer` / `GgufExpertSource` in `glm5_next_bridge.{h,cpp}`, an `ExpertSource` seam and a grouped per-expert visit in `glm5_next_moe.{h,cpp}`, a `LayerWeightSource` overload of `TextModelForward` in `glm5_next_layer.{h,cpp}`, the new `glm5_next_forward.{h,cpp}`, the registry hook, the new `tests/vllm/models/test_glm5_next_forward.cpp`, four appended cases in `test_glm5_next_bridge.cpp`, the MOVED refusal pin in `test_glm5_next_scaffold.cpp`, two CMake registrations, one `scripts/runner-routing-allowlist.txt` entry, this claim, one appended `.agents/issue-index.md` row, and the spec's §W5b-2b outcome plus `## Owed` O27 and `## Now`. **EXCLUDES the numerics of every primitive it now reaches** — W2's KDA, W3's DSA indexer, W4's mHC, W5's MoE and W5b-1's attention are CALLED and not modified, and their seven suites are rerun unchanged. EXCLUDES the vision tower (W6), the converter (W7), the safetensors arm, the MTP head, any parity-pin advance and the `.agents/model-matrix.md` row: the row's lifecycle state does not move, because O1 does not. EXCLUDES routing the experts through `layers::MlpGateUpMethodBase` / `vt::MergedGemmGroup`: O19 / [#2260](https://github.com/mudler/vllm.cpp/issues/2260) records that doing so makes `MoeGateUpSwiGLUGroupedCuda` throw for this model, and the per-expert source keeps that unreachable BY CONSTRUCTION because it hands the block host floats. EXCLUDES ragged batching and the device arm, both REFUSED BY NAME rather than approximated, and both recorded under O27 | `ACTIVE` | 2026-08-30 — **`ModelRegistry::Forward` reaches this model, and the reachability mutation is the deliverable.** Deleting the `Glm5NextHostForward` call in the registry hook and returning an empty `ForwardLogits{}` REDS `test_glm5_next_forward` at 11 of 118 assertions; O26 said there was no production call site to delete, and there is one now. **RED CAPTURED FIRST from an EXISTING gate**: `test_glm5_next_scaffold`'s case pinning "the forward REFUSES BY NAME" went red at 8 assertions in 1 case the moment the hook ran, and the pin MOVED with the change the way W3's `MlaBlockDims` pin moved rather than being deleted by it. **The residency decision, with its arithmetic**: one sparse layer's three expert banks are 27.0 GiB in f32 and the 42 sparse layers together are 1,134 GiB against ~119.63 GiB usable — 9.5x over — so the banks are NEVER bridged; `num_experts_per_tok` is 8 of 288, one expert is 100,663,296 bytes (0.09375 GiB) and one bank ROW is 33,554,432 (32x UNDER the unchanged 1 GiB ceiling, where the bank is 9x over), and `MoeForward` visits each HIT expert ONCE. The forward's f32 peak is one layer plus one expert plus one 64 MiB `lm_head` chunk, under 0.75 GiB, against a 426.72 GiB resident tower. **Thirteen negative mutations on tree `91354df62`, each proved applied by a diff hash, each BUILD rc=0, each restored byte-for-byte; ALL THIRTEEN now kill their gate — but TWO SURVIVED the first suite and both repairs are in this branch.** M12, swapping the two mHC sites inside the layer source, left every logit BIT-IDENTICAL at 86 of 86: the fixture's mHC `fn` payloads are ramps in the hundreds and thousands, so the sigmoid gates saturate and the Sinkhorn projection converges to the same matrix from either site, and a gate that could only see the swap through the logits is a mute switch at that geometry — the mapping is now asserted STRUCTURALLY with the two tensors asserted to DIFFER, and M12 reds 24. M13, removing the per-expert grouping, left 32030 of 32030 green because the case ran ONE token, where every selected expert is hit once whatever the code does; a six-token case now fills 12 slots from 2 distinct experts and M13 reds 3. M5 kills by SIGSEGV (rc=139) rather than by an assertion, which corrected the refusal's own message: without it the loop dereferences a null source, it does not read zeros. **Green:** forward 9 cases / 118 assertions, bridge 19 / 32228, scaffold 38 / 2652, with the six sibling suites unchanged and green (layer 10/1656, moe 8/1614, gguf_load 16/8731, attn 14/160, kda 28/342, dsa 10/1934, mhc 5/98). **No token, no load and no speed number is claimed for the 321.32B model**, and none was gated: what runs in CI is a synthetic 4-layer `glm5next` miniature at `hidden_size` 32, and O1 stands unchanged. The `dgx:gpu0` attempt at the staged 101.25 GiB artifact is reported as exactly what it printed, in the pull request body |
Loading