Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
8e4c579
fix(LTX25-RESIDENCY-W0): the anchors are counted by the RENDER, and a…
mudler Aug 20, 2026
940fd2e
fix(LTX25-RESIDENCY-W0): the PPM writer gets its own anchor, because …
mudler Aug 20, 2026
f47c2e0
docs(LTX25-RESIDENCY-W0): what the third review found, what closed it…
mudler Aug 20, 2026
f7e7f4c
merge: take origin/main under the third review's anchors
mudler Aug 20, 2026
165db63
record(LTX25-RESIDENCY-W0): the third review's findings get an issue,…
mudler Aug 20, 2026
821ef54
fix(LTX25-RESIDENCY-W0): the two anchors a fourth review could still …
mudler Aug 20, 2026
6447d55
wip(LTX25-RESIDENCY-W0): checkpoint the N1/N6 anchor repairs before t…
mudler Aug 20, 2026
8c66ffa
merge: take origin/main under the fourth review's anchors
mudler Aug 20, 2026
b4b75af
record(LTX25-RESIDENCY-W0): the ten-mutation set re-run on the merged…
mudler Aug 20, 2026
99703ce
merge: take origin/main again, and re-measure rather than inherit
mudler Aug 20, 2026
1609e1d
record(LTX25-RESIDENCY-W0): the mutation table takes the second merge…
mudler Aug 20, 2026
8b82ad5
test(LTX25-RESIDENCY-W0): the SIBLING BOUNDARY moves 100% of a leaf o…
mudler Aug 20, 2026
17eba8c
merge: take origin/main a fourth time, and re-measure rather than inh…
mudler Aug 20, 2026
9e40c68
record(LTX25-RESIDENCY-W0): twelve mutations on the fourth merge head…
mudler Aug 20, 2026
1160d04
merge: take origin/main a fifth time, and re-run the set on a delta t…
mudler Aug 20, 2026
9d38ebd
record(LTX25-RESIDENCY-W0): the set reproduces byte-identically acros…
mudler Aug 20, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 3 additions & 0 deletions .agents/issue-index.md

Large diffs are not rendered by default.

453 changes: 436 additions & 17 deletions .agents/specs/ltx25-device-residency.md

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions benchmarks/demo/ltx25_phase_log_fixture_cpu.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
{
"_source": "MEASURED 2026-08-20 on mudler-ubuntu-box (Linux 6.8.0-136-generic x86_64, 20 cores, CONTENDED -- load average 31 to 55 while this ran, other sessions compiling throughout). Produced by `build/examples/ltx2-gen`, the shipped example, which is a client of `vllm.h` and nothing else: it calls `vllm_video_engine_load` + `vllm_video_generate` and asks `vllm_video_last_phase_log` for this file's path. Geometry `--frames 9 --width 64 --height 64 --seed 7 --max-phase 0 --device cpu` over the reduced fixture. Built from 3dc2ae98b on row/LTX25-RESIDENCY-W0 with `cmake -DCMAKE_BUILD_TYPE=Release -G Ninja`, gcc 13.3.0. Every byte below `_source`, `_caveat`, `_headline` and `_footnotes` is the render's own output. It REPLACES the artifact taken at e3f46560e, which predates the `denoise.step` and `decode.video.chunk` anchors and therefore could not show them.",
"_caveat": "NOT A BENCHMARK, AND NOT A PROFILE OF ANY SHIPPED CHECKPOINT. The weights are the REDUCED two-block fixture `tests/vllm/multimodal/ltx2_video_fixture.h` writes, not Lightricks' 21.00B DiT, so no duration here describes a real render. The host was contended, and the sharpest evidence for that is in this file's own history: the run taken ONE MINUTE before this one, from the same binary at the same geometry, measured 0.147 s of wall against this run's 4.463 s. Across five such runs the wall has read 0.147 s, 0.158 s, 1.676 s, 4.463 s, 6.138 s and 12.030 s, and the RANK of the two largest phases has reversed between them. Nothing in this file may be quoted as a speed, a ratio between phases, or an ordering. What it supports is the SHAPE of the table -- which phases exist, that they do not overlap, that each carrying phase CONTAINS its own sub-scopes, and that they add up. The `notice` field says the same thing qualitatively and travels with every phase log the library writes.",
"_source": "MEASURED 2026-08-20 on mudler-ubuntu-box (Linux 6.8.0-136-generic x86_64, 20 cores, CONTENDED -- load average 31 to 55 while this ran, other sessions compiling throughout). Produced by `build/examples/ltx2-gen`, the shipped example, which is a client of `vllm.h` and nothing else: it calls `vllm_video_engine_load` + `vllm_video_generate` and asks `vllm_video_last_phase_log` for this file's path. Geometry `--frames 9 --width 64 --height 64 --seed 7 --max-phase 0 --device cpu` over the reduced fixture. Built from 3dc2ae98b on row/LTX25-RESIDENCY-W0 with `cmake -DCMAKE_BUILD_TYPE=Release -G Ninja`, gcc 13.3.0. Every byte below `_source`, `_caveat`, `_headline` and `_footnotes` is the render's own output. It REPLACES the artifact taken at e3f46560e, which predates the `denoise.step` and `decode.video.chunk` anchors and therefore could not show them. This capture in turn PREDATES the third review's two further anchors -- `decode.video.vae`, one nested record per temporal group the tiled VAE decodes, and `artifacts.frames.ppm`, one per write callback -- so a table emitted by a later build carries two record kinds this one does not. Nothing else about the shape changed, and no duration here may be compared with one from any other run in either direction.",
"_caveat": "NOT A BENCHMARK, AND NOT A PROFILE OF ANY SHIPPED CHECKPOINT. The weights are the REDUCED two-block fixture `tests/vllm/multimodal/ltx2_video_fixture.h` writes, not Lightricks' 21.00B DiT, so no duration here describes a real render. The host was contended, and the sharpest evidence for that is in this file's own history: the run taken ONE MINUTE before this one, from the same binary at the same geometry, measured 0.147 s of wall against this run's 4.463 s. Across six such runs the wall has read 0.147 s, 0.158 s, 1.676 s, 4.463 s, 6.138 s and 12.030 s, and the RANK of the two largest phases has reversed between them. Nothing in this file may be quoted as a speed, a ratio between phases, or an ordering. What it supports is the SHAPE of the table -- which phases exist, that they do not overlap, that each carrying phase CONTAINS its own sub-scopes, and that they add up. The `notice` field says the same thing qualitatively and travels with every phase log the library writes.",
"_headline": "99.94% of an LTX-2.5 render's wall now has a phase name on it, over 33 entries, and each of the three phases that carry the render encloses the sub-scopes that name its work.",
"_footnotes": [
"CONTAINMENT, which is what this artifact is re-taken to show. Each carrying leaf encloses its own nested sub-scopes and they cover nearly all of it: denoise.step covers 99.993% of denoise over 8 denoiser evaluations, decode.video.chunk 99.982% of decode.video over 2 leaf records, and decode.audio.mel + decode.audio.vocoder 99.998% of decode.audio. A leaf whose name has been moved onto a neighbour's seconds stops containing its own sub-scopes, which is the one defect a sum cannot see.",
Expand Down
53 changes: 40 additions & 13 deletions docs/USAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -1134,19 +1134,46 @@ count. These fields say how complete it is, and how far it carries:

Some phases are **decomposed rather than partitioned**. `denoise` carries one
`denoise.step` per denoiser evaluation, `decode.video` carries
`decode.video.chunk` per streamed chunk, `decode.audio` carries
`decode.audio.mel` and `decode.audio.vocoder`, and a two-stage recipe's
`phase.prepare` carries `phase.upsample_latent`. Those records are marked
`nested`, are printed for the reader, and are **excluded from
`sum_leaf_seconds`** — they are inside a leaf that is already counted, so adding
them would make `unaccounted_seconds` the residue of double counting instead of
time nobody named.

A nested record is also what makes a phase NAME checkable. A leaf that claims to
cover the denoise must enclose its own `denoise.step` records; one that stops
short of the loop, or that hands the back half of it to a neighbouring name, no
longer does. The three phases that carry a render each carry such an anchor for
that reason.
`decode.video.chunk` per streamed chunk and `decode.video.vae` per temporal
group the tiled VAE decodes, `decode.audio` carries `decode.audio.mel` and
`decode.audio.vocoder`, `artifacts.frames` carries `artifacts.frames.ppm` per
write callback, and a two-stage recipe's `phase.prepare` carries
`phase.upsample_latent`. Those records are marked `nested`, are printed for the
reader, and are **excluded from `sum_leaf_seconds`** — they are inside a leaf
that is already counted, so adding them would make `unaccounted_seconds` the
residue of double counting instead of time nobody named.

A nested record is also what makes a phase NAME checkable, and it is checkable in
four ways rather than one. A leaf that claims to cover the denoise must ENCLOSE
its own `denoise.step` records; one that stops short of the loop, or that hands
the back half of it to a neighbouring name, no longer does. The anchor must
appear once per unit of work the RENDER counted — one `denoise.step` per
denoiser evaluation, one `decode.video.chunk` per streamed chunk plus the reopen
after the last one — which is the only check that is not a ratio against the
leaf, and therefore the only one an anchor that moves WITH its leaf cannot
satisfy. Two SIBLING anchors under one leaf must appear in the order the render
runs them, because nothing else distinguishes them: swapping the
`decode.audio.mel` and `decode.audio.vocoder` names moves 96.8% of that leaf's
decomposed seconds onto the wrong model and changes nothing any ratio can see.
And `decode.video.vae` must cover most of the `decode.video.chunk` seconds,
because a cardinality alone permits the anchor to sit beside the tile decode
rather than on it. The four phases that carry a render each carry such an anchor.

**What the anchors do NOT prove**, stated here because the table invites the
opposite reading. An anchor proves that the name is where the work is, in the
order the render runs it; it does not prove that a name covers the call it is
named after — that needs a scope inside the callee, and `load.dit` and the two
`conditioning.*` leaves do not have one. A leaf may also grow over adjacent time
that NOBODY named, up to its coverage slack: 5.3% for `denoise`, 11% for
`decode.video`, 1% for `decode.audio` and up to 100% for `artifacts.frames`,
whose threshold is loose because the leaf is sub-millisecond. Growing over a
NEIGHBOUR is caught, because the neighbour turns `nested`; growing over
`unaccounted_seconds` is not.

A record that is `nested` and is NOT one of those anchors means a leaf was left
open across a phase it does not name: the neighbour turns `nested`, leaves
`sum_leaf_seconds`, and overlaps nothing, so a table can lose a whole phase to
its neighbour without any two intervals crossing.

**Do not read a duration here as a measurement of this machine.** Every number
is wall clock under whatever else the box was doing, which the file does not
Expand Down
20 changes: 20 additions & 0 deletions include/vllm/multimodal/ltx2_video.h
Original file line number Diff line number Diff line change
Expand Up @@ -973,6 +973,26 @@ struct Ltx2ConditioningTrace {
// re-noises a keyframe.
std::vector<float> video_first_timesteps;

// ── THE VIDEO DECODE, counted by the RENDER (row LTX25-DEVICE-RESIDENCY W0)
//
// How many chunks the streaming video VAE handed back, incremented in the
// driver's own sink beside `rendered_frames`. It is here because W0's phase
// table needs ONE number about the decode that the phase table did not
// produce.
//
// WHY THAT MATTERS AND WHY A COUNTER RATHER THAN A LONGER COMMENT. Every
// assertion W0's containment case makes about `decode.video` — containment,
// coverage, exclusivity, non-overlap — is a RATIO against the leaf, so an
// instrument defect that moves the leaf and its sub-scope TOGETHER satisfies
// all four at once. A count taken by the render is the one quantity such a
// defect cannot move: emit the chunk scope once instead of once per chunk and
// the count disagrees, whatever the clock did.
//
// ONE PER `emit`, so it is the group count `Ltx2GroupTilesByTemporalSlice`
// produces for this request's tiling — which the gate re-derives from that
// function rather than trusting this field.
int64_t video_decode_chunks = 0;

// True only once the `Generate` that produced this conditioning RETURNED. The
// trace is filled immediately after the connector and BEFORE the denoise loop,
// because that is the only point at which the exact buffers cross-attention
Expand Down
34 changes: 31 additions & 3 deletions src/vllm/model_executor/models/ltx2_video_vae_tiled.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -41,6 +41,16 @@

#include "vllm/model_executor/models/ltx2_tiling.h"
#include "vllm/model_executor/models/ltx2_video_vae.h"
// W0 of LTX25-DEVICE-RESIDENCY (#1010). The one include in this file that is not
// about decoding video, and it is here deliberately: `decode.video` is the leaf
// W5's lever is measured against, and every other assertion the gate makes about
// that leaf is a RATIO taken against the leaf itself. A sub-scope opened in the
// DRIVER, two statements from the leaf's own `Open`, moves with the leaf and
// constrains nothing; a sub-scope opened HERE, around the tile decode the render
// actually spends the seconds in, does not. `render_phase_log.h` is a
// process-wide instrument rather than a multimodal model, and it is the first
// thing under `multimodal/` this directory includes.
#include "vllm/multimodal/render_phase_log.h"
#include "vt/dtype.h"

namespace vllm {
Expand Down Expand Up @@ -235,9 +245,27 @@ void Ltx2ConvVideoDecodeTiled(const Ltx2ConvVideoDecoderConfig& config,

ChunkBuffer buffer;
buffer.Allocate(config.out_channels, curr_stop - curr_start, full_h, full_w);
std::vector<float> curr_weights =
AccumulateTemporalGroup(config, weights, latent, latent_channels, latent_t, latent_h,
latent_w, noise, timestep, group, &buffer, complementary);
// W0 (#1010): THE VIDEO DECODE'S OWN WORK, bounded by production events on
// both ends. One record per temporal group, which is one record per chunk
// the sink is handed, so the count is a quantity the instrument cannot move.
// Nested, so the table's sum does not change.
//
// THE PLACEMENT IS GATED BY A COVERAGE FLOOR AND NOT BY THE COUNT ALONE.
// Moving this scope one statement up, onto `buffer.Allocate`, keeps the
// count, the `nested` flag and the containment in a chunk window — and a
// fourth fresh review measured what that costs: `decode.video.vae = 0.000 s`
// beside a five-millisecond `decode.video`, with the tile decode inside no
// sub-scope, both gates green. `test_ltx2_video`'s containment case now
// requires these records to cover at least half of the render's
// `decode.video.chunk` seconds (measured 91.1%-98.6%), so an anchor that
// sits BESIDE the work rather than ON it is a red.
std::vector<float> curr_weights;
{
const ::vllm::multimodal::phase::Scope vae_phase("decode.video.vae");
curr_weights =
AccumulateTemporalGroup(config, weights, latent, latent_channels, latent_t, latent_h,
latent_w, noise, timestep, group, &buffer, complementary);
}

if (have_previous) {
if (previous_stop > curr_start) {
Expand Down
36 changes: 36 additions & 0 deletions src/vllm/multimodal/ltx2_video.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -4571,6 +4571,24 @@ VideoResult Ltx2VideoEngine::Generate(const VideoGenParams& gen) {
// statement. A `decode.video` leaf that closes before its chunk arrives, or
// that is re-labelled after one, no longer contains the chunk it produced.
// Nested, so the sum does not move.
//
// ITS OPEN IS STILL THE STATEMENT NEXT DOOR, and the third fresh review showed
// what that costs: move the two `Close`s below to after the PPM write and the
// leaf AND its anchor grow together, so containment and coverage both hold
// while the whole writer is charged to `decode.video`. Two things close it.
// `decode.video.vae` is opened inside `Ltx2ConvVideoDecodeTiled` around
// `AccumulateTemporalGroup`, so `decode.video` has one sub-scope whose ends
// are both production events; and the gate asserts that nothing but an anchor
// is emitted NESTED, which is what a swallowed `artifacts.frames` becomes.
//
// AND THE VAE SUB-SCOPE IS HELD BY A COVERAGE FLOOR, not by its existence. A
// fourth fresh review moved that scope off `AccumulateTemporalGroup` and onto
// the `buffer.Allocate` beside it — seven lines — and the table reported
// `decode.video.vae = 0.000 s` beside a five-millisecond `decode.video`, with
// the tile decode inside no sub-scope at all and both gates green, because the
// vae anchor was checked for cardinality, `nested` and containment and its
// DURATION was compared against nothing. The gate now requires it to cover at
// least half of this render's `decode.video.chunk` seconds.
size_t chunk_handle =
phase::PhaseLog::Instance().Open("decode.video.chunk", /*span=*/false);
Ltx2VideoDecodeStreaming(
Expand All @@ -4586,6 +4604,18 @@ VideoResult Ltx2VideoEngine::Generate(const VideoGenParams& gen) {
shape.t = chunk.frames.frames;
shape.h = chunk.frames.height;
shape.w = chunk.frames.width;
// W0 repair (#1010, third fresh review): THE WRITER'S OWN ANCHOR, bound
// to the `WriteFileBytes` loop rather than to the scope beside it.
//
// Without it the writer is held by nothing. Empty this leaf and leave
// the `write(2)`s inside `decode.video` — which is what happens if the
// two `Close`s above move below the loop — and the table charges W5's
// lever with the cost of the writes while `artifacts.frames` reports
// four microseconds for nine files. Every interval assertion still
// holds, because the leaves stay disjoint and nothing nests. This scope
// moves with the WRITES, so a leaf that stops covering them stops
// containing it. Nested, so the sum does not move.
phase::Scope ppm_phase("artifacts.frames.ppm");
for (int64_t f = 0; f < chunk.frames.frames; ++f) {
char name[64];
// The GLOBAL frame index, which the chunk carries so the writer does not
Expand All @@ -4599,8 +4629,14 @@ VideoResult Ltx2VideoEngine::Generate(const VideoGenParams& gen) {
WriteFileBytes(JoinPath(gen.output_dir, name),
MiniMaxH3WritePpmFrame(chunk.frames.data, shape, f));
}
ppm_phase.Close();
rendered_frames += chunk.frames.frames;
rendered_channels = chunk.frames.channels;
// COUNTED BY THE RENDER, not by the instrument (#1010, third fresh
// review). W0's containment case needs one number about this decode that
// no phase-scope placement can move; this is it, and it sits beside the
// frame count for the same reason that one does.
im.trace.video_decode_chunks += 1;
write_phase.Close();
decode_handle = phase::PhaseLog::Instance().Open("decode.video", /*span=*/false);
chunk_handle = phase::PhaseLog::Instance().Open("decode.video.chunk", /*span=*/false);
Expand Down
Loading
Loading