Skip to content

Plant the row a delete removed from the run's own earlier image of it - #252

Merged
maverox merged 2 commits into
mainfrom
work/plant-the-row-a-delete-removed
Sep 25, 2026
Merged

maverox merged 2 commits into
mainfrom
work/plant-the-row-a-delete-removed

Conversation

@maverox

@maverox maverox commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

A recorded generic_delete answers Ok(true): it asserts that a row existed and carries none of it. Since #225 the seeder names that skip RecordedPresence, but naming it does not make the replay pass. The replayed DELETE finds no row, the router answers 404 where the recording answered 200, and the request fails on its response. #251 did not change that, because the response is compared and stays blocking.

This change plants the row.

Mechanism

The delete names its row. Diesel's debug binds are recorded as plain text, for example:

DELETE FROM "business_profile" WHERE (("business_profile"."profile_id" = $1) AND ("business_profile"."merchant_id" = $2)) -- binds: [ProfileId("pro_…"), MerchantId("…")]

Row identity is registered from the replay database's catalog before planning, so the planner can turn that statement into a row key. It then looks for images of that row among every event of the run's scoped recording, in any correlation. If it finds one, it attaches the image to the delete's seed entry, and the seeder plants it into the correlation's own schema exactly as it plants any other typed image.

The borrow applies only when the entry has no image of its own and its recorded result is the same presence the seeder already calls RecordedPresence: an Ok envelope holding true, or a bare true.

The correlation's own image comes first. If the correlation imaged the row itself before deleting it, that image is used, with origin Recording. Its own earlier read plants that image anyway, so borrowing a later one from another correlation would give the plan two entries for one row that disagree, and the certificate would name the one that lost. The borrow applies only to rows the correlation never saw.

A borrowed entry is recording-derived. So the plan keeps it against a later ambient or recorded entry for the same key, exactly as it keeps a recorded one.

The choice rule and why

The image with the greatest global sequence strictly below the delete's.

  • Deterministic. Global sequence is a total order over the tape, so the same recording always chooses the same image.
  • The right answer, not a tiebreak. The latest image before the delete is the state the row was in when the delete ran. It is what the recording's own earlier calls saw, so it is also the value least likely to disturb anything downstream.
  • Why not synthesize from the catalog. business_profile has 13 NOT NULL columns, one of them an enum the catalog does not load, plus CHECK and UNIQUE constraints. A borrowed row already satisfied all of them when it was recorded.

Images recorded after the delete, images from errored events, and deletes that do not name exactly one whole row (a partial composite key, or two rows) borrow nothing. They stay skipped as RecordedPresence, as before.

Provenance

A real row cannot be marked as planted, so the record carries it instead. The entry's origin is a new SeedOrigin::Borrowed { global_sequence }, and the seed certificate writes it as {"borrowed": {"global_sequence": N}}, naming the event whose image was planted. Existing certificates still deserialize.

This differs from the design note, which also carried correlation_id. The sequence alone names the source event uniquely, and leaving out the string keeps SeedOrigin Copy.

Scope

The planner still reads only the run's ScopedRecording, and that seam is not widened here. A row whose only earlier images belong to other runs' scopes stays skipped. Recounted on the cycle-2 ledgers, in the run's scope:

correlation table delete at latest image before
01a0ce1d-f554 merchant_key_store 1667 317
01a0ce20-461b merchant_key_store 13466 11971
01a0ce29-2142 business_profile 108135 108123
01a0ce2f-bbf5 business_profile 118395 118382
01a0ce32-0838 business_profile 76674 none in scope

So this can move between 0 and 4 of the five from FAIL, and 01a0ce32-0838 is expected to stay FAIL. Its row has 8 earlier images on the tape, all in other runs' scopes.

For the four, the row is planted and the replayed DELETE should find it. Each then either passes or diverges downstream on a named, planted row. The count above was taken on recorded result text. That each of those events also carries a typed row image the seeder can plant is what the tests below establish for the shape, not something measured on this tape.

A falsifier for the scope claim: replaying 01a0ce32-0838 in the same run as 01a0ce30-20c7 should produce a borrowed entry from sequence 72143.

Shared bind mapping

The column-to-bind mapping (equality predicates on the key columns, matched to bind positions) moves from the recorder's binds_read_keys into deja_runtime::replay::row_keys_for_binds, so both readers share one implementation.

  • Recorder: keeps its strict JSON bind parser and produces the same keys, so recordings are unchanged. This is checked by a test, not just read off the code; see below.
  • Planner: uses a tolerant reader that unwraps newtype debug renderings (ProfileId("pro_1") becomes "pro_1"). It refuses the whole list if any item is not a scalar, including a Vec bind, whose debug output is a JSON array.

Seed planning runs only in the orchestrator, so the router needs no pin bump for this.

How "recordings unchanged" is established. The move is not a pure relocation. It makes three edits:

  • The identity lookup now runs after the bind parse.
  • The early return on an unbound key column is gone. That column now contributes Null, which db_row_key_for_table refuses.
  • Keys are deduplicated twice.

Each is equivalent on reading, but the recorder's output feeds every recording, so reading is not the evidence. crates/deja/tests/binds_read_keys_unchanged.rs keeps the implementation before the move verbatim as an oracle, and asserts the new binds_read_keys returns identical keys on a grid of 25,191 statements:

  • single and composite keys, and unregistered tables
  • key columns unbound (312 cases), a leading column bound several times, and a key column whose name ends another column's ("id" in "merchant_id")
  • out-of-range and $0 positions
  • bind lists the strict parser accepts or refuses

2,648 of those produce keys, so the comparison is not vacuous. The test also asserts coverage of the unbound-column path, and four mutations of the shared mapping each fail it.

It optionally reads real statements from DEJA_BINDS_CORPUS. Run once over 6,524 distinct recorded statements quoted in our replay notes, it agreed on every one. Only 330 of those have bind lists that parse as strict JSON, though: the recorder refuses newtype binds such as MerchantId("…") before and after this change alike, so for most real statements both sides give no keys.

Verification

  • just verify green (fmt, clippy -D warnings, workspace tests).
  • The runtime crates check on 1.85.0.
  • New tests:
    • in deja-runtime, covering the borrow, the choice rule, each refusal, the tolerant reader and the certificate's serialized form
    • one in the orchestrator, showing a presence entry carrying a borrowed image is planted rather than skipped
  • The planner call site passes the whole scoped event list, and identity is registered in load_db_catalog before the per-correlation planning loop.

Each mutation was run against the whole workspace (--no-fail-fast):

mutation killed by
no borrow at all (red check) the two positive borrow tests
images after the delete eligible latest-before test
earliest image instead of latest latest-before test
errored events lend images errored-event test
any result borrows, not only presence presence-only test
bare true not recognised bare-true test
first of several row keys taken whole-row test
newtypes not unwrapped reader test, plus three borrow tests
non-scalar binds accepted reader test (Vec bind case)
own earlier image ignored own-image test
a later entry overwrites a borrowed one upsert test

Two guards were removed because no mutant of them could be killed:

  • a table filter, because the row key already names its table
  • an early return on a partial key, because db_row_state_key already refuses a partial row

The partial-key test still passes with db_row_state_key as the only guard.

A recorded DELETE answers Ok(true): it asserts a row existed and carries
none of it. The seeder has named that skip RecordedPresence since #225, but
naming it does not make the replay pass. The replayed DELETE finds nothing,
the router answers 404 where the recording answered 200, and the request
fails on its response.

The delete does name its row. Diesel's debug binds are recorded in plain
text, and the table's key is registered from the catalog at replay, so the
planner can derive the row key from the statement. When the same row was
imaged earlier in the run, by this correlation or another, the planner now
attaches the latest such image to the delete's entry, and the seeder plants
it.

The choice rule is the image with the greatest global sequence strictly
below the delete's. Global sequence is a total order over the tape, so the
choice is deterministic, and it is the state the row was in when the delete
ran rather than an arbitrary tiebreak. Images after the delete, images from
errored events, and deletes that do not name exactly one whole row borrow
nothing and stay skipped as before.

A real row cannot be marked as planted, so its provenance is recorded
instead. The entry's origin is Borrowed with the global sequence of the
event whose image was used, and the seed certificate carries it.

The planner reads only the run's scoped recording, and that seam is not
widened for this. A row whose only earlier images belong to other runs'
scopes stays skipped.

The column-to-bind mapping moves from the recorder's binds parser into the
runtime so both readers share it. The recorder keeps its strict JSON parser
and produces the same keys. The planner uses a tolerant reader that unwraps
newtype debug renderings such as ProfileId("pro_1").
…own row

Moving the column-to-bind mapping into the runtime was not a pure
relocation: the identity lookup now runs after the bind parse, the early
return on an unbound key column is gone, and keys are deduplicated twice.
Each is equivalent on reading, but the recorder's output feeds every
recording, so reading is not enough. The implementation before the move is
now kept verbatim in a test as an oracle, and the new one must agree with it
on a grid of every statement shape the mapping distinguishes, including a
key column left unbound, a leading column bound several times, and a key
column whose name ends another column's.

A correlation that already imaged the row before deleting it now plants its
own image rather than a later one from another correlation. Its own earlier
read plants that image anyway, so borrowing a rival would put two entries
for one row in the plan that disagree, with the certificate naming the one
that lost. The borrow applies only to rows the correlation never saw.

A borrowed entry is recording-derived, so the plan now keeps it against a
later ambient or recorded entry for the same key, as it already did for a
recorded one. The run-wide image index is built on the first presence-only
delete rather than for every correlation, and a doc comment the new
function had separated from its own function is back in place.
@maverox
maverox merged commit 4f66118 into main Sep 25, 2026
10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant