perf(streaming): incremental tokenize + link-ref scanning via safe resume boundaries (fixes #30) - #40
Merged
Merged
Conversation
…sume boundaries (#30) The last O(prefix)-per-update scans (issue #21 limitation K): tokenizeBlocks over the full content on every update, and parseLinkReferenceDefinitions over the committed prefix on every commit — measured at 4%/17%/32% of wall-clock at 9/37/74KB. Both now resume from a saved offset; per-update parse work is O(tail) and DOM-path streaming growth is essentially linear across the benched range (74KB total: 4,026ms -> 3,021ms; 37->74KB doubling: 2.40x -> 2.03x). New src/incremental-scan.ts: IncrementalSourceScanner caches tokens and the link-ref map up to a SAFE BOUNDARY — the end of a complete blank token whose nearest preceding non-blank token is not a list_item, indented_code, or blockquote (the three kinds whose single token can extend backward across blank lines; everything else is sealed by a blank at the tokenizer level, and an open fence swallows blanks into its own token so no blank token can exist inside one). Tokenizing the suffix from such a boundary is byte-identical to tokenizing the whole and slicing; link-ref definitions cannot span a blank line, so the map is the cached prefix map merged first-definition-wins (#544) with a suffix scan. Non-append-only snapshots (rewrites, retreating splits) are caught by a startsWith memcmp on the retained safe prefix — same guard family as frozenSource — and reset the cache. Wiring: StreamingMarkdownRenderer owns two scanners (raw content per update, committed prefix per commit); splitForStreamingFrom(content, blocks) lets the split consume scanner output; FrozenTailRenderer.update accepts the threaded link-ref map (computing it itself only when absent, for tests). The stateless string path (renderStreamingMarkdown) is unchanged by design. Equivalence is the oracle: scanner output deep-equals fresh tokenizeBlocks/parseLinkReferenceDefinitions at EVERY prefix of every baseline corpus example, plus a hazard document exercising each blank-crossing kind (list continuation, indented-code merge, blockquote blank-run, multi-line ref def, fenced fake def), rewrite/retreat resets, and a deterministic scannedChars CI guard (re-tokenized chars grow ~2x per input doubling, <3x threshold). Remaining O(prefix)-per-update work is byte comparison only (the append-only guards' memcmps and the includes('|') gate) — hardware-speed; documented in the module headers. Closes #30. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LnAFvAgDY3RVYT6zJqN1fy
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #30 — the last O(prefix)-per-update scans from issue #21 limitation K:
tokenizeBlocksover the full content on every update, andparseLinkReferenceDefinitionsover the committed prefix on every commit. Measured at 4% / 17% / 32% of wall-clock at 9 / 37 / 74 KB.Result: DOM-path streaming growth is essentially linear across the benched range. 74 KB total: 4,026ms → 3,021ms (−25%); the 37→74 KB doubling: 2.40× → 2.03×. Per-update parse work is now O(tail).
The safe resume boundary
New
src/incremental-scan.ts:IncrementalSourceScannercaches tokens and the link-ref map up to a safe boundary — the end of a completeblanktoken whose nearest preceding non-blank token is not alist_item,indented_code, orblockquote. Those are exactly the three kinds whose single token can extend backward across blank lines (indented continuation joins the previous item; indented code merges across blanks; blockquote tokens continue across blank runs — verified against the tokenizer). Everything else is sealed by a blank at the tokenizer level, and an open fence swallows blank lines into its own token, so no blank token can exist inside one.Tokenizing the suffix from such a boundary is byte-identical to tokenizing the whole string and slicing. Link-ref definitions cannot span a blank line, and the parser treats a slice start as a line start, so the definition map is the cached prefix map merged first-definition-wins (#544) with a suffix scan. The safe pointer only moves forward (amortized O(1) maintenance).
Rewrites/retreats (message regeneration, split jitter) are caught by a
startsWithmemcmp on the retained safe prefix — the same guard family asfrozenSourcefrom the #23 review — and reset the cache: correct, just not incremental for that update.Wiring
StreamingMarkdownRendererowns two scanners (raw content per update; committed prefix per commit).splitForStreamingFrom(content, blocks)lets the split consume scanner output;splitForStreamingis unchanged.FrozenTailRenderer.updateaccepts the threaded link-ref map (computes it itself only when absent).renderStreamingMarkdown) is unchanged by design.Verification
tokenizeBlocks/parseLinkReferenceDefinitionsat every prefix of every baseline corpus example, plus a hazard document exercising each blank-crossing kind (list continuation across a blank, indented-code merge, blockquote blank-run, multi-line ref def, fenced fake def), and rewrite/retreat resets.scannedChars(characters actually re-tokenized) grows ~2× per input doubling (< 3× threshold; ~4× = resume regression).What remains O(prefix), knowingly
Byte comparison only: the append-only guards' memcmps and the
includes('|')gate — hardware-speed (~GB/s), documented in the module headers. This closes out the #21 performance thread: committed-DOM work (#23), pending queries (#23 follow-up), list shapes (#39), and parse scans (this PR) are all O(tail) per update.🤖 Generated with Claude Code
https://claude.ai/code/session_01LnAFvAgDY3RVYT6zJqN1fy
Generated by Claude Code