Skip to content

perf(streaming): incremental tokenize + link-ref scanning via safe resume boundaries (fixes #30) - #40

Merged
jonathanKingston merged 1 commit into
mainfrom
claude/issue-21-ccvlpo
Jul 5, 2026
Merged

jonathanKingston merged 1 commit into
mainfrom
claude/issue-21-ccvlpo

Conversation

@jonathanKingston

Copy link
Copy Markdown
Collaborator

Fixes #30 — the last O(prefix)-per-update scans from issue #21 limitation K: tokenizeBlocks over the full content on every update, and parseLinkReferenceDefinitions over the committed prefix on every commit. Measured at 4% / 17% / 32% of wall-clock at 9 / 37 / 74 KB.

Result: DOM-path streaming growth is essentially linear across the benched range. 74 KB total: 4,026ms → 3,021ms (−25%); the 37→74 KB doubling: 2.40× → 2.03×. Per-update parse work is now O(tail).

The safe resume boundary

New src/incremental-scan.ts: IncrementalSourceScanner caches tokens and the link-ref map up to a safe boundary — the end of a complete blank token whose nearest preceding non-blank token is not a list_item, indented_code, or blockquote. Those are exactly the three kinds whose single token can extend backward across blank lines (indented continuation joins the previous item; indented code merges across blanks; blockquote tokens continue across blank runs — verified against the tokenizer). Everything else is sealed by a blank at the tokenizer level, and an open fence swallows blank lines into its own token, so no blank token can exist inside one.

Tokenizing the suffix from such a boundary is byte-identical to tokenizing the whole string and slicing. Link-ref definitions cannot span a blank line, and the parser treats a slice start as a line start, so the definition map is the cached prefix map merged first-definition-wins (#544) with a suffix scan. The safe pointer only moves forward (amortized O(1) maintenance).

Rewrites/retreats (message regeneration, split jitter) are caught by a startsWith memcmp on the retained safe prefix — the same guard family as frozenSource from the #23 review — and reset the cache: correct, just not incremental for that update.

Wiring

  • StreamingMarkdownRenderer owns two scanners (raw content per update; committed prefix per commit).
  • splitForStreamingFrom(content, blocks) lets the split consume scanner output; splitForStreaming is unchanged.
  • FrozenTailRenderer.update accepts the threaded link-ref map (computes it itself only when absent).
  • The stateless string path (renderStreamingMarkdown) is unchanged by design.

Verification

  • Equivalence is the oracle: scanner output deep-equals fresh tokenizeBlocks / parseLinkReferenceDefinitions at every prefix of every baseline corpus example, plus a hazard document exercising each blank-crossing kind (list continuation across a blank, indented-code merge, blockquote blank-run, multi-line ref def, fenced fake def), and rewrite/retreat resets.
  • Deterministic CI guard: scannedChars (characters actually re-tokenized) grows ~2× per input doubling (< 3× threshold; ~4× = resume regression).
  • 476 tests + CommonMark conformance + exhaustive convergence fuzz green.

What remains O(prefix), knowingly

Byte comparison only: the append-only guards' memcmps and the includes('|') gate — hardware-speed (~GB/s), documented in the module headers. This closes out the #21 performance thread: committed-DOM work (#23), pending queries (#23 follow-up), list shapes (#39), and parse scans (this PR) are all O(tail) per update.

🤖 Generated with Claude Code

https://claude.ai/code/session_01LnAFvAgDY3RVYT6zJqN1fy


Generated by Claude Code

…sume boundaries (#30)

The last O(prefix)-per-update scans (issue #21 limitation K): tokenizeBlocks
over the full content on every update, and parseLinkReferenceDefinitions over
the committed prefix on every commit — measured at 4%/17%/32% of wall-clock at
9/37/74KB. Both now resume from a saved offset; per-update parse work is
O(tail) and DOM-path streaming growth is essentially linear across the benched
range (74KB total: 4,026ms -> 3,021ms; 37->74KB doubling: 2.40x -> 2.03x).

New src/incremental-scan.ts: IncrementalSourceScanner caches tokens and the
link-ref map up to a SAFE BOUNDARY — the end of a complete blank token whose
nearest preceding non-blank token is not a list_item, indented_code, or
blockquote (the three kinds whose single token can extend backward across
blank lines; everything else is sealed by a blank at the tokenizer level, and
an open fence swallows blanks into its own token so no blank token can exist
inside one). Tokenizing the suffix from such a boundary is byte-identical to
tokenizing the whole and slicing; link-ref definitions cannot span a blank
line, so the map is the cached prefix map merged first-definition-wins (#544)
with a suffix scan. Non-append-only snapshots (rewrites, retreating splits)
are caught by a startsWith memcmp on the retained safe prefix — same guard
family as frozenSource — and reset the cache.

Wiring: StreamingMarkdownRenderer owns two scanners (raw content per update,
committed prefix per commit); splitForStreamingFrom(content, blocks) lets the
split consume scanner output; FrozenTailRenderer.update accepts the threaded
link-ref map (computing it itself only when absent, for tests). The stateless
string path (renderStreamingMarkdown) is unchanged by design.

Equivalence is the oracle: scanner output deep-equals fresh
tokenizeBlocks/parseLinkReferenceDefinitions at EVERY prefix of every baseline
corpus example, plus a hazard document exercising each blank-crossing kind
(list continuation, indented-code merge, blockquote blank-run, multi-line ref
def, fenced fake def), rewrite/retreat resets, and a deterministic scannedChars
CI guard (re-tokenized chars grow ~2x per input doubling, <3x threshold).

Remaining O(prefix)-per-update work is byte comparison only (the append-only
guards' memcmps and the includes('|') gate) — hardware-speed; documented in
the module headers.

Closes #30.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LnAFvAgDY3RVYT6zJqN1fy
@jonathanKingston
jonathanKingston merged commit b0791d0 into main Jul 5, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Incremental tokenizer + link-ref scan to reach true O(n) streaming (limitation K)

2 participants