Skip to content

agent: no tool holds the runtime thread; MCP Events start after a session's question gates - #154

Merged
jaredLunde merged 7 commits into
mainfrom
jared/runtime-blocking
Oct 7, 2026
Merged

jaredLunde merged 7 commits into
mainfrom
jared/runtime-blocking

Conversation

@jaredLunde

@jaredLunde jaredLunde commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

What

Code Mode no longer holds the async runtime thread, and a second stall probe catches tools that block the thread without using CPU; the tools it caught are fixed. Three events tests that timed out under load are also fixed, and so is the product race behind one of them.

Each fix has a test that fails without it. For every fix except the events-gate ordering, I checked this by reverting the fix in place and running its test. That race shows up only under load, so its evidence is a load comparison instead (see its row).

# problem fix proving test
1 Code Mode ran QuickJS's synchronous evaluation on the async executor. while (true) {} held the runtime thread until the 30 s interrupt deadline, freezing every other task on that worker: other sessions, progress, aborts. Evaluation runs on the blocking pool under a local executor. The JS slot permit travels with the evaluation and is released when the evaluation actually ends. A drop guard trips the interrupt however execute ends: returned, cancelled, or its future dropped by an abort. Nested host calls are spawned back onto the session's runtime and raced against cancellation. tool_reactor_stall::code_mode_keeps_the_runtime_responsive_while_a_script_spins (both probes). code_mode_abort_interrupts_a_spinning_script_promptly: by token and by drop, and the next execute runs at once.
2 The stall probe only counted CPU, so a tool that blocks without CPU passed. A second probe: a ticker on the same current_thread runtime. Each gap between ticks has the thread's run-queue wait (/proc/thread-self/schedstat) taken out, so a sleep, blocking I/O or a contended std mutex is caught, and host load is not. The module doc now states exactly what each probe checks. The probe is run across every tool that needs no external service: read, write, edit, ls, grep, find, bash, todo, memory, web, and Code Mode. the_responsiveness_probe_catches_a_tool_that_blocks_without_cpu: the CPU probe passes a tool that yields once and then sleeps 150 ms; the new probe fails it.
2a bash: the stale-spill sweep listed and stat'ed the whole system temp dir on the runtime thread (about 10 ms for this host's 7.5k entries, more over NFS). The sweep runs on the blocking pool. output::the_stale_sweep_runs_off_the_runtime_thread (deterministic: it records which thread swept)
2b memory: merging, rendering and freeing a large search hit list happened on the runtime thread. That work runs on the blocking pool, and rendering stops at the listing cap. memory_keeps_the_runtime_responsive
2c web: the first call built the reqwest client inline (TLS provider and root certs, about 50 ms of CPU). Every parsing call also spawned its isolated parser child and waited on it inline, for up to the parse timeout. Both run on the blocking pool. Web::with_parser_binary lets an in-process test point the parse at the agent binary. web_keeps_the_runtime_responsive: the first call and a warm call
2d edit: per-byte work on the runtime thread. The #148 probe's hold grew from about 0.3 ms to about 1.3 ms with an 8× input. Two passes over every byte were the cause: UTF-8 validation of the read, and FsBackend::write_if_unchanged copying the whole new contents (bytes.to_vec()) in the spawned write task. Validation runs on the blocking pool together with the match. write_if_unchanged now takes the buffer owned, like write_bytes, and LocalFs moves it straight to the blocking pool. edit_does_no_per_byte_work_on_the_runtime_thread measures all of the runtime thread's CPU during an edit, including the task it spawns: 0.40 ms on a 3 MB file vs 0.48 ms on a 30 MB one. Putting back the validation fails it (0.58 vs 1.73 ms), as does the copy (1.14 vs 6.79 ms).
3 Events gate race (the hub started before the session's question gates) Fixed in #152, which is now merged. This PR's duplicate of that fix was dropped in the rebase. —
4 Daemon nested test relied on its client attaching within a 1.5 s delay. With no client attached, an elicitation is correctly declined. The fixture now holds the request until the test signals that both clients are attached (MCP_FIXTURE_NESTED_ON_SIGNAL, POST /control/raise_nested), after each client has answered a get_state. in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session
5 Refused-subscription test waited for the one-shot refused frame, which is sent at boot, before a late client attaches. The test polls mcp_events_list for the recorded refusal. a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session

Rebased onto #152. #152 fixed item 3, so this PR no longer changes serve.rs. Items 4 and 5 had other causes, fixed here: a client attaching later than the fixture's fixed delay, and a one-shot frame sent before the client attaches.

Load proof for items 3–5, rerun on top of #152. I ran 12 concurrent copies of each of the three tests next to 16 CPU-burning processes, for three rounds: 108 of 108 passed. Before the fixes, the same load produced timeouts in one or two of every 36 runs.

cargo test vs nextest. The stall tests take a process-wide lock for their whole run. Under cargo test, one test's large allocations make another test's runtime thread wait in the kernel on the shared address space, which would read as a stall. nextest, which CI uses, already isolates each test in its own process.

Not changed: moving bash's posix_spawn off the runtime thread. The spawn's vfork costs about 2 ms, below what the probe can attribute, so no test catches it and I left it as it was.

Audit fixes (at c795815)

Each fix's test fails with that fix reverted in place: 6 mutations, all fail.

# finding fix proving test
A1 Security. If building the hardened web client panicked (new in this PR, via the JoinError) or returned an error (pre-existing), client() fell back to reqwest::Client::new(). That client has no SSRF resolver, no timeout, and follows redirects itself. The tool now fails closed: the call returns an error and nothing is sent. The builder is a field, so a test can inject a failing one. A grep found one more fallback: the MCP Events direct-HTTP client's .build().unwrap_or_default(), whose default client would follow redirects past the checked URL. It fails closed too. (exec_endpoint's plain client is deliberate for an operator-set URL; it is not a fallback.) web::a_client_that_cannot_be_built_fails_the_call_and_sends_nothing (an injected error and an injected panic; the listener sees no connection), mcp_events::an_events_http_client_that_cannot_be_built_fails_closed
A2/A3 A nested host call wasn't aborted when the run was cancelled, or when the execute future was dropped without its token being cancelled. In the drop case the JS slot also stayed held until the call finished. Nested calls are registered with execute's drop guard, which aborts every one still in flight however the run ends. The guard is now the one place this happens; the bridge's own abort on cancel was redundant and is gone. code_mode::cancelling_a_run_aborts_its_nested_call, dropping_execute_aborts_its_nested_call_and_frees_the_slot (no side effect after the cancel or drop; the next execute runs at once)
A4 The script's deadline wasn't enforced while awaiting a nested call (pre-existing). The host call runs under the deadline. Past it, the call is dropped and the run times out. code_mode::the_deadline_holds_across_a_nested_call
A5 Off Linux, the wall-clock probe silently fell back to plain wall clock and could flake. The probe now reports "unavailable", and its tests skip with a message saying why. the_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clock

Checks

🤖 Generated with Claude Code

https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

@jaredLunde
jaredLunde force-pushed the jared/runtime-blocking branch from 77b0e69 to c795815 Compare October 7, 2026 12:48
jaredLunde added a commit that referenced this pull request Oct 7, 2026
… calls aborted and under the deadline

1. SECURITY: web's client() fell back to reqwest::Client::new() — no SSRF resolver, no timeout,
   default redirect-following — when building the hardened client panicked (the spawn_blocking
   JoinError) or erred (pre-existing). It now fails closed: the call errors and nothing is sent.
   The builder is a field, so a test injects a failing (and a panicking) one. The grep found one
   more fallback: the MCP Events direct-HTTP client's `.build().unwrap_or_default()` (a default
   client follows redirects past the checked URL) — it fails closed too.
2./3. Code Mode's nested host calls are registered with execute's drop guard, which aborts every one
   still in flight however the run ends — cancelled through its token or its future dropped — so
   no host tool runs on (and leaves a side effect) after it, and the JS slot is freed at once. (The
   bridge's own abort on cancel is gone: the guard is the one place.)
4. The script's deadline holds across a nested await: the host call runs under it, and past it is
   dropped and the run times out.
5. tool_reactor_stall's responsiveness probe is skipped, saying so, where /proc/thread-self/schedstat
   is unavailable, rather than silently measuring plain wall clock.

Tests, each failing with its fix reverted (6 mutations):
a_client_that_cannot_be_built_fails_the_call_and_sends_nothing (error and panic),
an_events_http_client_that_cannot_be_built_fails_closed, cancelling_a_run_aborts_its_nested_call,
dropping_execute_aborts_its_nested_call_and_frees_the_slot, the_deadline_holds_across_a_nested_call,
the_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde and others added 6 commits October 7, 2026 06:49
… a wall-clock probe, and the tools it caught

1. Code Mode ran QuickJS's synchronous evaluation on the async executor: `while (true) {}` held
   the runtime thread until the 30 s interrupt deadline, freezing everything else on that worker.
   Evaluation now runs on the blocking pool under a local executor; the JS slot permit travels
   with it (released when it really ends); a drop guard trips the interrupt however `execute`
   ends (cancelled, or its future dropped by an abort); host calls are spawned back onto the
   session's runtime and raced against cancellation.
2. tool_reactor_stall's CPU probe could not see a tool that blocks without CPU. A second probe — a
   ticker on the same current_thread runtime, with run-queue wait (schedstat) subtracted so host
   load cannot fake a stall — catches wall-clock blocking too; the module doc says exactly what each
   probe checks. Run across every tool that needs no external service, it caught:
   - bash: the stale-spill sweep listed and stat'ed the whole system temp dir on the runtime thread
     (~10 ms for 7.5k entries here) — now on the blocking pool;
   - memory: rendering (and freeing) a large search hit list on the runtime thread — now merged,
     rendered and dropped on the blocking pool, and rendering stops at the listing cap;
   - web: the first call built the reqwest client (TLS provider, root certs: ~50 ms of CPU) inline,
     and every parsing call spawned and waited on its isolated parser child (up to its timeout)
     inline — both now on the blocking pool (Web::with_parser_binary lets an in-process test point
     the parse at the agent binary).

Tests, each failing with its fix reverted: code_mode_keeps_the_runtime_responsive_while_a_script_spins,
code_mode_abort_interrupts_a_spinning_script_promptly (token and drop), the
*_keeps_the_runtime_responsive probes, the_stale_sweep_runs_off_the_runtime_thread;
the_responsiveness_probe_catches_a_tool_that_blocks_without_cpu keeps the new probe's teeth.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
#152 fixed the third (the hub started before the session's question gates, now after them); these
two time out under load for different reasons, which #152 does not cover:

- in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its
  client attaching within a 1.5 s delay before the fixture raised the request (with no client an
  elicitation is rightly declined). The fixture now holds it until the test says both clients are
  attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a
  get_state on its own session.
- a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited
  for the one-shot `refused` frame, which goes out when the events session starts at boot — under
  load before the client attached. It now polls mcp_events_list for the recorded refusal (state
  and last_error).
- ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the
  gate ordering (#152's fix) noted where nested requests are described.

All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
… spin is calibrated

A tool that works or blocks inline stalls the runtime on every run; the host's own one-off stalls
of the runtime thread (a page fault served from disk on a busy machine — one 45 ms bash reading in
a full sweep) do not repeat. Each responsiveness probe now asserts on the least of up to three
runs, per-run setup outside the measurement. Every fix's revert still fails its probe. The Code
Mode spin doubles its iteration count until a run takes 200 ms, whatever this build's QuickJS speed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
The #148 probe's edit hold grew with the input (~0.3 ms → ~1.3 ms at 8x). Two passes over every
byte ran on the runtime thread: validating the read as UTF-8 (now done on the blocking pool with
the match), and FsBackend::write_if_unchanged copying the whole new contents (bytes.to_vec()) in the
spawned write task before handing them to its blocking writer. write_if_unchanged now takes the
buffer owned, like write_bytes; LocalFs moves it straight to the blocking pool.

Test: edit_does_no_per_byte_work_on_the_runtime_thread measures all of the runtime thread's CPU
during an edit (its polls and the task it spawns) on a ~3 MB and a ~30 MB file: 0.40 vs 0.48 ms.
Either per-byte pass put back fails it (UTF-8: 0.58 vs 1.73 ms; the copy: 1.14 vs 6.79 ms).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
… calls aborted and under the deadline

1. SECURITY: web's client() fell back to reqwest::Client::new() — no SSRF resolver, no timeout,
   default redirect-following — when building the hardened client panicked (the spawn_blocking
   JoinError) or erred (pre-existing). It now fails closed: the call errors and nothing is sent.
   The builder is a field, so a test injects a failing (and a panicking) one. The grep found one
   more fallback: the MCP Events direct-HTTP client's `.build().unwrap_or_default()` (a default
   client follows redirects past the checked URL) — it fails closed too.
2./3. Code Mode's nested host calls are registered with execute's drop guard, which aborts every one
   still in flight however the run ends — cancelled through its token or its future dropped — so
   no host tool runs on (and leaves a side effect) after it, and the JS slot is freed at once. (The
   bridge's own abort on cancel is gone: the guard is the one place.)
4. The script's deadline holds across a nested await: the host call runs under it, and past it is
   dropped and the run times out.
5. tool_reactor_stall's responsiveness probe is skipped, saying so, where /proc/thread-self/schedstat
   is unavailable, rather than silently measuring plain wall clock.

Tests, each failing with its fix reverted (6 mutations):
a_client_that_cannot_be_built_fails_the_call_and_sends_nothing (error and panic),
an_events_http_client_that_cannot_be_built_fails_closed, cancelling_a_run_aborts_its_nested_call,
dropping_execute_aborts_its_nested_call_and_frees_the_slot, the_deadline_holds_across_a_nested_call,
the_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clock.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
…pport::ports (#155)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde
jaredLunde force-pushed the jared/runtime-blocking branch from f19443a to b6b208b Compare October 7, 2026 14:01
@jaredLunde
jaredLunde enabled auto-merge (squash) October 7, 2026 14:43
@jaredLunde
jaredLunde merged commit b87a29c into main Oct 7, 2026
21 checks passed
@jaredLunde
jaredLunde deleted the jared/runtime-blocking branch October 7, 2026 15:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant