Repository navigation
agent: no tool holds the runtime thread; MCP Events start after a session's question gates - #154
Merged
Merged
Conversation
jaredLunde
force-pushed
the
jared/runtime-blocking
branch
from
October 7, 2026 12:48
77b0e69 to
c795815
Compare
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
… calls aborted and under the deadline 1. SECURITY: web's client() fell back to reqwest::Client::new() — no SSRF resolver, no timeout, default redirect-following — when building the hardened client panicked (the spawn_blocking JoinError) or erred (pre-existing). It now fails closed: the call errors and nothing is sent. The builder is a field, so a test injects a failing (and a panicking) one. The grep found one more fallback: the MCP Events direct-HTTP client's `.build().unwrap_or_default()` (a default client follows redirects past the checked URL) — it fails closed too. 2./3. Code Mode's nested host calls are registered with execute's drop guard, which aborts every one still in flight however the run ends — cancelled through its token or its future dropped — so no host tool runs on (and leaves a side effect) after it, and the JS slot is freed at once. (The bridge's own abort on cancel is gone: the guard is the one place.) 4. The script's deadline holds across a nested await: the host call runs under it, and past it is dropped and the run times out. 5. tool_reactor_stall's responsiveness probe is skipped, saying so, where /proc/thread-self/schedstat is unavailable, rather than silently measuring plain wall clock. Tests, each failing with its fix reverted (6 mutations): a_client_that_cannot_be_built_fails_the_call_and_sends_nothing (error and panic), an_events_http_client_that_cannot_be_built_fails_closed, cancelling_a_run_aborts_its_nested_call, dropping_execute_aborts_its_nested_call_and_frees_the_slot, the_deadline_holds_across_a_nested_call, the_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
… a wall-clock probe, and the tools it caught
1. Code Mode ran QuickJS's synchronous evaluation on the async executor: `while (true) {}` held
the runtime thread until the 30 s interrupt deadline, freezing everything else on that worker.
Evaluation now runs on the blocking pool under a local executor; the JS slot permit travels
with it (released when it really ends); a drop guard trips the interrupt however `execute`
ends (cancelled, or its future dropped by an abort); host calls are spawned back onto the
session's runtime and raced against cancellation.
2. tool_reactor_stall's CPU probe could not see a tool that blocks without CPU. A second probe — a
ticker on the same current_thread runtime, with run-queue wait (schedstat) subtracted so host
load cannot fake a stall — catches wall-clock blocking too; the module doc says exactly what each
probe checks. Run across every tool that needs no external service, it caught:
- bash: the stale-spill sweep listed and stat'ed the whole system temp dir on the runtime thread
(~10 ms for 7.5k entries here) — now on the blocking pool;
- memory: rendering (and freeing) a large search hit list on the runtime thread — now merged,
rendered and dropped on the blocking pool, and rendering stops at the listing cap;
- web: the first call built the reqwest client (TLS provider, root certs: ~50 ms of CPU) inline,
and every parsing call spawned and waited on its isolated parser child (up to its timeout)
inline — both now on the blocking pool (Web::with_parser_binary lets an in-process test point
the parse at the agent binary).
Tests, each failing with its fix reverted: code_mode_keeps_the_runtime_responsive_while_a_script_spins,
code_mode_abort_interrupts_a_spinning_script_promptly (token and drop), the
*_keeps_the_runtime_responsive probes, the_stale_sweep_runs_off_the_runtime_thread;
the_responsiveness_probe_catches_a_tool_that_blocks_without_cpu keeps the new probe's teeth.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
#152 fixed the third (the hub started before the session's question gates, now after them); these two time out under load for different reasons, which #152 does not cover: - in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its client attaching within a 1.5 s delay before the fixture raised the request (with no client an elicitation is rightly declined). The fixture now holds it until the test says both clients are attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a get_state on its own session. - a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited for the one-shot `refused` frame, which goes out when the events session starts at boot — under load before the client attached. It now polls mcp_events_list for the recorded refusal (state and last_error). - ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the gate ordering (#152's fix) noted where nested requests are described. All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
… spin is calibrated A tool that works or blocks inline stalls the runtime on every run; the host's own one-off stalls of the runtime thread (a page fault served from disk on a busy machine — one 45 ms bash reading in a full sweep) do not repeat. Each responsiveness probe now asserts on the least of up to three runs, per-run setup outside the measurement. Every fix's revert still fails its probe. The Code Mode spin doubles its iteration count until a run takes 200 ms, whatever this build's QuickJS speed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
The #148 probe's edit hold grew with the input (~0.3 ms → ~1.3 ms at 8x). Two passes over every byte ran on the runtime thread: validating the read as UTF-8 (now done on the blocking pool with the match), and FsBackend::write_if_unchanged copying the whole new contents (bytes.to_vec()) in the spawned write task before handing them to its blocking writer. write_if_unchanged now takes the buffer owned, like write_bytes; LocalFs moves it straight to the blocking pool. Test: edit_does_no_per_byte_work_on_the_runtime_thread measures all of the runtime thread's CPU during an edit (its polls and the task it spawns) on a ~3 MB and a ~30 MB file: 0.40 vs 0.48 ms. Either per-byte pass put back fails it (UTF-8: 0.58 vs 1.73 ms; the copy: 1.14 vs 6.79 ms). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
… calls aborted and under the deadline 1. SECURITY: web's client() fell back to reqwest::Client::new() — no SSRF resolver, no timeout, default redirect-following — when building the hardened client panicked (the spawn_blocking JoinError) or erred (pre-existing). It now fails closed: the call errors and nothing is sent. The builder is a field, so a test injects a failing (and a panicking) one. The grep found one more fallback: the MCP Events direct-HTTP client's `.build().unwrap_or_default()` (a default client follows redirects past the checked URL) — it fails closed too. 2./3. Code Mode's nested host calls are registered with execute's drop guard, which aborts every one still in flight however the run ends — cancelled through its token or its future dropped — so no host tool runs on (and leaves a side effect) after it, and the JS slot is freed at once. (The bridge's own abort on cancel is gone: the guard is the one place.) 4. The script's deadline holds across a nested await: the host call runs under it, and past it is dropped and the run times out. 5. tool_reactor_stall's responsiveness probe is skipped, saying so, where /proc/thread-self/schedstat is unavailable, rather than silently measuring plain wall clock. Tests, each failing with its fix reverted (6 mutations): a_client_that_cannot_be_built_fails_the_call_and_sends_nothing (error and panic), an_events_http_client_that_cannot_be_built_fails_closed, cancelling_a_run_aborts_its_nested_call, dropping_execute_aborts_its_nested_call_and_frees_the_slot, the_deadline_holds_across_a_nested_call, the_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clock. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
…pport::ports (#155) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
force-pushed
the
jared/runtime-blocking
branch
from
October 7, 2026 14:01
f19443a to
b6b208b
Compare
jaredLunde
enabled auto-merge (squash)
October 7, 2026 14:43
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Code Mode no longer holds the async runtime thread, and a second stall probe catches tools that block the thread without using CPU; the tools it caught are fixed. Three events tests that timed out under load are also fixed, and so is the product race behind one of them.
Each fix has a test that fails without it. For every fix except the events-gate ordering, I checked this by reverting the fix in place and running its test. That race shows up only under load, so its evidence is a load comparison instead (see its row).
while (true) {}held the runtime thread until the 30 s interrupt deadline, freezing every other task on that worker: other sessions, progress, aborts.executeends: returned, cancelled, or its future dropped by an abort. Nested host calls are spawned back onto the session's runtime and raced against cancellation.tool_reactor_stall::code_mode_keeps_the_runtime_responsive_while_a_script_spins(both probes).code_mode_abort_interrupts_a_spinning_script_promptly: by token and by drop, and the nextexecuteruns at once.current_threadruntime. Each gap between ticks has the thread's run-queue wait (/proc/thread-self/schedstat) taken out, so a sleep, blocking I/O or a contendedstdmutex is caught, and host load is not. The module doc now states exactly what each probe checks. The probe is run across every tool that needs no external service: read, write, edit, ls, grep, find, bash, todo, memory, web, and Code Mode.the_responsiveness_probe_catches_a_tool_that_blocks_without_cpu: the CPU probe passes a tool that yields once and then sleeps 150 ms; the new probe fails it.output::the_stale_sweep_runs_off_the_runtime_thread(deterministic: it records which thread swept)memory_keeps_the_runtime_responsiveWeb::with_parser_binarylets an in-process test point the parse at the agent binary.web_keeps_the_runtime_responsive: the first call and a warm callFsBackend::write_if_unchangedcopying the whole new contents (bytes.to_vec()) in the spawned write task.write_if_unchangednow takes the buffer owned, likewrite_bytes, andLocalFsmoves it straight to the blocking pool.edit_does_no_per_byte_work_on_the_runtime_threadmeasures all of the runtime thread's CPU during an edit, including the task it spawns: 0.40 ms on a 3 MB file vs 0.48 ms on a 30 MB one. Putting back the validation fails it (0.58 vs 1.73 ms), as does the copy (1.14 vs 6.79 ms).MCP_FIXTURE_NESTED_ON_SIGNAL,POST /control/raise_nested), after each client has answered aget_state.in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_sessionrefusedframe, which is sent at boot, before a late client attaches.mcp_events_listfor the recorded refusal.a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_sessionRebased onto #152. #152 fixed item 3, so this PR no longer changes
serve.rs. Items 4 and 5 had other causes, fixed here: a client attaching later than the fixture's fixed delay, and a one-shot frame sent before the client attaches.Load proof for items 3–5, rerun on top of #152. I ran 12 concurrent copies of each of the three tests next to 16 CPU-burning processes, for three rounds: 108 of 108 passed. Before the fixes, the same load produced timeouts in one or two of every 36 runs.
cargo testvsnextest. The stall tests take a process-wide lock for their whole run. Undercargo test, one test's large allocations make another test's runtime thread wait in the kernel on the shared address space, which would read as a stall.nextest, which CI uses, already isolates each test in its own process.Not changed: moving
bash'sposix_spawnoff the runtime thread. The spawn's vfork costs about 2 ms, below what the probe can attribute, so no test catches it and I left it as it was.Audit fixes (at c795815)
Each fix's test fails with that fix reverted in place: 6 mutations, all fail.
webclient panicked (new in this PR, via theJoinError) or returned an error (pre-existing),client()fell back toreqwest::Client::new(). That client has no SSRF resolver, no timeout, and follows redirects itself..build().unwrap_or_default(), whose default client would follow redirects past the checked URL. It fails closed too. (exec_endpoint's plain client is deliberate for an operator-set URL; it is not a fallback.)web::a_client_that_cannot_be_built_fails_the_call_and_sends_nothing(an injected error and an injected panic; the listener sees no connection),mcp_events::an_events_http_client_that_cannot_be_built_fails_closedexecutefuture was dropped without its token being cancelled. In the drop case the JS slot also stayed held until the call finished.execute's drop guard, which aborts every one still in flight however the run ends. The guard is now the one place this happens; the bridge's own abort on cancel was redundant and is gone.code_mode::cancelling_a_run_aborts_its_nested_call,dropping_execute_aborts_its_nested_call_and_frees_the_slot(no side effect after the cancel or drop; the nextexecuteruns at once)code_mode::the_deadline_holds_across_a_nested_callthe_responsiveness_probe_is_unavailable_without_schedstat_not_wall_clockChecks
code_mode_eval,code_mode_density,tool_reactor_stall).mcp_*suite, the serve suites,serve_harness_deadlines,tool_reactor_stall, and theagentandagent-corelib tests.cargo clippywith-D warningsis clean, both for the agent with code-mode and for the whole workspace. fmt and dprint are clean.check.pypasses 12/12 over HTTP and 12/12 over stdio.🤖 Generated with Claude Code
https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk