Repository navigation
agent: load-only flakes fixed at the cause (two real races); #137 follow-ups - #152
Merged
Merged
Conversation
Tests that failed only under heavy parallel load, each made independent of host timing: - mcp_events_nested (a real race): serve started the MCP Events hub before installing the session's elicitation and sampling gates, so a server asking a question during the very first events/poll was declined as "no client". The hub now starts after the gates. A debug-only seam (BEYOND_AI_AGENT_TEST_SLOW_GATE_INSTALL_MS) holds the window open; the new test fails on the old order. - mcp_events_receiver a_retry_that_beats_the_resubscribe (a real race): a restarted daemon answered 410 (stop) to a delivery for a persisted callback that arrived before its events session had read its state and reserved the token. Unknown tokens are now 503 from daemon start until that session has reserved its tokens (bounded). A debug-only seam (BEYOND_AI_AGENT_TEST_SLOW_EVENTS_RESTORE_MS) holds the window open, and the test no longer sleeps 800 ms. - serve_uds uds_stale_socket_file_is_reclaimed: the "stale" node was a listener bound and dropped in the test process, which a child forked by another test in that instant keeps alive until its exec; a connect then succeeded against it. The node is now a datagram socket, which no stream connect can reach. - mcp_stdio a_panic_unwinding_past_the_exit_guard_still_sweeps: sweep_before_exit's wait is bounded, and a loaded host could stall a sweep thread past it. The exiting thread now kills every server still running itself after the wait. The test reads "killed" from the kernel (gone, a zombie, or SIGKILL pending), not from scheduling. - mcp_events wire many_events_in_one_chunk_drain_in_linear_time: counts the bytes scanned and moved (must stay within 2x the input) instead of timing the drain. - serve_lifecycle a_slow_lifecycle_consumer_does_not_delay_the_prompt_response: the collector holds every POST unanswered until released, and the prompt response must arrive first: proved by order, not by a 1.5 s bound. - serve_health: /readyz callers wait out the one transient answer, "shard probe timed out" (the probe's real answer lands in the memo); any other answer, a 503 for another reason included, is returned at once. - mcp_skills_integrity a_listing_that_never_ends_is_cut_off_and_said_so: the listing stays fresh for the run, so the turn does not page through all 64 again under the refresh's 30 s bound; asserts exactly 64 skills/list. - tool_reactor_stall (#148's probe) failed on CI: a 4 MB write took under a millisecond on a CI runner, and the tool's fixed per-call overhead (~0.26 ms in a debug build) alone was a third of that baseline. The edit and write subjects are now ~36 MB, so the baseline is milliseconds everywhere and the bar (25%) is unchanged. #137 follow-ups: - The journal key's maker re-reads it through the locked descriptor (FileLock::file); with #144's OFD locks another descriptor's close no longer releases the lock, NFS included. - An old-name key (not yet carried over) trashes, restores and hard-deletes with its session. - Why a prompt needs no drain of queued journal writes is documented where the prompt reads the journal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
force-pushed
the
jared/load-flakes-2
branch
from
October 7, 2026 11:48
6e2595a to
f4a5624
Compare
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
#152 fixed the third (the hub started before the session's question gates, now after them); these two time out under load for different reasons, which #152 does not cover: - in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its client attaching within a 1.5 s delay before the fixture raised the request (with no client an elicitation is rightly declined). The fixture now holds it until the test says both clients are attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a get_state on its own session. - a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited for the one-shot `refused` frame, which goes out when the events session starts at boot — under load before the client attached. It now polls mcp_events_list for the recorded refusal (state and last_error). - ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the gate ordering (#152's fix) noted where nested requests are described. All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
… restore window covers runtime-only daemons (#156) Audit of #152: - A runtime subscription restored when its session starts (boot restore, and the restart after a panic) was started from the restore task spawned in `McpEventsHub::attach`, before the session's elicitation and sampling gates existed, so a question on its first poll could still be declined as "no client". One signal now gates them all: `Hub::questions_ready`, set by `McpEventsHub::start` once the gates are installed, and awaited by every subscription task (configured, restored, resubscribing) before its first request. Test: a restored runtime subscription asks on its first poll while the slow-gate seam holds the gates off; the question reaches the client. Fails (declined) without the wait. - The 503-while-restoring window was armed only when configured subscriptions existed, so a daemon whose only subscriptions are runtime ones still answered 410 to early retries. The window now tracks every session the daemon starts to restore callbacks (the events session and each runtime-subscription session) and closes when the last has reserved its tokens (still bounded at 10 min). Test: a runtime-only daemon answers an early retry 503, not 410. - Once restore completes, a token nothing holds is 410 at once, asserted in both restart tests, so a daemon stuck at 503 (restored() dropped) fails them. - release_seams: the serve_ws seams named by the audit already sit in `#[cfg(debug_assertions)]` functions (a release build carries none of the names: checked with grep on the release binary). The checker's hole was elsewhere: any attribute within three lines above counted as gating, so `#[cfg(debug_assertions)] let x = 1;` gated the call site after it. It now requires the attribute directly above the start of the statement or arm the name is in; pinned with that case. The restore seam moved into a gated function of its own. Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
jaredLunde
added a commit
that referenced
this pull request
Oct 7, 2026
#152 fixed the third (the hub started before the session's question gates, now after them); these two time out under load for different reasons, which #152 does not cover: - in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its client attaching within a 1.5 s delay before the fixture raised the request (with no client an elicitation is rightly declined). The fixture now holds it until the test says both clients are attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a get_state on its own session. - a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited for the one-shot `refused` frame, which goes out when the events session starts at boot — under load before the client attached. It now polls mcp_events_list for the recorded refusal (state and last_error). - ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the gate ordering (#152's fix) noted where nested requests are described. All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR fixes every test I saw fail only under heavy parallel load, each at its cause, not by retrying it or lengthening a timeout. Two of the eight turned out to be real product races. It also carries the three follow-ups from #137's verification.
How the flakes were found
beyond-ai-agentsuites, run over and over, plus up to 40 CPU-burner processes (load average 50–75 on 16 cores).tools::outputsweep.mcp_multiple_servers_connect_concurrently_not_sequentially,*_does_not_stall_the_runtime.serve_authlogin,serve_service_isolation::a_failed_probe_means_no_tool_ever_runs_on_the_replica(agent: a tool waits for its checkpoint; a closed session's reason always arrives #150's close-reason fix covers it: 0/96 failures on their branch), and their own two.Flakes fixed
mcp_events_nested::a_nested_elicitation_during_an_events_poll_reaches_the_owning_sessionservestarted the MCP Events hub before installing the session's elicitation and sampling gates. A server that asked during the very firstevents/pollwas declined as "no client" (seen in a debug log:MCP elicitation unanswered; declining err=NoClient).a_nested_elicitation_on_the_first_poll_waits_for_the_session_to_take_questions: a debug-only seam (BEYOND_AI_AGENT_TEST_SLOW_GATE_INSTALL_MS) holds the window open. It fails every time on the old order and passes on the new one.mcp_events_receiver::a_retry_that_beats_the_resubscribe_after_a_restart_is_told_to_retry_not_to_stop410, which tells the server to stop for good, to a retried delivery for a persisted callback. That happened when the delivery arrived before the events session had read its state and reserved the token.503(retry). The window is bounded byRESERVE_FOR.BEYOND_AI_AGENT_TEST_SLOW_EVENTS_RESTORE_MS) holds the state read off, and the test's 800 ms sleep is gone. Without the fix it gets410.serve_uds::uds_stale_socket_file_is_reclaimedexec. A connect in that window succeeds, so the test talked to a socket the daemon had not yet replaced.mcp_stdio::a_panic_unwinding_past_the_exit_guard_still_sweepssweep_before_exit's wait is bounded by grace + 2 s, and a loaded host can stall a sweep thread past it.SigPnd/ShdPnd. When the process is reaped is up to the scheduler.mcp_events::wire::many_events_in_one_chunk_drain_in_linear_timeserve_lifecycle::a_slow_lifecycle_consumer_does_not_delay_the_prompt_responseanswered == 0; afterwards the events are confirmed delivered. Ifemitawaited the network, the read would stall and fail.tool_reactor_stall::write_does_not_stall_the_runtime(#148's probe; failed on CI run 37614530412)write_bytes(Vec<u8>)move. On the CI runner a 4 MBwritetook 0.69–0.87 ms of thread CPU inline. The tool's fixed per-call overhead (stat,create_dir_all, the hand-off: ~0.26 ms in a debug build) was a third of that baseline.editandwritesubjects are now ~36 MB, so the baseline is milliseconds on any machine. Locally: write 0.20 ms of 24 ms, edit 1.3 ms of 37 ms, ls 0.34 ms of 7.8 ms. The 25% bar is unchanged.edit's hold grew from ~0.3 ms to ~1.3 ms with the 8× larger input. That suggests some per-byte work still runs on the executor, still under 4% of the baseline; worth a separate look.serve_health(readyz)readyz()waits out exactly that one answer. Any other answer, a503for any other reason included, is returned at once.mcp_skills_integrity::a_listing_that_never_ends_is_cut_off_and_said_soskills/listcalls.#137 follow-ups
FileLock::file()). With agent: OAuth fixture coverage (DELETE, streamed in-POST question); session lock is a record lock (serve_ws flake root cause) #144's OFD locks (merged since), no other descriptor's open or close can release the lock, NFS included, so the extra registry and test I had first written became moot and were dropped in the rebase. agent: OAuth fixture coverage (DELETE, streamed in-POST question); session lock is a record lock (serve_ws flake root cause) #144's own tests cover the OFD property.resume_into_turnadds the resumer's own placeholder ids. I couldn't build a test for this one. No observable behaviour depends on the drain, and the race can't be forced through serve's surface.a_trashed_session_takes_an_old_name_key_along: gone from the session dir and present in.trashafter delete, back and keyed after restore.Stability
Verification
-D warnings, all targets),cargo fmtanddprint checkare clean.tool_reactor_stall,mcp_events_nested,mcp_events_receiver,serve_lifecycle,serve_health,serve_uds,mcp_skills_integrity,serve_service_isolation,serve_service_session_end.🤖 Generated with Claude Code
https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk