Skip to content

agent: load-only flakes fixed at the cause (two real races); #137 follow-ups - #152

Merged
jaredLunde merged 1 commit into
mainfrom
jared/load-flakes-2
Oct 7, 2026
Merged

jaredLunde merged 1 commit into
mainfrom
jared/load-flakes-2

Conversation

@jaredLunde

@jaredLunde jaredLunde commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

This PR fixes every test I saw fail only under heavy parallel load, each at its cause, not by retrying it or lengthening a timeout. Two of the eight turned out to be real product races. It also carries the three follow-ups from #137's verification.

How the flakes were found

Flakes fixed

Test Cause Fix Proof it fails without the fix
mcp_events_nested::a_nested_elicitation_during_an_events_poll_reaches_the_owning_session A real race. serve started the MCP Events hub before installing the session's elicitation and sampling gates. A server that asked during the very first events/poll was declined as "no client" (seen in a debug log: MCP elicitation unanswered; declining err=NoClient). The hub starts after the gates are installed. New test a_nested_elicitation_on_the_first_poll_waits_for_the_session_to_take_questions: a debug-only seam (BEYOND_AI_AGENT_TEST_SLOW_GATE_INSTALL_MS) holds the window open. It fails every time on the old order and passes on the new one.
mcp_events_receiver::a_retry_that_beats_the_resubscribe_after_a_restart_is_told_to_retry_not_to_stop A real race. A restarted daemon answered 410, which tells the server to stop for good, to a retried delivery for a persisted callback. That happened when the delivery arrived before the events session had read its state and reserved the token. From daemon start until the events session has reserved its persisted tokens, every unknown token is answered 503 (retry). The window is bounded by RESERVE_FOR. A debug-only seam (BEYOND_AI_AGENT_TEST_SLOW_EVENTS_RESTORE_MS) holds the state read off, and the test's 800 ms sleep is gone. Without the fix it gets 410.
serve_uds::uds_stale_socket_file_is_reclaimed The "stale" socket was a listener bound and then dropped inside the multi-threaded test process. A child forked by another test at that instant keeps the descriptor until its exec. A connect in that window succeeds, so the test talked to a socket the daemon had not yet replaced. The stale node is bound as a datagram socket. No stream connect can reach it. Structural: nothing can make a stream connect to a datagram node succeed.
mcp_stdio::a_panic_unwinding_past_the_exit_guard_still_sweeps sweep_before_exit's wait is bounded by grace + 2 s, and a loaded host can stall a sweep thread past it. After the wait, the exiting thread itself kills every server still running: by pid only if not reaped, by group otherwise. The test now reads "killed" from the kernel: gone, a zombie, or SIGKILL pending in SigPnd/ShdPnd. When the process is reaped is up to the scheduler. Structural.
mcp_events::wire::many_events_in_one_chunk_drain_in_linear_time The test bounded wall-clock time at 3 s. It now counts bytes instead (test-only counter): bytes scanned for separators plus bytes moved by compaction must stay within 2× the input. A per-event drain or a rescan from the event start makes the count quadratic. By construction.
serve_lifecycle::a_slow_lifecycle_consumer_does_not_delay_the_prompt_response The test bounded the prompt response at 1.5 s against a 2 s slow collector. The collector now holds every POST unanswered until the test releases it. The prompt response must arrive while answered == 0; afterwards the events are confirmed delivered. If emit awaited the network, the read would stall and fail. By construction.
tool_reactor_stall::write_does_not_stall_the_runtime (#148's probe; failed on CI run 37614530412) The probe, not the code. This branch already had #148's write_bytes(Vec<u8>) move. On the CI runner a 4 MB write took 0.69–0.87 ms of thread CPU inline. The tool's fixed per-call overhead (stat, create_dir_all, the hand-off: ~0.26 ms in a debug build) was a third of that baseline. The edit and write subjects are now ~36 MB, so the baseline is milliseconds on any machine. Locally: write 0.20 ms of 24 ms, edit 1.3 ms of 37 ms, ls 0.34 ms of 7.8 ms. The 25% bar is unchanged. 60/60 under 40 CPU burners. Note: edit's hold grew from ~0.3 ms to ~1.3 ms with the 8× larger input. That suggests some per-byte work still runs on the executor, still under 4% of the baseline; worth a separate look.
serve_health (readyz) The shard probe has a 1 s deadline. Under load, a caller can get "shard probe timed out" while the probe is still running; its real answer then lands in the memo. readyz() waits out exactly that one answer. Any other answer, a 503 for any other reason included, is returned at once. —
mcp_skills_integrity::a_listing_that_never_ends_is_cut_off_and_said_so The fixture's listing had TTL 0, so the turn paged through all 64 pages a second time under the refresh's 30 s wall-clock bound. The listing stays fresh for the run, so the cut-off is the page budget alone. The test also asserts exactly 64 skills/list calls. —

#137 follow-ups

  1. NFS lock safety. The key's maker re-reads it through the locked descriptor (FileLock::file()). With agent: OAuth fixture coverage (DELETE, streamed in-POST question); session lock is a record lock (serve_ws flake root cause) #144's OFD locks (merged since), no other descriptor's open or close can release the lock, NFS included, so the extra registry and test I had first written became moot and were dropped in the rebase. agent: OAuth fixture coverage (DELETE, streamed in-POST question); session lock is a record lock (serve_ws flake root cause) #144's own tests cover the OFD property.
  2. The flush that was dropped. I documented, in place, why the prompt needs no drain: carried results cover a resumer restart; a placeholder reaches the path only during a run, whose end drains its mark; and resume_into_turn adds the resumer's own placeholder ids. I couldn't build a test for this one. No observable behaviour depends on the drain, and the race can't be forced through serve's surface.
  3. Old-name key and the trash. A key still under the old name is trashed, restored and hard-deleted with its session.
    • Test: a_trashed_session_takes_an_old_name_key_along: gone from the session dir and present in .trash after delete, back and keyed after restore.

Stability

Verification

🤖 Generated with Claude Code

https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

Tests that failed only under heavy parallel load, each made independent of
host timing:

- mcp_events_nested (a real race): serve started the MCP Events hub before
  installing the session's elicitation and sampling gates, so a server
  asking a question during the very first events/poll was declined as "no
  client". The hub now starts after the gates. A debug-only seam
  (BEYOND_AI_AGENT_TEST_SLOW_GATE_INSTALL_MS) holds the window open; the
  new test fails on the old order.
- mcp_events_receiver a_retry_that_beats_the_resubscribe (a real race): a
  restarted daemon answered 410 (stop) to a delivery for a persisted
  callback that arrived before its events session had read its state and
  reserved the token. Unknown tokens are now 503 from daemon start until
  that session has reserved its tokens (bounded). A debug-only seam
  (BEYOND_AI_AGENT_TEST_SLOW_EVENTS_RESTORE_MS) holds the window open, and
  the test no longer sleeps 800 ms.
- serve_uds uds_stale_socket_file_is_reclaimed: the "stale" node was a
  listener bound and dropped in the test process, which a child forked by
  another test in that instant keeps alive until its exec; a connect then
  succeeded against it. The node is now a datagram socket, which no stream
  connect can reach.
- mcp_stdio a_panic_unwinding_past_the_exit_guard_still_sweeps:
  sweep_before_exit's wait is bounded, and a loaded host could stall a
  sweep thread past it. The exiting thread now kills every server still
  running itself after the wait. The test reads "killed" from the kernel
  (gone, a zombie, or SIGKILL pending), not from scheduling.
- mcp_events wire many_events_in_one_chunk_drain_in_linear_time: counts the
  bytes scanned and moved (must stay within 2x the input) instead of
  timing the drain.
- serve_lifecycle a_slow_lifecycle_consumer_does_not_delay_the_prompt_response:
  the collector holds every POST unanswered until released, and the prompt
  response must arrive first: proved by order, not by a 1.5 s bound.
- serve_health: /readyz callers wait out the one transient answer, "shard
  probe timed out" (the probe's real answer lands in the memo); any other
  answer, a 503 for another reason included, is returned at once.
- mcp_skills_integrity a_listing_that_never_ends_is_cut_off_and_said_so: the
  listing stays fresh for the run, so the turn does not page through all 64
  again under the refresh's 30 s bound; asserts exactly 64 skills/list.

- tool_reactor_stall (#148's probe) failed on CI: a 4 MB write took under
  a millisecond on a CI runner, and the tool's fixed per-call overhead
  (~0.26 ms in a debug build) alone was a third of that baseline. The edit
  and write subjects are now ~36 MB, so the baseline is milliseconds
  everywhere and the bar (25%) is unchanged.

#137 follow-ups:
- The journal key's maker re-reads it through the locked descriptor
  (FileLock::file); with #144's OFD locks another descriptor's close no
  longer releases the lock, NFS included.
- An old-name key (not yet carried over) trashes, restores and hard-deletes
  with its session.
- Why a prompt needs no drain of queued journal writes is documented where
  the prompt reads the journal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
@jaredLunde
jaredLunde force-pushed the jared/load-flakes-2 branch from 6e2595a to f4a5624 Compare October 7, 2026 11:48
@jaredLunde
jaredLunde merged commit 1981cc3 into main Oct 7, 2026
21 checks passed
@jaredLunde
jaredLunde deleted the jared/load-flakes-2 branch October 7, 2026 12:20
jaredLunde added a commit that referenced this pull request Oct 7, 2026
#152 fixed the third (the hub started before the session's question gates, now after them); these
two time out under load for different reasons, which #152 does not cover:

- in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its
  client attaching within a 1.5 s delay before the fixture raised the request (with no client an
  elicitation is rightly declined). The fixture now holds it until the test says both clients are
  attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a
  get_state on its own session.
- a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited
  for the one-shot `refused` frame, which goes out when the events session starts at boot — under
  load before the client attached. It now polls mcp_events_list for the recorded refusal (state
  and last_error).
- ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the
  gate ordering (#152's fix) noted where nested requests are described.

All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
jaredLunde added a commit that referenced this pull request Oct 7, 2026
… restore window covers runtime-only daemons (#156)

Audit of #152:

- A runtime subscription restored when its session starts (boot restore,
  and the restart after a panic) was started from the restore task spawned
  in `McpEventsHub::attach`, before the session's elicitation and sampling
  gates existed, so a question on its first poll could still be declined as
  "no client". One signal now gates them all: `Hub::questions_ready`, set by
  `McpEventsHub::start` once the gates are installed, and awaited by every
  subscription task (configured, restored, resubscribing) before its first
  request. Test: a restored runtime subscription asks on its first poll
  while the slow-gate seam holds the gates off; the question reaches the
  client. Fails (declined) without the wait.
- The 503-while-restoring window was armed only when configured
  subscriptions existed, so a daemon whose only subscriptions are runtime
  ones still answered 410 to early retries. The window now tracks every
  session the daemon starts to restore callbacks (the events session and
  each runtime-subscription session) and closes when the last has reserved
  its tokens (still bounded at 10 min). Test: a runtime-only daemon answers
  an early retry 503, not 410.
- Once restore completes, a token nothing holds is 410 at once, asserted in
  both restart tests, so a daemon stuck at 503 (restored() dropped) fails
  them.
- release_seams: the serve_ws seams named by the audit already sit in
  `#[cfg(debug_assertions)]` functions (a release build carries none of
  the names: checked with grep on the release binary). The checker's hole
  was elsewhere: any attribute within three lines above counted as gating,
  so `#[cfg(debug_assertions)] let x = 1;` gated the call site after it.
  It now requires the attribute directly above the start of the statement
  or arm the name is in; pinned with that case. The restore seam moved into
  a gated function of its own.


Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
jaredLunde added a commit that referenced this pull request Oct 7, 2026
#152 fixed the third (the hub started before the session's question gates, now after them); these
two time out under load for different reasons, which #152 does not cover:

- in_a_daemon_a_nested_elicitation_during_an_events_poll_reaches_the_events_session relied on its
  client attaching within a 1.5 s delay before the fixture raised the request (with no client an
  elicitation is rightly declined). The fixture now holds it until the test says both clients are
  attached (MCP_FIXTURE_NESTED_ON_SIGNAL + POST /control/raise_nested), after each has answered a
  get_state on its own session.
- a_permanently_refused_configured_subscription_is_reported_and_does_not_keep_the_session waited
  for the one-shot `refused` frame, which goes out when the events session starts at boot — under
  load before the client attached. It now polls mcp_events_list for the recorded refusal (state
  and last_error).
- ARCHITECTURE.md: the blocking-pool rule and the two stall probes; Code Mode off the executor; the
  gate ordering (#152's fix) noted where nested requests are described.

All three events tests: 108/108 under 12 concurrent copies each with 16 CPU hogs, three rounds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JimHGjsfk2Ktm5GxyZJKKk
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant