This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject: MSQWorkerFaultToleranceTest.test_cancelledWorker_isRetried_ifFaultToleranceIsEnabled (embedded-tests, unit tests (25, M*,A*,V*,W*))
Failures: 1 · First seen: 2026-10-10 · Last seen: 2026-10-10
Root cause
The test cancels the MSQ worker on the faulty Indexer and then waits for the controller task to succeed with waitForTaskToSucceed(taskId, overlord.latchableEmitter()) (line 141), which uses LatchableEmitter's default 10 s timeout. In that window the controller has to detect the failed worker, relaunch it on the newly started functional Indexer (which must first be discovered and announce capacity), re-read the input and publish segments. On a busy runner this sequence took longer than 10 s, so the wait timed out even though the earlier steps finished in about 1.5 s. The failing commit (#20535) is a Dependabot bump of async-http-client, and the same job passed on the neighbouring master commits.
Suggested fix
Wait for the controller's task/run/time event with an explicit longer timeout, using the LatchableEmitter.waitForEvent(Predicate<Event>, long timeoutMillis) overload with about 60 s, then check the task status. Alternatively, first wait (with a longer timeout) for the relaunched worker's task/run/time SUCCESS event.
Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
| Date |
Commit |
Job |
Failure log |
Detail |
Reported in |
| 2026-10-10 |
7921967 (#20535) |
unit tests / unit tests(main) (25, M*,A*,V*,W*) / test-jdk25-[M*,A*,V*,W*] |
job 114107083909 |
ISE: Timed out waiting for event after [10,000]ms at line 141 (controller did not finish after worker relaunch); single attempt, no surefire retry |
|
This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject:
MSQWorkerFaultToleranceTest.test_cancelledWorker_isRetried_ifFaultToleranceIsEnabled(embedded-tests,unit tests (25, M*,A*,V*,W*))Failures: 1 · First seen: 2026-10-10 · Last seen: 2026-10-10
Root cause
The test cancels the MSQ worker on the faulty Indexer and then waits for the controller task to succeed with
waitForTaskToSucceed(taskId, overlord.latchableEmitter())(line 141), which usesLatchableEmitter's default 10 s timeout. In that window the controller has to detect the failed worker, relaunch it on the newly started functional Indexer (which must first be discovered and announce capacity), re-read the input and publish segments. On a busy runner this sequence took longer than 10 s, so the wait timed out even though the earlier steps finished in about 1.5 s. The failing commit (#20535) is a Dependabot bump ofasync-http-client, and the same job passed on the neighbouring master commits.Suggested fix
Wait for the controller's
task/run/timeevent with an explicit longer timeout, using theLatchableEmitter.waitForEvent(Predicate<Event>, long timeoutMillis)overload with about 60 s, then check the task status. Alternatively, first wait (with a longer timeout) for the relaunched worker'stask/run/timeSUCCESS event.Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
unit tests / unit tests(main) (25, M*,A*,V*,W*) / test-jdk25-[M*,A*,V*,W*]ISE: Timed out waiting for event after [10,000]msat line 141 (controller did not finish after worker relaunch); single attempt, no surefire retry