Skip to content

Flaky test: MSQWorkerFaultToleranceTest.test_cancelledWorker_isRetried_ifFaultToleranceIsEnabled #20554

Description

@FrankChen021

This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.

Status: Open: no fix PR
Subject: MSQWorkerFaultToleranceTest.test_cancelledWorker_isRetried_ifFaultToleranceIsEnabled (embedded-tests, unit tests (25, M*,A*,V*,W*))
Failures: 1 · First seen: 2026-10-10 · Last seen: 2026-10-10

Root cause

The test cancels the MSQ worker on the faulty Indexer and then waits for the controller task to succeed with waitForTaskToSucceed(taskId, overlord.latchableEmitter()) (line 141), which uses LatchableEmitter's default 10 s timeout. In that window the controller has to detect the failed worker, relaunch it on the newly started functional Indexer (which must first be discovered and announce capacity), re-read the input and publish segments. On a busy runner this sequence took longer than 10 s, so the wait timed out even though the earlier steps finished in about 1.5 s. The failing commit (#20535) is a Dependabot bump of async-http-client, and the same job passed on the neighbouring master commits.

Suggested fix

Wait for the controller's task/run/time event with an explicit longer timeout, using the LatchableEmitter.waitForEvent(Predicate<Event>, long timeoutMillis) overload with about 60 s, then check the task status. Alternatively, first wait (with a longer timeout) for the relaunched worker's task/run/time SUCCESS event.

Occurrences

Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.

Date Commit Job Failure log Detail Reported in
2026-10-10 7921967 (#20535) unit tests / unit tests(main) (25, M*,A*,V*,W*) / test-jdk25-[M*,A*,V*,W*] job 114107083909 ISE: Timed out waiting for event after [10,000]ms at line 141 (controller did not finish after worker relaunch); single attempt, no surefire retry

Activity

  1. akshat-lakhera commented on Oct 11, 2026

    @akshat-lakhera

    Hey @FrankChen021, I'd like to take this flaky test.

    The wait in MSQWorkerFaultToleranceTest uses the default timeout, which looks too short when the controller has to relaunch the worker on a new Indexer. I'd wait on the relaunch event or use an explicit longer timeout, no sleeps, and run it repeatedly under load to show it's stable.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions