Skip to content

Flaky test: EmbeddedKafkaSupervisorTest (test_runKafkaSupervisor, test_runSupervisor_withEmptyDimension) #20478

Description

@FrankChen021

This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.

Status: Resolved by #20479
Subject: EmbeddedKafkaSupervisorTest.test_runKafkaSupervisor, EmbeddedKafkaSupervisorTest.test_runSupervisor_withEmptyDimension (extensions-core/kafka-indexing-service, unit tests (25, E*,G*,F*))
Failures: 0 (3 passed on retry) · First seen: 2026-10-01 · Last seen: 2026-10-02

Tracked although every occurrence so far passed on retry: the flake recurs daily, each surefire rerun restarts an embedded cluster (~30 s added to the job), and the fix is test-only.

Root cause

  • test_runKafkaSupervisor line 125 (expected: <1> but was: <2>): the supervisor uses taskDuration=PT0.5S, but the test asserts exactly one active task about 1 s after the first task starts. In job 110571049503 the first task started at 20:52:35.758, its duration expired at 36.800 (end offsets set, task still active while publishing), a second task was created at 36.859, and the test listed tasks at 36.951 and saw two. Because the test fails before suspending its supervisor, that supervisor keeps rolling over tasks for the rest of the class.
  • test_runSupervisor_withEmptyDimension line 210 (timeout) / line 219 (expected: <100> but was: <83>): with maxRowsPerSegment=1 and one row per day, the 100 rows create 100 one-row segments. The test waits on ingest/events/processed, but the class does not set druid.monitoring.emissionPeriod (default PT1M), so the task emits that metric only from removeMonitor when it exits, which happens after 100 segment handoffs. In job 110571049503 all 100 rows were read by +4.7 s, but only 77 segments had loaded on the Historical when the 10 s wait expired, while the leaked supervisor from the failed test above competed for the same Indexer and Historical. The passing rerun took 7.97 s, so the margin is thin even without contention. The line 219 variant is the same slow handoff, observed by the row-count query.

None of the commits involved touch the Kafka indexing service.

Suggested fix

Test-only changes in EmbeddedKafkaSupervisorTest: (1) assert that at least one task exists and that all are RUNNING instead of exactly one; (2) add an @AfterEach that suspends or terminates the test's supervisor so one failure cannot slow the next test; (3) set indexer.addProperty("druid.monitoring.emissionPeriod", "PT0.1s") (as KafkaClusterMetricsTest does) and drop maxRowsPerSegment(1) for test_runSupervisor_withEmptyDimension, which only needs to verify empty-dimension handling. Fix PR #20479 implements these changes.

Occurrences

Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.

Date Commit Job Failure log Detail Reported in
2026-10-01 631f05d (#20458) unit tests (25, E*,G*,F*) job 110375258572 test_runSupervisor_withEmptyDimension:219 expected 100 rows but was 83; 1 of 2 attempts failed (passed on retry) #20437
2026-10-02 e855cd9 (#20448) unit tests (25, E*,G*,F*) job 110571049503 test_runKafkaSupervisor:125 expected 1 task but was 2; 1 of 2 attempts failed (passed on retry) #20437
2026-10-02 e855cd9 (#20448) unit tests (25, E*,G*,F*) job 110571049503 test_runSupervisor_withEmptyDimension:210 timed out after 10 s waiting for ingest/events/processed; 1 of 2 attempts failed (passed on retry) #20437

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions