This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Resolved by #20479
Subject: EmbeddedKafkaSupervisorTest.test_runKafkaSupervisor, EmbeddedKafkaSupervisorTest.test_runSupervisor_withEmptyDimension (extensions-core/kafka-indexing-service, unit tests (25, E*,G*,F*))
Failures: 0 (3 passed on retry) · First seen: 2026-10-01 · Last seen: 2026-10-02
Tracked although every occurrence so far passed on retry: the flake recurs daily, each surefire rerun restarts an embedded cluster (~30 s added to the job), and the fix is test-only.
Root cause
test_runKafkaSupervisor line 125 (expected: <1> but was: <2>): the supervisor uses taskDuration=PT0.5S, but the test asserts exactly one active task about 1 s after the first task starts. In job 110571049503 the first task started at 20:52:35.758, its duration expired at 36.800 (end offsets set, task still active while publishing), a second task was created at 36.859, and the test listed tasks at 36.951 and saw two. Because the test fails before suspending its supervisor, that supervisor keeps rolling over tasks for the rest of the class.
test_runSupervisor_withEmptyDimension line 210 (timeout) / line 219 (expected: <100> but was: <83>): with maxRowsPerSegment=1 and one row per day, the 100 rows create 100 one-row segments. The test waits on ingest/events/processed, but the class does not set druid.monitoring.emissionPeriod (default PT1M), so the task emits that metric only from removeMonitor when it exits, which happens after 100 segment handoffs. In job 110571049503 all 100 rows were read by +4.7 s, but only 77 segments had loaded on the Historical when the 10 s wait expired, while the leaked supervisor from the failed test above competed for the same Indexer and Historical. The passing rerun took 7.97 s, so the margin is thin even without contention. The line 219 variant is the same slow handoff, observed by the row-count query.
None of the commits involved touch the Kafka indexing service.
Suggested fix
Test-only changes in EmbeddedKafkaSupervisorTest: (1) assert that at least one task exists and that all are RUNNING instead of exactly one; (2) add an @AfterEach that suspends or terminates the test's supervisor so one failure cannot slow the next test; (3) set indexer.addProperty("druid.monitoring.emissionPeriod", "PT0.1s") (as KafkaClusterMetricsTest does) and drop maxRowsPerSegment(1) for test_runSupervisor_withEmptyDimension, which only needs to verify empty-dimension handling. Fix PR #20479 implements these changes.
Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
| Date |
Commit |
Job |
Failure log |
Detail |
Reported in |
| 2026-10-01 |
631f05d (#20458) |
unit tests (25, E*,G*,F*) |
job 110375258572 |
test_runSupervisor_withEmptyDimension:219 expected 100 rows but was 83; 1 of 2 attempts failed (passed on retry) |
#20437 |
| 2026-10-02 |
e855cd9 (#20448) |
unit tests (25, E*,G*,F*) |
job 110571049503 |
test_runKafkaSupervisor:125 expected 1 task but was 2; 1 of 2 attempts failed (passed on retry) |
#20437 |
| 2026-10-02 |
e855cd9 (#20448) |
unit tests (25, E*,G*,F*) |
job 110571049503 |
test_runSupervisor_withEmptyDimension:210 timed out after 10 s waiting for ingest/events/processed; 1 of 2 attempts failed (passed on retry) |
#20437 |
This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Resolved by #20479
Subject:
EmbeddedKafkaSupervisorTest.test_runKafkaSupervisor,EmbeddedKafkaSupervisorTest.test_runSupervisor_withEmptyDimension(extensions-core/kafka-indexing-service,unit tests (25, E*,G*,F*))Failures: 0 (3 passed on retry) · First seen: 2026-10-01 · Last seen: 2026-10-02
Tracked although every occurrence so far passed on retry: the flake recurs daily, each surefire rerun restarts an embedded cluster (~30 s added to the job), and the fix is test-only.
Root cause
test_runKafkaSupervisorline 125 (expected: <1> but was: <2>): the supervisor usestaskDuration=PT0.5S, but the test asserts exactly one active task about 1 s after the first task starts. In job 110571049503 the first task started at 20:52:35.758, its duration expired at 36.800 (end offsets set, task still active while publishing), a second task was created at 36.859, and the test listed tasks at 36.951 and saw two. Because the test fails before suspending its supervisor, that supervisor keeps rolling over tasks for the rest of the class.test_runSupervisor_withEmptyDimensionline 210 (timeout) / line 219 (expected: <100> but was: <83>): withmaxRowsPerSegment=1and one row per day, the 100 rows create 100 one-row segments. The test waits oningest/events/processed, but the class does not setdruid.monitoring.emissionPeriod(defaultPT1M), so the task emits that metric only fromremoveMonitorwhen it exits, which happens after 100 segment handoffs. In job 110571049503 all 100 rows were read by +4.7 s, but only 77 segments had loaded on the Historical when the 10 s wait expired, while the leaked supervisor from the failed test above competed for the same Indexer and Historical. The passing rerun took 7.97 s, so the margin is thin even without contention. The line 219 variant is the same slow handoff, observed by the row-count query.None of the commits involved touch the Kafka indexing service.
Suggested fix
Test-only changes in
EmbeddedKafkaSupervisorTest: (1) assert that at least one task exists and that all areRUNNINGinstead of exactly one; (2) add an@AfterEachthat suspends or terminates the test's supervisor so one failure cannot slow the next test; (3) setindexer.addProperty("druid.monitoring.emissionPeriod", "PT0.1s")(asKafkaClusterMetricsTestdoes) and dropmaxRowsPerSegment(1)fortest_runSupervisor_withEmptyDimension, which only needs to verify empty-dimension handling. Fix PR #20479 implements these changes.Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
unit tests (25, E*,G*,F*)test_runSupervisor_withEmptyDimension:219 expected 100 rows but was 83; 1 of 2 attempts failed (passed on retry)unit tests (25, E*,G*,F*)test_runKafkaSupervisor:125 expected 1 task but was 2; 1 of 2 attempts failed (passed on retry)unit tests (25, E*,G*,F*)test_runSupervisor_withEmptyDimension:210 timed out after 10 s waiting foringest/events/processed; 1 of 2 attempts failed (passed on retry)