This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject: KinesisFaultToleranceTest.test_supervisorRecovers_afterChangeInTopicPartitions_withEmptyShards (embedded-tests, unit tests (25, K*,U*,Z*,Y*,X*))
Failures: 2 · First seen: 2026-09-23 · Last seen: 2026-10-10
Root cause
After the stream is resharded from 2 to 4 shards, the supervisor must finish the closed shards, publish, and start tasks on the child shards. On a busy runner, ingest/rows/published did not reach 2000 within 240 s, and the timeout message does not say whether the tasks were stuck or just slow. The 2026-10-10 failure was on a Dependabot bump of amazon-kinesis-client (#20533), but the same job passed on the two later master commits that include that bump.
Suggested fix
Wait for the supervisor to report 4 partitions before waiting on the row aggregate, and log the supervisor status and the current row sum on timeout.
Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
| Date |
Commit |
Job |
Failure log |
Detail |
Reported in |
| 2026-09-23 |
4eede66 (#20380) |
unit tests (25, K*,U*,Z*,Y*,X*) |
job 106878085304 |
timed out after 240 s waiting for 2000 published rows; 1/1 attempt |
#20414 |
| 2026-10-10 |
2144e34 (#20533) |
unit tests / unit tests(main) (25, K*,U*,Z*,Y*,X*) / test-jdk25-[K*,U*,Z*,Y*,X*] |
job 114106560713 |
timed out after 240 s in verifyAndTearDown waiting for published rows; 1/1 attempt |
|
This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject:
KinesisFaultToleranceTest.test_supervisorRecovers_afterChangeInTopicPartitions_withEmptyShards(embedded-tests,unit tests (25, K*,U*,Z*,Y*,X*))Failures: 2 · First seen: 2026-09-23 · Last seen: 2026-10-10
Root cause
After the stream is resharded from 2 to 4 shards, the supervisor must finish the closed shards, publish, and start tasks on the child shards. On a busy runner,
ingest/rows/publisheddid not reach 2000 within 240 s, and the timeout message does not say whether the tasks were stuck or just slow. The 2026-10-10 failure was on a Dependabot bump ofamazon-kinesis-client(#20533), but the same job passed on the two later master commits that include that bump.Suggested fix
Wait for the supervisor to report 4 partitions before waiting on the row aggregate, and log the supervisor status and the current row sum on timeout.
Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
unit tests (25, K*,U*,Z*,Y*,X*)unit tests / unit tests(main) (25, K*,U*,Z*,Y*,X*) / test-jdk25-[K*,U*,Z*,Y*,X*]verifyAndTearDownwaiting for published rows; 1/1 attempt