This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject: KubernetesClusterDockerTest class setup (embedded-tests, docker-tests)
Failures: 2 · First seen: 2026-09-10 · Last seen: 2026-09-18
Root cause
K3sClusterResource.waitUntilPodIsReady waits 300 s for each pod's Ready=True condition. On timeout it throws an ISE that contains only the pod name. The pods have no readiness probe and run with DRUID_XMX=128m, so a container restart or a node NotReady blip on a busy runner leaves the pod unready, and the cause is not visible in the log.
Suggested fix
On timeout, include the pod status (conditions, restart count, last termination reason), its events and the log tail in the exception, and dump pod logs to druid-container-logs. Add a /status/health readinessProbe with an initial delay to manifests/druid-service.yaml.
Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
Status: Open: no fix PR
Subject:
KubernetesClusterDockerTestclass setup (embedded-tests,docker-tests)Failures: 2 · First seen: 2026-09-10 · Last seen: 2026-09-18
Root cause
K3sClusterResource.waitUntilPodIsReadywaits 300 s for each pod'sReady=Truecondition. On timeout it throws an ISE that contains only the pod name. The pods have no readiness probe and run withDRUID_XMX=128m, so a container restart or a nodeNotReadyblip on a busy runner leaves the pod unready, and the cause is not visible in the log.Suggested fix
On timeout, include the pod status (conditions, restart count, last termination reason), its events and the log tail in the exception, and dump pod logs to
druid-container-logs. Add a/status/healthreadinessProbewith an initial delay tomanifests/druid-service.yaml.Occurrences
Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.
docker-testsdocker-tests