This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
This triage covers the 3 commits merged to master on 2026-09-18, all Dependabot version bumps merged within 3 minutes of each other. One commit (3cbd00d, #20363, the last of the three) had a failed job: the push-triggered docker-tests job. The commit did not cause it; it is a one-line json-flattener version bump in benchmarks/pom.xml. KubernetesClusterDockerTest timed out waiting for the k3s druid-coordinator pod to become Ready. The same test passed in the two docker-tests jobs that ran at the same time for the commits merged 1 and 3 minutes earlier (b0b1784, 3248d73), and on the PR's own pre-merge run (5da3638). It also passed on every other master push since 2026-09-10, when the same setup failure last hit the router pod (#20312). No re-runs had been triggered, and embedded-tests runs with surefire.rerunFailingTestsCount=0, so the docker test failed on its single attempt.
Summary
| Commit |
Failed job |
Failure log |
Root cause |
Verdict |
| 3cbd00d (#20363, build(deps): bump json-flattener) |
docker-tests |
job 105449382571 |
KubernetesClusterDockerTest setup: ISE: Timed out waiting for pod[druid-coordinator-7fdb8bb578-l2jwd] to be ready after the 300 s POD_READY_TIMEOUT_SECONDS in K3sClusterResource.waitUntilPodIsReady, although the coordinator JVM had announced itself in ZooKeeper 21 s after the manifests were applied; single attempt (no surefire retries in embedded-tests); the other 101 docker tests in the job passed |
Flaky / infra |
Analysis and suggested fixes
Status of each item updated on 2026-09-26.
Open
1. KubernetesClusterDockerTest: k3s pod never becomes Ready (embedded-tests, docker-tests)
K3sClusterResource.waitUntilPodIsReady waits 300 s for Ready=True and then throws an ISE that includes only the pod name, so the cause (container restart, 128 MB heap, node NotReady) is not visible. It hit the router pod on 2026-09-10 (#20312) and the coordinator pod on 2026-09-18 (#20385). No fix PR is open.
Suggested fix: on timeout, include the pod status, events and log tail in the exception and dump the pod logs to druid-container-logs. Also add a /status/health readinessProbe with an initial delay to manifests/druid-service.yaml.
This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.
This triage covers the 3 commits merged to master on 2026-09-18, all Dependabot version bumps merged within 3 minutes of each other. One commit (3cbd00d, #20363, the last of the three) had a failed job: the push-triggered
docker-testsjob. The commit did not cause it; it is a one-linejson-flattenerversion bump inbenchmarks/pom.xml.KubernetesClusterDockerTesttimed out waiting for the k3sdruid-coordinatorpod to become Ready. The same test passed in the twodocker-testsjobs that ran at the same time for the commits merged 1 and 3 minutes earlier (b0b1784, 3248d73), and on the PR's own pre-merge run (5da3638). It also passed on every other master push since 2026-09-10, when the same setup failure last hit the router pod (#20312). No re-runs had been triggered, andembedded-testsruns withsurefire.rerunFailingTestsCount=0, so the docker test failed on its single attempt.Summary
docker-testsKubernetesClusterDockerTestsetup:ISE: Timed out waiting for pod[druid-coordinator-7fdb8bb578-l2jwd] to be readyafter the 300 sPOD_READY_TIMEOUT_SECONDSinK3sClusterResource.waitUntilPodIsReady, although the coordinator JVM had announced itself in ZooKeeper 21 s after the manifests were applied; single attempt (no surefire retries inembedded-tests); the other 101 docker tests in the job passedAnalysis and suggested fixes
Status of each item updated on 2026-09-26.
Open
1.
KubernetesClusterDockerTest: k3s pod never becomes Ready (embedded-tests,docker-tests)K3sClusterResource.waitUntilPodIsReadywaits 300 s forReady=Trueand then throws an ISE that includes only the pod name, so the cause (container restart, 128 MB heap, nodeNotReady) is not visible. It hit the router pod on 2026-09-10 (#20312) and the coordinator pod on 2026-09-18 (#20385). No fix PR is open.Suggested fix: on timeout, include the pod status, events and log tail in the exception and dump the pod logs to
druid-container-logs. Also add a/status/healthreadinessProbewith an initial delay tomanifests/druid-service.yaml.