Skip to content

Flaky test: KinesisFaultToleranceTest.test_supervisorRecovers_afterChangeInTopicPartitions_withEmptyShards #20434

Description

@FrankChen021

This issue was generated automatically by Claude Code (Anthropic's AI coding agent) running a scheduled CI-triage routine on behalf of @FrankChen021. Analysis and suggested fixes are AI-produced; please verify before acting on them.

Status: Open: no fix PR
Subject: KinesisFaultToleranceTest.test_supervisorRecovers_afterChangeInTopicPartitions_withEmptyShards (embedded-tests, unit tests (25, K*,U*,Z*,Y*,X*))
Failures: 1 · First seen: 2026-09-23 · Last seen: 2026-09-23

Root cause

After the stream is resharded from 2 to 4 shards, the supervisor must finish the closed shards, publish, and start tasks on the child shards. On a busy runner, ingest/rows/published did not reach 2000 within 240 s, and the timeout message does not say whether the tasks were stuck or just slow.

Suggested fix

Wait for the supervisor to report 4 partitions before waiting on the row aggregate, and log the supervisor status and the current row sum on timeout.

Occurrences

Failed push-triggered master jobs only. The daily triage routine adds one row per new failed job.

Date Commit Job Failure log Detail Reported in
2026-09-23 4eede66 (#20380) unit tests (25, K*,U*,Z*,Y*,X*) job 106878085304 timed out after 240 s waiting for 2000 published rows; 1/1 attempt #20414

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions