Skip to content

fix: Fix flaky TaskQueueScaleTest.doMassLaunchAndExit - #20426

Closed
Akshat-Jain wants to merge 1 commit into
apache:masterfrom
Akshat-Jain:akshat/fix-flaky-TaskQueueScaleTest
Closed

Akshat-Jain wants to merge 1 commit into
apache:masterfrom
Akshat-Jain:akshat/fix-flaky-TaskQueueScaleTest

Conversation

@Akshat-Jain

Copy link
Copy Markdown
Contributor

Description

This PR fixes flaky TaskQueueScaleTest.doMassLaunchAndExit. Sample failure: https://github.com/apache/druid/actions/runs/36084362652/job/107912839489

   🧪 - indexing-service/target/test-classes/org/apache/druid/indexing/overlord/TaskQueueScaleTest.class | all tasks should be known ==> expected: <1000> but was: <973>
  Error: all tasks should be known ==> expected: <1000> but was: <973>

TaskQueueScaleTest.doMassLaunchAndExit reads the running, pending and waiting counts at three separate instants while the TaskQueue-Manager thread is still submitting tasks to the runner. Since getWaitingTaskCount() counts active tasks the runner does not know about, a task that crosses taskRunner.run() between the pending read and the waiting read is counted by neither, so the sum comes up short (expected: <1000> but was: <988>).

The test had a high failure rate without this PR change. With this PR change, I ran the test 100 times locally, it passes everytime.


This PR has:

  • been self-reviewed.

@FrankChen021

Copy link
Copy Markdown
Member

Thanks for the fix. However, a simple while-loop is not enough. It has been fixed by #20291
Let's see if we have similar flaky test after that.

BTW, under the GitHub issue section, there's a series issues created by Claude Code everyday that triages CI failures in previous day, some of those are still open, hope someone can fix them.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants