Skip to content

perf: avoid rebuilding the eligible-worker snapshot per pending task in HttpRemoteTaskRunner - #20299

Open
Anubhav-Roy wants to merge 2 commits into
apache:masterfrom
Anubhav-Roy:hrtr-avoid-rebuilding-worker-snapshot
Open

Anubhav-Roy wants to merge 2 commits into
apache:masterfrom
Anubhav-Roy:hrtr-avoid-rebuilding-worker-snapshot

Conversation

@Anubhav-Roy

Copy link
Copy Markdown

Fixes #20295.

Description

Under a large pending-task backlog with saturated workers, the Overlord's httpRemote
task runner stalls: one pending-task-runner thread holds statusLock while deep in
worker-snapshot reconstruction, blocking new task submission, status updates, and
worker-sync operations cluster-wide. Restarting doesn't help because the active task set is
reloaded from metadata, the backlog reappears, and the loop re-enters the same
lock-holding scan.

Fixed the Overlord stall on a large pending-task backlog

In HttpRemoteTaskRunner.pendingTasksExecutionLoop(), the loop holds the single
statusLock while iterating every pending task, and for each task calls
findWorkerToRunTask(Task), which rebuilds a full immutable snapshot of all workers via
getWorkersEligibleToRunTasks(). This makes the loop cost
O(pendingTasks × workers × tasksAnnouncedPerWorker), while it holds the lock.

This change computes the snapshot once per pass,
inside synchronized (statusLock) before iterating, and passes it into a new overload
findWorkerToRunTask(Task, ImmutableMap<String, ImmutableWorkerInfo> eligibleWorkers).
The existing findWorkerToRunTask(Task) is retained and now delegates to the overload,
so no other call site changes behavior.

This reduces per-pass cost to O(workers × tasksPerWorker) with no behavioral change:
worker selection within a pass is identical because the input snapshot is identical to
what each per-task call would have recomputed.

Release note

Fixed an issue where the Overlord using the httpRemote task runner could stall for
extended periods (holding statusLock) when a large backlog of pending tasks
accumulated while workers were saturated, blocking task submission and status updates.


Key changed/added classes in this PR
  • HttpRemoteTaskRunner

This PR has:

  • been self-reviewed.
  • added comments explaining the "why" and the intent of the code wherever would not be obvious for an unfamiliar reader.

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity Findings
P0 0
P1 0
P2 2
P3 0
Total 2

Reviewed 1 of 1 changed files.


This is an automated review by Codex GPT-5.6-Luna(max)

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up assessment

The first finding, concerning an unnecessary eligible-worker snapshot when the pending queue is empty, is resolved in the current head. The remaining stale-capacity finding was handled in inline reply 4004995846; no additional inline finding is needed from this follow-up. Reviewed 1 of 1 changed files.


This is an automated review by Codex GPT-5.6-Luna(max)

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The current head still has a race between fresh worker selection and reservation: worker task announcements can update the worker snapshot outside statusLock, and the assignment path does not revalidate capacity after the selection. The new pre-filter also leaves the cached snapshot stale after a failed fresh selection, which can restore the original repeated snapshot-rebuild lock stall during a worker-capacity transition. Both issues are included as inline findings.

Reviewed 1 of 1 changed files. Static review covered the changed scheduling loop, WorkerHolder snapshot publication, worker-selection strategies, the reservation and assignment path, and the existing review discussion. No tests or build commands were run.

Severity Findings
P0 0
P1 0
P2 2
P3 0
Total 2

This is an automated review by Codex GPT-5.6 Luna(Max)

After addressing the findings or replying to the comments, you can request another review from me to trigger a new automated review.

@Anubhav-Roy

Copy link
Copy Markdown
Author

@kfaraz this fixes an Overlord statusLock stall under a large pending-task backlog, that we faced in Prod.

It touches only the HttpRemoteTaskRunner worker selection.

Would appreciate a review when you have a chance. Automated feedback from @FrankChen021 has been addressed.

(Adding you since I could see you had touched these files.)

@FrankChen021

Copy link
Copy Markdown
Member

Thanks for reporting and submitting a patch. Could you check if above two findings are true or not? If they are not worthy of fixing, you can reply to these comments and AI will re-check the source code with your justification to make a judge.

@Anubhav-Roy Anubhav-Roy changed the title Avoid rebuilding the eligible-worker snapshot per pending task in Htt… avoid rebuilding the eligible-worker snapshot per pending task in HttpRemoteTaskRunner Sep 16, 2026
@Anubhav-Roy

Copy link
Copy Markdown
Author

They are not regressions from this code. Addressed and resolved

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Follow-up assessment

I reviewed 1 of 1 changed files at the current head, with the scheduling loop checked against WorkerHolder snapshot publication, all findWorkerToRunTask call sites, selection strategies, reservation, assignment, and lifecycle notifications. No new PR-caused findings remain.

The selection-to-reservation capacity race in the first thread is unchanged from master. The stale-cache path in the second thread can still fall back to master's existing per-task rebuild after a worker-state transition, but this PR does not introduce that behavior; in its targeted saturated-worker case it avoids the repeated snapshot reconstruction. Both supplied threads are therefore resolved for this PR, and no inline reply is warranted.

Validation: git diff --check passed. No tests, builds, dependency installs, formatters, or long validations were run.


This is an automated review by Codex GPT-5.6-Luna(max)

@Anubhav-Roy Anubhav-Roy changed the title avoid rebuilding the eligible-worker snapshot per pending task in HttpRemoteTaskRunner perf: avoid rebuilding the eligible-worker snapshot per pending task in HttpRemoteTaskRunner Sep 22, 2026
@Anubhav-Roy

Copy link
Copy Markdown
Author

Only the docker-tests fails, but I built the branch locally on JDK 25 and ran the embedded test that actually covers this code path, HttpRemoteTaskRunnerWorkerFailTest, and it passes (1 run, 0 failures). The build is clean too.

@Anubhav-Roy Anubhav-Roy reopened this Sep 23, 2026
@Anubhav-Roy
Anubhav-Roy marked this pull request as draft September 23, 2026 13:08
@Anubhav-Roy
Anubhav-Roy marked this pull request as ready for review September 23, 2026 13:09
@FrankChen021

Copy link
Copy Markdown
Member

looks like unit tests were not triggered, can you rebase the branch to the latest master branch?

@Anubhav-Roy
Anubhav-Roy force-pushed the hrtr-avoid-rebuilding-worker-snapshot branch from a326c03 to 30ceba0 Compare September 30, 2026 09:20
@Anubhav-Roy

Copy link
Copy Markdown
Author

rebased and updated the PR

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The lazy per-pass snapshot removes the repeated worker/task-announcement reconstruction from the saturated pending-task scan while preserving the existing fresh worker selection immediately before reservation. Empty and non-schedulable queues do not materialize the snapshot, and the current worker-selection call sites, WorkerHolder snapshot publication, reservation path, and lifecycle notifications were reviewed. The previously reported PR-specific issues are addressed; the remaining selection-to-reservation race is unchanged from the base implementation.

Reviewed 1 of 1 changed files. Static review covered the changed scheduling loop, all findWorkerToRunTask call sites, built-in and JavaScript worker-selection strategies, WorkerHolder snapshot publication, reservation/assignment, and the existing review discussion.

Validation: git diff --check 0ad366d5d4a2bf3b1d35dd7f515c2302c5731e7e..30ceba0cdd930895b420f63e21564df9b1777f77 passed; the local git merge-tree --write-tree conflict check against origin/master passed. No tests, builds, dependency installs, formatters, or long validations were run.


This is an automated review by Codex GPT-5.6 Luna(Max)

@FrankChen021

Copy link
Copy Markdown
Member

I think aboves findings from the AI should be addressed. Under the case that when a worker disappears after the snapshot is built, current code fallback to previous behaviour, that's true, but here we are going to solve the performance problem here right?
and because the snapshot is built before each iteration, I think compared to the original code, it increases the probability that we will encounter above problem.

2nd, after this change, getWorkersEligibleToRunTasks can be called twice, it increases the overhead for a pending task scan that can find a worker.

Instead of patching the current pending executor path, another way is to track the tasks of a worker when its announcement changes. we can do it in the WorkerHolder#fullSync and WorkerHolder#deltaSync, and update the WorkerHolder#toImmutable to return the computed result directly instead of computing in the fly as current implementation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

HttpRemoteTaskRunner: Overlord stalling on huge backlog

2 participants