Skip to content

perf(semaphore): avoid the queue lock for positive-balance releases - #316

Open
orthur2 wants to merge 4 commits into
apache:mainfrom
orthur2:perf/semaphore-positive-balance-release
Open

orthur2 wants to merge 4 commits into
apache:mainfrom
orthur2:perf/semaphore-positive-balance-release

Conversation

@orthur2

@orthur2 orthur2 commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Semaphore::release adds permits with one compare-exchange while the balance is positive and takes the waiter mutex only when the balance is zero or the exchange fails. A positive balance proves that no waiter is linked, so uncontended RwLock read-guard drops, multi-permit releases with permits to spare, and bounded pool returns below capacity no longer lock. Mutex guards keep the locked path through an internal release_all_held.
  • The second commit adds a contention flag: once a locked release loses a race on the balance, other releases take the lock until it finishes. Without it, two threads that only acquire and release one primitive take more than twice as long as on main on an Apple M5 under macOS. With it, the remaining losses are tight loops with no work between operations on the M5, 0% to 26% at two threads under macOS and 8% to 19% at three or four in a Linux VM; every other measured cell is faster than main or too noisy to rank. The flag costs 8 bytes per Semaphore, Mutex, and RwLock and is up to 14% (Hygon) or 25% (M5) slower than the first commit in some contended cells, so it is a separate commit that can be dropped.
  • Tests cover the lock-free release, the zero-balance hand-off, the locked path against a positive balance, the flag, releases racing acquisitions across threads, two releases racing at usize::MAX - 1, and, under Miri, that an acquire still synchronizes with earlier releases.

Design Notes

A waiter is linked only after acquired_or_enqueue has drained the balance to zero, in an exchange it performs while holding the queue lock, and insert_permits_with_lock adds permits to the balance only once no waiter is linked. A positive balance therefore means the queue is empty. If an acquisition drains the balance while a release attempts its exchange, the acquirer already holds the lock, so the release's exchange fails and the release takes the locked path, which runs after the waiter is linked and hands the permits to it.

I attempt the exchange once, in its strong form. Retrying it in a loop was several times slower than main once several threads released together, and a spurious failure would send an uncontended release to the lock.

The fast path can add to a positive balance while another thread holds the queue lock, so the locked path can no longer check for overflow and then call fetch_add. For a positive balance it uses the fast path's checked compare-exchange, and overflow still panics before any permit is added. A zero balance cannot change under the lock or overflow, so it keeps fetch_add, and every write to the balance stays a read-modify-write. I tried a plain store for a zero balance, and Miri reported a data race: the store ended the release sequence of the previous release, so a later acquire no longer synchronized with it. semaphore_ordering_test covers that, and it fails on the plain-store revision.

Mutex guards hold the only permit, so their balance is always zero and they call release_all_held to skip the probe. RwLock write guards qualify too; I leave them out because #351 moves the guard permits into access tokens.

The flag targets the two-thread loss on the M5. There a failed compare-exchange costs about 12 ns against 1.75 ns for a successful one, and in a single-word model of the two-thread loop 83% of the release exchanges failed; on the Hygon both cost about 11 ns and almost none failed. A locked release that still loses the race sets the flag, other releases see it and queue on the lock instead of adding to the traffic, and the locked release clears it once its addition lands. I tried a flag that trips only on a second loss, which did not help two threads, and one that stays set until a later uncontended addition, which cost 30% to 50% at four to eight threads.

Benchmarks

main is c9dff9e; "PR" is the first commit and "+flag" both. Linux: Hygon C86 7360, 2 sockets, 8 NUMA nodes. macOS: Apple M5, 4 performance and 6 efficiency cores, background load 3 to 6. Rust 1.96.0 on both, binaries in shuffled order each round.

Contention rows come from an out-of-tree probe: N threads run the same loop on one primitive after a ready barrier, and the value is wall-clock nanoseconds per iteration of one thread over 200,000 iterations, median of 15 rounds. On Linux each thread is pinned to its own core, on one NUMA node for up to four threads and on two for eight. A short or long gap is a dependent multiply chain of 50 or 200 steps between a release and the next acquire, about 18 and 140 ns on the M5 and 80 and 320 ns on the Hygon. Every Linux round below beats every main round of its cell.

Contention, ns per iteration Linux main PR +flag macOS main PR +flag
RwLock read, 2 threads 268.1 141.3 141.8 51.8 84.9 65.3
RwLock read, 4 threads 916.7 423.7 470.6 155.8 117.0 85.2
RwLock read, 8 threads 2486 999 997 377.3 236.7 201.7
RwLock read, 2 threads, short gap 238.1 168.1 169.1 146.1 114.5 118.2
RwLock read, 8 threads, long gap 3943 1670 1878 2746 1019 1235
RwLock read held across a short gap, 4 threads 926.5 549.9 590.0 378.8 188.3 235.6
Semaphore 64 permits, take 2, 2 threads 277.8 168.9 180.3 32.0 69.4 28.3
Semaphore 64 permits, take 2, 8 threads 1963 445 444 239.9 105.4 112.3
Semaphore 2 permits, 2 threads 266.6 213.1 242.9 47.9 47.8 49.4

All measured losses are on the M5 with no work between operations. Under macOS, two threads on RwLock read guards are 64% to 77% slower than main with the first commit and 0% to 26% with the flag across five runs, and the flag's two-thread Semaphore cells overlap main. To separate the OS from the CPU, I also ran the probe in a Linux VM on the same M5 (colima, 6 vCPUs). There main is about three times slower than under macOS on the two-thread loops without work, and both commits beat it in every two-thread cell, by 31% to 52%, but four threads on RwLock read guards are 8% to 18% slower than main with the first commit and 8% to 11% with the flag across three pinnings, and three threads 27% and 19%. I have no native ARM Linux host to tell how much of that is the VM. With either gap, no cell on any platform is more than 2% slower than main. The two-permit loops at four and eight threads park on nearly every acquisition and vary up to twofold between runs.

Single-thread rows run cargo x bench --bench primitives -- --sample-count 500 --sample-size 5000 with name filters, median of 20 rounds, pinned to one core on Linux. Three rows are new so that each affected path has a direct measurement. The flag adds one relaxed load to the fast path.

cargo x bench, ns Linux main PR +flag macOS main PR +flag
rwlock::read::read_reuse 88.0 49.0 48.8 13.16 7.65 7.59
rwlock::read::eight_reads_then_one_write 775.8 492.4 493.3 117.90 76.45 76.19
semaphore::acquire::owned_try_acquire_release 69.5 35.5 38.7 8.80 5.00 5.00
semaphore::acquire::try_acquire_release_with_spare_permits 58.3 20.2 20.2 7.27 3.95 3.98
pool::bounded::bounded_warm_get_and_return_with_spare_capacity 282.6 248.6 246.9 51.38 44.86 44.84
semaphore::acquire::try_acquire_release 58.3 56.7 56.7 7.26 7.55 7.53
mutex::lock::uncontended_reuse 94.1 93.9 93.8 13.09 12.98 12.90

On the M5, try_acquire_release is 0.3 ns (4%) slower with either commit: a single-permit release finds the balance at zero and reads it once more before taking the lock. The Hygon shows 3% the other way, and Mutex skips that read through release_all_held.

Validation

  • cargo x check, cargo x test, cargo x lint, and cargo x miri on each commit
  • cargo x miri now runs semaphore_ordering_test

@orthur2
orthur2 force-pushed the perf/semaphore-positive-balance-release branch from 9c09ca1 to ad5b6c2 Compare October 8, 2026 08:07
@orthur2

orthur2 commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

I looked into where the M5 losses come from by adding per-thread CPU time and getrusage context-switch counts to the probe. They come from main's queue lock rather than from the exchange. In a loop with no work, main wins where its lock puts threads to sleep and makes them take turns, so each thread runs alone with its cache lines; the first commit keeps the threads running concurrently, at about 100 ns per iteration with two threads on the M5 under either OS. On the M5 that turn-taking happens at two threads under macOS, whose pthread_mutex defaults to the first-fit policy, and at three or four threads under Linux once the futex mutex stops spinning. On the Hygon, main only gets slower once its threads start sleeping. The flag brings back part of the turn-taking.

RwLock read with no work between operations. Each cell is ns per iteration, CPU utilization of the threads, and context switches per 1,000 iterations (voluntary plus involuntary), median of 7 rounds:

Host, threads main PR +flag
M5 macOS, 2 56, 78%, 21 102, 99%, 0.1 59, 90%, 6.8
M5 macOS, 3 256, 74%, 95 80, 91%, 11 70, 74%, 15
M5 Linux VM, 2 167, 99%, 0.1 98, 99%, 0 88, 99%, 0
M5 Linux VM, 3 119, 81%, 13 168, 95%, 0.4 161, 94%, 1.2
M5 Linux VM, 4 163, 77%, 32 206, 85%, 21 182, 84%, 21
Hygon Linux, 2 271, 99%, 0 147, 97%, 0 156, 94%, 0
Hygon Linux, 4 998, 83%, 116 381, 81%, 21 424, 89%, 30

The Semaphore loop taking two of 64 permits splits the same way at two threads. In the VM at three and four threads, main sleeps there too but stays behind both commits.

@orthur2
orthur2 marked this pull request as ready for review October 8, 2026 11:58
@orthur2
orthur2 requested review from QwQBiG and tisonkun October 8, 2026 11:59

@QwQBiG QwQBiG left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Local tests passed on Linux/WSL with stable and Rust 1.86.

One small clarification in the PR description: the zero-balance locked path still uses fetch_add, so “both paths share one checked compare-exchange” needs that exception. :3

Comment thread xtask/src/main.rs Outdated
Comment thread CHANGELOG.md Outdated
@orthur2

orthur2 commented Oct 10, 2026

Copy link
Copy Markdown
Contributor Author

Pushed the changes and updated the description.

@QwQBiG

QwQBiG commented Oct 10, 2026

Copy link
Copy Markdown
Member

Thanks for the updates!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants