Repository navigation
Conversation
9c09ca1 to
ad5b6c2
Compare
|
I looked into where the M5 losses come from by adding per-thread CPU time and
The |
QwQBiG
left a comment
There was a problem hiding this comment.
Local tests passed on Linux/WSL with stable and Rust 1.86.
One small clarification in the PR description: the zero-balance locked path still uses fetch_add, so “both paths share one checked compare-exchange” needs that exception. :3
|
Pushed the changes and updated the description. |
|
Thanks for the updates! |
Summary
Semaphore::releaseadds permits with one compare-exchange while the balance is positive and takes the waiter mutex only when the balance is zero or the exchange fails. A positive balance proves that no waiter is linked, so uncontendedRwLockread-guard drops, multi-permit releases with permits to spare, and bounded pool returns below capacity no longer lock.Mutexguards keep the locked path through an internalrelease_all_held.mainon an Apple M5 under macOS. With it, the remaining losses are tight loops with no work between operations on the M5, 0% to 26% at two threads under macOS and 8% to 19% at three or four in a Linux VM; every other measured cell is faster thanmainor too noisy to rank. The flag costs 8 bytes perSemaphore,Mutex, andRwLockand is up to 14% (Hygon) or 25% (M5) slower than the first commit in some contended cells, so it is a separate commit that can be dropped.usize::MAX - 1, and, under Miri, that an acquire still synchronizes with earlier releases.Design Notes
A waiter is linked only after
acquired_or_enqueuehas drained the balance to zero, in an exchange it performs while holding the queue lock, andinsert_permits_with_lockadds permits to the balance only once no waiter is linked. A positive balance therefore means the queue is empty. If an acquisition drains the balance while a release attempts its exchange, the acquirer already holds the lock, so the release's exchange fails and the release takes the locked path, which runs after the waiter is linked and hands the permits to it.I attempt the exchange once, in its strong form. Retrying it in a loop was several times slower than
mainonce several threads released together, and a spurious failure would send an uncontended release to the lock.The fast path can add to a positive balance while another thread holds the queue lock, so the locked path can no longer check for overflow and then call
fetch_add. For a positive balance it uses the fast path's checked compare-exchange, and overflow still panics before any permit is added. A zero balance cannot change under the lock or overflow, so it keepsfetch_add, and every write to the balance stays a read-modify-write. I tried a plain store for a zero balance, and Miri reported a data race: the store ended the release sequence of the previous release, so a later acquire no longer synchronized with it.semaphore_ordering_testcovers that, and it fails on the plain-store revision.Mutexguards hold the only permit, so their balance is always zero and they callrelease_all_heldto skip the probe.RwLockwrite guards qualify too; I leave them out because #351 moves the guard permits into access tokens.The flag targets the two-thread loss on the M5. There a failed compare-exchange costs about 12 ns against 1.75 ns for a successful one, and in a single-word model of the two-thread loop 83% of the release exchanges failed; on the Hygon both cost about 11 ns and almost none failed. A locked release that still loses the race sets the flag, other releases see it and queue on the lock instead of adding to the traffic, and the locked release clears it once its addition lands. I tried a flag that trips only on a second loss, which did not help two threads, and one that stays set until a later uncontended addition, which cost 30% to 50% at four to eight threads.
Benchmarks
mainisc9dff9e; "PR" is the first commit and "+flag" both. Linux: Hygon C86 7360, 2 sockets, 8 NUMA nodes. macOS: Apple M5, 4 performance and 6 efficiency cores, background load 3 to 6. Rust 1.96.0 on both, binaries in shuffled order each round.Contention rows come from an out-of-tree probe: N threads run the same loop on one primitive after a ready barrier, and the value is wall-clock nanoseconds per iteration of one thread over 200,000 iterations, median of 15 rounds. On Linux each thread is pinned to its own core, on one NUMA node for up to four threads and on two for eight. A short or long gap is a dependent multiply chain of 50 or 200 steps between a release and the next acquire, about 18 and 140 ns on the M5 and 80 and 320 ns on the Hygon. Every Linux round below beats every
mainround of its cell.RwLockread, 2 threadsRwLockread, 4 threadsRwLockread, 8 threadsRwLockread, 2 threads, short gapRwLockread, 8 threads, long gapRwLockread held across a short gap, 4 threadsSemaphore64 permits, take 2, 2 threadsSemaphore64 permits, take 2, 8 threadsSemaphore2 permits, 2 threadsAll measured losses are on the M5 with no work between operations. Under macOS, two threads on
RwLockread guards are 64% to 77% slower thanmainwith the first commit and 0% to 26% with the flag across five runs, and the flag's two-threadSemaphorecells overlapmain. To separate the OS from the CPU, I also ran the probe in a Linux VM on the same M5 (colima, 6 vCPUs). Theremainis about three times slower than under macOS on the two-thread loops without work, and both commits beat it in every two-thread cell, by 31% to 52%, but four threads onRwLockread guards are 8% to 18% slower thanmainwith the first commit and 8% to 11% with the flag across three pinnings, and three threads 27% and 19%. I have no native ARM Linux host to tell how much of that is the VM. With either gap, no cell on any platform is more than 2% slower thanmain. The two-permit loops at four and eight threads park on nearly every acquisition and vary up to twofold between runs.Single-thread rows run
cargo x bench --bench primitives -- --sample-count 500 --sample-size 5000with name filters, median of 20 rounds, pinned to one core on Linux. Three rows are new so that each affected path has a direct measurement. The flag adds one relaxed load to the fast path.cargo x bench, nsrwlock::read::read_reuserwlock::read::eight_reads_then_one_writesemaphore::acquire::owned_try_acquire_releasesemaphore::acquire::try_acquire_release_with_spare_permitspool::bounded::bounded_warm_get_and_return_with_spare_capacitysemaphore::acquire::try_acquire_releasemutex::lock::uncontended_reuseOn the M5,
try_acquire_releaseis 0.3 ns (4%) slower with either commit: a single-permit release finds the balance at zero and reads it once more before taking the lock. The Hygon shows 3% the other way, andMutexskips that read throughrelease_all_held.Validation
cargo x check,cargo x test,cargo x lint, andcargo x mirion each commitcargo x mirinow runssemaphore_ordering_test