Repository navigation
Explore ManualResetEvent waiter storage and state fast paths #252
Description
Activity
Research notes
Conclusion
Both ideas are viable, but they are not equally independent or equally cheap:
- An atomic signal bit is the best first experiment.
is_set()and an unregistered wait that observes the set state can use an acquire load without taking the waiter lock. The bit does not replace synchronization aroundset,reset, and waiter registration; those transitions should initially remain serialized by the waiter lock. WaitSetcan replaceWaitList, but not as a drop-in substitution. The event needs a newWaitSetoperation that turns a stale/drained token into a persistent "this wait was committed" result. Without that operation, the currentregisterbehavior would re-register a stale token after a rapidset/resetand lose the commitment made byset.- A fully lock-free transition path is possible, but comparable implementations need packed flags, waiter generations, and more elaborate per-future state. That does not look like the right first optimization without measurements showing transition contention.
1. Replacing
WaitListwithWaitSetThe event does not need
WaitListfor FIFO ordering. Its important property is that an unlinked node remains addressable:setmarks each waiternotified, removes its waker, and the future can still observe that notification afterreset.WaitSetalready has most of an alternative representation. A non-emptydrainadvances its epoch, so a future whose token belongs to the previous epoch can infer that its waker was drained. However, that fact is currently only used to prevent slot aliasing. It is not exposed as completion, and callingregisterwith such a token creates a new registration in the current epoch.A viable event-specific use therefore needs an API equivalent to
take_drained(&mut Option<WakerToken>) -> bool, called while holding theWaitSetlock. Event polling would use this order:- If an existing token belongs to an earlier epoch, clear it and return
Ready, regardless of the current signal bit. - If there is no token and the event is currently set, return
Ready. - Otherwise register or update the waker in the current epoch.
setmust change the signal state and drain the current epoch in one critical section;resetand the registration recheck must be ordered by that same lock. Waker clone, drop, and wake callbacks remain outside the lock.This would remove the event-specific
Waiter { notified, waker }, the unlink loop, and detached nodes that remain allocated until their futures are polled or dropped. A queuedWaitSetentry stores only aWaker, andsetcan release all arena slots immediately. The costs are less obvious at the future boundary: on a 64-bit target, the currentWaiterIdis one word whileWakerToken { epoch, slot }is two words. More importantly, the sharedWaitSetAPI gains a new semantic contract, and its checkedu64epoch increment introduces a theoretical overflow panic into a primitive intended for unlimited reuse. That overflow policy should be resolved explicitly before adopting this representation.My assessment: this is a credible simplification experiment, not an obvious win. It should be evaluated separately from the atomic-state change so code-size, future-size, waiter-memory, and fan-out effects remain attributable.
2. Making the signal state atomic
A simple and conservative shape is:
struct ManualResetEvent { is_set: AtomicBool, waiters: Mutex<WaitList<Waiter>>, // or the separately evaluated WaitSet design }
is_set()can use an acquire load. A wait with no existing registration can use the same load as its ready fast path. After observingfalseand cloning a waker outside the lock, it must acquire the waiter lock and recheck the atomic state before registering. The unset-to-set transition should publishtruewith release ordering.For the first version, both state-changing operations should still acquire the waiter lock and recheck the bit there. In particular, a design where
setstorestrue, obtains the lock later to drain, andresetindependently storesfalseis incorrect for the current contract:- an old
setstorestruebut has not drained yet; resetstoresfalse;- a new-generation wait registers;
- the old
setdrains and incorrectly commits that new wait.
Serializing the state transition, cohort detach, reset, and registration recheck avoids that ABA-style race while still making the two common observation paths lock-free. Asyncband's
CountdownStatealready demonstrates the atomic-probe plus lock-and-recheck pattern withAtomicU32andWaitSet, but countdown completion is monotonic; reusable reset is the additional constraint here.A packed state word is not needed merely to optimize
is_set. It becomes useful only if we also want a lock-freesetfast path with aHAS_WAITERSbit or encode a generation. That optimization has materially more proof and test surface and should follow evidence, not precede it.Ecosystem comparison
Implementation State and waiters Relevant tradeoff .NET ManualResetEventSlimPacks the signaled bit and waiter count into an integer; IsSetis a volatile read,Setuses an interlocked update and pulses a monitor when waiters exist.Strong precedent for an atomic observation path, but its Resetdocumentation explicitly disallows concurrent use with other members. Asyncband promises useful concurrentset/resetbehavior, so its transition algorithm cannot be copied directly.Microsoft's AsyncManualResetEventsketchEach generation is a TaskCompletionSource;setcompletes it andresetatomically swaps in a new one.Old tasks remain complete naturally after reset. The price is a new completion object per generation and continuation-scheduling concerns during set. AnArc<Generation>design would be the analogous Rust alternative.CPython threading.EventA boolean flag plus a condition lock; set/clearhold that lock, whileis_setdirectly reads the flag.Its waiter returns the condition notification result instead of re-reading the flag, so a later cleardoes not retract that notification. The direct Python read is not a Rust memory-ordering recipe, but the split between a cheap snapshot and locked transitions is relevant.CPython asyncio.EventA boolean plus a deque of per-waiter futures; setcompletes all queued futures andclearonly changes the boolean.This closely matches the committed-wait contract: completed waiter futures stay complete after clear. It can avoid thread synchronization because asyncio primitives are single-event-loop and explicitly not thread-safe.Rust events::ManualResetEventand its sourceUses an AtomicU8containingIS_SETandHAS_WAITERS, plus a mutex-protected intrusive/pinnedAwaiterSet, waiter notification state, and generations.It implements a directly comparable committed-wait contract and shows that more aggressive atomic transitions are possible. It also shows the amount of machinery required; this is substantially more complex than replacing the current mutex-protected boolean with AtomicBool.Tokio Notify/watchNotifystores at most one permit or wakes only the current cohort;watchretains only the latest value.Neither is a direct manual-reset event: Notifyhas different permit/cohort semantics, and a watcher may miss an intermediatetruefollowed byfalse.Suggested experiment order
- Prototype
AtomicBool + Mutex<WaitList<_>>, limiting the atomic tois_set, the no-token ready path, and publication of the set transition. Keep transitions and registration serialized by the waiter lock. - Add focused schedules for false-probe versus registration,
set/resetbefore repoll, cancellation before and after detach, and reentrant clone/drop/wake callbacks; preserve the existing publication test. - Benchmark
is_set, already-set waits, pending registration/cancellation, set/reset reuse, and fan-out against the merged implementation. - Separately prototype
WaitSetwith an explicit drained-token observation API and resolve the epoch-overflow/API question. Compare code complexity and both per-future and per-waiter memory, not only throughput. - Consider a packed atomic/generation design only if the first two experiments show that transition locking, rather than state observation or waiter storage, is a meaningful bottleneck.
So the two ideas should not be bundled initially. The atomic read path directly addresses a concrete cost with a small design delta;
WaitSetis an independent representation choice whose main expected benefit is simpler and smaller queued waiter state.- An atomic signal bit is the best first experiment.
- addedhelp wantedExtra attention is neededExtra attention is neededgood first issueGood for newcomersGood for newcomers
on Aug 31, 2026
Background
PR #243 added
ManualResetEventwith a mutex-protected boolean state andWaitList. Its approval review left the internal representation as a follow-up exploration: the ideas may improve the implementation, but keeping the current design is valid if the alternatives do not preserve its contracts or reduce overall cost.Questions to explore
ManualResetEventreplaceWaitListand its per-waiternotifiedstate withWaitSetwhile preserving the rule that every wait registered beforesetremains committed even ifresethappens before its next poll?is_setand the already-set wait path use an atomic boolean, or another compact atomic state, instead of acquiring the state mutex on every observation?Requirements
set, andreset;setfollowed byreset;is_set, already-set waits, pending registration/cancellation, set/reset reuse, and fan-out;Follow-up to #221 and #243. The originating approval review is here.