Found while investigating why Bug_189_fails_if_there_are_many_messages_in_queue_on_startup has been a chronic CIRabbitMQ failure. The test was the messenger; the finding underneath is worth its own issue.
What happens
When the broker rejects a basic.ack — PRECONDITION_FAILED - unknown delivery tag — it closes that channel. If deliveries are still in flight on the channel being torn down, RabbitMQ.Client races itself:
System.Collections.Generic.Dictionary`2.get_Item(TKey key)
RabbitMQ.Client.Impl.SessionManager.Lookup(Int32 number)
RabbitMQ.Client.Framing.Connection.ProcessFrameAsync(InboundFrame frame, CancellationToken)
RabbitMQ.Client.Framing.Connection.ReceiveLoopAsync(CancellationToken)
RabbitMQ.Client.Framing.Connection.MainLoop()
---- System.Collections.Generic.KeyNotFoundException : The given key '4' was not present in the dictionary.
An inbound frame arrives for the channel number that was just removed from the session map, Lookup throws, and the client escalates that into a library-initiated close of the entire connection:
AlreadyClosedException : Already closed: The AMQP operation was interrupted:
AMQP close-reason, initiated by Library, code=541,
text='The given key '4' was not present in the dictionary.'
So the blast radius of one bad delivery tag is not the channel — it is every listener and sender on that connection.
Wolverine cannot catch this. It is thrown on the client's own MainLoop thread, inside the receive loop, after the ack has already gone out. Wolverine's settle paths are already hardened for the tag rejection itself (RabbitMqChannelCallback.Complete and RabbitMqInteropFriendlyCallback.MoveToErrorsAsync both swallow AlreadyClosedException when the message matches PRECONDITION_FAILED - unknown delivery tag), and that hardening is not the gap — the connection dies regardless.
Measurements
All local, Wolverine.RabbitMQ.Tests, RabbitMQ from docker compose, client 7.1.2:
| Scenario |
KeyNotFoundException |
code=541 |
| Bug_189 as written (~5% of 1000 deliveries poisoned) |
every run (3/3) |
every run |
| Bug_189, poisoning removed entirely |
0/3 |
0/3 |
| Bug_189, exactly one poisoned tag |
5/5 runs |
4/5 runs |
5 messages, one poisoned tag (stale_delivery_tag_settling) |
0/5 |
0/5 |
The third row is the important one: a single rejected ack is enough, provided the channel has deliveries in flight. It is the in-flight traffic, not the number of bad acks, that decides whether the connection dies.
Same signature in CI on both a PR (run 31894510951) and on main (run 31741860971).
Why it matters beyond the test
A stale delivery tag is not purely synthetic. Wolverine already guards the case it knows about — RabbitMqListener.CanSettle refuses to settle against a channel that has been replaced, precisely because tags restart at 1 on a new channel. But any path that does reach the broker with a tag it no longer recognises will, on a busy channel, cost the whole connection rather than one message.
Suggestions
- Report upstream to rabbitmq-dotnet-client.
SessionManager.Lookup doing an indexer read on a dictionary it may have just mutated is the actual defect; a TryGetValue returning "frame for a dead channel, drop it" would contain the damage to the channel. This is the real fix and it is not ours.
- Confirm Wolverine recovers. Locally the connection dies and the affected tests still pass, so
ConnectionMonitor does appear to reconnect — but that is observed, not asserted anywhere. A test that kills a connection mid-flight and asserts listeners resume would turn an assumption into a guarantee.
- Consider whether a rejected settle should proactively quiesce the channel rather than leave in-flight deliveries racing the teardown.
Related
The chaos injection that surfaced this has been moved out of Bug_189 into a focused low-volume test — Bug_189 is about starting up against a full queue and does not need chaos to prove it.
Found while investigating why
Bug_189_fails_if_there_are_many_messages_in_queue_on_startuphas been a chronic CIRabbitMQ failure. The test was the messenger; the finding underneath is worth its own issue.What happens
When the broker rejects a
basic.ack—PRECONDITION_FAILED - unknown delivery tag— it closes that channel. If deliveries are still in flight on the channel being torn down, RabbitMQ.Client races itself:An inbound frame arrives for the channel number that was just removed from the session map,
Lookupthrows, and the client escalates that into a library-initiated close of the entire connection:So the blast radius of one bad delivery tag is not the channel — it is every listener and sender on that connection.
Wolverine cannot catch this. It is thrown on the client's own
MainLoopthread, inside the receive loop, after the ack has already gone out. Wolverine's settle paths are already hardened for the tag rejection itself (RabbitMqChannelCallback.CompleteandRabbitMqInteropFriendlyCallback.MoveToErrorsAsyncboth swallowAlreadyClosedExceptionwhen the message matchesPRECONDITION_FAILED - unknown delivery tag), and that hardening is not the gap — the connection dies regardless.Measurements
All local,
Wolverine.RabbitMQ.Tests, RabbitMQ fromdocker compose, client 7.1.2:KeyNotFoundExceptioncode=541stale_delivery_tag_settling)The third row is the important one: a single rejected ack is enough, provided the channel has deliveries in flight. It is the in-flight traffic, not the number of bad acks, that decides whether the connection dies.
Same signature in CI on both a PR (run 31894510951) and on
main(run 31741860971).Why it matters beyond the test
A stale delivery tag is not purely synthetic. Wolverine already guards the case it knows about —
RabbitMqListener.CanSettlerefuses to settle against a channel that has been replaced, precisely because tags restart at 1 on a new channel. But any path that does reach the broker with a tag it no longer recognises will, on a busy channel, cost the whole connection rather than one message.Suggestions
SessionManager.Lookupdoing an indexer read on a dictionary it may have just mutated is the actual defect; aTryGetValuereturning "frame for a dead channel, drop it" would contain the damage to the channel. This is the real fix and it is not ours.ConnectionMonitordoes appear to reconnect — but that is observed, not asserted anywhere. A test that kills a connection mid-flight and asserts listeners resume would turn an assumption into a guarantee.Related
The chaos injection that surfaced this has been moved out of Bug_189 into a focused low-volume test — Bug_189 is about starting up against a full queue and does not need chaos to prove it.