Version: WolverineFx.CosmosDb 6.39.1 (same on main).
CosmosDbMessageStore.MarkIncomingEnvelopeAsHandledAsync stamps a keepUntil, but nothing in the provider ever reads it back:
message.Status = EnvelopeStatus.Handled;
message.KeepUntil = DateTimeOffset.UtcNow.Add(_options.Durability.KeepAfterMessageHandling);
CosmosDbDurabilityAgent's only sweep is tryDeleteExpiredDeadLetters(). There is no equivalent of the RDBMS DeleteExpiredHandledEnvelopesCommand, and no ttl is written (container TTL is off in any case — MigrateAsync creates the container with default ContainerProperties). So handled inbox documents are kept forever, bodies included, since mark-handled is a read/modify/replace that leaves body alone. RavenDb doesn't have this problem: it sets @metadata.@expires when marking handled.
Why that bites harder on Cosmos than on a SQL store: IncomingMessage.PartitionKey is envelope.Destination, so every handled envelope for a listening endpoint piles into one logical partition — 20 GB hard ceiling, 10k RU/s cap, i.e. the same failure shape as #3415 for sagas. FetchCountsAsync, which CheckHealthAsync calls on every health check, also counts that pile cross-partition.
Nor is there a way to clean up after the fact: DeleteAllHandledAsync() throws NotSupportedException("This function is not yet supported by CosmosDb"), so the built-in clear-handled command fails on this provider. We run an out-of-band sweeper over the wolverine container instead.
Worth noting too that DurabilitySettings.HandledMessageCleanupPollingTime/BatchSize/MaxBatchesPerCycle read as general durability knobs but are referenced only from Wolverine.RDBMS.DurabilityAgent, so on Cosmos they do nothing.
The fix looks like a tryDeleteExpiredHandledEnvelopes mirroring the dead-letter sweep — docType = 'incoming' AND status = @handled AND keepUntil < @now, paged and throttled — plus DeleteAllHandledAsync implemented with the same query minus the keepUntil predicate. Happy to send a PR with a regression test next to Bug_4286_dead_letter_expiration if that's useful.
Version:
WolverineFx.CosmosDb6.39.1 (same onmain).CosmosDbMessageStore.MarkIncomingEnvelopeAsHandledAsyncstamps akeepUntil, but nothing in the provider ever reads it back:CosmosDbDurabilityAgent's only sweep istryDeleteExpiredDeadLetters(). There is no equivalent of the RDBMSDeleteExpiredHandledEnvelopesCommand, and nottlis written (container TTL is off in any case —MigrateAsynccreates the container with defaultContainerProperties). So handled inbox documents are kept forever, bodies included, since mark-handled is a read/modify/replace that leavesbodyalone. RavenDb doesn't have this problem: it sets@metadata.@expireswhen marking handled.Why that bites harder on Cosmos than on a SQL store:
IncomingMessage.PartitionKeyisenvelope.Destination, so every handled envelope for a listening endpoint piles into one logical partition — 20 GB hard ceiling, 10k RU/s cap, i.e. the same failure shape as #3415 for sagas.FetchCountsAsync, whichCheckHealthAsynccalls on every health check, also counts that pile cross-partition.Nor is there a way to clean up after the fact:
DeleteAllHandledAsync()throwsNotSupportedException("This function is not yet supported by CosmosDb"), so the built-inclear-handledcommand fails on this provider. We run an out-of-band sweeper over thewolverinecontainer instead.Worth noting too that
DurabilitySettings.HandledMessageCleanupPollingTime/BatchSize/MaxBatchesPerCycleread as general durability knobs but are referenced only fromWolverine.RDBMS.DurabilityAgent, so on Cosmos they do nothing.The fix looks like a
tryDeleteExpiredHandledEnvelopesmirroring the dead-letter sweep —docType = 'incoming' AND status = @handled AND keepUntil < @now, paged and throttled — plusDeleteAllHandledAsyncimplemented with the same query minus thekeepUntilpredicate. Happy to send a PR with a regression test next toBug_4286_dead_letter_expirationif that's useful.