Found reviewing #3410. Pre-existing, affects C# and F# alike.
What's happening
The Cosmos saga frames read and write with PartitionKey.None. Saga documents carry no partitionKey property, and the container (DocumentTypes.cs:16) is keyed on /partitionKey.
To be fair to the current code: this is self-consistent. Read and write agree on the same "undefined" partition, so there is no cross-partition query and no silent miss — I checked, because that was my first worry. It is not a correctness bug.
Why it still matters
It means every saga in the application shares one logical partition. In Cosmos, a logical partition is hard-capped at 20 GB and 10,000 RU/s. So an app's entire saga workload is pinned to a single partition's throughput budget and storage ceiling, no matter how many physical partitions the container has.
For a low-volume saga workload nobody will ever notice. For anyone using sagas at the scale Cosmos is normally chosen for, it is a wall they will hit without any indication of why — the symptom is 429 throttling that doesn't improve when you scale the container up, because the constraint is per-logical-partition, not per-container.
Ask
Decide deliberately what the partition key for a saga document should be, rather than inheriting "none" by default. The saga id is the obvious candidate (it's the point-read key, so a saga-id partition key keeps every operation a single-partition point read, which is the ideal Cosmos access pattern). That would be a breaking storage change for anyone with existing Cosmos saga documents, so it needs a migration story or an opt-in.
At absolute minimum, document the current ceiling so it isn't discovered as a production 429 storm.
Refs #3410.
Found reviewing #3410. Pre-existing, affects C# and F# alike.
What's happening
The Cosmos saga frames read and write with
PartitionKey.None. Saga documents carry nopartitionKeyproperty, and the container (DocumentTypes.cs:16) is keyed on/partitionKey.To be fair to the current code: this is self-consistent. Read and write agree on the same "undefined" partition, so there is no cross-partition query and no silent miss — I checked, because that was my first worry. It is not a correctness bug.
Why it still matters
It means every saga in the application shares one logical partition. In Cosmos, a logical partition is hard-capped at 20 GB and 10,000 RU/s. So an app's entire saga workload is pinned to a single partition's throughput budget and storage ceiling, no matter how many physical partitions the container has.
For a low-volume saga workload nobody will ever notice. For anyone using sagas at the scale Cosmos is normally chosen for, it is a wall they will hit without any indication of why — the symptom is 429 throttling that doesn't improve when you scale the container up, because the constraint is per-logical-partition, not per-container.
Ask
Decide deliberately what the partition key for a saga document should be, rather than inheriting "none" by default. The saga id is the obvious candidate (it's the point-read key, so a saga-id partition key keeps every operation a single-partition point read, which is the ideal Cosmos access pattern). That would be a breaking storage change for anyone with existing Cosmos saga documents, so it needs a migration story or an opt-in.
At absolute minimum, document the current ceiling so it isn't discovered as a production 429 storm.
Refs #3410.