Split out of #3752, which was closed because its premise — "CIMarten and CIAzureServiceBus set the wall clock" — had gone stale. One observation in it did not go stale, and it needs a title that describes the real constraint rather than a wall-clock one.
Wall clock is fine. Runner-minutes are the constraint.
Measured on the last fully-completed main run, 32649119089 (c3e1c24d7, 34/34 green):
- 35 jobs, 211 runner-minutes, longest job 827s, median job 350s.
The 20-minute cap is comfortable — the longest job sits at 69% of it — and there is no pole: the top four are within 66 seconds of each other. Optimising any single job still buys ~nothing in wall clock, exactly as #3752 concluded.
But every merge costs 211 runner-minutes, and that is what bit today.
What happened on 2026-08-23
Five merges landed in about an hour (#4016, #4020, #4021, #4023, #4024, then #4025 and #4027). The queue reached:
- 26 runs queued, 0 in progress, for roughly two and a half hours.
- Of the queued runs, 11 were
tests.yml — about 2,300 runner-minutes of backlog.
Consequence: three merges to main were unverifiable for hours. main was confirmed green only through c3e1c24d7; the runs covering #4024, #4025 and #4027 had not started a single job. When a change merges before its own matrix finishes — which happened with #4024 — that verification gap is the only thing standing behind it.
This is a different failure mode from the one the CI effort has been optimising. Latency per run is healthy; throughput under concurrent merges is not, and no issue tracks it.
Where the minutes actually go
Same run, jobs by duration:
| job |
s |
|
job |
s |
| CIRabbitMQ |
827 |
|
CIRedis |
584 |
| CISqlServer |
790 |
|
CIPulsar |
477 |
| CIAzureServiceBus |
763 |
|
CIAzureServiceBusRouting |
460 |
| CIPubsub |
761 |
|
CIRavenDb |
443 |
| CIPersistence |
699 |
|
CIAzureServiceBusLeader |
422 |
| CIEfCore |
644 |
|
CIPolecat |
413 |
| CIKafka |
615 |
|
…21 more, 15s–384s |
|
The distribution is flat — no single job is more than 7% of the total — so there is nothing here to shard. Sharding moves minutes between jobs and adds them (#3818 added 11 runner-minutes of fixed per-job overhead: checkout, build, container warm-up). Against a runner-minutes budget, sharding is a cost, not a win. That inverts the recommendation that made sense when wall clock was the target.
The lead worth measuring
#3752's surviving observation: the shared compliance battery dominates the transport suites. Measured on Wolverine.Redis.Tests, two compliance classes held 61% of the suite's total test time (366.7s of 599.4s), and those same classes exist across eight transport suites. If that generalises, the compliance battery is a large fraction of the 211 minutes — and unlike sharding, shortening it removes minutes outright, from many jobs at once.
Also from that measurement: the Redis suite used 5% CPU over 7m38s. It is idle-dominated — waiting on brokers and polls — and ~37s of it was unconditional Task.Delay. Fixed sleeps are the kind of cost that polling would remove and make more reliable.
Suggested first step — measure, do not act
The per-class profile that exists was taken to answer a wall-clock question. Retake it against runner-minutes:
- For the top ~8 jobs by duration, sum per-class test time from a TRX and rank by total minutes contributed, not by "does this class set the floor".
- Report what fraction of the 211 minutes is the shared compliance battery, and what fraction is unconditional sleeps.
- Only then decide whether to shorten the battery, replace sleeps with polling, or accept the cost and raise concurrency instead.
Nothing here should be acted on before that measurement. The last time this was acted on from the wrong framing, it produced a sharding change that helped wall clock and hurt runner-minutes.
Explicitly not proposed
- Sharding anything further. See above.
- Cutting coverage. The compliance battery is what makes the transports comparable; the question is whether it costs what it should, not whether to run it.
Split out of #3752, which was closed because its premise — "CIMarten and CIAzureServiceBus set the wall clock" — had gone stale. One observation in it did not go stale, and it needs a title that describes the real constraint rather than a wall-clock one.
Wall clock is fine. Runner-minutes are the constraint.
Measured on the last fully-completed
mainrun, 32649119089 (c3e1c24d7, 34/34 green):The 20-minute cap is comfortable — the longest job sits at 69% of it — and there is no pole: the top four are within 66 seconds of each other. Optimising any single job still buys ~nothing in wall clock, exactly as #3752 concluded.
But every merge costs 211 runner-minutes, and that is what bit today.
What happened on 2026-08-23
Five merges landed in about an hour (#4016, #4020, #4021, #4023, #4024, then #4025 and #4027). The queue reached:
tests.yml— about 2,300 runner-minutes of backlog.Consequence: three merges to
mainwere unverifiable for hours.mainwas confirmed green only throughc3e1c24d7; the runs covering #4024, #4025 and #4027 had not started a single job. When a change merges before its own matrix finishes — which happened with #4024 — that verification gap is the only thing standing behind it.This is a different failure mode from the one the CI effort has been optimising. Latency per run is healthy; throughput under concurrent merges is not, and no issue tracks it.
Where the minutes actually go
Same run, jobs by duration:
The distribution is flat — no single job is more than 7% of the total — so there is nothing here to shard. Sharding moves minutes between jobs and adds them (#3818 added 11 runner-minutes of fixed per-job overhead: checkout, build, container warm-up). Against a runner-minutes budget, sharding is a cost, not a win. That inverts the recommendation that made sense when wall clock was the target.
The lead worth measuring
#3752's surviving observation: the shared compliance battery dominates the transport suites. Measured on
Wolverine.Redis.Tests, two compliance classes held 61% of the suite's total test time (366.7s of 599.4s), and those same classes exist across eight transport suites. If that generalises, the compliance battery is a large fraction of the 211 minutes — and unlike sharding, shortening it removes minutes outright, from many jobs at once.Also from that measurement: the Redis suite used 5% CPU over 7m38s. It is idle-dominated — waiting on brokers and polls — and ~37s of it was unconditional
Task.Delay. Fixed sleeps are the kind of cost that polling would remove and make more reliable.Suggested first step — measure, do not act
The per-class profile that exists was taken to answer a wall-clock question. Retake it against runner-minutes:
Nothing here should be acted on before that measurement. The last time this was acted on from the wrong framing, it produced a sharding change that helped wall clock and hurt runner-minutes.
Explicitly not proposed