Skip to content

CI throughput: 211 runner-minutes per run makes concurrent merges unverifiable for hours #4033

Description

@jeremydmiller

Split out of #3752, which was closed because its premise — "CIMarten and CIAzureServiceBus set the wall clock" — had gone stale. One observation in it did not go stale, and it needs a title that describes the real constraint rather than a wall-clock one.

Wall clock is fine. Runner-minutes are the constraint.

Measured on the last fully-completed main run, 32649119089 (c3e1c24d7, 34/34 green):

  • 35 jobs, 211 runner-minutes, longest job 827s, median job 350s.

The 20-minute cap is comfortable — the longest job sits at 69% of it — and there is no pole: the top four are within 66 seconds of each other. Optimising any single job still buys ~nothing in wall clock, exactly as #3752 concluded.

But every merge costs 211 runner-minutes, and that is what bit today.

What happened on 2026-08-23

Five merges landed in about an hour (#4016, #4020, #4021, #4023, #4024, then #4025 and #4027). The queue reached:

  • 26 runs queued, 0 in progress, for roughly two and a half hours.
  • Of the queued runs, 11 were tests.yml — about 2,300 runner-minutes of backlog.

Consequence: three merges to main were unverifiable for hours. main was confirmed green only through c3e1c24d7; the runs covering #4024, #4025 and #4027 had not started a single job. When a change merges before its own matrix finishes — which happened with #4024 — that verification gap is the only thing standing behind it.

This is a different failure mode from the one the CI effort has been optimising. Latency per run is healthy; throughput under concurrent merges is not, and no issue tracks it.

Where the minutes actually go

Same run, jobs by duration:

job s job s
CIRabbitMQ 827 CIRedis 584
CISqlServer 790 CIPulsar 477
CIAzureServiceBus 763 CIAzureServiceBusRouting 460
CIPubsub 761 CIRavenDb 443
CIPersistence 699 CIAzureServiceBusLeader 422
CIEfCore 644 CIPolecat 413
CIKafka 615 …21 more, 15s–384s

The distribution is flat — no single job is more than 7% of the total — so there is nothing here to shard. Sharding moves minutes between jobs and adds them (#3818 added 11 runner-minutes of fixed per-job overhead: checkout, build, container warm-up). Against a runner-minutes budget, sharding is a cost, not a win. That inverts the recommendation that made sense when wall clock was the target.

The lead worth measuring

#3752's surviving observation: the shared compliance battery dominates the transport suites. Measured on Wolverine.Redis.Tests, two compliance classes held 61% of the suite's total test time (366.7s of 599.4s), and those same classes exist across eight transport suites. If that generalises, the compliance battery is a large fraction of the 211 minutes — and unlike sharding, shortening it removes minutes outright, from many jobs at once.

Also from that measurement: the Redis suite used 5% CPU over 7m38s. It is idle-dominated — waiting on brokers and polls — and ~37s of it was unconditional Task.Delay. Fixed sleeps are the kind of cost that polling would remove and make more reliable.

Suggested first step — measure, do not act

The per-class profile that exists was taken to answer a wall-clock question. Retake it against runner-minutes:

  1. For the top ~8 jobs by duration, sum per-class test time from a TRX and rank by total minutes contributed, not by "does this class set the floor".
  2. Report what fraction of the 211 minutes is the shared compliance battery, and what fraction is unconditional sleeps.
  3. Only then decide whether to shorten the battery, replace sleeps with polling, or accept the cost and raise concurrency instead.

Nothing here should be acted on before that measurement. The last time this was acted on from the wrong framing, it produced a sharding change that helped wall clock and hurt runner-minutes.

Explicitly not proposed

  • Sharding anything further. See above.
  • Cutting coverage. The compliance battery is what makes the transports comparable; the question is whether it costs what it should, not whether to run it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    wontfixThis will not be worked on

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions