Skip to content

CI wall clock is set by CIMarten and CIAzureServiceBus; compliance classes dominate transport suites #3752

Description

@jeremydmiller

Where the CI wall clock actually goes

The tests.yml matrix runs ~30 jobs concurrently, so the workflow's wall clock is the slowest single job, not the total. From run 30634981923:

Job Duration
CIMarten 18m
CIAzureServiceBus 18m
CISqlServer 14m
CIPubsub 12m
CIPersistence 12m
CIKafka 12m
CIRabbitMQ 11m
CIEfCore 11m
CIRedis 9m
CIRavenDb 9m
CICosmosDb 6m
CIPolecat / CICircuitBreaking / CIAWSSqs 5m
CIPulsar 3m

Two jobs set the number. Any work on a job below ~14m changes the workflow's wall clock by zero. That is worth stating explicitly because it is easy to optimise the suite you happen to be looking at.

What a measurement of one suite showed

Wolverine.Redis.Tests (144 tests) was measured in detail as a pilot — running it across worker processes, partitioned by test class:

Workers Wall Speed-up Result
1 451.8s 1.00x 144/144
2 243.1s 1.86x 144/144
4 205.8s 2.20x 144/144
8 199.0s 2.27x 144/144

Green at every size, with retries disabled so nothing could mask instability. But it plateaus hard at ~199s, and the reason is the useful part:

Class Total Tests
BufferedSendingAndReceivingCompliance 188.1s 22
InlineSendingAndReceivingCompliance 178.6s 22
everything else (29 classes) 233s 100

Two compliance classes hold 61% of the suite's total test time (366.7s of 599.4s). Since a test class cannot be split across processes without breaking its isolation contract, the largest class is a hard floor: 599.4 / 188.1 = 3.19x is the ceiling, and no amount of parallelism goes below ~188s.

That floor is reachable from a single run — sum the per-test durations, divide by the largest class's total. It predicted 188s; the four-run sweep measured 199s. Within 6%. So the "how much would parallelism buy this job" question is answerable per suite without a bisection.

Why this probably generalises to the jobs that matter

BufferedSendingAndReceivingCompliance / InlineSendingAndReceivingCompliance are not Redis-specific — they derive from the shared Wolverine.ComplianceTests base, and the same classes exist across the transports:

  • Wolverine.AzureServiceBus.Tests — 6 compliance classes (InlineSendingAndReceivingCompliance, BufferedSendingAndReceivingCompliance, sending_compliance_with_prefixes, two topic/subscription suites, …)
  • Wolverine.RabbitMQ.Tests — 6 (durable_compliance, quorum_queue_compliance, stream_queue_compliance, …)
  • Wolverine.Pubsub.Tests — 4
  • Wolverine.Kafka.Tests, Wolverine.Nats.Tests — 1 each

CIAzureServiceBus is an 18m long pole and has six of them. If its time distribution looks like Redis's, it is both the most valuable job to attack and a better-shaped one — six large partitions parallelise further than two do.

This is unmeasured. The claim here is only that the shape is worth checking on CIAzureServiceBus and CIMarten first, because that is where the wall clock is.

Possible directions, roughly in order of cost

  1. Measure CIAzureServiceBus and CIMarten the way Redis was measured: per-class time totals and the sum / largest-class ceiling. One run each. This decides whether anything below is worth doing.
  2. Shard the long poles across CI jobs, the way CIPolecat / CIPolecatWorkflow / CIPolecatSagas already shard one project by namespace (CI: CIAWS and CIPolecat jobs chronically hit the 20-minute execution timeout #3350). Needs no new tooling, uses the existing testFilter parameter on RunTestProject, and turns one 18m job into several shorter concurrent ones. Most likely the cheapest real win.
  3. Attack the compliance base itself. If two classes are 61% of a transport suite, and the same classes exist in eight transport suites, then shortening the compliance tests shortens many jobs at once. The slowest individual tests in the Redis run were 12–15s each (schedule_send, can_stop_receiving_when_too_busy_and_restart_listeners, can_schedule_retry) — worth checking whether those are inherently slow or waiting on fixed timeouts.
  4. Multi-process parallelism within a job. Real (2.27x measured on Redis) but the largest-class floor limits it, and sharding across CI jobs gets much of the same benefit with machinery that already exists here.

Caveats on the numbers

  • Measured on an Apple Silicon dev machine, not a CI runner. Absolute times will differ.
  • The Redis suite is idle-dominated — the sequential run used 5% CPU over 7m38s — so it is waiting on brokers and polls, not computing. That is why parallelism helped despite limited cores, and it is a reason to expect similar behaviour on a CI runner. It may not hold for CPU-bound suites.
  • ~37s of the Redis suite is unconditional Task.Delay (15s in ScheduledMessageTests, 6.5s in RetryLimitTests, 8s in DeadLetterQueueTests). Only ~8% of the run, so sleeps are not the main cost here — but they are the kind of cost that replacing with polling would remove and make more reliable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions