You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The tests.yml matrix runs ~30 jobs concurrently, so the workflow's wall clock is the slowest single job, not the total. From run 30634981923:
Job
Duration
CIMarten
18m
CIAzureServiceBus
18m
CISqlServer
14m
CIPubsub
12m
CIPersistence
12m
CIKafka
12m
CIRabbitMQ
11m
CIEfCore
11m
CIRedis
9m
CIRavenDb
9m
CICosmosDb
6m
CIPolecat / CICircuitBreaking / CIAWSSqs
5m
CIPulsar
3m
Two jobs set the number. Any work on a job below ~14m changes the workflow's wall clock by zero. That is worth stating explicitly because it is easy to optimise the suite you happen to be looking at.
What a measurement of one suite showed
Wolverine.Redis.Tests (144 tests) was measured in detail as a pilot — running it across worker processes, partitioned by test class:
Workers
Wall
Speed-up
Result
1
451.8s
1.00x
144/144
2
243.1s
1.86x
144/144
4
205.8s
2.20x
144/144
8
199.0s
2.27x
144/144
Green at every size, with retries disabled so nothing could mask instability. But it plateaus hard at ~199s, and the reason is the useful part:
Class
Total
Tests
BufferedSendingAndReceivingCompliance
188.1s
22
InlineSendingAndReceivingCompliance
178.6s
22
everything else (29 classes)
233s
100
Two compliance classes hold 61% of the suite's total test time (366.7s of 599.4s). Since a test class cannot be split across processes without breaking its isolation contract, the largest class is a hard floor: 599.4 / 188.1 = 3.19x is the ceiling, and no amount of parallelism goes below ~188s.
That floor is reachable from a single run — sum the per-test durations, divide by the largest class's total. It predicted 188s; the four-run sweep measured 199s. Within 6%. So the "how much would parallelism buy this job" question is answerable per suite without a bisection.
Why this probably generalises to the jobs that matter
BufferedSendingAndReceivingCompliance / InlineSendingAndReceivingCompliance are not Redis-specific — they derive from the shared Wolverine.ComplianceTests base, and the same classes exist across the transports:
Wolverine.Kafka.Tests, Wolverine.Nats.Tests — 1 each
CIAzureServiceBus is an 18m long pole and has six of them. If its time distribution looks like Redis's, it is both the most valuable job to attack and a better-shaped one — six large partitions parallelise further than two do.
This is unmeasured. The claim here is only that the shape is worth checking on CIAzureServiceBus and CIMarten first, because that is where the wall clock is.
Possible directions, roughly in order of cost
Measure CIAzureServiceBus and CIMarten the way Redis was measured: per-class time totals and the sum / largest-class ceiling. One run each. This decides whether anything below is worth doing.
Shard the long poles across CI jobs, the way CIPolecat / CIPolecatWorkflow / CIPolecatSagas already shard one project by namespace (CI: CIAWS and CIPolecat jobs chronically hit the 20-minute execution timeout #3350). Needs no new tooling, uses the existing testFilter parameter on RunTestProject, and turns one 18m job into several shorter concurrent ones. Most likely the cheapest real win.
Attack the compliance base itself. If two classes are 61% of a transport suite, and the same classes exist in eight transport suites, then shortening the compliance tests shortens many jobs at once. The slowest individual tests in the Redis run were 12–15s each (schedule_send, can_stop_receiving_when_too_busy_and_restart_listeners, can_schedule_retry) — worth checking whether those are inherently slow or waiting on fixed timeouts.
Multi-process parallelism within a job. Real (2.27x measured on Redis) but the largest-class floor limits it, and sharding across CI jobs gets much of the same benefit with machinery that already exists here.
Caveats on the numbers
Measured on an Apple Silicon dev machine, not a CI runner. Absolute times will differ.
The Redis suite is idle-dominated — the sequential run used 5% CPU over 7m38s — so it is waiting on brokers and polls, not computing. That is why parallelism helped despite limited cores, and it is a reason to expect similar behaviour on a CI runner. It may not hold for CPU-bound suites.
~37s of the Redis suite is unconditional Task.Delay (15s in ScheduledMessageTests, 6.5s in RetryLimitTests, 8s in DeadLetterQueueTests). Only ~8% of the run, so sleeps are not the main cost here — but they are the kind of cost that replacing with polling would remove and make more reliable.
Where the CI wall clock actually goes
The
tests.ymlmatrix runs ~30 jobs concurrently, so the workflow's wall clock is the slowest single job, not the total. From run 30634981923:Two jobs set the number. Any work on a job below ~14m changes the workflow's wall clock by zero. That is worth stating explicitly because it is easy to optimise the suite you happen to be looking at.
What a measurement of one suite showed
Wolverine.Redis.Tests(144 tests) was measured in detail as a pilot — running it across worker processes, partitioned by test class:Green at every size, with retries disabled so nothing could mask instability. But it plateaus hard at ~199s, and the reason is the useful part:
BufferedSendingAndReceivingComplianceInlineSendingAndReceivingComplianceTwo compliance classes hold 61% of the suite's total test time (366.7s of 599.4s). Since a test class cannot be split across processes without breaking its isolation contract, the largest class is a hard floor:
599.4 / 188.1 = 3.19xis the ceiling, and no amount of parallelism goes below ~188s.That floor is reachable from a single run — sum the per-test durations, divide by the largest class's total. It predicted 188s; the four-run sweep measured 199s. Within 6%. So the "how much would parallelism buy this job" question is answerable per suite without a bisection.
Why this probably generalises to the jobs that matter
BufferedSendingAndReceivingCompliance/InlineSendingAndReceivingComplianceare not Redis-specific — they derive from the sharedWolverine.ComplianceTestsbase, and the same classes exist across the transports:Wolverine.AzureServiceBus.Tests— 6 compliance classes (InlineSendingAndReceivingCompliance,BufferedSendingAndReceivingCompliance,sending_compliance_with_prefixes, two topic/subscription suites, …)Wolverine.RabbitMQ.Tests— 6 (durable_compliance,quorum_queue_compliance,stream_queue_compliance, …)Wolverine.Pubsub.Tests— 4Wolverine.Kafka.Tests,Wolverine.Nats.Tests— 1 eachCIAzureServiceBus is an 18m long pole and has six of them. If its time distribution looks like Redis's, it is both the most valuable job to attack and a better-shaped one — six large partitions parallelise further than two do.
This is unmeasured. The claim here is only that the shape is worth checking on CIAzureServiceBus and CIMarten first, because that is where the wall clock is.
Possible directions, roughly in order of cost
sum / largest-classceiling. One run each. This decides whether anything below is worth doing.CIPolecat/CIPolecatWorkflow/CIPolecatSagasalready shard one project by namespace (CI: CIAWS and CIPolecat jobs chronically hit the 20-minute execution timeout #3350). Needs no new tooling, uses the existingtestFilterparameter onRunTestProject, and turns one 18m job into several shorter concurrent ones. Most likely the cheapest real win.schedule_send,can_stop_receiving_when_too_busy_and_restart_listeners,can_schedule_retry) — worth checking whether those are inherently slow or waiting on fixed timeouts.Caveats on the numbers
Task.Delay(15s inScheduledMessageTests, 6.5s inRetryLimitTests, 8s inDeadLetterQueueTests). Only ~8% of the run, so sleeps are not the main cost here — but they are the kind of cost that replacing with polling would remove and make more reliable.