Skip to content

perf(ci): run TXE-backed noir tests in chunks, not one nargo process per test - #307

Draft
fcarreiro wants to merge 2 commits into
mainfrom
fc/batch-noir-txe-tests
Draft

fcarreiro wants to merge 2 commits into
mainfrom
fc/batch-noir-txe-tests

Conversation

@fcarreiro

@fcarreiro fcarreiro commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Runs the TXE-backed noir tests (aztec-nr and noir-contracts) as 45 chunks of up to 64 tests each, one nargo process per chunk, instead of 1,040 nargo processes with one test each. Also raises CI's TXE count from 1 to 4.

Why

Each nargo test --exact <test> process parses and type-checks the whole package before running its single test. That costs ~13s of the ~15s average per test (a trivial test that makes no oracle calls takes 12.9s locally). On a PR that touches anything in the TXE's import closure, every one of these tests reruns: on a recent PR run (1790153506230403) they took 13,779 CPU-s of the test engine, ~28% of its total, and ran at the end of the run after the build. In a single process nargo elaborates the package once per test thread, not once per test.

What changes

  • noir-projects/scripts/test_chunks.sh groups nargo test --list-tests output by (package, kind), where kind is oracle for __oracle_test__ tests and txe for the rest, and splits each group into as few chunks as keep each at or under 64 tests.
  • noir-projects/scripts/run_test_chunk.sh lists the package's tests, selects the chunk's slice (sorted, dealt round-robin), and runs them in one nargo test --exact ... --test-threads min(CPUS, n) process. CPUS is the test engine's per-command budget (default 2, with one command per 2 CPUs), so a chunk stays within the CPU share a single-test command already had. It fails if the slice is empty.
  • test_cmds in aztec-nr and noir-contracts emit one command per chunk. Cache keys are unchanged: they were already shared per package (noir-contracts) or across all tests (aztec-nr). They now also cover noir-projects/scripts/.
  • Every command carries the TXE port range as <base_port> <num_ports>, and the runner picks a TXE at random. The command line is the cache key, so a per-chunk port assignment would re-key chunks whenever the chunk list changes.
  • NUM_TXES=4 in the root bootstrap. Batched runs load the TXE much more heavily than per-test runs. With one TXE the batched run peaked at ~4.8 cores in the TXE and was TXE-bound (279s vs 223s with 4 TXEs). The CI box has ~430 GB free at its lowest point, and a TXE peaked at ~7 GB RSS.

Measurements

Both layouts run through ci3/parallelize 16 → exec_test on the same 32 cores (taskset and CPU_LIST), with the TXEs and resolver pinned to those cores too, so neither layout gets more CPU than the other:

Layout Wall nargo CPU TXE + resolver CPU Busy cores (avg / p90 / max of 32)
Per test, 1 TXE (main) 836s 12,454s 543s 18.1 / 21.8 / 28.7
Chunked, 2 threads, 4 TXEs (this PR), two runs 132s / 160s 1,160s / 1,166s 508s / 730s 15.7-18.9 / 25-27 / 29

At equal budget the chunked layout is 5-6x faster and does ~10x less nargo work, with the same peak core usage. Its critical path is the largest contract packages, each a single chunk (amm_contract and token_contract, ~105-120s). Most chunks finish well before them.

In earlier exploratory runs on 64 cores, plain per-package runs (no chunking) were limited by noir_aztec (623 TXE tests in one job, 222s even with 4 TXEs), which is why chunks are capped at 64 tests. With a single TXE, batched runs were TXE-bound.

Running the runner against a stub NARGO shows the chunks select exactly the 1,041 tests nargo test --list-tests reports, each once, with oracle tests only in oracle chunks.

Trade-offs

  • A flake retry reruns a whole chunk (up to ~120s) instead of one test (~15s).
  • 4 TXEs is a fixed count. It only applies in the ci-* modes (build_and_test); the subproject ./bootstrap.sh test paths still start one TXE. A TXE peaked at ~7 GB RSS here. No .test_patterns.yml entry targets a noir test today.
  • Test history and failure reports are per chunk. The failing test's name is in the chunk's log, not in the command.
  • Tests in a chunk share a package elaboration per thread (see nargo test --no-context-reuse), so they are less isolated than one-process-per-test. ./bootstrap.sh test-one still runs a single test in its own process.
  • Chunk membership shifts when tests are added to a group of more than one chunk. That only affects history, since every chunk of a group shares one cache key.

…per test

Each `nargo test --exact <test>` process elaborates the whole package before
running its one test, which is most of the ~15s every aztec-nr and
noir-contracts test takes. Run the tests as chunks of up to 64 per
(package, kind) in one nargo process each, so the package is elaborated once
per test thread instead of once per test: 45 commands instead of 1,040.

Batched runs load the TXE far more heavily, so CI runs 4 TXEs. Commands carry
the TXE port range and the runner picks a TXE at random, keeping the port out
of the cache key. The chunking scripts are part of the test hashes.
…CPU budget

Run each chunk on CPUS test threads, which the test engine defaults to 2 with
one command per 2 CPUs, instead of 4. A chunk then uses no more CPU than the
single-test command it replaces, so the chunked layout does not rely on cores
the engine does not budget for. On an equal 32-core budget it still runs the
suite in 132-160s against 836s for one process per test.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant