Skip to content

perf(api,bench): derive concurrency from the machine; drop the adjacent-RSS growth gate - #1466

Merged
DecisionNerd merged 3 commits into
mainfrom
perf/unpin-compute-drop-rss-growth-gate
Sep 19, 2026
Merged

DecisionNerd merged 3 commits into
mainfrom
perf/unpin-compute-drop-rss-growth-gate

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Two independent commits that unblock the measurement work in plan rev 8. Both were sitting on a local-only branch; neither had been through CI.

1287a2f8 — derive instance concurrency from the machine (#1463)

ExecutionResourcePolicy::default() pinned a fixed two-worker facade — tokio_worker_threads: Some(2), target_partitions: Some(2), io_concurrency: Some(2), compute_threads: Some(2) — preserving pre-#337 behaviour. On an 8-core / 16-thread host that is simply wrong, and every default-constructed instance took two workers whatever the machine had.

The default is now mode: Automatic with those four knobs None, deferring to the mode. Explicit policies are untouched.

Why it matters beyond the obvious: the G500 ladder runs at 0.85–0.89 effective cores across a 64x edge range on a 16-thread host. Until this lands, a serial-fraction instrument (#1462) measures the facade rather than the engine, so it is a precondition for the whole phase-0 measurement step, not merely independent cleanup.

Both pinning tests were rewritten to assert the invariant rather than the observation — defaults_preserve_fixed_two_worker_baseline becomes defaults_derive_concurrency_from_machine_parallelism, checking the derived value instead of the constant 2.

5bbb8a71 — drop the adjacent-RSS growth gate, keep the 4 GiB limit

Removes rss_bounded_or_plateaued from admission and max_adjacent_rss_growth_fraction from the harness, all seven G500 profiles and the profile schema. rss_growth_fraction is still measured and reported in the evidence, and the qualification test is rewritten to assert exactly that, so the architectural signal survives the gate's removal. RSS is governed by the absolute 4 GiB rung limit alone.

Rationale, recorded on #1387:

  1. Headroom. S25 projected 757 MiB against a 4 GiB rung limit on a 125 GB host. A 10% adjacent-rung growth ceiling is not measuring a resource risk at that ratio.
  2. It asserted an observation, not a property. What matters is that memory stays inside budget and does not grow without bound.
  3. It is incompatible with the planned refactor. epic(storage): scale complete ingest across cores and reach 1M edges/s #1387's own 4 GiB decision records that the redesign's parallel shard lanes each need a sort buffer, and that a comparable DataFusion change measured +52%. A 10% adjacent-growth ceiling would refuse epic(storage): stop maintaining a hand-rolled dataflow engine; reuse Arrow/DataFusion where determinism and durability allow #1456 for working as designed.

What is given up: this was the only check refusing "RSS scales with edge count rather than with the streaming window" before it becomes an absolute problem, and #1387 records that shape as binding at S28/S30. It will now surface as a breached limit at a higher rung instead of a refused admission at a lower one. The fraction stays in the evidence, so it should be watched rather than assumed. #1439 keeps its place on the other merit — it remains the precondition for any SortExec adoption.

Consequence for the ladder

S25 admission had two failing checks. With this one gone, io_reader_publication_headroom is the only remaining refusal, so the ladder is now gated on byte amplification alone.

Testing

  • cargo test -p graphforge-api --lib — 744 passed, 0 failed, 2 ignored.
  • Benchmark harness, CI's invocation (PYTHONPATH=harness uv run --locked python -m unittest) across test_progressive_qualification, test_progressive_run, test_progressive_provider_plan, test_progressive_provider_run, test_native_ladder_controller, test_lifecycle_runtime — 97 passed.
  • TCK (cargo test -p graphforge-api --test bdd) — API BDD 118 passed, 0 failed; openCypher TCK 3897/3897, 0 regressed, 0 xpass.

On the TCK perf warning

The run emits TCK PERF WARNING: openCypher TCK total: 141,248 ms (baseline 66,505 ms, threshold 83,131 ms). Since unpinning worker threads could plausibly oversubscribe a parallel test harness, it was measured both ways on the same quiet host:

openCypher TCK total
with 1287a2f8 (machine-derived) 141,248 ms
with the two files reverted (fixed 2 workers) 139,752 ms
difference +1.1%, noise

The overrun is pre-existing and not caused by this change. The recorded 66,505 ms baseline was captured on a GitHub-hosted runner (runner = ubuntu-latest in tests/tck/performance_baseline.json) and the comparison never checks that, so it warns on every local run. Filed as #1467, with the two per-scenario outliers (29.7x and 20.6x against a 2.1x aggregate) flagged there as probably real and not to be absorbed into a refreshed baseline.

Note for anyone reproducing locally: the graphforge-api suite fails 243 tests under the default TMPDIR on this host because /tmp is tmpfs and filesystem admission fail-closes with UnsupportedFilesystem: phase=CLASSIFY cause=filesystem_class_unproven. Point TMPDIR at an ext4 path and it is green. That is environmental, unrelated to these commits.

Closes #1463

🤖 Generated with Claude Code


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

DecisionNerd and others added 2 commits September 18, 2026 02:41
…fixed pair

`ExecutionResourcePolicy::default()` requested `Explicit` mode with
`tokio_worker_threads`, `target_partitions`, `io_concurrency` and
`compute_threads` all hard-pinned to `2`, commented as preserving the
pre-#337 two-worker facade. Every default-constructed instance therefore
took two workers whatever the host had.

Measured consequence on the G500 ladder (`f80f69fe`, 8 cores / 16 threads):

| rung | edges | wall | CPU-s | effective cores |
|------|-------|------|-------|-----------------|
| S18  | 4.19M |   69 |  58.5 | 0.85 |
| S20  | 16.8M |  270 | 231.5 | 0.86 |
| S22  | 67.1M | 1088 | 960.8 | 0.88 |
| S24  |  268M | 4725 |4201.5 | 0.89 |

Flat across a 64x edge range: wall time is CPU time and fifteen threads
idle. Default to `Automatic`, which derives a bounded count from observed
parallelism (`observed.div_ceil(2).clamp(MIN_THREADS, 8)`) — 8 here rather
than 2, still bounded, still fail-closed against the combined-budget and
reserve caps. Callers passing explicit knobs are unaffected.

The `<= 2` CPU machines stay serial by the existing Automatic rule, so
small hosts keep the low-overhead path.

Refs #1387

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…iB limit

`rss_bounded_or_plateaued` refused a rung whenever peak RSS grew more than
10% between the two projection sources. It refused S25 on the `f80f69fe`
ladder at a growth fraction of 1.5613 — while absolute RSS was 470 MiB at
S24 and 757 MiB projected at S25, against a 4 GiB limit that passed with
over four times' margin on a 125 GB host.

The absolute budget is the gate that governs memory now. `RSS_LIMIT_BYTES`
stays at 4 GiB and `rss_headroom` is unchanged.

`rss_growth_fraction` is still computed and still recorded in the evidence
document, so the architectural signal the gate was watching for survives
its removal as an observation rather than a refusal.

Removed from: the harness check set, both schemas, the certify gate
assertion and its `ProgressiveHeadroom` field (`deny_unknown_fields`, so
the struct and the profile JSONs must agree), and the
`max_adjacent_rss_growth_fraction` parameter in all seven G500 profiles.

The evidence schema still *accepts* `rss_bounded_or_plateaued` as an
optional boolean: `checks` is `additionalProperties: false`, and recorded
runs carry the key. Without this, `read_native_rung` fails to validate
every historical projection it reconciles against.

Contract tests updated rather than deleted: the refusal test now asserts
the fraction is recorded and does not refuse, and the S24 historical
admission in `lifecycle-runtime-1279-baseline.json` flips to `admitted`
because the growth gate was its only failing check. Regenerating that
baseline moved exactly two lines.

Refs #1387

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository: CurateLabs/graphforge/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 83a157a2-3be6-47ed-b07a-b0a64c7dcceb

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Warning

Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Do not merge yet: this change makes the S18 rung time out

A/B on one host, quiet, same harness, same certify and generator binaries. The only difference is whether gf carries 1287a2f8.

gf build direct validate (no cgroup) S18 rung under benchexec (16 cores, 4 GiB memory cap)
pinned two workers (origin/main) 35.90 s ~55 s, all ten phases passed
with this change (Automatic) 35.12 s killed at 420 s

Reproduced twice. The first observation was a --maximum-scale 20 run that sat at 100% CPU for 17 minutes on S18 before I stopped it; the second was a dedicated --rung S18 run that hit a 420 s timeout (exit=124).

What it is not

It is not a slowdown of the work itself. Driven directly, outside benchexec, the whole import sequence is unchanged:

begin 0.10s  register-nodes 0.10s  register-edges 0.40s  validate 35.22s  commit 2.00s

versus 35.90 s for validate on the pinned build. I checked this first, because my initial reading was "this change is a severe regression" and the direct measurement did not support it. It only appears under the ladder.

Likely mechanism, stated as a hypothesis rather than a result

The ladder runs inside a cgroup with cores: 16 and memory_bytes: 4294967296. This change takes tokio_worker_threads, target_partitions, io_concurrency and compute_threads from a fixed 2 to a machine-derived value — about 8 on this host. More concurrent workers means more simultaneous buffers, sort scratch and cache windows against a hard 4 GiB ceiling that the unconstrained direct run never touches.

I have not confirmed that memory is the binding resource — the killed runs left no evidence behind. It could also be thread oversubscription interacting with the cgroup's CPU controller. The reproduction is solid; the explanation is not.

Why CI is green

Nothing in the crate suites, the harness suites, clippy or make pre-push-fast runs an import at S18 scale inside a 4 GiB cgroup. The change is correct by every gate we have, and it still stops the ladder.

This is the third fail-closed surface this work has found that only a real rung exercises, after the certify validator (#1462) and the evidence schema. A green PR touching concurrency defaults should be rung-tested before it merges, and that is now cheap: --rung S18 is about a minute on a quiet host.

Suggested next step

Before merging, re-run --rung S18 with this commit and capture the peak RSS, or raise the rung's memory limit temporarily to see whether it completes. If memory is the binding resource, this change needs a memory-aware bound on the derived worker count rather than a plain logical_cpus() derivation — which is also what #1387's own 4 GiB budget decision implies, since it says to budget per shard rather than treat 4 GiB as a target.

The change is still right in principle: a two-worker facade cannot show multi-core behaviour whatever the engine does (#1462 depends on this landing). It needs a bound, not reverting.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Diagnosed: this exposes a livelock, not a resource limit

Correcting my previous comment. I hypothesised the 4 GiB cgroup cap was the binding resource. It is not. Caught the hung S18 rung in the act:

VmRSS: 183 MB                      <- nowhere near the 4 GiB cap
Threads: 17
voluntary_ctxt_switches: 602       <- in 100 seconds
read_bytes: 242 MB  write_bytes: 510 MB

    TID   %CPU STAT WCHAN
2166093    0.3 SNl  futex_do_wait   <- main thread, waiting
2166096..  0.0 SNl  futex_do_wait   <- 15 more, all waiting
2166106   99.8 RNl  -               <- ONE thread, running, no wchan

Sampled utime + stime ten seconds apart: exactly 1000 ticks, precisely one full core, continuously.

One thread spins indefinitely while all sixteen others block on futexes. That is a livelock, not a deadlock — a deadlock parks every thread at 0% CPU. And it is not memory: RSS is 183 MB against a 4096 MB ceiling.

It is non-deterministic

Driven directly outside benchexec with the same derived worker count, the identical import completes in 35.12 s. Under benchexec it hangs. Same binary, same data, same host. So this is timing-dependent — a race that the direct run wins and the containerised run loses, not a property of the worker count alone.

That also means the direct measurement I used to exonerate this change was luck, and a green CI run proves nothing here. A concurrency default cannot be validated by a single passing execution.

What this means for the change

The change is not wrong. It raises tokio_worker_threads, target_partitions, io_concurrency and compute_threads from a fixed 2 to about 8, and that exposes a pre-existing race in the ingest path that a two-worker facade has been hiding.

Two consequences worth separating:

  1. This PR should not merge until the race is found. Landing it turns an intermittent hang loose on every default-configured ingest, and the ladder is where it will show up first.
  2. The race is the more important finding. It is in code that runs today; the facade merely makes it improbable. epic(storage): stop maintaining a hand-rolled dataflow engine; reuse Arrow/DataFusion where determinism and durability allow #1456 argues the hand-rolled dataflow machinery is where the defects live, and this is that argument with a reproduction attached.

Reproducing it

--rung S18 under benchexec, gf built with 1287a2f8

Hung in 2 of 2 attempts (17 minutes before I stopped the first; 420 s timeout on the second; a third reproduced within 100 s and was inspected above). The pinned build passes the same rung in ~55 s, and S18/S19/S20 all pass on it.

The spinning thread's stack would name the loop. gdb -p <pid> -batch -ex "thread apply all bt" on a live hang is the next step, and the hang reproduces inside two minutes, so it is cheap to catch.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Correcting myself again: it is not a livelock. It is pathological parquet re-decompression.

I called this a livelock because one thread sat at 100% CPU while sixteen blocked on futexes. A profile of the running thread shows it is doing real work, not spinning on a lock. perf record on the thread in R state, 12 s, 2377 samples:

45.92%  ZSTD_decompressSequences_bmi2      tokio-rt-worker
15.12%  core::hash::BuildHasher::hash_one
 6.34%  core::iter::Map::fold
 5.41%  core::hash::sip::Hasher::write
 1.72%  core::slice::sort::unstable::quicksort
 1.59%  ZSTD_buildFSETable_body
 1.14%  arrow_array::BooleanArray::from_trusted_len_iter
 0.75%  parquet::encodings::rle::RleDecoder::get_batch_with_dict
 0.25%  parquet::column::reader::GenericColumnReader::skip_records
 0.17%  parquet::arrow::ParquetRecordBatchReader::next
 0.13%  parquet::arrow::array_reader::builder::ArrayReaderBuilder::build_reader

Zstd decompression of parquet, plus SipHash hashing, on a tokio worker — continuously, for minutes, where the pinned build completes the whole S18 rung in ~55 s. The other threads are blocked because everything is serialised behind this one.

Hypothesis, and why I think it, stated as a hypothesis

ArrayReaderBuilder::build_reader and GenericColumnReader::skip_records are both in the profile. Building a reader is setup, and skip_records is seeking; neither belongs in the steady state of a single linear scan. Their presence alongside 46% zstd decompression is the signature of readers being constructed repeatedly and each skipping forward through the file, re-decompressing to reach its offset — linear work per batch, quadratic overall.

Raising target_partitions from 2 to 8 would multiply exactly that. It is consistent with every observation: one worker saturated, the rest idle, 242 MB read but 510 MB written, and a 20x+ blowup that does not appear when the same binary runs the same import outside the container.

I cannot prove causality from this profile. The release build has no frame pointers, so the call graph is too shallow to name the caller. What is established: the thread is doing parquet decode work, not waiting, and the volume is pathological.

Corrections to my earlier comments on this PR

Three, and they are worth stating plainly because each was confidently wrong:

  1. "Severe regression" — the direct measurement showed no difference (35.12 s vs 35.90 s).
  2. "The 4 GiB cgroup cap is the binding resource" — RSS was 183 MB. Not memory.
  3. "Livelock" — the thread is running real code, not spinning on a futex.

What has held throughout: the S18 rung hangs with this change and passes without it, reproduced 4 of 4 times, and S18/S19/S20 all pass on the pinned build.

Suggested next step

Rebuild with debug = true in the release profile (or force-frame-pointers), reproduce, and re-profile. The call graph will then name the caller of build_reader, which is where the repeated scan originates. The hang reproduces inside two minutes, so this is a short loop.

The change should still land eventually — a two-worker facade cannot show multi-core behaviour, and #1462 depends on it. It should land after whatever this is, because the defect it exposes is in code that runs today.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Located: it is the one-hop query, not ingest

A frame-pointer build resolves the call graph completely. This is the whole of the hot path, 2410 samples over 12 s on the one running thread:

<graphforge_exec::demand::RssProbeStream as Stream>::poll_next
  graphforge_exec::expand_exec::expand_single_hop_chunk
    |--43.63%-- graphforge_storage::catalog::filtered_parquet::read_edges_filtered_projected_from_inventory
    |             read_required_edge_filtered
    |               read_parquet_filtered_u64_attempt
    |                 |--29.91%-- StructArrayReader::skip_records
    |                 |             |--15.88%-- array_reader::skip_records -> ZSTDCodec::decompress
    |                 |              --13.95%-- array_reader::skip_records -> ZSTD decompress
    |                  --13.72%-- ArrowReaderBuilder::build -> ZSTD decompress
     --1.93%--  read_nodes_filtered_projected_observed

expand_single_hop_chunk is a query, not ingest. The S18 rung runs ten phases and query is one of them; that is where it hangs.

This explains the observation I could not reconcile

I reported that driving the import directly showed no difference — 35.12 s with this change against 35.90 s without — and used that to argue the change was not a regression. That measurement ran begin, register-parquet, validate and commit. It never ran a query. The rung does. So the direct comparison was measuring a path this defect does not touch, and my conclusion from it was worthless.

What the defect is

read_parquet_filtered_u64_attempt spends 43.6% of the thread in two places that should not be hot:

  • skip_records (29.9%), each call decompressing zstd to seek forward.
  • ArrowReaderBuilder::build (13.7%), also decompressing — building a reader is setup, and setup is not a steady-state cost.

That is the signature of readers being constructed repeatedly and each seeking forward through compressed data. Raising target_partitions from 2 to 8 multiplies it, which is why the pinned build completes the rung in ~55 s and this one does not finish in 400 s.

Why this matters beyond this PR

The defect is in code that runs today. The two-worker default keeps the multiplier small enough to hide it. Two open issues are already circling it without naming it:

This profile is the same query path, and it says the cost is repeated seeking and reader construction over zstd-compressed parquet. A filtered read that decompresses in order to skip is doing the decompression twice: once to find the rows, once to read them.

Status of this PR

Still should not merge as it stands, but the reason has changed: it is not an ingest regression, it is a query-path defect that this change amplifies. Fixing the query path is the prerequisite, and it is worth doing regardless of this PR, since it is on the critical path for #1388 as well.

Reproduction: --rung S18 under benchexec with gf built from 1287a2f8. Hung 5 of 5 attempts; the frame-pointer build reproduces within 110 s and profiles cleanly.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

The profile says PR #1453 is the fix for this, and that it also unblocks PR #1466

Following the call path posted above to its cause, and ruling out the obvious suspect first.

It is not a missing page index. permanent_parquet::writer_properties() already sets EnabledStatistics::Page, set_offset_index_disabled(false) and a column index truncate length, and the reader asks for PageIndexPolicy::Optional. The index is written and it is read.

It is the access pattern against the page granularity. Pages are PAGE_ROWS = 20_000 rows / PAGE_BYTES = 1 MiB. At S18's 4.19M edges that is roughly 210 pages per column. A one-hop query asks for a scattered set of edge ids, and a scattered selection lands at least one wanted row in nearly every page. A page index can only skip pages that contain nothing wanted, so it skips almost nothing here, and skip_records walks and zstd-decompresses essentially the whole column either way.

That is why 43.6% of the thread is in skip_records and ArrowReaderBuilder::build, both decompressing.

The consequence

This is a layout-versus-access-pattern mismatch, not a coding defect. Filtered parquet scanning is the wrong mechanism for a scattered adjacency lookup, at any page size — shrinking pages trades decompression for index size and per-page overhead.

Which is exactly what this issue proposes to remove. PR #1453 publishes the adjacency CSR with the generation instead of rebuilding it per process, so a hop query opens a CSR rather than scanning filtered parquet. If the hop path stops going through read_edges_filtered_projected_from_inventory, this cost disappears rather than being tuned.

The dependency worth recording

PR #1453 plausibly unblocks PR #1466. #1466 (unpin the default resource policy) currently hangs the S18 rung — 5 of 5 attempts — and the profile puts the hang in this query path, amplified by target_partitions going from 2 to 8. If #1453 takes the hop query off filtered parquet, the amplified cost has nothing to amplify.

That is a hypothesis, not a result: I have not built #1453 and re-run the rung. It is a cheap test — --rung S18 is about a minute on a quiet host when it passes, and hangs inside two minutes when it does not.

Supporting measurement

#1449 records that a one-hop at S20 costs 14 s user CPU and 6.4 GB of disk reads while read_path_scan attributes 579 KB. The 6.4 GB is consistent with decompressing whole columns to serve a scattered selection, which is what this profile shows happening.

Reproduction for anyone picking this up: --rung S18 under benchexec with gf built from 1287a2f8, frame-pointer release build, perf record --call-graph fp on the thread in R state. Hangs within 110 s, profiles cleanly.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Tested: #1453 removes the hang that blocks #1466

I posted this as a hypothesis. It now has a measurement.

build S18 rung under benchexec
#1466 alone hung, 5 of 5 attempts — >400 s, one thread spinning in expand_single_hop_chunk
#1466 + #1453 work completes in 37.6 s

37.6 s is ordinary S18 territory; the hang was unbounded. Taking the hop query off filtered parquet removes the cost that #1466's target_partitions increase was multiplying, exactly as the profile predicted.

Suggested merge order: #1453 before #1466.

Two honest caveats

The combined run does not pass. It completes the work and then fails evidence_invalid at ingest. The receipt key sets are identical to a passing run — no missing or extra fields — so it is a value-level rejection rather than a shape change. I have not chased it.

And it may be my mess rather than #1453's. This was a local merge across three worktrees carrying my own in-flight fixes, and I tripped over that twice while testing: first a certify/gf binary mismatch, then finding that perf/1464-chain-split forked from observe/1462-wire-ingest-phases before I committed the certify validator fix on it. Neither was a defect in #1453.

The timing result is robust to that — 37.6 s versus >400 s is not a subtle difference and does not depend on which certify binary validated it. The evidence_invalid is not robust to it, and should be reproduced from clean branches before anyone treats it as a property of #1453.

Reproduction

# hangs, 5/5
--rung S18, gf built from 1287a2f8

# completes in 37.6s
--rung S18, gf built from 1287a2f8 merged with origin/fix/adjacency-csr-persist

Both on a quiet host, benchexec, same generator and data.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Merge-order constraint, measured: #1453 must land before this PR.

On this branch alone, S18 hangs 5/5 attempts. With #1453 applied, the same work completes in 37.6 s. The cause is on the query path rather than in this change: expand_single_hop_chunk spends 43.6% of its time in skip_records plus ArrowReaderBuilder::build, both zstd-decompressing, and #1453's persisted adjacency CSR is what removes that per-process rebuild.

So this is a sequencing constraint, not a conflict — the two merge cleanly in either order, but landing this one first leaves main unable to complete a rung. Holding it out of the merge queue until #1453 is on main.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core source code changes documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(api): the default resource policy pins two workers, so nothing can show multi-core behaviour

1 participant