Skip to content

perf(exec): parallelize exact cosine KNN by source (#342) - #494

Merged
cursor[bot] merged 4 commits into
mainfrom
cursor/342-parallel-cosine-knn-c6a4
Aug 10, 2026
Merged

cursor[bot] merged 4 commits into
mainfrom
cursor/342-parallel-cosine-knn-c6a4

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Aug 10, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Parallelize exact / filtered cosine KNN and all-score cosine similarity by canonical source through the instance-owned Rayon ComputePool from #337, preserving bit-for-bit serial numeric order and merge semantics.

Closes #342.

Changes

  • Add graphforge-exec::compute_pool (instance-owned Rayon pool sized by compute_threads)
  • Parallelize exact cosine KNN / similarity / filtered KNN by source ordinal above measured crossover (COSINE_PARALLEL_CROSSOVER_OPS = 32_768)
  • Worker-local checkpoint counters (no shared atomic per MAC) so the parallel path wins above crossover
  • Deterministic lowest-index chunk error selection; strengthened pool/path tests

Test plan

  • cargo test -p graphforge-exec --lib algorithm_similar_knn
  • cargo test -p graphforge-exec --lib compute_pool
  • cargo clippy -p graphforge-exec --lib -- -D warnings
  • Release-mode crossover measurement (measure_cosine_parallel_crossover --ignored)
  • Exact-head CI / CI Gate green

Notes

Land before #343 / #344 so those PRs can reuse ComputePool without duplicating the module.

Open in Web Open in Cursor 

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Note

Parallelize exact cosine KNN by source using a private Rayon thread pool

  • Introduces a ComputePool wrapper around a private Rayon ThreadPool in compute_pool.rs, sized by resource_policy.compute_threads and owned per GraphForge instance, avoiding the global Rayon pool.
  • Adds a COSINE_PARALLEL_CROSSOVER_OPS threshold in algorithm_similar_knn.rs; when estimated work exceeds it and a pool is attached, sources are chunked and processed in parallel.
  • Adds similar_algorithm_with_compute in algorithm_similar.rs and threads the pool through AlgorithmControl so the KNN kernel can use it; existing similar_algorithm and similar_algorithm_with_limits callers are unaffected.
  • Parallel execution preserves deterministic output ordering and enforces row limits and cancellation atomically via AtomicUsize.
  • Risk: worker panics are caught and mapped to an execution error rather than propagating; parallel path activates automatically when compute_threads > 1 and work exceeds the crossover, changing runtime behavior for existing callers who set compute_threads.

Macroscope summarized 903ac87.

Summary by CodeRabbit

  • New Features

    • Added configurable CPU parallelism for similarity and cosine KNN searches.
    • Improved KNN performance with workload-based parallel execution while preserving deterministic results.
    • Added instance-level compute limits and cancellation-aware processing.
  • Bug Fixes

    • Improved handling of filtered searches, output limits, worker failures, and execution errors.
    • Ensured serial and parallel search results remain equivalent.
    • Added safer processing for invalid candidates and computation failures.

@coderabbitai

coderabbitai Bot commented Aug 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

Walkthrough

The change adds instance-owned Rayon compute pools, propagates compute-thread limits through similarity execution, and parallelizes exact and filtered cosine KNN by canonical source chunks while preserving deterministic results and execution limits.

Changes

Deterministic cosine parallelization

Layer / File(s) Summary
Compute pool and algorithm control
crates/graphforge-exec/Cargo.toml, tools/bazel/drift/cargo_feature_fingerprint.json, cargo-bazel-lock.json, crates/graphforge-exec/src/compute_pool.rs, crates/graphforge-exec/src/algorithm_dispatch.rs, crates/graphforge-exec/src/algorithm_analyze_conductance.rs, crates/graphforge-exec/src/algorithm_analyze_modularity.rs, crates/graphforge-exec/src/lib.rs
Adds bounded ComputePool and SharedComputePool APIs. Adds AlgorithmLimits.compute_threads and compute-pool access to AlgorithmControl. Updates dependency metadata and test initializers.
Compute-aware similarity entry point
crates/graphforge-exec/src/algorithm_similar.rs, crates/graphforge-api/src/lib.rs, crates/graphforge-api/src/embedding_refresh.rs
Adds similar_algorithm_with_compute and routes similarity execution through configured limits and the shared compute pool. Refresh workers retain the pool.
Canonical serial and parallel cosine KNN
crates/graphforge-exec/src/algorithm_similar_knn.rs
Adds workload-based path selection, source chunking, serial and private-pool execution, deterministic merging, validation, cancellation, output-limit handling, panic conversion, and equivalence tests.
GraphForge pool ownership and policy documentation
crates/graphforge-api/src/resource_policy.rs
Documents compute_threads as the budget for the instance-owned private CPU pool, including serial behavior and crossover selection.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant GraphForge
  participant similar_algorithm_with_compute
  participant AlgorithmControl
  participant ComputePool
  participant CosineKNN
  GraphForge->>similar_algorithm_with_compute: submit similarity request with limits and pool
  similar_algorithm_with_compute->>AlgorithmControl: attach shared compute pool
  AlgorithmControl->>CosineKNN: provide compute-thread budget
  CosineKNN->>ComputePool: execute canonical source chunks
  ComputePool-->>CosineKNN: return per-source scores
  CosineKNN-->>GraphForge: return deterministic shaped results
Loading

Possibly related issues

  • #343: Extends the same instance-owned bounded CPU pool and deterministic parallel execution framework to another graph algorithm.

Suggested labels: testing

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (1 warning, 1 inconclusive)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 75.47% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
Linked Issues check ❓ Inconclusive The code evidence covers the pool, crossover, deterministic execution, limits, and tests, but documentation requirements cannot be verified because one relevant file was excluded. Review docs/development/execution-resource-policy.md, excluded by !/*.md and !/docs/**, and confirm it documents exactness, crossover, resource policy, and performance methodology.
✅ Passed checks (3 passed)
Check name Status Explanation
Out of Scope Changes check ✅ Passed All reviewed changes support the linked performance objective, including the compute pool, dispatch wiring, KNN execution, tests, documentation, and dependency metadata.
Title check ✅ Passed The title clearly identifies the primary change: parallelizing exact cosine KNN in graphforge-exec.
Description check ✅ Passed The description clearly covers the implementation scope, linked issue, testing performed, performance intent, and sequencing, despite not using every template heading.
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch cursor/342-parallel-cosine-knn-c6a4

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added executor Changes to query executor core Core source code changes documentation Improvements or additions to documentation labels Aug 10, 2026
@cursor
cursor Bot force-pushed the cursor/341-bounded-arrow-shaping-c6a4 branch 3 times, most recently from 94eb2cc to ae473c3 Compare August 10, 2026 01:48
@cursor
cursor Bot deleted the branch main August 10, 2026 01:58
cursoragent and others added 2 commits August 10, 2026 01:58
Partition independent source rows across the instance-owned private
compute pool (#337 budget) while preserving serial per-source dot-product
order, deterministic merge, fingerprints, and structured limit/cancel
outcomes. Keep a serial path below the documented crossover or when the
policy provides one compute thread.

Co-authored-by: David Spencer <DecisionNerd@users.noreply.github.com>
Allow the compute-pool similar entry arity and iterate filtered candidate
chunks without a needless range index.

Co-authored-by: David Spencer <DecisionNerd@users.noreply.github.com>
@cursor
cursor Bot changed the base branch from cursor/341-bounded-arrow-shaping-c6a4 to main August 10, 2026 02:01
@blacksmith-sh

This comment has been minimized.

Bazel Bootstrap failed once on checkpoint_pin_and_open_lease_control_recovery_cleanup
(WriterBusy) — outside this PR's exec/KNN diff; local cargo+bazel storage suites pass.

Co-authored-by: David Spencer <DecisionNerd@users.noreply.github.com>
@DecisionNerd
DecisionNerd marked this pull request as ready for review August 10, 2026 02:13

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (4)
crates/graphforge-exec/src/algorithm_dispatch.rs (2)

173-207: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Consider marking AlgorithmLimits as #[non_exhaustive].

This change adds a public field to a public struct. Every external struct-literal construction breaks. The three in-repo test sites that this PR had to update confirm the literal pattern is in use. If you mark the struct #[non_exhaustive], later field additions stay source-compatible for downstream crates, and callers move to AlgorithmLimits::default() plus builder methods.

The .max(1) clamp and the doc comment match the existing with_batch_size contract, so the builder itself is consistent.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_dispatch.rs` around lines 173 - 207,
Mark the public AlgorithmLimits struct as #[non_exhaustive] so adding future
fields does not break external struct-literal construction. Keep its existing
Default implementation and with_batch_size/with_compute_threads builder methods
unchanged.

247-253: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low value

Document that the pool overrides the declared thread budget.

with_compute_pool overwrites limits.compute_threads with pool.num_threads(). A caller that chains with_compute_threads(n) and then with_compute_pool(pool) loses n without any signal. The override direction is the safe one, because select_cosine_path and source_chunks must never request more chunks than the pool has workers. State that invariant in the doc comment so a later refactor does not reverse the assignment order.

📝 Proposed doc clarification
     /// Attach the instance-owned private CPU pool (`#337` / `#342`).
+    ///
+    /// The pool is authoritative: this replaces any previously configured
+    /// `compute_threads` with `pool.num_threads()`. Chunk planning depends on
+    /// `compute_threads` never exceeding the pool's real worker count.
     #[must_use]
     pub(crate) fn with_compute_pool(mut self, pool: crate::SharedComputePool) -> Self {
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_dispatch.rs` around lines 247 - 253,
Update the doc comment for with_compute_pool to explicitly state that attaching
the pool overrides the declared compute thread budget with the pool’s worker
count, preserving the invariant that select_cosine_path and source_chunks never
request more chunks than available pool workers.
crates/graphforge-exec/src/algorithm_similar_knn.rs (2)

441-459: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

Consider more chunks than workers so Rayon can balance the load.

source_chunks creates exactly one contiguous range per worker. After the split no work stealing is possible, so the slowest chunk sets total latency.

For exact_filtered_cosine_knn_parallel the candidate-set size varies per source. Equal source counts do not mean equal work. One source with a large candidate set serializes the tail while the other workers idle.

If you emit a small multiple of threads chunks, Rayon balances the ranges through work stealing. Determinism is unaffected, because merge_chunk_pairs concatenates by chunk index and the ranges stay in ascending source order.

Note that source_chunks_cover_canonical_ranges at Lines 773-779 and the chunks: 4 assertion at Lines 762-769 both pin the current one-chunk-per-worker shape, so those tests need updating with any change here.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_similar_knn.rs` around lines 441 - 459,
Update source_chunks to produce a small multiple of threads contiguous,
ascending source ranges rather than exactly one range per worker, while
preserving empty-input handling and deterministic ordering. Adjust
source_chunks_cover_canonical_ranges and the chunks: 4 assertion to validate the
new chunking shape and coverage, and ensure exact_filtered_cosine_knn_parallel
continues merging chunks by index.

272-273: 🚀 Performance & Scalability | 🔵 Trivial | 🏗️ Heavy lift

Avoid the per-source seen allocation in the filtered path.

score_source_filtered allocates and zeroes a vec![false; vectors.len()] for every source. The cost is proportional to the total vector count, not to the candidate-set size. For n sources this is O(n²) allocation and zeroing work even when each source has only a handful of candidates.

Two options preserve exact iteration order, and therefore preserve bit-for-bit results:

  • Reuse one generation-stamped Vec<u32> per chunk. Compare each entry against the current source ordinal instead of clearing the buffer.
  • Deduplicate source_candidates directly when the candidate set is much smaller than vectors.len().

The bounds check must stay, because it produces the filtered KNN candidate is outside vector selection error that the tests assert.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_similar_knn.rs` around lines 272 - 273,
Update score_source_filtered to avoid allocating and clearing a vec![false;
vectors.len()] for each source by reusing a generation-stamped buffer per chunk
or deduplicating source_candidates when appropriate. Preserve the existing
candidate iteration order and bit-for-bit results, retain the bounds check that
reports “filtered KNN candidate is outside vector selection,” and keep
deduplication semantics unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/graphforge-exec/src/algorithm_similar_knn.rs`:
- Around line 18-23: Validate COSINE_PARALLEL_CROSSOVER_OPS with benchmarks
comparing serial and source-parallel execution, including Rayon scheduling,
per-chunk allocation, and merge overhead on the target hosts. Replace the
coincidental 16_384 value with the smallest measured workload where parallel
execution wins, and document the supporting benchmark results alongside the
constant.
- Around line 824-857: Update
parallel_output_limits_and_cancellation_remain_atomic to use workload dimensions
that exceed COSINE_PARALLEL_CROSSOVER_OPS, such as adversarial_vectors(32, 32),
while preserving the expected errors. Add an explicit select_cosine_path
assertion confirming the parallel path is selected, following the guard used by
thread_matrix_preserves_exact_knn_cosine_and_filtered_fingerprints.
- Around line 119-138: Update the parallel chunk collection in the current KNN
implementation and exact_filtered_cosine_knn_parallel to collect all chunk
results without short-circuiting, preserving their range order. After
collection, return the first error by chunk index, and otherwise unwrap the
successful chunk results for downstream aggregation.

In `@crates/graphforge-exec/src/compute_pool.rs`:
- Around line 85-95: Update multi_thread_pool_runs_on_private_workers to inspect
each parallel iterator task’s current thread name and assert it uses the
configured private-worker name prefix, while retaining the sum assertion. Ensure
the check occurs inside the install closure so execution on Rayon’s global pool
cannot satisfy the test.

---

Nitpick comments:
In `@crates/graphforge-exec/src/algorithm_dispatch.rs`:
- Around line 173-207: Mark the public AlgorithmLimits struct as
#[non_exhaustive] so adding future fields does not break external struct-literal
construction. Keep its existing Default implementation and
with_batch_size/with_compute_threads builder methods unchanged.
- Around line 247-253: Update the doc comment for with_compute_pool to
explicitly state that attaching the pool overrides the declared compute thread
budget with the pool’s worker count, preserving the invariant that
select_cosine_path and source_chunks never request more chunks than available
pool workers.

In `@crates/graphforge-exec/src/algorithm_similar_knn.rs`:
- Around line 441-459: Update source_chunks to produce a small multiple of
threads contiguous, ascending source ranges rather than exactly one range per
worker, while preserving empty-input handling and deterministic ordering. Adjust
source_chunks_cover_canonical_ranges and the chunks: 4 assertion to validate the
new chunking shape and coverage, and ensure exact_filtered_cosine_knn_parallel
continues merging chunks by index.
- Around line 272-273: Update score_source_filtered to avoid allocating and
clearing a vec![false; vectors.len()] for each source by reusing a
generation-stamped buffer per chunk or deduplicating source_candidates when
appropriate. Preserve the existing candidate iteration order and bit-for-bit
results, retain the bounds check that reports “filtered KNN candidate is outside
vector selection,” and keep deduplication semantics unchanged.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: f58cd092-57c5-4348-aabe-543c1561059b

📥 Commits

Reviewing files that changed from the base of the PR and between e869833 and 6e6af62.

⛔ Files ignored due to path filters (3)
  • Cargo.lock is excluded by !**/*.lock, !**/*.lock
  • docs/development/execution-resource-policy.md is excluded by !**/*.md, !**/docs/**
  • docs/development/m4-entry-baseline.md is excluded by !**/*.md, !**/docs/**
📒 Files selected for processing (13)
  • cargo-bazel-lock.json
  • crates/graphforge-api/src/embedding_refresh.rs
  • crates/graphforge-api/src/lib.rs
  • crates/graphforge-api/src/resource_policy.rs
  • crates/graphforge-exec/Cargo.toml
  • crates/graphforge-exec/src/algorithm_analyze_conductance.rs
  • crates/graphforge-exec/src/algorithm_analyze_modularity.rs
  • crates/graphforge-exec/src/algorithm_dispatch.rs
  • crates/graphforge-exec/src/algorithm_similar.rs
  • crates/graphforge-exec/src/algorithm_similar_knn.rs
  • crates/graphforge-exec/src/compute_pool.rs
  • crates/graphforge-exec/src/lib.rs
  • tools/bazel/drift/cargo_feature_fingerprint.json

Comment thread crates/graphforge-exec/src/algorithm_similar_knn.rs Outdated
Comment thread crates/graphforge-exec/src/algorithm_similar_knn.rs Outdated
Comment thread crates/graphforge-exec/src/algorithm_similar_knn.rs
Comment thread crates/graphforge-exec/src/compute_pool.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
crates/graphforge-exec/src/algorithm_similar_knn.rs (2)

272-273: 🚀 Performance & Scalability | 🔵 Trivial | ⚡ Quick win

The per-source seen buffer scales with the vector count, not the candidate count.

score_source_filtered allocates and zeroes vec![false; vectors.len()] for every source. Total cost is O(sources × vectors), even when each candidate list holds only a few entries. For filtered KNN the candidate lists are normally short, so this allocation dominates the actual scoring work on large selections.

Size the deduplication structure by source_candidates.len() instead.

♻️ Proposed change to size deduplication by candidate count
-    let mut seen = vec![false; vectors.len()];
     let mut candidates = Vec::with_capacity(source_candidates.len());
+    let mut seen = std::collections::HashSet::with_capacity(source_candidates.len());
     for &target_index in source_candidates {
         checkpoint(control, work)?;
-        let Some(target_seen) = seen.get_mut(target_index) else {
+        if target_index >= vectors.len() {
             return Err(execution(
                 "filtered KNN candidate is outside vector selection",
             ));
-        };
-        if source_index == target_index || std::mem::replace(target_seen, true) {
+        }
+        if source_index == target_index || !seen.insert(target_index) {
             continue;
         }

Confirm that the traversal order and the resulting candidate order stay identical, because #342 requires bit-for-bit parity with the serial path. The order is unchanged here, since seen only gates insertion.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_similar_knn.rs` around lines 272 - 273,
Update the per-source `seen` buffer in `score_source_filtered` to use
`source_candidates.len()` rather than `vectors.len()`, reducing allocation to
the candidate count. Preserve the existing traversal and candidate insertion
order so filtered KNN results remain bit-for-bit identical with the serial path.

321-325: 🚀 Performance & Scalability | 🔵 Trivial | 💤 Low value

Full sort per source costs more than a partial selection.

top_k_pairs sorts every candidate, then keeps k. The exact path produces sources - 1 candidates per source, so the cost is O(n log n) per source while only k results survive. select_nth_unstable_by followed by a sort of the first k elements reduces this to O(n + k log k).

This change is optional. If you apply it, verify the tie order stays identical, because #342 requires bit-for-bit parity with the current output.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/graphforge-exec/src/algorithm_similar_knn.rs` around lines 321 - 325,
Optimize top_k_pairs by replacing the full candidates sort with
select_nth_unstable_by to partition around the top-k boundary, then sort only
the retained first k candidates using the existing score-descending and
index-ascending comparator. Preserve the current tie ordering exactly to
maintain bit-for-bit parity.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@crates/graphforge-exec/src/algorithm_similar_knn.rs`:
- Around line 272-273: Update the per-source `seen` buffer in
`score_source_filtered` to use `source_candidates.len()` rather than
`vectors.len()`, reducing allocation to the candidate count. Preserve the
existing traversal and candidate insertion order so filtered KNN results remain
bit-for-bit identical with the serial path.
- Around line 321-325: Optimize top_k_pairs by replacing the full candidates
sort with select_nth_unstable_by to partition around the top-k boundary, then
sort only the retained first k candidates using the existing score-descending
and index-ascending comparator. Preserve the current tie ordering exactly to
maintain bit-for-bit parity.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 98285860-7738-4677-807c-c9a801bf55f9

📥 Commits

Reviewing files that changed from the base of the PR and between e869833 and 6e6af62.

⛔ Files ignored due to path filters (3)
  • Cargo.lock is excluded by !**/*.lock, !**/*.lock
  • docs/development/execution-resource-policy.md is excluded by !**/*.md, !**/docs/**
  • docs/development/m4-entry-baseline.md is excluded by !**/*.md, !**/docs/**
📒 Files selected for processing (13)
  • cargo-bazel-lock.json
  • crates/graphforge-api/src/embedding_refresh.rs
  • crates/graphforge-api/src/lib.rs
  • crates/graphforge-api/src/resource_policy.rs
  • crates/graphforge-exec/Cargo.toml
  • crates/graphforge-exec/src/algorithm_analyze_conductance.rs
  • crates/graphforge-exec/src/algorithm_analyze_modularity.rs
  • crates/graphforge-exec/src/algorithm_dispatch.rs
  • crates/graphforge-exec/src/algorithm_similar.rs
  • crates/graphforge-exec/src/algorithm_similar_knn.rs
  • crates/graphforge-exec/src/compute_pool.rs
  • crates/graphforge-exec/src/lib.rs
  • tools/bazel/drift/cargo_feature_fingerprint.json
🚧 Files skipped from review as they are similar to previous changes (12)
  • crates/graphforge-exec/Cargo.toml
  • crates/graphforge-exec/src/algorithm_analyze_conductance.rs
  • cargo-bazel-lock.json
  • crates/graphforge-api/src/embedding_refresh.rs
  • crates/graphforge-exec/src/algorithm_dispatch.rs
  • tools/bazel/drift/cargo_feature_fingerprint.json
  • crates/graphforge-exec/src/algorithm_similar.rs
  • crates/graphforge-exec/src/algorithm_analyze_modularity.rs
  • crates/graphforge-exec/src/lib.rs
  • crates/graphforge-api/src/resource_policy.rs
  • crates/graphforge-exec/src/compute_pool.rs
  • crates/graphforge-api/src/lib.rs

Raise COSINE_PARALLEL_CROSSOVER_OPS to the measured 32_768 win boundary,
use worker-local checkpoint counters (no shared atomic per MAC), prefer
lowest-index chunk errors, and strengthen pool/path tests per review.

Co-authored-by: David Spencer <DecisionNerd@users.noreply.github.com>
This was referenced Aug 10, 2026
@DecisionNerd
DecisionNerd deleted the cursor/342-parallel-cosine-knn-c6a4 branch August 13, 2026 02:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core source code changes documentation Improvements or additions to documentation executor Changes to query executor

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(exec): parallelize exact cosine KNN by canonical source

2 participants