Skip to content

build(bazel): verify Blacksmith remote cache hits and collect perf pairs (#5) - #426

Merged
DecisionNerd merged 12 commits into
mainfrom
build/5-blacksmith-cache-evidence
Aug 7, 2026
Merged

DecisionNerd merged 12 commits into
mainfrom
build/5-blacksmith-cache-evidence

Conversation

@DecisionNerd

@DecisionNerd DecisionNerd commented Aug 6, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Verify Blacksmith Bazel remote cache hits on identical-SHA warm observation (distinct --output_base prime/warm).
  • Collect and check in ≥10 cold/warm representative pairs in perf-sample.json; strict evaluate passes (warm p50 speedup ≈97%, compute reduction ≈98%, cold regression negative).
  • Prove cache-unavailable cold correctness and affected-input isolation; keep no in-repo --remote_cache.

Test plan

  • Bazel Bootstrap: observe-warm: remote cache hits observed
  • collect-pairs ≥10 pairs → checked-in perf-sample.json status complete
  • Strict evaluate without --allow-pending passes
  • Test Suite + CI Gate green on head SHA

Closes #5

Same-output-base warm re-runs hide Blacksmith remote cache hits behind
the local action cache. Prime and warm across fresh output bases, and
collect ≥10 cold/warm pairs once hits are observed.

Co-authored-by: Cursor <cursoragent@cursor.com>
@coderabbitai

coderabbitai Bot commented Aug 6, 2026 •

Copy link
Copy Markdown

Review Change Stack

Walkthrough

The Bazel cache performance harness now isolates cold and warm output bases, measures representative targets, collects cold/warm pairs, validates remote-cache evidence, and writes updated evidence JSON. A new collect-pairs CLI mode controls collection.

Changes

Bazel cache performance collection

Layer / File(s) Summary
Isolated warm observation and affected-input validation
scripts/ci/bazel-cache-perf.py, scripts/ci/test-bazel-cache-perf.py
Warm and mutated builds use distinct output bases. The harness records remote-cache announcements, remote hits, local actions, and protocol metadata. Tests verify the isolation protocol.
Representative cold/warm collection
scripts/ci/bazel-cache-perf.py, scripts/ci/test-ci-storage-policy.py
measure_representative combines representative test, build, and binding measurements. mode_collect_pairs collects isolated runs, evaluates evidence gates, and writes collection metadata and evidence JSON. The collected evidence artifact is approved by the storage policy.
Collect-pairs CLI wiring
scripts/ci/bazel-cache-perf.py
The CLI adds the collect-pairs mode, the --pairs option, command documentation, and dispatch requiring --write.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant Operator
  participant collect_pairs
  participant measure_representative
  participant Bazel
  participant Evidence_JSON

  Operator->>collect_pairs: request paired collection
  collect_pairs->>measure_representative: run isolated cold and warm measurements
  measure_representative->>Bazel: execute representative targets
  Bazel-->>measure_representative: timing and remote-hit evidence
  measure_representative-->>collect_pairs: measurement result
  collect_pairs->>Evidence_JSON: write pairs, metadata, gates, and status
Loading

Possibly related issues

  • CurateLabs/graphforge-x issue 1: Related to Bazel cache performance measurement and remote-cache observability for cold and warm CI builds.

Possibly related PRs

🚥 Pre-merge checks | ✅ 3 | ❌ 2

❌ Failed checks (2 warnings)

Check name Status Explanation Resolution
Linked Issues check ⚠️ Warning The PR adds the collection mechanisms, but the required 10-pair evidence and strict performance-gate evaluation remain incomplete. Collect at least 10 valid pairs, verify remote hits, affected-input isolation, cold-build correctness, and all performance gates, then attach complete evidence with the exact SHA.
Description check ⚠️ Warning The description states the goals, outcomes, testing plan, and related issue, but omits most required template sections and checklist details. Use the repository template and add change type, detailed changes, test commands, checklist status, performance notes, breaking-change status, and reviewer context.
✅ Passed checks (3 passed)
Check name Status Explanation
Out of Scope Changes check ✅ Passed The reviewed changes support remote-cache verification, performance collection, evidence upload, testing, and CI execution for issue #5.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Title check ✅ Passed The title clearly summarizes the main changes: verifying Bazel remote cache hits and collecting performance pairs.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch build/5-blacksmith-cache-evidence

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/ci/bazel-cache-perf.py`:
- Around line 450-468: Update the pair-record construction to persist explicit
cold and warm compute evidence, including CPU time, action counts, and storage
metrics from each leg’s existing results. Replace the current
warm["wall_seconds"] assignment for compute_proxy_seconds with a value derived
from those compute metrics, while retaining wall-time fields for timing context.
- Around line 476-512: Reset evidence["status"] to "collecting" unconditionally
before calling evaluate_evidence() in the collection flow. Keep the existing
transition to "complete" only inside the passing gate branch, so failed
evaluations cannot preserve a prior completed status.

In `@scripts/ci/test-bazel-cache-perf.py`:
- Around line 78-91: Update test_observe_warm_uses_distinct_output_bases to mock
or capture measure_bazel, invoke mode_observe_warm, and assert that the two
recorded --output_base values are distinct. Remove the source-text checks while
preserving the existing Blacksmith Bazel Bootstrap execution and
bazel-warm-observation.json upload as integration evidence.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 699f662f-c598-42d6-8ccf-261e7285930f

📥 Commits

Reviewing files that changed from the base of the PR and between c52be21 and 40e9141.

⛔ Files ignored due to path filters (2)
  • .github/workflows/test.yml is excluded by !**/.github/**
  • docs/development/bazel-migration-perf.md is excluded by !**/*.md, !**/docs/**
📒 Files selected for processing (2)
  • scripts/ci/bazel-cache-perf.py
  • scripts/ci/test-bazel-cache-perf.py

Comment thread scripts/ci/bazel-cache-perf.py
Comment thread scripts/ci/bazel-cache-perf.py
Comment thread scripts/ci/test-bazel-cache-perf.py
DecisionNerd and others added 2 commits August 6, 2026 14:22
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Blocked on GitHub Actions major outage (no github-actions check suite for this head). Will empty-commit / re-sync when https://www.githubstatus.com shows Actions recovered, then confirm Blacksmith remote cache hit via observe-warm before collecting ≥10 pairs.

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Blocked on GitHub Actions major outage (no github-actions check suite for this head). Will empty-commit / re-sync when https://www.githubstatus.com shows Actions recovered, then confirm Blacksmith remote cache hit via observe-warm before collecting ≥10 pairs.

Co-authored-by: Cursor <cursoragent@cursor.com>
@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Temporarily closing to retrigger Test Suite after GitHub Actions outage (webhook throttle).

@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Reopening to retrigger CI for #5 Blacksmith cache evidence.

@DecisionNerd DecisionNerd reopened this Aug 6, 2026
DecisionNerd and others added 3 commits August 6, 2026 16:38
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation ci-cd CI/CD configuration changes tooling Developer tooling and automation release:none No release note or version impact labels Aug 6, 2026
Blacksmith remote hits are confirmed; make affected-inputs use distinct
output bases so isolation is measurable, allow the collected sample
artifact, and satisfy ruff format/check.

Co-authored-by: Cursor <cursoragent@cursor.com>
@DecisionNerd

Copy link
Copy Markdown
Contributor Author

Update

Remote cache hits are observed on Blacksmith (observe-warm: remote cache hits observed (3); remote_cache_announced: true) in run https://github.com/CurateLabs/graphforge/actions/runs/31130077088.

That run still failed on:

  • affected-inputs (same-output-base + remote cache hid the isolation signal)
  • Python Quality (ruff format)
  • Repository Policy (missing dist/perf-sample-collected.json allowlist entry)

Pushed a fix for those three so Bazel Bootstrap can proceed to collect-pairs.

Record 10 paired runs with remote cache hits from Bazel Bootstrap, mark
perf-sample complete, and update harness tests for the close gate.

Co-authored-by: Cursor <cursoragent@cursor.com>

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
scripts/ci/bazel-cache-perf.py (1)

476-484: 🗄️ Data Integrity & Integration | 🟠 Major | 🏗️ Heavy lift

Bind reused observations to the measured commit.

Line 476 retains prior observation booleans without checking their commit SHA. Line 484 then records the current SHA. A new commit can therefore reuse old cache_unavailable_cold_correct and affected_inputs_isolation proofs, while evaluate_evidence() accepts those booleans as current evidence.

Store provenance for each observation, or clear reused observations when the measured SHA changes.

Proposed fix
     evidence = json.loads(evidence_path.read_text(encoding="utf-8"))
+    measured_sha = git_sha(root)
     pairs: list[dict[str, Any]] = []
     remote_hits_seen = False
...
     obs = evidence.setdefault("observations", {})
+    if evidence.get("git_sha_measured") != measured_sha:
+        for key in (
+            "remote_cache_hits_on_identical_sha",
+            "cache_unavailable_cold_correct",
+            "affected_inputs_isolation",
+        ):
+            obs.pop(key, None)
...
-                "git_sha": git_sha(root),
+                "git_sha": measured_sha,
...
-    evidence["git_sha_measured"] = git_sha(root)
+    evidence["git_sha_measured"] = measured_sha
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/ci/bazel-cache-perf.py` around lines 476 - 484, Update the evidence
reuse flow around observations and git_sha(root) so
cache_unavailable_cold_correct, affected_inputs_isolation, and related proof
booleans cannot carry across measured commits. Track each observation’s
originating SHA and only preserve it when it matches the current measured
commit; otherwise clear the stale observation values before evaluate_evidence()
consumes them.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@scripts/ci/bazel-cache-perf.py`:
- Around line 604-613: Remove the final local-work/process-count fallback from
the `ok` computation. Rely on `local_fallback_ok` for runs without an execution
log, ensuring every accepted path still rejects runs where `core_rebuilt` is
true.

---

Outside diff comments:
In `@scripts/ci/bazel-cache-perf.py`:
- Around line 476-484: Update the evidence reuse flow around observations and
git_sha(root) so cache_unavailable_cold_correct, affected_inputs_isolation, and
related proof booleans cannot carry across measured commits. Track each
observation’s originating SHA and only preserve it when it matches the current
measured commit; otherwise clear the stale observation values before
evaluate_evidence() consumes them.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: 8e4b79bd-e675-415f-aee5-e80432d9854b

📥 Commits

Reviewing files that changed from the base of the PR and between 40e9141 and d94a697.

⛔ Files ignored due to path filters (1)
  • .github/workflows/test.yml is excluded by !**/.github/**
📒 Files selected for processing (3)
  • scripts/ci/bazel-cache-perf.py
  • scripts/ci/test-bazel-cache-perf.py
  • scripts/ci/test-ci-storage-policy.py
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/ci/test-bazel-cache-perf.py

Comment thread scripts/ci/bazel-cache-perf.py Outdated
DecisionNerd and others added 3 commits August 6, 2026 18:09
Distinct output bases under remote cache masked source mutations as remote
hits. Warm-then-mutate on one output_base again proves local rebuild of the
changed crate only.

Co-authored-by: Cursor <cursoragent@cursor.com>
Match the probe that already passed in Bootstrap: warm then mutate against
the job default output base populated by earlier smoke builds.

Co-authored-by: Cursor <cursoragent@cursor.com>
A fixed comment marker was uploaded by earlier probe runs, so warm→mutate
got a remote cache hit instead of a local sandbox rebuild.

Co-authored-by: Cursor <cursoragent@cursor.com>
@DecisionNerd
DecisionNerd merged commit bad99f9 into main Aug 7, 2026
21 checks passed
@DecisionNerd
DecisionNerd deleted the build/5-blacksmith-cache-evidence branch August 13, 2026 02:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-cd CI/CD configuration changes documentation Improvements or additions to documentation release:none No release note or version impact tooling Developer tooling and automation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bazel: Blacksmith remote cache enablement and cold/warm performance gates

1 participant