Skip to content

ci: run perf jobs on bamboo - #36

Merged
jdx merged 6 commits into
mainfrom
agent/bamboo-perf-runner
Aug 2, 2026
Merged

jdx merged 6 commits into
mainfrom
agent/bamboo-perf-runner

Conversation

@jdx

@jdx jdx commented Aug 1, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • add perf and perf:record mise tasks for Tak itself
  • add main and pull-request performance workflows on the isolated bamboo-perf runner
  • keep pull-request reporting on a GitHub-hosted runner with a separate write token

Validation

  • actionlint (with bamboo-perf registered as the expected custom label)
  • mise tasks ls
  • git diff --check
  • trusted main baseline queued on the disposable runner

AI-assisted — Tool: Codex; model: unavailable; version: unavailable.


Note

Medium Risk
Introduces self-hosted CI that writes git notes on main and executes PR-supplied code on bamboo-perf, though PR measure jobs are read-only and commenting is isolated on GitHub-hosted runners.

Overview
Adds instruction-count performance CI on the dedicated bamboo-perf runner, with mise tasks perf and perf:record (release build + tak run / tak run --record).

perf on main measures each merge (and optional workflow_dispatch), records via perf:record, and pushes results to refs/notes/tak only on main. Builds run without target/ cache and install valgrind so counts stay comparable on one runner class (TAK_RUNNER pinned).

perf-pr compares the PR head (not the merge commit) against the merge-base with tak compare, fails the check on regressions, and posts/updates a sticky PR comment from a separate report job on ubuntu-latest so PR code never runs with write tokens. PR measurements stay local (no notes push). Workflows are currently gated to PRs from user jdx; actionlint registers the bamboo-perf label.

Reviewed by Cursor Bugbot for commit 4e13307. Bugbot is set up for automated code reviews on this repo. Configure here.

Summary by CodeRabbit

  • New Features
    • Added automated performance checks for pull requests and pushes to the main branch, with optional manual runs.
    • Performance results are compared against the merge-base commit to detect regressions and summarized in workflow reports.
    • Added commands for running benchmarks and recording results for historical tracking.
    • Reports are uploaded for review and may be posted to pull requests automatically.
    • Performance runs now use a consistent, dedicated execution environment for more reliable comparisons.

@coderabbitai

coderabbitai Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

The changes add local performance tasks and two GitHub Actions workflows. The workflows record Valgrind instruction counts, compare pull-request changes with the merge base, publish main-branch history, upload reports, update pull-request comments, and enforce regression gates.

Changes

Performance automation

Layer / File(s) Summary
Local performance tasks
mise.toml
Adds release-mode perf and perf:record tasks for benchmarks and git-note recording.
Main-branch performance history
.github/actionlint.yaml, .github/workflows/perf.yml
Configures the performance runner. Runs serialized measurements on main pushes or manual dispatches. Publishes tak notes from main and writes history to the workflow summary.
Pull-request performance gate
.github/workflows/perf-pr.yml
Measures the PR head and merge base, uploads the report, updates a same-repository sticky comment, and fails on measurement errors or instruction-count regressions.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

  • jdx/tak#35: Describes the performance workflow and mise task integrations added here.

Sequence Diagram(s)

sequenceDiagram
  participant PullRequest
  participant GitHubActions
  participant ValgrindTak
  participant ArtifactStorage
  participant PullRequestComment
  PullRequest->>GitHubActions: Trigger perf-pr workflow
  GitHubActions->>ValgrindTak: Measure PR head and merge base
  ValgrindTak-->>GitHubActions: Return report and gate status
  GitHubActions->>ArtifactStorage: Upload report and status
  ArtifactStorage-->>GitHubActions: Download report artifact
  GitHubActions->>PullRequestComment: Create or update sticky comment
  GitHubActions-->>PullRequest: Report success or fail the gate
Loading

Poem

A rabbit watched the benchmarks run,
While Valgrind counted every one.
Notes hopped onto main with care,
PR gates checked the numbers there.
“No regressions!” the bunny sings.
🐇📈 Fresh speed in measured things!

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: running performance jobs on the Bamboo runner.

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Comment thread .github/workflows/perf-pr.yml Outdated
Comment thread .github/workflows/perf-pr.yml
@greptile-apps

greptile-apps Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

Greptile Summary

The PR adds dedicated performance workflows for main and eligible pull requests, backed by reusable mise tasks.

  • Records and publishes main-branch instruction-count measurements from the isolated Bamboo runner.
  • Measures PR heads locally, compares them with their merge-base baseline, and transfers the report to a GitHub-hosted reporting job.
  • Registers the custom runner label with actionlint.

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Important Files Changed

Filename Overview
.github/workflows/perf-pr.yml Adds a split measure/report workflow that executes performance benchmarks without write credentials and reports the resulting gate from a GitHub-hosted runner.
.github/workflows/perf.yml Adds serialized main-branch measurement and publication of the git-notes performance history.
mise.toml Adds release-mode performance measurement and recording tasks consistent with the existing CLI contract.
.github/actionlint.yaml Registers the Bamboo performance runner label for workflow validation.

Reviews (6): Last reviewed commit: "ci: build perf binaries on bamboo" | Re-trigger Greptile

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (2)
.github/workflows/perf.yml (2)

32-42: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Register bamboo-perf with actionlint in both workflows. Both files trigger the same actionlint runner-label false positive because bamboo-perf is a self-hosted label that isn't declared to actionlint. Add it once to actionlint.yaml to fix both findings.

  • .github/workflows/perf.yml#L32-L42: no per-file change needed once actionlint.yaml lists bamboo-perf; this is the root-cause site to reference when adding the config.
  • .github/workflows/perf-pr.yml#L39-L39: same fix applies here once actionlint.yaml is updated.
🔧 Proposed actionlint.yaml addition
self-hosted-runner:
  labels:
    - bamboo-perf

As per static analysis hints, actionlint flagged label "bamboo-perf" is unknown on both files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/perf.yml around lines 32 - 42, Add the self-hosted runner
label bamboo-perf to actionlint.yaml under self-hosted-runner.labels. No direct
changes are needed at .github/workflows/perf.yml lines 32-42 or
.github/workflows/perf-pr.yml line 39; both runner-label findings are resolved
by the shared actionlint configuration.

Source: Linters/SAST tools


43-65: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Share the checkout/toolchain/instrumentation setup between the two workflows. Both files duplicate the same checkout, Swatinem/rust-cache, jdx/mise-action, valgrind install, and TAK_RUNNER value. Because this PR's comparisons depend on main and pull-request runs using identical instrumentation, letting these two blocks drift (for example, a valgrind version bump landing in only one file) would silently invalidate the instruction-count comparison instead of failing loudly.

  • .github/workflows/perf.yml#L43-L65: extract this setup sequence into a reusable composite action (or a shared workflow) that both files call.
  • .github/workflows/perf-pr.yml#L46-L73: replace this duplicated sequence with a call to the same composite action.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/perf.yml around lines 43 - 65, Extract the duplicated
checkout, Swatinem/rust-cache, jdx/mise-action, valgrind installation, and
TAK_RUNNER setup into one reusable composite action or shared workflow. Update
.github/workflows/perf.yml lines 43-65 and .github/workflows/perf-pr.yml lines
46-73 to invoke that shared setup, preserving identical instrumentation and
configuration in both workflows.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/perf-pr.yml:
- Around line 33-58: Before allowing the measure job to run external pull
requests on bamboo-perf, verify maintainer approval, ephemeral per-job runner
cleanup, and isolation from trusted workflows and secrets; if these controls are
unavailable, gate or remove the untrusted PR trigger or move measurement to a
GitHub-hosted runner.

---

Nitpick comments:
In @.github/workflows/perf.yml:
- Around line 32-42: Add the self-hosted runner label bamboo-perf to
actionlint.yaml under self-hosted-runner.labels. No direct changes are needed at
.github/workflows/perf.yml lines 32-42 or .github/workflows/perf-pr.yml line 39;
both runner-label findings are resolved by the shared actionlint configuration.
- Around line 43-65: Extract the duplicated checkout, Swatinem/rust-cache,
jdx/mise-action, valgrind installation, and TAK_RUNNER setup into one reusable
composite action or shared workflow. Update .github/workflows/perf.yml lines
43-65 and .github/workflows/perf-pr.yml lines 46-73 to invoke that shared setup,
preserving identical instrumentation and configuration in both workflows.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Central YAML (base), Organization UI (inherited)

Review profile: CHILL

Plan: Pro Plus

Run ID: fcf76300-b107-4e24-a46f-9fe58e59dcce

📥 Commits

Reviewing files that changed from the base of the PR and between 672657d and 422156e.

📒 Files selected for processing (3)
  • .github/workflows/perf-pr.yml
  • .github/workflows/perf.yml
  • mise.toml

Comment thread .github/workflows/perf-pr.yml
@github-actions

github-actions Bot commented Aug 1, 2026 •

Copy link
Copy Markdown

Instruction counts

benchmark trend instructions Δ wall (min) Δ
help ▁█ 932,979 → 943,490 +1.13% ⚠️ 0.67 → 0.69ms +2.76%
version ▁█ 673,662 → 684,185 +1.56% ⚠️ 0.66 → 0.74ms +12.20%

2 benchmark(s) above the 1% gate: help +1.13%, version +1.56%

Measured on the base but not here — a benchmark that stops running also stops gating: help on bamboo-v1-ubuntu24.04-x64-rust1.97.1, version on bamboo-v1-ubuntu24.04-x64-rust1.97.1

Only instruction counts gate. Wall clock is shown for context — on identical hardware it moves 4-20% run to run.

Measured by tak — instruction-counted CLI benchmarks, stored in this repository's git notes.

4e13307ce879 vs 672657dceca4 · measured on the runner, not pushed to the history.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 3ac139e. Configure here.

report:
name: Report and gate
needs: measure
if: always() && github.event.pull_request.user.login == 'jdx'

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Report runs after cancelled measure

Medium Severity

With cancel-in-progress: true, a superseded measure run is cancelled while report still starts because of always(). That job then tries to download the tak-report artifact from a run that often never uploaded it, so the check fails even though a newer workflow run may succeed.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 3ac139e. Configure here.

@jdx
jdx merged commit 963dad0 into main Aug 2, 2026
12 of 13 checks passed
@jdx
jdx deleted the agent/bamboo-perf-runner branch August 2, 2026 01:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant