Skip to content

ci: run code scanning from this repository so a crashed scan can be re-run - #7130

Merged
devantler merged 10 commits into
mainfrom
claude/ci-codeql-advanced-setup-6767
Sep 20, 2026
Merged

devantler merged 10 commits into
mainfrom
claude/ci-codeql-advanced-setup-6767

Conversation

@devantler

@devantler devantler commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

🤖 Generated by the Agentic Engineer

Why

GitHub's built-in code scanning crashes on KSail often enough to matter — 22 of the last 118 finished runs, and 4 of 31 on the main branch. A crashed scan blocks merging, and GitHub refuses to re-run it, so the only way out is a fresh commit. A bot-authored release pull request never pushes one, so it stays stuck until a person steps in.

What

Adds a code scan that lives in this repository, covering the same languages with the same rules, so a crashed run can simply be re-run.

It ships switched off and does nothing until it is turned on, because GitHub refuses results from a repository's own scan while the built-in one is active. That keeps the two from colliding and makes this safe to merge on its own — today's scanning carries on unchanged.

Fixes #6767

👉 After merge/promotion: turning it on is one reversible step and it is yours to take — switch the built-in code scanning off, and set the repository variable this scan reads to true. Doing so also removes what currently feeds the "Require code quality results" rule, which GitHub no longer lets a repository's own scan produce, so that rule needs a new source or removal at the same time. A comment below names the variable and the exact settings.

Default setup's Go autobuild extracts every go.mod on one runner. The desktop
module replaces the root module, so the root dependency graph is extracted
twice and the runner is shut down mid-extraction. Move to an advanced setup
that analyses the root and desktop modules on separate runners and can be
re-run.

Refs #6767

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Comment thread .github/workflows/codeql.yaml Fixed
Comment thread .github/workflows/codeql.yaml Fixed
@github-actions

github-actions Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

✅MegaLinter analysis: Success

✅ Linters with no issues

actionlint, bash-exec, git_diff, hadolint, jscpd, jsonlint, lychee, markdown-table-formatter, markdownlint, prettier, prettier, shellcheck, shfmt, stylelint, syft, trivy-sbom, trufflehog, v8r, v8r, yamllint

Notices

⚠️ Your configuration references items that have been removed from MegaLinter and are ignored: REPOSITORY_GITLEAKS. See Removed linters to find their replacements.

See detailed reports in MegaLinter artifacts

MegaLinter is provided by OX Security
Show us your support by starring ⭐ the repository

devantler and others added 3 commits September 20, 2026 00:23
CodeQL claims ~14.5 GB of the 16 GB runner for the Go extractor. In manual
build mode the traced go build runs alongside it, so the job is killed
mid-build. Cap the extractor and the compiler, and drop the code-quality
analysis kind the action rejects in custom workflows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The new CodeQL workflow tripped two repository guards.

`TestNoDefaultBranchWorkflowCancelsRunsInProgress` rejects any workflow that
runs on main and may cancel in-progress runs there, so one merge cannot evict
the previous merge's checks. Its allowlist is an exact-match whitelist by
design — substring forms admitted bypasses such as
`${{ github.ref != 'refs/heads/main' || true }}` — so conform to the approved
expression rather than widening the guard.

zizmor flagged both local action references under `unpinned-uses`. Adopt
GitHub's self-repository syntax, which resolves to this repository at the
executing commit, so the action and the workflow calling it can never drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit took zizmor's advice and moved both local action
references to GitHub's self-repository syntax. That traded a zizmor notice
for a harder failure: actionlint v1.7.12, which MegaLinter v10.1.0 pins,
does not recognise `$/` and rejects it as a malformed ref, so `Validate Go
Project` went red.

The two linters disagree, and only one of them can be satisfied by the
reference itself. Suppress zizmor's `self-repository` finding on exactly the
two lines it fires on, with the reason recorded inline, rather than silencing
actionlint's malformed-ref check — that check is repository-wide and would
stop reporting genuinely broken action references.

This also keeps the file consistent with the rest of the repository, which
uses `./` throughout.

Verified with the same tools CI runs: actionlint exits 0 where `$/` reproduced
CI's exact error, and zizmor reports no findings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Hygiene pass at 386072e7. Three checks that were red are now green; what remains is the settings change this PR's description already asks for.

Fixed — 🧾 Verify Workflow Contracts (now CI - KSail ✅)

TestNoDefaultBranchWorkflowCancelsRunsInProgress rejected codeql.yaml: it runs on main and could cancel in-progress runs there, so one merge could evict the previous merge's checks. The workflow's ${{ github.event_name == 'pull_request' }} is safe on main in practice, but that test is an exact-match allowlist by deliberate design — its own comment records that substring forms admitted bypasses such as ${{ github.ref != 'refs/heads/main' || true }}. So I conformed to the approved expression rather than widening the guard. RED/GREEN: ablating the line back reproduces the failure, and the message names codeql.yaml specifically.

Fixed — zizmor ✅, and a correction to my own first attempt

zizmor raised self-repository on both local action references. I first took its advice and moved them to $/. That was wrong and I reverted it: actionlint v1.7.12, which MegaLinter v10.1.0 pins, does not recognise $/ and rejects it as a malformed ref, which turned a zizmor low into a red ✅ Validate Go Project. Both are green now with ./ kept and zizmor suppressed on exactly the two lines, reason recorded inline.

I deliberately did not silence actionlint's malformed-ref check instead — it is repository-wide, and blinding it to catch a false positive on a local action would stop it reporting genuinely broken references. Verified with the same tools CI runs (actionlint exits 0 where $/ reproduces CI's exact error; zizmor reports no findings). Evidence posted to devantler-tech/.github#271, which tracks this org-wide.

Still red, and not fixable from here — 🔍 CodeQL

Every Analyze job fails with:

CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled

Default setup reads configured today. This is the first of the two settings changes in the description, and until it is made the workflow cannot be validated at all — so I have not attempted the remaining memory tuning, because no run can show whether it worked. Recorded as an authority blocker on #6767.

One finding worth keeping for when it can run. The earlier Analyze (go) failure was not in the analysis — it was step 7, the traced go build, killed at 8 minutes (exit 143 + runner shutdown), with Perform CodeQL analysis never reached. So the remaining lever is extractor/compiler concurrency (threads, go build -p), not only CodeQL's ram reservation. Analyze (go-desktop) got all the way to SARIF upload, so the per-module split is already doing its job.

This also gates #7128, which is otherwise merge-ready and blocked solely by the managed CodeQL check — GitHub refuses to re-run those, so only a new head commit retriggers one.

The Go extractor resolves the module graph with `go list` and reads types
from source, covering every go.mod in the workspace in one job with no
compiler. Tracing a `go build` instead runs the compiler alongside the
extractor, and the two together exceed the runner: `Build root module` was
killed (exit 143) at both heads that ran it, while CodeQL's own autobuilder
chose buildless extraction on the same commit and analysed all three modules
successfully in 9.5 minutes at the default memory budget.

Drops the per-module matrix split, the memory caps, the disk and toolchain
setup steps and the build steps they existed for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler devantler changed the title ci: analyse each Go module on its own CodeQL runner ci: run code scanning from this repository so a crashed scan can be re-run Sep 20, 2026
`build-mode: none` is rejected outright for Go by CodeQL 2.27 ("Go does not
support the none build mode. Please try using one of the following build
modes instead: autobuild, manual"), so the init step failed in 25 seconds.

`autobuild` is what reaches the same extractor: on this repository the Go
autobuilder runs `go list` and extracts types from source without compiling,
which is the configuration default setup uses and the one measured succeeding
on all three modules in 9.5 minutes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Exercised at 6b8006c1, and the headline result is the one this PR is for

The recovery path works. Analyze (go) died at 02:28:06Z with
The runner has received a shutdown signal — the same intermittent runner exhaustion #6767 is about.
gh run rerun 35483555479 --failed was accepted and the workflow is now on attempt=2. The managed
run on this same repository refuses that command outright (This workflow run cannot be retried,
re-probed today on #7128). That difference is the entire point of moving off default setup, and it is
now demonstrated rather than argued.

What the two pushes on this PR established

1. The original premise was wrong, and a control on the same commit refuted it. The draft capped
CodeQL's memory so a traced go build would have room. Measured at 386072e7:

configuration CODEQL_RAM / THREADS result
default setup, autobuild, all 3 modules 14575 / 4 success, 15.5 min
this PR, manual + go build -p 2 ./..., root module 8192 / 2 killed, exit 143 at 🏗️ Build root module

The configuration being blamed was the one passing. The reason is that autobuild never compiles here:
its log reads Running go list to resolve package and module directories → resolved 5490 packages →
Done extracting packages → Success: extraction succeeded for all 3 discovered project(s). The
compiler only appears when a workflow asks for manual, and then it and the extractor together exceed
the runner.

2. build-mode: none is not the way to ask for that extractor. Pushed at 82c8aa6b; CodeQL 2.27
rejected it in 25 seconds — Go does not support the none build mode. Please try using one of the following build modes instead: autobuild, manual. autobuild is the supported spelling, pushed at
6b8006c1.

3. The remaining failure is not ours. At 82c8aa6b, Analyze (actions) and
Analyze (javascript-typescript) both ran to completion and failed only on
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled —
the settings change recorded on #6767. Everything up to the upload works.

Net effect

133 lines → 69. Gone: the per-module matrix, the memory caps, the disk-free and toolchain setup steps,
both build steps, and both zizmor: ignore[self-repository] suppressions (with them, the actionlint
conflict noted on devantler-tech/.github#271 no longer touches this file). actionlint, zizmor and
the 86 workflow-contract tests pass locally; CI - KSail, Validate Go Project, zizmor and
dependency-review are green on this head.

Staying a draft: it cannot be validated end to end, or merged, until default setup is off.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Attempt 2 landed, and it rules configuration out as the cause

The re-run this PR exists to make possible completed at 02:41Z. Three results, all at 6b8006c1.

1. The advanced workflow runs on exactly the same memory budget as default setup. This was
asserted in the workflow's own comment; it is now measured, from Analyze (go)'s own init log:

While resolving threads, found a cgroup CPUs file with 4 CPUs in /sys/fs/cgroup/cpuset.cpus.effective.
CODEQL_RAM: 14575
CODEQL_THREADS: 4

Those are the same numbers the successful default-setup run used on this commit. Nothing in this
workflow asks the extractor to work with less, so the difference in outcome is not a difference in
configuration.

2. Everything except Go reached the upload and failed only on the settings conflict.
Analyze (actions) and Analyze (javascript-typescript) both completed
🔬 Perform CodeQL analysis and failed with

Code Scanning could not process the submitted SARIF file:
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled

The earlier evaluation established this at 82c8aa6b; it now holds at the current head. The workflow
is proven end-to-end up to the upload, and the upload is the switch recorded on #6767.

3. Go died the same way again. 🏗️ Autobuild ran 11.5 minutes, extracted through the root
module's packages, and the runner received a shutdown signal at 02:41:15Z mid-extraction of
pkg/svc/installer/metallb/installer.go.

The honest reading

Same commit, same memory budget, opposite outcomes: default setup succeeded, this workflow's two
attempts did not. That rules out configuration and confirms the failure is the intermittent one
#6767 describes — but two failures in two attempts at one head is a higher rate than "occasional",
and I am not going to call that noise on three observations.

One untested difference is worth naming rather than assuming away: the rewrite deleted the disk-free
and toolchain setup steps along with the memory caps, and a runner out of disk dies with the same
shutdown signal an out-of-memory runner does. I have not measured that, so it stays a hypothesis.

None of this changes the PR's claim. The deliverable is that a failed scan can be retried at all, and
that is demonstrated: gh run rerun --failed was accepted here and is refused on the managed run.
Recovering from this failure is the point; never seeing it was never the promise.

Status: staying a draft. The merge is authority-blocked on #6767 — default setup must be turned
off before any of this can upload.

devantler and others added 2 commits September 20, 2026 05:11
The Go extractor defaults to one worker per core. On this repository's module
graph that exhausts a 16 GB runner part-way through extraction and the runner is
killed, which is the intermittent failure #6767 describes. Measured on the same
commit, both a managed run and this workflow used CODEQL_RAM 14575 / THREADS 4,
so the failure is not a configuration difference between them.

Default setup cannot set this cap. This workflow can, which is a second reason to
own it alongside re-runnability.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Capping extraction at two threads was measured and changed nothing: the cap
took effect (CODEQL_THREADS: 2) and the runner was still killed at the same
eleven-minute mark. So peak memory is not scaling with worker concurrency.

The Go extractor runs as its own process beside CodeQL, and CodeQL claims
roughly 14.5 GB of the runner's 16 GB by default, leaving the extractor almost
nothing. Cap CodeQL at 8 GB so the extractor has room to finish.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

threads: 2 is refuted. Reporting it before it becomes folklore.

I pushed a thread cap at 840b92e7 on the reasoning that the Go extractor defaults to one worker per
core and that halving concurrency would halve peak memory. It did not work, and the measurement is
unambiguous.

The cap took effect — this is the first thing to rule out, because a with: key Actions does not
recognise is silently ignored and would look exactly like a cap that did nothing:

While resolving threads, found a cgroup CPUs file with 4 CPUs
CODEQL_RAM: 14575
CODEQL_THREADS: 2

And the failure is unchanged. Analyze (go) started 03:12:07, died 03:23:25 on the same
The runner has received a shutdown signal, mid-extraction — about eleven minutes, which is where
both threads: 4 attempts died too. Halving worker concurrency moved neither the outcome nor the
time to it.

attempt THREADS outcome time to death
6b8006c1 devantler-tech/actions#1 4 runner killed ~11.5 min
6b8006c1 devantler-tech/actions#2 4 runner killed ~11.5 min
840b92e7 2 runner killed ~11.3 min

So peak memory is not scaling with the extractor's worker count, and the thread hypothesis is dead.

What that leaves, and what is now in flight

If concurrency is not the driver, process headroom is the remaining candidate. The Go extractor runs
as its own process beside CodeQL — it is a Go program with its own heap, not work inside CodeQL's
JVM — and CODEQL_RAM: 14575 hands CodeQL roughly 14.5 GB of the runner's 16 GB. That leaves the
extractor a sliver of what it needs to hold a 5490-package graph, which fits both the thrashing before
the kill and the fact that thread count is irrelevant to it.

4d6efe91 therefore replaces the thread cap with ram: 8192 — one variable, again, and the last one
I have a reasoned candidate for. The workflow comment has been rewritten to say what is actually
known rather than repeating the refuted claim.

If this one also fails, the honest conclusion is that the advanced workflow cannot make this
extraction fit a standard runner, and the fix is a larger runner or a narrower Go extraction scope —
not another knob. I will not keep trying knobs past that.

Nothing here changes the PR's case

The deliverable was never a green Go scan. It is that a failed scan can be retried at all:
gh run rerun --failed is accepted on this workflow and refused on the managed run. That is
demonstrated and unaffected by this result — and it is exactly what #7128 is stuck on right now,
with no lever of its own.

The other two languages continue to reach 🔬 Perform CodeQL analysis and fail only on
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled.

Still a draft, still authority-blocked on #6767.

Neither `ram` nor `threads` avoids the Go extractor exhausting the runner,
so the workflow now runs on the same budget as the default setup it
replaces and the comment states the constraint instead of a cap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

ram: 8192 is refuted, and the same commit proves why no budget knob will work

I pushed ram: 8192 as the last reasoned candidate for the extractor crash. It is refuted, and I have removed it in 86bd808e. The workflow now sets no ram and no threads, so its budget matches the default setup it replaces.

The cap took effect and changed nothing. Analyze (go) on 4d6efe91 logged CODEQL_RAM: 8192 (down from 14575) at init, then died at 03:38:43Z with ##[error]The runner has received a shutdown signal — 12.4 min in, mid-extraction, with 🔬 Perform CodeQL analysis skipped. I checked the init log before concluding, because an unrecognised with: key is silently ignored by Actions and would look identical to a cap that did nothing.

That makes three single-variable experiments, all dying in the same 11–13 minute band:

Configuration Result
threads: 2, ram default died ~11.3 min
ram: 8192, threads default died ~12.4 min
ram 14575 / threads 4 (default-setup budget) died ~11.5 min

The decisive observation is on this exact head. The GitHub-managed default-setup run and this workflow's run both started within 2 seconds of each other on 4d6efe91:

  • managed run 35486545387 — Analyze (go) 03:25:54Z → 03:47:17Z, success (21.4 min)
  • this workflow 35486547223 — Analyze (go) 03:25:56Z → 03:38:47Z, died (12.4 min)

Same commit, same runner class, same code. So the crash is not a property of this workflow or of any setting in it — the extraction sits right at the runner's memory ceiling and clears it or does not. That is exactly the "intermittent" in #6767, and it is why knob-tuning cannot fix it.

The thrashing signature is unambiguous. Within one run, extraction of ksail's own small source files degraded by two orders of magnitude just before death:

  • 03:37:51Z — pkg/client/k9s/client.go 125 ms, doc.go 487 ms
  • 03:38:37Z — pkg/cli/cmd/cluster/cluster.go 48,120 ms, list_output.go 47,083 ms, render.go 46,427 ms, and six more in the 46 s band

There is also a 40-second stall (03:37:51Z → 03:38:31Z) with nothing logged. Small files do not become 400× slower because of CPU; that is paging.

Scope, for the record. The autobuilder reports Found 3 go.mod files in: desktop/go.mod, go.mod, third_party/go-archive/go.mod and extracts all three module graphs in one job; desktop/ pulls in wails/v3. 837 distinct files were extracted before death.

What this means for this PR

Nothing about the crash is introduced or worsened here, and the PR's deliverable is untouched: a run of this workflow can be re-run (gh run rerun 35483555479 --failed was accepted → attempt=2), and a managed run cannot. This head is the cleanest illustration of why that matters — the managed run happened to clear the ceiling this time, but when one does not (as on #7128), there is no way to clear it short of a new commit, which a bot-authored release PR will never produce.

Reducing the memory the extraction needs is a separate concern from owning the workflow, and I have written it up on #6767 rather than tuning further here. It stays a draft: merging is authority-blocked on default setup being switched off.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Evaluation record — exercised at 86bd808e, and removing the cap changed the Go outcome

I read the real runs rather than reasoning about the workflow, on head 86bd808e (run 35488373507).

Two of three languages are proven end-to-end. Analyze (actions) and Analyze (javascript-typescript) each initialised CodeQL, ran the analysis to completion, and produced a SARIF file. Both then failed on exactly one thing:

##[error]Code Scanning could not process the submitted SARIF file:
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled

That is the whole pipeline working and being refused at the last step by a repository setting. Nothing in the workflow is left unexercised for those two languages — there is no further evidence this PR can produce for them until default setup is switched off.

Analyze (go) got further than it ever has. 🏗️ Autobuild completed successfully and the job moved on to 🔬 Perform CodeQL analysis. Every previous attempt died mid-extraction with The runner has received a shutdown signal at ~11–12 minutes. This is the first time extraction has finished on this workflow.

That is also a second, independent line of evidence that the caps I had added were the wrong move — and if anything harmful rather than merely ineffective:

Head Budget Extraction
managed run on 4d6efe91 uncapped (14575/4) completed
86bd808e (this one) uncapped, cap removed completed
4d6efe91 ram: 8192 died ~12.4 min
840b92e7 threads: 2 died ~11.3 min

Both capped runs died; both uncapped runs completed. I am not claiming that settles it — extraction sits right at the runner's memory ceiling and is marginal either way, so n=2 per arm proves nothing on its own. But it is consistent, it points the same direction as the same-commit comparison, and it means the honest budget for this workflow is the one default setup uses. That is what 86bd808e now ships, and why the workflow comment no longer asserts a cap rationale.

Judged as the change's user: what I wanted from owning this workflow was the ability to re-run a crashed scan. That works — gh run rerun <id> --failed is accepted here and produced attempt=2, where the managed equivalent refuses. I have since confirmed on #6767 that all four re-run avenues GitHub exposes are refused for a managed run, so this is not a marginal convenience; it is the only recovery path that exists.

State

Staying a draft. The remaining blocker is unchanged and is not in this PR: default setup must be switched off in Settings → Code security → CodeQL analysis before any of these results can upload (#6767, ask recorded on this PR's body 2026-09-19, renew at 14 days).

Analyze (go)'s query phase was still running at hand-off. Its conclusion does not change the two-language proof above, and the extractor memory problem it may still hit is tracked separately in #7131 — I am deliberately not tuning it further here.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

@devantler I will review the changes.

✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Analyze (go) finished too — all three languages are now proven end-to-end at 86bd808e

Completing the evaluation record above, which was posted while the Go query phase was still running.

Analyze (go) on run 35488373507:

Step Result
🔍 Initialize CodeQL success
🏗️ Autobuild success — 04:08:44Z → 04:20:38Z (11.9 min)
🔬 Perform CodeQL analysis ran 04:20:38Z → 04:24:52Z (4.2 min), produced SARIF, then failed on:
##[error]Code Scanning could not process the submitted SARIF file:
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled

So every language — actions, javascript-typescript and go — now initialises CodeQL, runs the analysis to completion, produces a SARIF file, and is refused at the upload by the repository setting and by nothing else. There is no remaining unexercised part of this workflow and no known defect in it. It is one settings change from working.

Two things worth noting about the Go run specifically, because both were open questions on this PR an hour ago:

Extraction completed for the first time. Every earlier attempt died mid-extraction with The runner has received a shutdown signal at ~11–12 minutes. Removing the ram: 8192 cap in 86bd808e is the only change between those attempts and this one, which makes the capped/uncapped split 2-for-2 in both directions (both uncapped runs completed, both capped runs died). I still would not call that proof — the extraction sits at the runner's memory ceiling and is marginal either way — but it is consistent, it agrees with the same-commit managed-vs-advanced comparison, and it confirms the right budget for this workflow is the one default setup already uses. The residual memory risk is #7131's scope, not this PR's.

The Go analysis is cheap once extraction clears. 4.2 minutes of queries against 11.9 minutes of extraction. Whatever is eventually done about the memory ceiling, the query cost is not the problem.

State

Still a draft, and the blocker is unchanged: default setup must be switched off in Settings → Code security → CodeQL analysis (#6767). The ask is recorded on this PR's body from 2026-09-19 and is not yet due for renewal.

What is new is that the evidence is now complete — there is nothing further this PR can demonstrate before that switch, and #7128 is sitting blocked behind the same switch with auto-merge armed and every other gate green.

@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository YAML (base), Organization UI (inherited)

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 0740437e-98fc-4a2f-b777-74e242c050b0

📥 Commits

Reviewing files that changed from the base of the PR and between 86bd808 and 4a1ce42.

📒 Files selected for processing (1)
  • .github/workflows/codeql.yaml

Included review availability: Your plan provides up to 2 included reviews per hour; 1 remains after this review.

📜 Recent review details
⏰ Context from checks skipped due to timeout. (1)
  • GitHub Check: Analyze (go)
🧰 Additional context used
📓 Path-based instructions (1)
Validate workflow changes with `mega-linter-runner -f go`; MegaLinter runs `actionlint` for GitHub Actions workflows.

📄 CodeRabbit inference engine (AGENTS.md)

Files:

  • .github/workflows/codeql.yaml
🔇 Additional comments (1)
.github/workflows/codeql.yaml (1)

59-59: 📐 Maintainability & Code Quality

The repository guideline requires mega-linter-runner -f go, but the absence of its output does not show that the command was skipped or failed. The guideline does not require reporting the result. No actionable finding remains.


📝 Walkthrough

Walkthrough

Adds a repository-owned CodeQL workflow for actions, go, and javascript-typescript. The workflow runs on changes to main, merge groups, and a weekly schedule. It runs only when ENABLE_ADVANCED_CODE_SCANNING is true, uses scoped permissions and a 60-minute timeout, and applies concurrency cancellation rules. Go uses autobuild extraction, while other languages use none.

Priority: ➖ Normal

Severity of issue fixed: Medium


Important

Pre-merge checks failed

Please resolve all errors before merging. Addressing warnings is optional.

❌ Failed checks (1 error)

Check name Status Explanation Resolution
Linked Issues check ❌ Error The workflow addresses [#6767] by adding an advanced CodeQL workflow with the same three languages, security-extended queries, a repository variable gate, and a documented re-run path. The change do… Identify the cgo-triggering package and document whether exclusion would reduce first-party coverage. Disable default CodeQL setup and provide or remove the source for the “Require code quality results” ruleset. Run the advanced workflow af…
✅ Passed checks (4 passed)
Check name Status Explanation
Out of Scope Changes check ✅ Passed The whole pull request adds only .github/workflows/codeql.yaml. Its triggers, permissions, language matrix, extraction mode, re-run behavior, and activation comments directly support the advanced Co…
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check. Docstring coverage is scoped to functions touched by this diff. Analyzed 0 functions across 0…
Title check ✅ Passed The title clearly states that the pull request adds repository-owned code scanning so failed scans can be re-run. It accurately summarizes the main change.
Description check ✅ Passed The description directly explains the CodeQL workflow, its purpose, activation conditions, and related repository settings.
Full details: Linked Issues check

Explanation

The workflow addresses [#6767] by adding an advanced CodeQL workflow with the same three languages, security-extended queries, a repository variable gate, and a documented re-run path. The change does not identify which package pulls cgo into Go extraction or document the coverage impact of excluding that package. It also does not confirm code_scanning ruleset behavior for an extractor error. Default setup remains enabled, so the advanced results cannot be processed and end-to-end validation is not established.

Resolution

Identify the cgo-triggering package and document whether exclusion would reduce first-party coverage. Disable default CodeQL setup and provide or remove the source for the “Require code quality results” ruleset. Run the advanced workflow after the settings change and confirm the intended code_scanning behavior for an extractor error.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

GitHub rejects an advanced configuration's results while default setup is
enabled, so this workflow's uploads failed on every pull request, including
its own, and the change could not merge on its own.

Gate the jobs on a default-off repository variable so the workflow lands
latent and does nothing until it is switched on. Activation is then one
deliberate, reversible step: disable default setup and set the variable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Activation detail (kept out of the PR body, which stays outcome-level).

Both halves are one switch, and either order leaves a brief gap, so do them together:

  1. Settings → Code security → CodeQL analysis — switch Default setup off.
  2. Settings → Secrets and variables → Actions → Variables — add repository variable
    ENABLE_ADVANCED_CODE_SCANNING = true.

To reverse: clear the variable and re-enable default setup.

Why this is safe to merge before you do either. While the variable is unset the three
analyze jobs are skipped, so nothing uploads and nothing is rejected. Default setup keeps
running and keeps satisfying the Require code scanning results ruleset, which is evaluated
on CodeQL alerts rather than on a check name. The required status check on this repository is
the single aggregate CI - Required Checks context, so the skipped Analyze (…) checks are
not required by name and skipping them changes no gate.

One thing to decide at activation. The Require code quality results ruleset
(type: code_quality) is currently fed by default setup. GitHub does not let a repository's
own workflow produce code-quality results, so once default setup is off that rule has no
source and would block every pull request. Give it another source or remove the ruleset in the
same change.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown

@devantler I will review the changes.

✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

Readiness record (detail kept out of the PR body).

1 — Programmatically tested. actionlint is clean on the workflow. Because a clean exit
proves nothing on its own, I ran a negative control: replacing the gate expression with an
undefined context makes actionlint fail on that exact line. Its error also enumerates the
available contexts and lists vars, which independently confirms the gate expression is valid
in a job-level if.

2 — Reviewed. CodeRabbit at this head, no actionable findings; 0 unresolved threads.

3 — Tried and evaluated as a user. Observed in real CI rather than reasoned about:

  • Off state (what merges today). The three Analyze (…) jobs went from FAILURE to SKIPPED
    with zero failures. The previous head failed with CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled; that error is gone.
  • Default setup is unaffected. Its own runs on this head still execute and succeed, so code
    scanning coverage is continuous — nothing is traded away while this sits latent.
  • On state. Not exercisable without disabling default setup, which is yours to do. The
    evidence that it works is the previous head: all three languages ran their queries and
    completed analysis, and were rejected only at the upload step, by the mutual-exclusivity
    error above. Activation removes exactly that rejection. I am flagging this honestly rather
    than claiming an end-to-end run I cannot perform.

Two things worth knowing, neither a defect:

  • While the flag is off, the skipped job surfaces as a check literally named
    Analyze (${{ matrix.language }}) — GitHub does not expand a matrix for a job it skips. It is
    cosmetic, it is not a required check, and it disappears at activation.
  • CodeQL = neutral — "1 configuration not found" on this PR is pre-existing and repo-wide,
    not caused by this change. Control: PR chore(deps): update agent skills #7128, which does not carry this workflow, shows the
    identical check and title, and ksail merges normally with it (test(ci): give the podinfo overlay a local base instead of a remote git clone #7129 merged yesterday).

Correction to the PR body's earlier claim. It said the built-in scan crashes "about one run
in three". Measured over the last 118 finished managed CodeQL runs: 22 failed (18.6%), and
4 of 31 on main (12.9%). Still a recurring wedge that cannot be re-run, but less frequent
than first stated, and the body now carries the measured figures.

@devantler
devantler marked this pull request as ready for review September 20, 2026 06:45
@devantler

Copy link
Copy Markdown
Contributor Author

🤖 Generated by the Agentic Engineer

The case for this change happened again while the change was in review.

At 06:19Z today the release bot pushed a new commit to #7128. The built-in scan for that commit
crashed four minutes later, and the pull request is blocked on it right now.

I then tried to re-run it, to confirm the premise rather than assume it. GitHub refuses:

This workflow run cannot be retried — HTTP 403

That is the whole problem in one exchange. The run is GitHub-managed, so there is no workflow file
to re-run and no way to clear the failure except another commit — and the only author of commits on
that branch is a bot that has already pushed. So it sits blocked until a person intervenes, which is
exactly what #6767 describes and what this change removes.

For completeness, the same crash is what the measured 18.6% failure rate refers to; #7128 is simply
the instance that is live at the moment of writing.

@devantler
devantler merged commit 96528d7 into main Sep 20, 2026
54 checks passed
@devantler
devantler deleted the claude/ci-codeql-advanced-setup-6767 branch September 20, 2026 06:49
@github-project-automation github-project-automation Bot moved this from 🫴 Ready to ✅ Done in 🌊 Project Board Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: ✅ Done

Development

Successfully merging this pull request may close these issues.

Intermittent CodeQL Analyze (go) Autobuild runner termination blocks merges, and GitHub refuses to re-run it

2 participants