ci: run code scanning from this repository so a crashed scan can be re-run - #7130
Conversation
Default setup's Go autobuild extracts every go.mod on one runner. The desktop module replaces the root module, so the root dependency graph is extracted twice and the runner is shut down mid-extraction. Move to an advanced setup that analyses the root and desktop modules on separate runners and can be re-run. Refs #6767 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
✅MegaLinter analysis: Success✅ Linters with no issuesactionlint, bash-exec, git_diff, hadolint, jscpd, jsonlint, lychee, markdown-table-formatter, markdownlint, prettier, prettier, shellcheck, shfmt, stylelint, syft, trivy-sbom, trufflehog, v8r, v8r, yamllint Notices
See detailed reports in MegaLinter artifacts
|
CodeQL claims ~14.5 GB of the 16 GB runner for the Go extractor. In manual build mode the traced go build runs alongside it, so the job is killed mid-build. Cap the extractor and the compiler, and drop the code-quality analysis kind the action rejects in custom workflows. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The new CodeQL workflow tripped two repository guards.
`TestNoDefaultBranchWorkflowCancelsRunsInProgress` rejects any workflow that
runs on main and may cancel in-progress runs there, so one merge cannot evict
the previous merge's checks. Its allowlist is an exact-match whitelist by
design — substring forms admitted bypasses such as
`${{ github.ref != 'refs/heads/main' || true }}` — so conform to the approved
expression rather than widening the guard.
zizmor flagged both local action references under `unpinned-uses`. Adopt
GitHub's self-repository syntax, which resolves to this repository at the
executing commit, so the action and the workflow calling it can never drift.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit took zizmor's advice and moved both local action references to GitHub's self-repository syntax. That traded a zizmor notice for a harder failure: actionlint v1.7.12, which MegaLinter v10.1.0 pins, does not recognise `$/` and rejects it as a malformed ref, so `Validate Go Project` went red. The two linters disagree, and only one of them can be satisfied by the reference itself. Suppress zizmor's `self-repository` finding on exactly the two lines it fires on, with the reason recorded inline, rather than silencing actionlint's malformed-ref check — that check is repository-wide and would stop reporting genuinely broken action references. This also keeps the file consistent with the rest of the repository, which uses `./` throughout. Verified with the same tools CI runs: actionlint exits 0 where `$/` reproduced CI's exact error, and zizmor reports no findings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hygiene pass at Fixed —
Fixed — zizmor raised I deliberately did not silence actionlint's malformed-ref check instead — it is repository-wide, and blinding it to catch a false positive on a local action would stop it reporting genuinely broken references. Verified with the same tools CI runs ( Still red, and not fixable from here — Every Default setup reads One finding worth keeping for when it can run. The earlier This also gates #7128, which is otherwise merge-ready and blocked solely by the managed CodeQL check — GitHub refuses to re-run those, so only a new head commit retriggers one. |
The Go extractor resolves the module graph with `go list` and reads types from source, covering every go.mod in the workspace in one job with no compiler. Tracing a `go build` instead runs the compiler alongside the extractor, and the two together exceed the runner: `Build root module` was killed (exit 143) at both heads that ran it, while CodeQL's own autobuilder chose buildless extraction on the same commit and analysed all three modules successfully in 9.5 minutes at the default memory budget. Drops the per-module matrix split, the memory caps, the disk and toolchain setup steps and the build steps they existed for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`build-mode: none` is rejected outright for Go by CodeQL 2.27 ("Go does not
support the none build mode. Please try using one of the following build
modes instead: autobuild, manual"), so the init step failed in 25 seconds.
`autobuild` is what reaches the same extractor: on this repository the Go
autobuilder runs `go list` and extracts types from source without compiling,
which is the configuration default setup uses and the one measured succeeding
on all three modules in 9.5 minutes.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Exercised at
|
| configuration | CODEQL_RAM / THREADS |
result |
|---|---|---|
default setup, autobuild, all 3 modules |
14575 / 4 | success, 15.5 min |
this PR, manual + go build -p 2 ./..., root module |
8192 / 2 | killed, exit 143 at 🏗️ Build root module |
The configuration being blamed was the one passing. The reason is that autobuild never compiles here:
its log reads Running go list to resolve package and module directories → resolved 5490 packages →
Done extracting packages → Success: extraction succeeded for all 3 discovered project(s). The
compiler only appears when a workflow asks for manual, and then it and the extractor together exceed
the runner.
2. build-mode: none is not the way to ask for that extractor. Pushed at 82c8aa6b; CodeQL 2.27
rejected it in 25 seconds — Go does not support the none build mode. Please try using one of the following build modes instead: autobuild, manual. autobuild is the supported spelling, pushed at
6b8006c1.
3. The remaining failure is not ours. At 82c8aa6b, Analyze (actions) and
Analyze (javascript-typescript) both ran to completion and failed only on
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled —
the settings change recorded on #6767. Everything up to the upload works.
Net effect
133 lines → 69. Gone: the per-module matrix, the memory caps, the disk-free and toolchain setup steps,
both build steps, and both zizmor: ignore[self-repository] suppressions (with them, the actionlint
conflict noted on devantler-tech/.github#271 no longer touches this file). actionlint, zizmor and
the 86 workflow-contract tests pass locally; CI - KSail, Validate Go Project, zizmor and
dependency-review are green on this head.
Staying a draft: it cannot be validated end to end, or merged, until default setup is off.
Attempt 2 landed, and it rules configuration out as the causeThe re-run this PR exists to make possible completed at 02:41Z. Three results, all at 1. The advanced workflow runs on exactly the same memory budget as default setup. This was Those are the same numbers the successful default-setup run used on this commit. Nothing in this 2. Everything except Go reached the upload and failed only on the settings conflict. The earlier evaluation established this at 3. Go died the same way again. The honest readingSame commit, same memory budget, opposite outcomes: default setup succeeded, this workflow's two One untested difference is worth naming rather than assuming away: the rewrite deleted the disk-free None of this changes the PR's claim. The deliverable is that a failed scan can be retried at all, and Status: staying a draft. The merge is authority-blocked on #6767 — default setup must be turned |
The Go extractor defaults to one worker per core. On this repository's module graph that exhausts a 16 GB runner part-way through extraction and the runner is killed, which is the intermittent failure #6767 describes. Measured on the same commit, both a managed run and this workflow used CODEQL_RAM 14575 / THREADS 4, so the failure is not a configuration difference between them. Default setup cannot set this cap. This workflow can, which is a second reason to own it alongside re-runnability. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Capping extraction at two threads was measured and changed nothing: the cap took effect (CODEQL_THREADS: 2) and the runner was still killed at the same eleven-minute mark. So peak memory is not scaling with worker concurrency. The Go extractor runs as its own process beside CodeQL, and CodeQL claims roughly 14.5 GB of the runner's 16 GB by default, leaving the extractor almost nothing. Cap CodeQL at 8 GB so the extractor has room to finish. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
| attempt | THREADS |
outcome | time to death |
|---|---|---|---|
6b8006c1 devantler-tech/actions#1 |
4 | runner killed | ~11.5 min |
6b8006c1 devantler-tech/actions#2 |
4 | runner killed | ~11.5 min |
840b92e7 |
2 | runner killed | ~11.3 min |
So peak memory is not scaling with the extractor's worker count, and the thread hypothesis is dead.
What that leaves, and what is now in flight
If concurrency is not the driver, process headroom is the remaining candidate. The Go extractor runs
as its own process beside CodeQL — it is a Go program with its own heap, not work inside CodeQL's
JVM — and CODEQL_RAM: 14575 hands CodeQL roughly 14.5 GB of the runner's 16 GB. That leaves the
extractor a sliver of what it needs to hold a 5490-package graph, which fits both the thrashing before
the kill and the fact that thread count is irrelevant to it.
4d6efe91 therefore replaces the thread cap with ram: 8192 — one variable, again, and the last one
I have a reasoned candidate for. The workflow comment has been rewritten to say what is actually
known rather than repeating the refuted claim.
If this one also fails, the honest conclusion is that the advanced workflow cannot make this
extraction fit a standard runner, and the fix is a larger runner or a narrower Go extraction scope —
not another knob. I will not keep trying knobs past that.
Nothing here changes the PR's case
The deliverable was never a green Go scan. It is that a failed scan can be retried at all:
gh run rerun --failed is accepted on this workflow and refused on the managed run. That is
demonstrated and unaffected by this result — and it is exactly what #7128 is stuck on right now,
with no lever of its own.
The other two languages continue to reach 🔬 Perform CodeQL analysis and fail only on
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled.
Still a draft, still authority-blocked on #6767.
Neither `ram` nor `threads` avoids the Go extractor exhausting the runner, so the workflow now runs on the same budget as the default setup it replaces and the comment states the constraint instead of a cap. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
| Configuration | Result |
|---|---|
threads: 2, ram default |
died ~11.3 min |
ram: 8192, threads default |
died ~12.4 min |
| ram 14575 / threads 4 (default-setup budget) | died ~11.5 min |
The decisive observation is on this exact head. The GitHub-managed default-setup run and this workflow's run both started within 2 seconds of each other on 4d6efe91:
- managed run
35486545387—Analyze (go)03:25:54Z → 03:47:17Z, success (21.4 min) - this workflow
35486547223—Analyze (go)03:25:56Z → 03:38:47Z, died (12.4 min)
Same commit, same runner class, same code. So the crash is not a property of this workflow or of any setting in it — the extraction sits right at the runner's memory ceiling and clears it or does not. That is exactly the "intermittent" in #6767, and it is why knob-tuning cannot fix it.
The thrashing signature is unambiguous. Within one run, extraction of ksail's own small source files degraded by two orders of magnitude just before death:
03:37:51Z—pkg/client/k9s/client.go125 ms,doc.go487 ms03:38:37Z—pkg/cli/cmd/cluster/cluster.go48,120 ms,list_output.go47,083 ms,render.go46,427 ms, and six more in the 46 s band
There is also a 40-second stall (03:37:51Z → 03:38:31Z) with nothing logged. Small files do not become 400× slower because of CPU; that is paging.
Scope, for the record. The autobuilder reports Found 3 go.mod files in: desktop/go.mod, go.mod, third_party/go-archive/go.mod and extracts all three module graphs in one job; desktop/ pulls in wails/v3. 837 distinct files were extracted before death.
What this means for this PR
Nothing about the crash is introduced or worsened here, and the PR's deliverable is untouched: a run of this workflow can be re-run (gh run rerun 35483555479 --failed was accepted → attempt=2), and a managed run cannot. This head is the cleanest illustration of why that matters — the managed run happened to clear the ceiling this time, but when one does not (as on #7128), there is no way to clear it short of a new commit, which a bot-authored release PR will never produce.
Reducing the memory the extraction needs is a separate concern from owning the workflow, and I have written it up on #6767 rather than tuning further here. It stays a draft: merging is authority-blocked on default setup being switched off.
Evaluation record — exercised at
|
| Head | Budget | Extraction |
|---|---|---|
managed run on 4d6efe91 |
uncapped (14575/4) | completed |
86bd808e (this one) |
uncapped, cap removed | completed |
4d6efe91 |
ram: 8192 |
died ~12.4 min |
840b92e7 |
threads: 2 |
died ~11.3 min |
Both capped runs died; both uncapped runs completed. I am not claiming that settles it — extraction sits right at the runner's memory ceiling and is marginal either way, so n=2 per arm proves nothing on its own. But it is consistent, it points the same direction as the same-commit comparison, and it means the honest budget for this workflow is the one default setup uses. That is what 86bd808e now ships, and why the workflow comment no longer asserts a cap rationale.
Judged as the change's user: what I wanted from owning this workflow was the ability to re-run a crashed scan. That works — gh run rerun <id> --failed is accepted here and produced attempt=2, where the managed equivalent refuses. I have since confirmed on #6767 that all four re-run avenues GitHub exposes are refused for a managed run, so this is not a marginal convenience; it is the only recovery path that exists.
State
Staying a draft. The remaining blocker is unchanged and is not in this PR: default setup must be switched off in Settings → Code security → CodeQL analysis before any of these results can upload (#6767, ask recorded on this PR's body 2026-09-19, renew at 14 days).
Analyze (go)'s query phase was still running at hand-off. Its conclusion does not change the two-language proof above, and the extractor memory problem it may still hit is tracked separately in #7131 — I am deliberately not tuning it further here.
@coderabbitai review |
|
✅ Action performedReview finished.
|
|
| Step | Result |
|---|---|
🔍 Initialize CodeQL |
success |
🏗️ Autobuild |
success — 04:08:44Z → 04:20:38Z (11.9 min) |
🔬 Perform CodeQL analysis |
ran 04:20:38Z → 04:24:52Z (4.2 min), produced SARIF, then failed on: |
##[error]Code Scanning could not process the submitted SARIF file:
CodeQL analyses from advanced configurations cannot be processed when the default setup is enabled
So every language — actions, javascript-typescript and go — now initialises CodeQL, runs the analysis to completion, produces a SARIF file, and is refused at the upload by the repository setting and by nothing else. There is no remaining unexercised part of this workflow and no known defect in it. It is one settings change from working.
Two things worth noting about the Go run specifically, because both were open questions on this PR an hour ago:
Extraction completed for the first time. Every earlier attempt died mid-extraction with The runner has received a shutdown signal at ~11–12 minutes. Removing the ram: 8192 cap in 86bd808e is the only change between those attempts and this one, which makes the capped/uncapped split 2-for-2 in both directions (both uncapped runs completed, both capped runs died). I still would not call that proof — the extraction sits at the runner's memory ceiling and is marginal either way — but it is consistent, it agrees with the same-commit managed-vs-advanced comparison, and it confirms the right budget for this workflow is the one default setup already uses. The residual memory risk is #7131's scope, not this PR's.
The Go analysis is cheap once extraction clears. 4.2 minutes of queries against 11.9 minutes of extraction. Whatever is eventually done about the memory ceiling, the query cost is not the problem.
State
Still a draft, and the blocker is unchanged: default setup must be switched off in Settings → Code security → CodeQL analysis (#6767). The ask is recorded on this PR's body from 2026-09-19 and is not yet due for renewal.
What is new is that the evidence is now complete — there is nothing further this PR can demonstrate before that switch, and #7128 is sitting blocked behind the same switch with auto-merge armed and every other gate green.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository YAML (base), Organization UI (inherited) Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (1)
Included review availability: Your plan provides up to 2 included reviews per hour; 1 remains after this review. 📜 Recent review details⏰ Context from checks skipped due to timeout. (1)
🧰 Additional context used📓 Path-based instructions (1)Validate workflow changes with `mega-linter-runner -f go`; MegaLinter runs `actionlint` for GitHub Actions workflows.📄 CodeRabbit inference engine (AGENTS.md) Files:
🔇 Additional comments (1)
📝 WalkthroughWalkthroughAdds a repository-owned CodeQL workflow for Priority: ➖ Normal Severity of issue fixed: Medium Important Pre-merge checks failedPlease resolve all errors before merging. Addressing warnings is optional. ❌ Failed checks (1 error)
✅ Passed checks (4 passed)
Full details: Linked Issues checkExplanation The workflow addresses [ Resolution Identify the cgo-triggering package and document whether exclusion would reduce first-party coverage. Disable default CodeQL setup and provide or remove the source for the “Require code quality results” ruleset. Run the advanced workflow after the settings change and confirm the intended Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
GitHub rejects an advanced configuration's results while default setup is enabled, so this workflow's uploads failed on every pull request, including its own, and the change could not merge on its own. Gate the jobs on a default-off repository variable so the workflow lands latent and does nothing until it is switched on. Activation is then one deliberate, reversible step: disable default setup and set the variable. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Activation detail (kept out of the PR body, which stays outcome-level). Both halves are one switch, and either order leaves a brief gap, so do them together:
To reverse: clear the variable and re-enable default setup. Why this is safe to merge before you do either. While the variable is unset the three One thing to decide at activation. The |
@coderabbitai review |
|
✅ Action performedReview finished.
|
Readiness record (detail kept out of the PR body). 1 — Programmatically tested. 2 — Reviewed. CodeRabbit at this head, no actionable findings; 0 unresolved threads. 3 — Tried and evaluated as a user. Observed in real CI rather than reasoned about:
Two things worth knowing, neither a defect:
Correction to the PR body's earlier claim. It said the built-in scan crashes "about one run |
The case for this change happened again while the change was in review. At 06:19Z today the release bot pushed a new commit to #7128. The built-in scan for that commit I then tried to re-run it, to confirm the premise rather than assume it. GitHub refuses:
That is the whole problem in one exchange. The run is GitHub-managed, so there is no workflow file For completeness, the same crash is what the measured 18.6% failure rate refers to; #7128 is simply |

Why
GitHub's built-in code scanning crashes on KSail often enough to matter — 22 of the last 118 finished runs, and 4 of 31 on the main branch. A crashed scan blocks merging, and GitHub refuses to re-run it, so the only way out is a fresh commit. A bot-authored release pull request never pushes one, so it stays stuck until a person steps in.
What
Adds a code scan that lives in this repository, covering the same languages with the same rules, so a crashed run can simply be re-run.
It ships switched off and does nothing until it is turned on, because GitHub refuses results from a repository's own scan while the built-in one is active. That keeps the two from colliding and makes this safe to merge on its own — today's scanning carries on unchanged.
Fixes #6767
👉 After merge/promotion: turning it on is one reversible step and it is yours to take — switch the built-in code scanning off, and set the repository variable this scan reads to true. Doing so also removes what currently feeds the "Require code quality results" rule, which GitHub no longer lets a repository's own scan produce, so that rule needs a new source or removal at the same time. A comment below names the variable and the exact settings.