Repository navigation
OpenCode: ACP hosting, gated unattended reviewer, and SessionStart memory injection - #484
Conversation
`opencode acp` speaks ACP protocolVersion 1 over stdio, so hosting is one
descriptor plus a factory registration. Everything asserted about the vendor
here was measured against opencode 1.18.9 rather than read off its
advertisement — `docs/probes/2026-08-07-opencode-acp/`, since the advertisement
demonstrably decides nothing: Kiro, Gemini and Copilot all publish the same
`mcpCapabilities {http, sse}` shape and disagree with each other about stdio.
Two findings drove the shape of this:
Model selection needed a second read shape. OpenCode publishes no `models`
object at all — its selectable models are `session/new`'s
`configOptions[id=="model"].options[]`. The shared resolver looked only at
`models.availableModels`, so a caller-requested model would have been silently
discarded on every OpenCode launch with nothing in the logs naming the reason.
`AcpSessionModelList` now reads both shapes, `models` first so a vendor growing
a mirror of the same fact cannot have the mirror reinterpreted as authority.
The write half (`session/set_config_option`) is verified at EFFECT level, not
on its echoed success: with `currentValue` moved to a free model the model
self-identified as that id and never named the previous one. That distinction
is why Gemini is still on `NoOpModelSelector`.
Capture precedence had to be decided, and OpenCode is the one vendor where two
kcap capture paths can be live at once. The operator's installed live-ingest
plugin loads INSIDE the `opencode acp` process and starts a second, top-level
recording of the very session the ACP mapper is already recording. Measured
with a controlled pair (equal 10s dwell, plugin confirmed installed, so
neither arm is vacuous): default env created the plugin's capture file,
`OPENCODE_PURE=1` did not. So the daemon passes it on EVERY hosted launch,
interactive included — scoping it to review launches would leave the common
case double-ingested, and the duplication is timing-dependent, which is worse
to operate than a deterministic one. Accepted cost, stated rather than
discovered later: a hosted agent does not carry the operator's other plugins.
For the reviewer, OpenCode's whole posture is env-shaped — `opencode acp`
accepts none of the global flags — so `UnattendedTrustArgv` is empty and
`OpenCodeLaunchEnvironment` owns it instead. The permission table is
deny-all plus the read family plus the injected result channel: no shell, no
write, no network. Two controls make that claim mean something. Flipping `bash`
to allow makes the shell actually run, so the denial is attributable to the
rule rather than to a model that chose not to shell out; and removing the
`{server}_*` entry takes the result channel out of the model's toolset
entirely, which rules out the alternative explanation that MCP tools bypass
permissions and the entry is decorative. Shell success is judged on a file
existing, never on the answer text — an earlier revision scored the denied arm
as having run the shell because the model quoted the command while explaining
it could not.
A denied tool is ABSENT from the model's surface rather than refused when
called, so a correct launch raises no interaction frame at all and `Fail` is
the honest policy. That also closes the skill-derail hazard structurally:
OpenCode reads skills from HOME, which the isolated config dir does not
suppress, but with the `skill` tool denied it omits the skills section from the
system prompt entirely.
Containment is source suppression, as for Kiro: an empty per-launch
`OPENCODE_CONFIG_DIR` removes the operator's global MCP servers (`kcap-flows`
among them, hence nested flows), and `OPENCODE_DISABLE_PROJECT_CONFIG` stops
the reviewed BRANCH's own config and AGENTS.md reaching the reviewer judging
it. Credentials live outside the config dir — measured, and the fact the whole
approach depends on.
Unattended is gated behind operator consent plus a version floor, because the
narrow tool surface bounds a WRITE and not a read: `read`/`grep`/`glob`/`list`
are whole-filesystem primitives under the daemon's uid. The affirm verb and the
service-unit env allowlist both had to learn the vendor, and the allowlist is
now DERIVED from the affirmable-reviewer registry rather than listed twice —
they were two hand-maintained lists of one set, and a vendor missing from the
allowlist means a supervised install silently drops its consent flag while the
refusal text points at a variable the unit never received.
The Kiro-only launch-budget and cleanup machinery is generalized to both
vendors through one `ReviewerLaunchTimeoutSeconds`, so the budget, the
first-output deadline and the cleanup hook cannot reach different conclusions —
a budget without its cleanup strands exactly the directory the budget fired to
reclaim.
Not done, and not claimed: no real unattended round has been run end to end.
The reviewer is off by default behind operator consent, so nothing is live
until someone opts in. `SupportsReconnectResume` stays false as UNPROBED,
which is a different claim from Kiro's and Gemini's measured-ineligible.
Fixes AI-1405
Fixes AI-1411
Tests (targeted): OpenCodeHostedLaunchTests 7/7, OpenCodeReviewerLaunchTests
12/12, OpenCodeReviewerCapabilityTests 14/14, AcpSessionModelListTests 17/17.
Neighbours green: AcpHostedAgentRuntimeFactoryTests 82/82,
KiroReviewerLaunchTests 10/10, KiroReviewerHomeTests 9/9,
KiroReviewerCapabilityTests 32/32, GeminiReviewerLaunchTests 31/31,
AcpVendorDescriptorTests 22/22, ServiceEnvironmentTests 17/17,
DaemonReviewerCommandTests 8/8, DaemonRunnerCursorAvailabilityTests 35/35.
AOT publish clean (no IL2026/IL3050).
Mutation-checked: scoping the plugin suppression to review launches fails 2
tests; adding `bash` to the read set fails 1; bypassing the consent gate fails
1. One survivor was found and fixed — the `configOptions` id filter, whose
test used a fixture where `model` happened to lead the array.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OpenCode has no shell hooks, so the generated plugin's existing shell-out to
`kcap hook --opencode --event session-start` becomes a small data bridge:
stdout carries the team-memory fragment (raw text, no envelope) and the plugin
appends it to the model's system prompt through
`experimental.chat.system.transform`. kcap keeps every decision behind it — auth,
scope, the durable once-per-session lease, the byte budget — so the plugin
never talks to the server and never renders anything.
The transform contract was read out of the opencode 1.18.9 binary rather than
guessed, and four of its properties are load-bearing:
1. The same hook name is ALSO triggered by OpenCode's agent-config generator,
with no `sessionID`. Injecting there would put the team-memory index into
every generated agent definition, so an absent id declines.
2. `system` arrives holding ONE pre-joined element and is rebuilt per LLM
request. So appending on every request is what keeps the index in context,
and each array still ends up with exactly one copy. The marker check keeps
that correct in the other direction too, if a future build ever retains
transformed entries instead of rebuilding.
3. Mutations are honoured, but `system[0]` is OpenCode's whole assembled
prompt and its own collapse of entries 1..n is conditional on `system[0]`
being unchanged — so this appends and never assigns.
4. OpenCode awaits the hook with no try/catch of its own, so a throw would
surface as a failed model request. Nothing here can throw.
The gated live cert earned its place immediately. It caught a race that the
whole unit suite could not see: the in-flight start marker was published inside
`start()`, which only runs after classification resolves, while OpenCode
publishes `session.created` and then issues the first LLM request. The
transform reliably landed in that window, saw no pending start, and declined to
inject — so a one-turn session got no index at all, intermittently rather than
visibly. `ensureStarted` now publishes the marker synchronously, before the
classify's first await, and both event paths go through it.
Cert result: positive case and negative control both pass against opencode
1.18.9, with the model reproducing a nonce that could only have reached it
through the injected index. Recorded in-code as the verified experimental-API
baseline, because that API is experimental and an upstream change would stop
delivering while every unit test still passed — this cert is the only thing
that would notice. The gate names the plugin-installed precondition (without
the plugin the negative control passes vacuously) and the requirement that the
`kcap` on PATH be the same build as the plugin (a released `kcap` emits no
fragment, and the positive case then fails for the wrong reason). It also takes
an optional model override: the account default answered 401 on the first real
run, failing the cert on the exit code rather than on the nonce.
Fixes AI-1466
Tests (targeted): OpenCodeSessionStartMemoryTests 16/16,
OpenCodeMemoryIndexLiveCertTests 2/2 live (2 skipped at ~1ms with the gate
unset, as CI sees them), OpenCodeExtensionInstallerTests 4/4,
PluginCommandOpenCodeTests 9/9, SessionStartMemoryFoundationTests 31/31,
MemoryIndexEmitterTests 13/13. AOT publish clean.
README carries the docs for all three OpenCode issues in this branch: the
hosted-agent section, the unattended-reviewer section, and this row of the
memory-injection matrix — one file, so it lands once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR Summary by QodoAdd OpenCode ACP hosting, gated unattended reviewer, and session-start memory injection
AI Description
Diagram
High-Level Assessment
Files changed (29)
|
Code Review by Qodo
1. OpenCodeLaunchEnvironment docs too verbose
|
| if (sessionNewResult.ValueKind != JsonValueKind.Object) | ||
| return []; |
There was a problem hiding this comment.
1. Acpsessionmodellist uses valuekind 📘 Rule violation ⚙ Maintainability
New JSON parsing code directly compares JsonElement.ValueKind instead of using the project-standard JsonElementExtensions, reducing consistency and increasing the chance of future kind-check mistakes.
Agent Prompt
## Issue description
`AcpSessionModelList` performs JSON kind checks via direct `JsonElement.ValueKind` comparisons. The compliance checklist requires using the project-provided `JsonElementExtensions` helpers for kind checks.
## Issue Context
`JsonElementExtensions` already provides `IsObject` and `Arr(...)` helpers intended to avoid scattered manual `ValueKind` comparisons.
## Fix Focus Areas
- src/Capacitor.Cli.Core/Acp/AcpSessionModelList.cs[38-63]
ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools
| /// <summary> | ||
| /// Applies the settings that every hosted OpenCode launch needs, interactive included. | ||
| /// | ||
| /// <para><b>Why <c>OPENCODE_PURE=1</c> is not optional.</b> OpenCode is the one vendor where two |
There was a problem hiding this comment.
2. Opencodelaunchenvironment docs too verbose 📘 Rule violation ⚙ Maintainability
New/modified comments are extremely long and contain extensive operational narrative, making them harder to maintain and obscuring the code’s intent.
Agent Prompt
## Issue description
Several newly added comment blocks are overly verbose (multi-paragraph operational/probe narrative) rather than concise context, which violates the project guideline to keep comments minimal and prefer self-explanatory code.
## Issue Context
The repository already has dedicated probe documentation under `docs/probes/...`; long-form measurement/provenance is better maintained there, while code comments should stay brief and point to the doc.
## Fix Focus Areas
- src/Capacitor.Cli.Daemon/Acp/OpenCodeLaunchEnvironment.cs[5-83]
- src/Capacitor.Cli.Daemon/Acp/AcpVendorDescriptor.cs[378-444]
ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools
| var availableModels = AcpSessionModelList.Extract(sessionNewResult); | ||
|
|
||
| var resolvedModelId = AcpModelResolver.Resolve(requestedModel, availableModels); | ||
| if (resolvedModelId is null) { |
There was a problem hiding this comment.
3. Missing model parse diagnostics 🐞 Bug ◔ Observability
AcpSessionModelList swallows JsonException and returns an empty model list, and SessionModelResolution now treats that as a normal “no selectable models” condition. This can downgrade a vendor/schema regression into a misleading generic warning (“requested model not found”) or no parse-specific signal, making model-selection failures significantly harder to diagnose.
Agent Prompt
## Issue description
`AcpSessionModelList` converts malformed `session/new` model payloads into an empty list by catching `JsonException` and returning `[]` without logging. The caller (`SessionModelResolution.ResolveOrNull`) then logs only a generic “requested model not found” warning (or nothing parse-specific), which hides the root cause when a vendor changes schema or returns malformed JSON.
## Issue Context
- `AcpSessionModelList.FromModelsObject` catches `JsonException` and returns `[]`.
- `AcpSessionModelList.FromConfigOptions` skips unreadable entries via `catch (JsonException) { continue; }`.
- `SessionModelResolution.ResolveOrNull` now only sees the empty list, so a parse regression becomes indistinguishable from “no models were published” and/or is misattributed as “requested model not found”.
## Fix Focus Areas
- src/Capacitor.Cli.Core/Acp/AcpSessionModelList.cs[47-72]
- src/Capacitor.Cli.Daemon/Acp/IAcpModelSelector.cs[72-90]
## Suggested fix
- Add an optional `ILogger? logger = null` parameter to `AcpSessionModelList.Extract(...)` (or a `TryExtract(..., out models, out parseError)` API).
- When catching `JsonException`, emit a `LogDebug` (or `LogWarning` if you want it operator-visible) that indicates which shape failed (`models` vs `configOptions`) and that model selection will be skipped.
- Consider adjusting the warning in `SessionModelResolution.ResolveOrNull` to differentiate between:
- “selectable-model list missing/unparseable” and
- “selectable-model list present but requested model not found”.
This keeps the fail-open behavior while restoring actionable diagnostics.
ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools
Found by running the thing this branch builds: an unattended `opencode` reviewer on an isolated dev daemon, reviewing this changeset. Four real issues, plus one version-reporting defect the dev daemon's own startup log exposed. Containment no longer depends on a merge rule that was never measured. What the probe established is that OPENCODE_PERMISSION overrides an operator config saying `"*": "ask"`. What it did NOT establish — the reviewer drew the distinction precisely — is whether a WILDCARD from the env beats a SPECIFIC key from a file: OpenCode merges per key and then resolves patterns specific-before-wildcard, so an operator's `bash: "allow"` could plausibly survive a bare `"*": "deny"`, leaving only the assumption that the isolated config dir stops that file loading at all between a reviewer and a shell. Each forbidden tool now carries its own `deny` key, so the merge order is irrelevant whether or not any file config loaded. That also makes `ForbiddenTools` load-bearing rather than the dead constant the reviewer noticed it was — one fix for both findings. An injected server name is interpolated into a GLOB, so a metacharacter in it would widen that entry past the server. Unreachable today (every injected name is a per-launch alias this daemon generated), which is why it is a cheap refusal rather than a redesign: it stops the day a name starts coming from somewhere else, instead of that day quietly granting a wider surface. The `models.availableModels` path did not filter junk entries while the `configOptions` path did, and that asymmetry was a crash. `ModelId` is non-nullable in C# but nothing stops an agent omitting it; `AcpModelResolver`'s prefix arm then calls `StartsWith` on null and throws a NullReferenceException past a caller whose only guard is `JsonException` — turning a malformed vendor response into a failed LAUNCH, for a feature documented as never being a launch precondition. The plugin's request-path healing is now bounded. A session whose classification keeps failing never reaches `start()`, so `started` never records it and the transform re-triggered `classify` on EVERY model request for the life of the process. The reviewer got there by tracing what populates `started` — it is set inside `start()`, which that path never reaches. The event path's own retry (session.idle) is deliberately unaffected; only the request-path healing, which exists for a plugin loaded mid-session, is capped. Separately, `ComputeUnattendedVendorCapabilities`'s per-vendor path map did not know `opencode`, so it fell to the generic arm and reported `CLI version unknown` while the capability gate had just admitted the vendor on a resolved version. Two answers about one build, and the WRONG one is what reaches the server and the operator's startup log — the first place anyone looks when a reviewer misbehaves. Named, exactly as the comment above the Antigravity entry says to. Deliberately NOT changed: the reviewer observed the MEMORY_MARKER dedupe may never fire, since OpenCode rebuilds the system array every request so a single push never reaches the length that triggers its own collapse. That is correct and already documented — the marker is defence for the OTHER behaviour (a future build retaining transformed entries), not for the measured one. Tests: OpenCodeReviewerLaunchTests 25/25 (13 new), AcpSessionModelListTests 18/18 (1 new — the assertion that threw before the filter), OpenCodeSessionStartMemoryTests 17/17 (1 new), OpenCodeHostedLaunchTests 7/7, OpenCodeReviewerCapabilityTests 14/14. AOT publish clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
3d995d2 to
5fc5458
Compare
…ng glob syntax
Round-2 review point. The guard was a list of glob metacharacters
(`* ? [ ] !`) and its error message claimed a name "cannot be admitted
verbatim" — but the list cannot be exhaustive for a glob engine we do not
control: extglobs (`?()`, `*()`, `!()`), braces, pipes, parens and backslash
were all plausible and all absent. Enumerating another implementation's syntax
is a losing game.
Inverted to a conservative character allowlist (`[A-Za-z0-9._-]`), which is
total by construction: nothing outside it can mean anything to any engine. Same
house rule as escaping — admit only what is verifiable, else refuse.
Worth stating that this was never a containment breach even while incomplete,
and review said so: a widened `{name}_*` still needs a literal `_` before the
wildcard, so it could only ever additionally match another injected server's
flattened tool names, which are all intended-allowed. Native tool names carry no
underscore. So this closes a stated-totality gap, not a reachable escalation.
Tests: OpenCodeReviewerLaunchTests 35/35 — the refusal cases now range past
`* ? [ ]` into extglob/brace/pipe/backslash/space/slash, plus a positive set
pinning that real per-launch alias shapes are still admitted (a guard that
refused the legitimate case would be worse than none). AOT publish clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… failure Two things, both found by looking rather than assuming. **The epic specifies a delivery contract this did not implement.** AI-1456's normative cross-runtime section pins `kcap hook --opencode … --memory-contract 1` and requires the extension to VALIDATE the marker before injecting. Both now hold: - The binary fetches the index only when the caller DECLARES it can consume the output. Absence means contract 0. This matters because the fetch spends the session's once-only injection lease: a new binary paired with an already-installed older plugin — which discards this command's stdout — would otherwise spend that lease on output nobody delivers. The reverse pairing needs nothing, since this command has always ignored unrecognised arguments. - The plugin treats stdout as a fragment only when it opens with the marker. stdout is a data channel here, and a data channel needs a shape; without this, any line some future code path prints there would land in the model's system prompt verbatim. **Deliberate deviation, stated rather than silent:** the marker is NOT stripped before injection, though the epic's text says to. It is the only way to recognise an already-appended fragment in a system array this plugin does not own — the guard that keeps injection correct if a future OpenCode retains transformed entries instead of rebuilding them — and it is an invisible HTML comment, so leaving it in costs a reader nothing. Pinned by a test so it stays deliberate. **The Windows-only failure was mine, not the known flake.** A hosted-launch test asserted the CONSENT refusal code, but the capability gate checks the platform first and short-circuits, so Windows answers `unsupported_platform` before consent is consulted. Correct behaviour, wrong assertion — and precisely the trap `KiroReviewerCapability.Decide` documents having been caught by once already. It passed locally on macOS. Both platforms now assert a coded refusal; only the consent-specific code is POSIX-scoped. **Re-certified, because the delivery path changed.** The live cert was re-run after the contract flag and marker validation landed: both cases pass against opencode 1.18.9. Two intermediate failures were investigated rather than retried — the readiness probe (new) proved the index carried the nonce, and an isolated turn reproduced a patterned nonce exactly, so the cause was a small free model mistranscribing 32 random hex characters. That is a model-capability flake, not a delivery defect. The assertion now carries the model's own answer in its failure message so the two are distinguishable, and the class doc says which response each calls for. The assertion itself was NOT weakened: an exact echo is the whole proof. Also added: a bounded readiness wait before spending a model turn, so an index-propagation race cannot be reported as a delivery defect. Tests: OpenCodeSessionStartMemoryTests 25/25 (8 new — contract negotiation, marker validation, marker retention), OpenCodeHostedLaunchTests 7/7, OpenCodeReviewerLaunchTests 35/35, OpenCodeReviewerCapabilityTests 14/14, AcpSessionModelListTests 18/18. Live cert 2/2. AOT publish clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The contract flag was ADDED to an existing argument array, so dropping `--session` or `--file` while editing it would break CAPTURE — the watcher spawn and the lifecycle POST — which no memory test would notice and which is a far worse regression than losing the index. Verified live as well: with no contract flag the hook writes zero bytes to stdout and still spawns the watcher, so the gate is scoped to memory and nothing else. Tests: OpenCodeSessionStartMemoryTests 25/25. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
OpenCode across three of the epics: hosted agent (AI-1405), unattended
review-flow reviewer (AI-1411), and SessionStart memory injection (AI-1466).
Everything asserted about the vendor is measured against opencode 1.18.9,
not read off its advertisement — the advertisement demonstrably decides
nothing, since Kiro, Gemini and Copilot all publish the same
mcpCapabilities {http, sse}shape and disagree with each other about stdio.Probe record:
docs/probes/2026-08-07-opencode-acp/findings.md.Why one PR
AI-1405 and AI-1411 are not separable at the file level — both live in
AcpVendorDescriptorandAcpHostedAgentRuntimeFactory, and the reviewer flipis a change to the same descriptor hosting introduces. AI-1466 is separable and
sits in disjoint files; it shares only
README.md. Happy to split it into itsown PR if you'd rather review it apart.
The findings that shaped this
Model selection needed a second read shape. OpenCode publishes no
modelsobject — its selectable models are
session/new'sconfigOptions[id=="model"].options[]. The shared resolver read onlymodels.availableModels, so a caller-requested model would have been silentlydiscarded on every OpenCode launch. The write half is verified at effect
level (the model self-identified as the requested id), not on its echoed
success — the distinction that keeps Gemini on
NoOpModelSelector.Capture precedence. OpenCode is the one vendor where two kcap capture paths
can be live at once: the operator's live-ingest plugin loads inside the
opencode acpprocess and starts a second recording of the session the ACPmapper is already recording. Controlled pair (equal dwell, plugin confirmed
installed): default env created the plugin's capture file,
OPENCODE_PURE=1did not. Applied on every hosted launch, interactive included — the
duplication is timing-dependent, which is worse to operate than a deterministic
one. Accepted cost: a hosted agent does not carry the operator's other plugins.
The reviewer's posture is env-shaped.
opencode acpaccepts none of theglobal flags, so
UnattendedTrustArgvis empty andOpenCodeLaunchEnvironmentowns it. Deny-all plus the read family plus the injected result channel: no
shell, no write, no network. Two controls make that non-vacuous — flipping
bashto allow makes the shell actually run, and removing the{server}_*entry takes the result channel out of the model's toolset entirely (ruling out
"MCP tools bypass permissions, so the entry is decorative"). Shell success is
judged on a file existing, never the answer text: an earlier revision scored
the denied arm as having run the shell because the model quoted the command
while explaining it could not.
A denied tool is absent from the model's surface rather than refused at call
time, so a correct launch raises no interaction frame and
Failis honest. Thatalso closes the skill-derail hazard structurally — with
skilldenied, OpenCodeomits its skills section from the system prompt entirely.
Verified live, not just unit-tested
start_review_flow(vendor="opencode")ranthree rounds —
findings→findings→clean— on an isolated devdaemon (own
--name, consent flag in its env; the operator's daemonuntouched), with the result arriving through the injected
submit_review_resultchannel and zero human-routed interactions.Both live runs found defects that everything green had missed, which is the
main argument for having done them:
resting on an unmeasured merge rule (env wildcard vs a file-specific key), an
unescaped glob interpolation, an NRE in the model list
(
ModelId.StartsWithon null, turning a malformed vendor response into afailed launch), and an unbounded per-request re-probe. All fixed; it then
returned
clean. Its metacharacter-list objection led to inverting that guardinto a character allowlist, total by construction.
in-flight marker was published after an async classify resolved, while
OpenCode publishes
session.createdand then issues the first LLM request —so a one-turn session got no index at all, intermittently.
Also in here
ComputeUnattendedVendorCapabilities's path map did not knowopencode, so alive dev daemon logged
CLI version unknownwhile the gate had just admitted thevendor on a resolved version. Named, per the comment above the Antigravity
entry. Flagging deliberately: this is adjacent to cancelled AI-1756 — I read
that as being about maintaining certified version sets, not a one-line path
entry, but say so if you disagree. Kiro and Gemini stay unnamed either way.
Epic AI-1456's normative cross-runtime section is now honoured
(
--memory-contract 1, so the binary spends the session's once-only injectionlease only for a caller that can deliver; plus marker validation before
injecting). One deliberate deviation: the marker is NOT stripped before
injection, because it is the only way to recognise an already-appended fragment
in a system array the plugin does not own, and it is an invisible HTML comment.
Pinned by a test so it stays a decision rather than a drift.
What is NOT claimed
SupportsReconnectResumeisfalseas unprobed — a different claim fromKiro's and Gemini's measured-ineligible.
reviewer in the caller's live checkout. Fails closed to an owned worktree.
*: denyversus afile-specific allow for a future native tool, and non-tool permission
categories. Both require a permission-bearing config to load despite the
isolated config dir, which is unmeasured.
small free model flakes on transcription. The failure message carries the
model's own answer so that is distinguishable from a delivery failure; the
assertion was deliberately not weakened.
Verification
Targeted runs only (no full-suite sweeps):
OpenCodeHostedLaunchTests7/7,OpenCodeReviewerLaunchTests35/35,OpenCodeReviewerCapabilityTests14/14,AcpSessionModelListTests18/18,OpenCodeSessionStartMemoryTests25/25, livecert 2/2. Neighbours green:
AcpHostedAgentRuntimeFactoryTests82/82,KiroReviewerLaunchTests10/10,KiroReviewerHomeTests9/9,KiroReviewerCapabilityTests32/32,GeminiReviewerLaunchTests31/31,AcpVendorDescriptorTests22/22,ServiceEnvironmentTests17/17,DaemonReviewerCommandTests8/8,DaemonRunnerCursorAvailabilityTests35/35.AOT publish clean. No Linear IDs in
.cs.Mutation-checked: scoping the plugin suppression to review launches fails 2
tests; adding
bashto the read set fails 1; bypassing the consent gate failsconfigOptionsid filter, whosetest used a fixture where
modelhappened to lead the array.A Windows-only failure in the first CI run was mine, not the known flake: a
test asserted the consent refusal code, but the capability gate checks platform
first and short-circuits, so Windows answers
unsupported_platform. Correctbehaviour, wrong assertion — and the exact trap
KiroReviewerCapability.Decide's own doc comment records. Fixed; both platformsnow assert a coded refusal and only the consent-specific code is POSIX-scoped.
Fixes AI-1405
Fixes AI-1411
Fixes AI-1466
🤖 Generated with Claude Code