Repository navigation
[AI-1391] Daemon sequenced-command settlement (B2-b, daemon half) - #347
Conversation
…w-round-1 applied Implementation plan for the daemon (kcap-cli) half of B2-b: 17 TDD tasks (wire DTOs, sequenced lanes+watermark, resolved-candidates ledger + 4 crash-consistent hooks, coverage-journal/boot-chain attestation, per-platform StartupReapComplete, daemon heal-barrier). Authored + adversarially reviewed + revised (genesis atomicity, compile defects, full crash-injection matrix, synthesized-item monotonicity, racing-liveness). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…artup discovery) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ks, RequestStatusReport) StopAgentV2 + CommandAck/CommandRejected (+ their reason/state/outcome/liveness enums), AckProcessedPrefix, AckResolvedCandidates (+ ResolvedCandidateAck), RequestStatusReport, and the additive LaunchAgentCommand Epoch?/Seq?/CommandId? sequencing fields. Registered on CapacitorJsonContext, snake_case, enum tokens pinned with safe zero-defaults; raw-JSON round-trip asserts. All additive. Phase B2-b (sequenced-settlement design). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Open the per-name lock file FileAccess.ReadWrite and, immediately after inspecting the leftover PID file (while holding the exclusive flock, so the prior holder is provably gone), read the previous holder's InstanceId from the lock-file content BEFORE truncating and rewriting our own. Exposed as the new DaemonLock.PriorInstanceId (null on a genuinely fresh lock / unreadable content; read failures fail-closed to null). Unlike the PID file (deleted on clean shutdown), the lock file's InstanceId is the persistent per-boot nonce every shipped version rewrites at boot and never deletes, so it witnesses the immediately-preceding boot even one by an unaware binary. Phase B2-b (sequenced-settlement design) uses it as the coverage boot-chain's chain-check. Additive and unused until later tasks consume it. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…testation) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ssible Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…pend-before-delete) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…es ledger Phase B2-b (sequenced-settlement design): wire the real ResolvedCandidatesLedger into AgentOrchestrator and add the two remaining confirmed-gone hooks, both append-before-delete and idempotent on the source-stable (AgentId, OldEpoch) key: - Hook B (quarantine drain, RetryQuarantineOnceAsync): emit (AgentId, _daemonEpoch, flow...) before deleting each drained entry's PID record. RetryAllAsync now returns IReadOnlyList<Entry> so the drain has the flow identity. - Hook C (StopAgent fallback, TryStopByPidRecordAsync): emit (agentId, record.DaemonEpoch, record.flow...) from the trusted record before the delete. Also wires the Task 7 OrphanReaper record-pass callback (Hook A) to the ledger and adds ResolvedLedgerSnapshotForTest / QuarantineForTest seams. Shipped teardown/ quarantine/orphan-reap semantics are unchanged. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
… into the ledger Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rune Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…vedStartupCandidates Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ontiguous watermark, synthesized-error item) Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…collisions Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…re (no unacked eviction) Backpressure guard in the accept path: once the per-epoch identity cache reaches its bound, further exact-next accepts are rejected with CommandRejectedReason.Backpressure rather than evicting an unacked entry. AckPrefix retires cache entries through a VALIDATED AckProcessedPrefix (current epoch, not over-ahead of LastProcessedSeq, strictly monotonic); stale-epoch / over-ahead / regressing acks are ignored without eviction. 2 tests. Phase B2-b. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…y/semantic rejections
HandleNonNextLocked now emits CommandRejected{wrong_next} for a gap/too-low Seq (never
accepted out of order; the server transport resyncs) — accept path + watermark untouched.
Execution-time terminal rejections (daemon_capacity/semantic) already flow via the lane's
LaunchRejected+RejectReason mapping and advance the watermark as a terminal item. 2 tests. Phase B2-b.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…lity, RequestStatusReport Phase B2-b (sequenced-settlement design): own an epoch-scoped SequencedCommandProcessor in AgentOrchestrator; route Seq'd LaunchAgentCommand/StopAgentV2 through its serial lane (the execute closure returns a CommandOutcome — LaunchRejected+daemon_capacity/semantic where the shipped launch would reject, LaunchExecuted/LaunchFailedCleaned/StopExecuted otherwise), keeping un-Seq'd commands on the legacy unsequenced lane. Add ReadLiveness (confirmed-death precedence over _agents ∪ _quarantine), advertise Epoch/HighestAcceptedSeq/ LastProcessedSeq counters on BuildStatusReport + the enriched DaemonConnect payload, set SupportsSequencedCommands=true, and serve StopAgentV2/AckProcessedPrefix/RequestStatusReport plus the one-way CommandAck/CommandRejected sends. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…r advances watermark Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
OrphanReaper.EmitAndClear looked up a durable record by env-derived agentId alone. A prior-epoch descendant that inherited a still-live leader X's KCAP_AGENT_ID env triple, reaped under a different pid, would match X's leader record — emitting a false trusted-flow death proof for the LIVE leader and deleting X's record (stranding it). The lookup was also unscoped by epoch. Now corroborate the reaped pid AND the prior epoch (r.Pid == c.Pid && r.DaemonEpoch == c.OldEpoch) before trusting/deleting the record; a non-corroborating record falls back to the recordless null-flow path with no record delete. The legitimate identity_unavailable case (record genuinely IS the reaped pid) is unaffected. Regression test added. Phase B2-b (sequenced-settlement design). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…donly, SafeInvoke - Fix 2: OrphanReaper.CurrentDiscovery is a multi-field struct written on the reap thread and read on the report/connect thread; guard get/set with a dedicated gate so the read can't tear (sweep is single-writer, so no wider critical section is needed). - Fix 3: advertise the DaemonConnect epoch from the orchestrator's own per-boot _daemonEpoch via a GetDaemonEpoch seam (the single source the processor is scoped to), falling back to _config.DaemonEpoch when unwired — removes the test-divergence footgun; prod behaviour unchanged (DaemonRunner pins config). - Fix 4: mark _resolvedLedger / _markerCandidates / _processor readonly. - Fix 5: route the RequestStatusReport receive through SafeInvoke like the other command handlers. - Fix 6: drop the dead report/expected locals in the startup-discovery test. Phase B2-b (sequenced-settlement design). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR Summary by QodoAI-1391 B2-b: daemon-side sequenced-command settlement (producer half)
AI Description
Diagram
High-Level Assessment
Files changed (23)
|
Code Review by Qodo
1.
|
…nt §5.5) OrphanReaper.BlockedCandidates() only listed IdentityUnavailable records and the macOS legacy-live case. A Present prior-epoch record the record pass spared (a transient ambiguous identity read on Linux, or a record-pass fault) stayed on disk yet unlisted, so ComputeStartupReapComplete() could return true while a prior-epoch process may still be alive — the paired server would treat the omission as proof of death and launch a duplicate. Add a third arm: any other prior-epoch record still on disk is unresolved (confirmed-dead records are deleted at reap time), so it blocks as identity_unresolvable. A record whose process dies between passes blocks only until the next reap tick deletes it — a false-incomplete only delays relaunch, never mints a duplicate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A capacity/semantic launch rejection is cached as LaunchRejected with its RejectReason, but the processed-duplicate CommandAck was built without RejectionReason. A retransmitted duplicate of a rejected launch therefore couldn't distinguish daemon_capacity (requeue) from semantic (fail) — breaking the per-reason ticket semantics for exactly the lost-rejection case the identity cache exists to answer. Pass the cached reason on the processed ack, serialized through the same CapacitorJsonContext the SignalR hub uses for CommandRejected.Reason, so the ack's RejectionReason string equals that enum's [JsonStringEnumMemberName] snake_case wire token (daemon_capacity / semantic). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
HandleLaunchAgent routed to the processor only when Epoch/Seq/CommandId were ALL present; any partial tuple fell through to the unwatermarked legacy lane, where a malformed/truncated capable-server command's retry could be re-accepted on the sequenced lane and double-create the generation (violating at-most-once-per-generation). Split routing three ways: none present -> legacy lane (old server); all three present -> sequenced lane; anything in between -> fail closed with a LaunchFailed, never the legacy lane. The watermark is untouched and nothing spawns. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rite (§5.5) Upsert mutated _entries and _nextGeneration in memory BEFORE Persist(). If Persist threw, memory led disk: the next sweep's Upsert returned the unpersisted in-memory entry via the idempotent short-circuit WITHOUT persisting, and the caller then deleted the durable source — a crash could then lose both (violating append-before-delete). Persist the prospective state (post-increment high-water) BEFORE committing the counter, and roll back the in-memory entry on a write failure. Refactor Persist() into PersistState(nextGeneration); Ack keeps the plain Persist() (its memory-behind-disk direction is benign — re-loads + re-acks idempotently). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…5.5) The report/connect payloads carried the resolved-candidates array but not the daemon-lifetime monotonic high-water, so once sparse acks prune entries the server lost the generation frontier. Add an optional long? HighestResolutionGeneration to DaemonStatusReport and DaemonConnect (additive, snake_case), a ledger high-water property (_nextGeneration - 1; persists across prunes/restarts), and wire it into BuildStatusReport plus a new GetHighestResolutionGeneration ServerConnection seam on the connect payload (mirroring the GetResolvedStartupCandidates pattern). DaemonConnect already advertises ResolvedStartupCandidates, so the high-water goes on connect too. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The env-marker scan discovers a recordless prior-epoch survivor, persists a durable marker-candidate source, then kills it. The per-PID try/catch logged and continued on ANY failure, then the pass unconditionally published MarkerScanState.Complete. If the source WRITE fails (disk-full / permission / I/O), the survivor is neither recorded (invisible to BlockedCandidates) nor killed — yet the scan reported Complete, so with the spared-record completion fix StartupReapComplete could go true beside a live prior-epoch survivor and the paired server would launch a duplicate. Split the write from the resolution: a WRITE failure sets captureFailed and skips the kill (never resolve without a durable source), leaving discovery Failed for the pass (retried next heartbeat, last-successful-scan time preserved). A failure AFTER a successful write is unchanged — the pending_marker source blocks via BlockedCandidates and the next boot re-derives + retries the kill. Linux-gated regression test: a MarkerCandidateStore whose state dir is a file (every Write throws) + a live prior-epoch survivor -> discovery stays Failed and the survivor is not killed. Phase B2-b (sequenced-settlement design §5.5). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…s (§4.2.3)
DaemonLock returned null PriorInstanceId for BOTH a genuinely-empty lock file
(a real first-ever/genesis slot) and a read failure on a non-empty file, and
CoverageJournal.RecordBoot treated `priorLockInstanceId is null` as genesis-
eligible. So an absent journal + a corrupt/unreadable non-empty lock file wrongly
attested RecordlessSurvivorsImpossible=true — violating the boot-chain's
fail-closed invariant (any missing/corrupt state ⇒ false).
- DaemonLock now exposes `PriorLockIndeterminate`: true iff the lock file was
non-empty but yielded no id (a read fault OR blank/garbage). A genuinely empty
file stays null + not-indeterminate (genesis-eligible).
- RecordBoot takes `priorLockReadFailed`; genesis is now
`!priorLockReadFailed && priorLockInstanceId is null`, so an indeterminate
prior fails closed instead of being mistaken for genesis.
- The lock-file read is now bounded (`new byte[(int)Math.Min(stream.Length,
4096)]`) so a corrupt/oversized file can't drive an unbounded allocation (the
content is a 32-char GUID + newline). ('new byte[long]' already compiled — the
reviewer's 'breaks build' note was a false positive; this addresses the valid
unbounded-allocation half.)
Tests: DaemonLock indeterminate flag (non-empty-unreadable ⇒ indeterminate;
fresh ⇒ not) + RecordBoot fail-closed on read-failed prior vs genesis on a
genuinely empty prior.
Phase B2-b (sequenced-settlement design §4.2.3).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
What & why
Part of AI-1391 B2-b — the bilateral sequenced daemon-command settlement protocol (parent AI-1313 §5.5/§6/§7/§8). This is the daemon (producer) half; the paired kcap-server PR consumes the wire. Rollout is daemon-first: everything here is advertised-but-inert — the
SupportsSequencedCommandscapability istrueonDaemonConnect, but no behaviour changes until the server PR lands and starts sending sequenced commands. All wire changes are additive (an old server ignores unknown fields/methods), so either merge order is safe.Spec:
docs/superpowers/specs/2026-07-17-ai1391-sequenced-settlement-design.md. Plan (in this PR):docs/superpowers/plans/2026-07-22-ai1391-b2b-daemon-sequenced-settlement.md.Scope (daemon producer side)
Cli.Core/Models.cs) —StopAgentV2/CommandAck/CommandRejected/AckProcessedPrefix/AckResolvedCandidates/RequestStatusReport;ResolvedStartupCandidate/UnresolvedStartupCandidate/StartupDiscovery; and additive fields onLaunchAgentCommand(Epoch/Seq/CommandId) +DaemonConnect/DaemonStatusReport(watermark + startup-completeness +SupportsSequencedCommands). Enums pinned with[JsonStringEnumMemberName]and zero = safe default.SequencedCommandProcessor— two sequenced lanes (Seq'dLaunchAgentCommand+StopAgentV2) executed strictly serially per epoch; exact-next acceptance is one atomic op under lock;LastProcessedSeqis the contiguous terminal-processed prefix; duplicates answered withCommandAck(never re-executed);wrong_next/duplicate_collision/backpressure/stale_epoch/internal_errorrejections;AckProcessedPrefixretires identity-cache entries (never evicts an unacked entry). Un-Seq'd commands stay on the legacy unsequenced lane (old-server compat) and never advance the watermark.ResolvedCandidatesLedger— durable positive-death-evidence outbox in the daemon state dir (atomic temp+rename, monotonicGeneration, append-before-source-delete,(AgentId, OldEpoch)reconcile key). Four hooks feed it: OrphanReaper record-pass, quarantine drain, StopAgent PID-record fallback, and the env-marker recordless resolution matrix. Pruned by a validatedAckResolvedCandidates.CoverageJournal(single-atomic-write genesis,cumulative_coveredfold) +DaemonLock.PriorInstanceIdchain-check → the fail-closed, sticky-falseRecordlessSurvivorsImpossibleflag advertised on connect.StartupReapComplete+StartupDiscovery(marker-scan state) +UnresolvedStartupCandidates; the daemon heal-barrier servesRequestStatusReportand emitsCommandRejectedper reason.Deferred (noted seams, not in this PR)
DaemonLaunchQueue(FIFO ticket state machine) — the spec frames it as a separate layer; anOnLaunchRequeueseam is provided.Marker-path safety (review follow-up)
OrphanReaper.EmitAndClearnow corroborates the reaped pid AND prior epoch (r.Pid == c.Pid && r.DaemonEpoch == c.OldEpoch) before trusting/deleting a durable record on the marker path — the env agentId is untrusted for role authority, so a prior-epoch descendant that inherited a still-live leader's env can no longer emit a false trusted-flow death proof for (or delete the record of) the live leader. Non-corroborating matches fall back to the recordless null-flow path with no record delete.Testing
Capacitor.Cli.Tests.Unitfull suite green — total 3663, failed 0, 2 skipped (a gated live-ACP E2E). New coverage across the processor, ledger + 4 hooks, coverage journal, DaemonLock prior-instance chain, heal-barrier report, marker-candidate resolution matrix, startup completeness, and the wire DTOs. Tests use isolated dummy processes only — no real daemon, no live flows. The Linux-only marker-scan matrix is exercised under Linux CI.Compatibility
Fully additive;
SupportsSequencedCommandsis the single gate. An old server ignores the new fields/methods and the daemon behaves exactly as before. No projection or read-model changes.🤖 Generated with Claude Code