Skip to content

fix(subagent): price the startup memory reserve at the learned cost - #13294

Merged
bolichen97 merged 1 commit into
mainfrom
fix/spawn-reserve-learned-cost
Sep 24, 2026
Merged

bolichen97 merged 1 commit into
mainfrom
fix/spawn-reserve-learned-cost

Conversation

@bolichen97

@bolichen97 bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Problem / Motivation

A fan-out of dedicated sub-agents can take a 64 GB host from ~25 GB free to under 1 GB in about 90 seconds, after which the gateway's event loop stalls and the watchdog hard-exits it. Every one of those starts passed the spawn memory guard: the guard priced each pending start at the 0.5 GB first-boot fallback (agent.subagent_cost_gb) while the host's own cost store already held a learned p90 of ~6 GB per run.

Why it matters

The cap (compute_max_subagents) is already sized from the learned cost, but the cap is a count, not a memory guard. The per-spawn reserve is the only thing that prices a start between admission and the reaper's first RSS sample 60 s later, and a runtime takes tens of seconds to reach its resident size. Under-pricing that window by ~12x lets several starts each clear a raw free-memory check and then grow into the same headroom together. That is an OOM, which is unrecoverable, on any host whose real per-run cost is well above 0.5 GB (a heavy MCP roster, a non-shared backend).

What changed (motivation → approach → change)

Root cause: spawn_impl's memory guard in subagent_manager/admission/gate.py read startup_cost = memory_cfg.subagent_cost_gb and never consulted the learned cost.

Two prices, in subagent.py. A WARMING start — the next one, a claim awaiting registration, a dedicated worker fewer than _RSS_SAMPLES_TO_SETTLE (2) sweeps have measured — owes _effective_next_start_gb (the larger of the configured fallback, the learned figure and every live dedicated peak) less the RSS it holds. A SETTLED worker owes only the gap between the larger of the configured cost and its OWN peak and its observed RSS, so a high learned p90 never becomes a reserve no sample can retire. _startup_cost_gb supplies the learned half as max(configured, learned). SubagentInfo._rss_samples counts the sweeps that measured a run (subagent_manager/monitoring.py); the cancel-recovery respawn in subagent_manager/cancellation.py resets that count and the last reading and bumps _rss_generation, which the sweep re-checks after its off-loop /proc read so a dead process's reading cannot settle its replacement (the peak stays); those fields are registered in DELIVERY_ROUTING_FIELDS.

The learned figure is per cost bucket and dedicated-only. _cost_bucket — the explicit agent, else the inherited template — is the one key the sample write and the gate's lookup share. In subagent_cost.py, a session-shared run records shared: true, read_learned_costs(dedicated_only=True) leaves those per-session shares out, compaction keeps one window per (agent, shared) so shared runs cannot evict dedicated history, and learned_cost_for prices a spawn at its own bucket only — a bucket without dedicated history answers None and the configured cost plus live peaks prices it, never another agent's figure, since on a sharing-default backend a share-eligible agent's own bucket never forms. The reserve's read leaves out samples older than 30 days (_SAMPLE_MAX_AGE_SECS), so a price learned under a removed workload expires on its own; the cap's reader applies no horizon. The log is streamed, never held whole (_iter_samples, also under compaction): a bounded deque per bucket, a parse ceiling on distinct buckets, keys over _BUCKET_KEY_CAP dropped, and the heaviest _MAX_BUCKETS returned and held (cap_buckets). Records written before the shared field existed read as dedicated until their window turns over. The map reaches the gate as SubagentManager._learned_costs_gb, published off-loop by the reaper sweep; it is cleared when the log is absent; a complete read is authoritative and replaces it, so a bucket that expired or fell below min_samples retires on the running gateway (read_learned_costs_checked), which also serves a replaced log (cost_log_identity: new inode or shrunk size; the identity is read on both sides of the parse and a mismatch keeps the prior state); only an incomplete read — a refused record ended the parse early, or the present log could not be opened or inspected — is merged, so an unreached bucket is not silently lowered.

A low-memory deferral names the effective per-start price, its source and the reset path in the log line, the SEL record (startup_cost_gb, learned_cost_gb) and the deferral reason; the session-memory sampled flag reads a counted sweep or a live reading, not the kept peak. docs/system-specs/modules/subagent.md, docs/system-specs/modules/adaptive-concurrency.md, src/kiro_crew/docs/dynamic-subagent-sizing.md and src/kiro_crew/docs/subagents.md state all of this.

Backwards compatibility

Compatible: no config key, API or file format changes; _startup_memory_reserve_gb's new argument defaults to the old behaviour, read_learned_cost keeps its signature and result (no horizon; the heaviest bucket is always among the _MAX_BUCKETS returned, so the cap's input is unchanged for any real log), and every existing pin on the reserve still holds. A host with no learned cost yet behaves exactly as before. A host with a learned cost above the fallback defers an unmeasured start to the durable queue until the reserve is satisfied and says so in the deferral.

Tests

  • test/test_admission_gate.py::TestSpawnAdmissionGate::test_startup_reserve_prices_the_pending_start_at_the_learned_cost — through the real spawn path, the guard's min_gb is floor + learned cost when the manager carries one (10.0, not 4.5), floor + fallback when it carries none, and keeps an operator pin above a lower learned value. Proven: reverting only the gate line fails it with [4.5] == [10.0].
  • ...::test_a_lightweight_agent_is_priced_by_its_own_history — with {kirocrew: 9.0, light: 1.0} learned, spawning light reserves 5.0 and an agent with no history reserves the configured 4.5, never the heavy agent's figure; ...::test_an_agentless_spawn_is_priced_by_the_template_it_inherits — an agent-less spawn under an inherited heavy template is priced from the heavy bucket, not the default one.
  • test/test_adaptive_startup_memory.py::test_cost_samples_are_written_under_the_bucket_the_gate_reads — samples land under the same key the gate reads: inherited template, named agent, or the default.
  • ...::test_low_memory_deferral_names_the_learned_price, ...::test_low_memory_deferral_names_the_configured_price_when_nothing_is_learned, ...::test_low_memory_deferral_reports_the_live_peak_that_set_the_price — the SEL record carries the effective startup_cost_gb / learned_cost_gb and the durable row's deferred event names the price and its real source (learned p90, configured pin, or a 7.5 GB live peak above the learned figure).
  • test/test_adaptive_startup_memory.py::test_startup_reserve_prices_unmeasured_starts_at_the_learned_cost — the reserve arithmetic: warming starts and once-sampled workers at the learned price less held RSS, a settled worker at its own gap only (a heavy sibling raises the next start, not the light worker's gap), an observed peak above the learned figure raising the next start, shared sessions adding nothing; ::test_effective_next_start_price_folds_in_live_peaks.
  • test/test_adaptive_startup_memory.py::test_dedicated_pricing_leaves_shared_session_shares_out, ::test_compaction_keeps_dedicated_history_under_a_flood_of_shared_runs, ::test_a_reset_recreated_within_one_sweep_is_read_fresh, ::test_a_sweep_that_straddles_a_respawn_does_not_settle_the_new_process, ::test_cost_log_identity_tells_absent_from_uninspectable — shared shares excluded from pricing and from evicting dedicated history, a re-created log read fresh, a straddling sweep discarded, and the identity probe; the refresh test also covers a partial read merging into the held map.
  • test/test_adaptive_startup_memory.py::test_expired_samples_do_not_price_a_start, ::test_an_expired_bucket_retires_on_a_long_lived_gateway, ::test_the_held_map_is_bounded_against_an_agent_writable_log, ::test_a_log_replaced_during_the_read_keeps_the_prior_map — the 30-day horizon (legacy records without ts still count) and its taking effect on a running gateway, the parse-time bounds, and a replacement landing mid-read keeping the prior map; the refresh test pins complete-read-replaces vs incomplete-read-merges.
  • test/test_adaptive_startup_memory.py::test_startup_cost_is_the_larger_of_configured_and_learned, ::test_learned_cost_for_prices_a_spawn_by_its_own_agent, ::test_reaper_sweep_publishes_the_learned_cost_off_loop — the price helper, the per-agent lookup with its heaviest-known fallback, the presence probe on a failed stat, and the sweep publishing the per-bucket map while a raising read and a present-but-empty read both keep the previous values and a deleted log drops them.
  • test/test_subagent_reap_race.py::test_recovery_respawn_is_priced_as_a_fresh_process — a respawned run starts over as warming (samples 0, last reading 0, peak kept) and the guard reserves the full learned price for it again. Proven: removing the reset fails it.
  • Every pre-existing pin on _startup_memory_reserve_gb (default next_start_gb) is unchanged and passes.

Manual verification

N/A — unit coverage sufficient: the reserve arithmetic is exercised through the real SubagentManager.spawn path with the memory reader stubbed, and the refresh is exercised against a real cost log.

Related Issues

no linked issue: found while investigating an internal report (Kiro Crew 0.7.0.8, a sub-agent fan-out exhausting a 64 GB host followed by a gateway watchdog restart); the report lives in an internal tracker, not a GitHub issue.

Pattern harvest

Rule candidate: review-prompt
Pattern: "two guards for the same resource read the estimate from different sources" — a sizing path consults the learned store while the admission path reads the static fallback for the same quantity.

Checklist

  • At most two commits (one is the norm), with a Conventional Commits title (feat|fix|docs|style|refactor|perf|test|chore|ci|build|revert: ...)
  • Existing tests pass and new tests added for new functionality
  • Self-review completed; code follows project style guidelines
  • Documentation updated (if applicable)
  • No secrets, credentials, or internal references in the diff

@bolichen97
bolichen97 requested review from a team and CrysisDeu September 24, 2026 08:16
@bolichen97

Copy link
Copy Markdown
Contributor Author

Intent: Make the per-spawn memory reserve price a pending dedicated start at the learned per-run cost the cap is already sized from, so a burst of admissions inside the pre-sample window can no longer over-commit host memory.
Not a goal: Changing how or when cost samples are recorded (short runs and runs that end in a gateway crash still never write a sample), changing the cap formula, or gating dashboard sessions created outside the SubagentManager.

@github-actions github-actions Bot added the readiness: checking Automated validation is still running label Sep 24, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

GPT 5.6 Review — ✅ human override accepted

Human judgment by @bolichen97 overrides the GPT 5.6 finding for da66f842e1f3cca4dad08b41d4a56e3e4b464d3e; the recorded reason is authoritative for this commit.

[GPT-OVERRIDE] da66f84

This comment is updated in place on each push.

The model was not re-run because an authorized human decision supersedes it.

False positive or not applicable? A repository writer can comment:
/ai-review override gpt da66f842e1f3cca4dad08b41d4a56e3e4b464d3e: <one-sentence reason>

@bolichen97
bolichen97 enabled auto-merge (squash) September 24, 2026 08:19
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

First Principles Review (Fable 5) — ✅ PASS

Premise-level review of da66f842e1f3cca4dad08b41d4a56e3e4b464d3e — why this exists and whether the shipped surface is the smallest honest version. Updated in place on each push. A BLOCK verdict blocks PR readiness; PASS/CONCERNS are advisory.

First-Principles-Verdict: PASS

Verify deferred starts re-attempt from the durable queue: a stale learned p90 now holds unmeasured spawns until the 30-day horizon or a manual log reset.

What this change ships

Inventory (10 items) — 10 justified

Intent: FIX — stop a subagent fan-out from OOMing the host by pricing the spawn guard's startup reserve at the learned per-run cost instead of the 0.5 GB first-boot fallback (provenance: the added gate test fails on base, [4.5] == [10.0]; the cap at subagent.py:1362 already used the learned figure).

  1. Each warming start is priced at the learned p90, never below the configured cost — justified
  2. A worker measured by two sweeps reserves only its own peak-vs-RSS gap — justified
  3. A spawn is priced by its own agent/template bucket, the same key its samples are written under — justified
  4. Shared-session samples are marked shared and left out of pricing and compaction eviction — justified
  5. Reserve pricing ignores samples older than 30 days; the cap's reader is unchanged — justified
  6. The agent-writable cost log is streamed with bounded buckets, keys and windows — justified
  7. Learned prices reach the gate off-loop each sweep, replaced/merged/cleared by log identity — justified
  8. A cancel-recovery respawn is re-priced as warming, with a generation guard against straddling sweeps — justified
  9. Low-memory deferrals name the per-start price, its source and the reset remedy — justified
  10. Dashboard memory rows show a respawned run as unmeasured, not "0 MB" — justified

[FIRST-PRINCIPLES-REVIEWED] da66f84

@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Design Review (Fable 5) — 🟡 CONCERNS

Design-level review of da66f842e1f3cca4dad08b41d4a56e3e4b464d3e — updated in place on each push. A BLOCK verdict blocks PR readiness; PASS/CONCERNS are advisory.

Design-Verdict: CONCERNS

A stale learned price can defer every spawn on a memory-tight host for up to 30 days, and deferred runs record no samples to correct it.

Watch

Self-locking deferral: the docstring claims "over-reserving here only defers a start until the next sample," but when the learned p90 outlives its workload (the description's own "deferred runs record no new samples to correct it"), no next sample exists — every start defers until the 30-day horizon or a manual cost_samples.jsonl deletion, so subagents on that host are unavailable for weeks with only a log line pointing at the remedy. Pre-PR the same host ran (and risked OOM); the trade is deliberate, but a human should own the recovery-time choice.
Clears when: a maintainer accepts the 30-day manual-remedy residual, or a follow-up adds a bounded re-probe (one admission at the configured price after sustained all-starts deferral) so a corrective sample can land.

[DESIGN-REVIEWED] da66f84

@github-actions github-actions Bot added readiness: action required A blocking check or review needs attention and removed readiness: checking Automated validation is still running labels Sep 24, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Opus 5 Review — ✅ no blocking findings

Reviewed da66f842e1f3cca4dad08b41d4a56e3e4b464d3e — this comment is updated in place on each push.

Review details

An uninspectable cost-log identity takes the replace branch the adjacent comment says is "additive at most", wiping held learned prices.

FINDING — src/kiro_crew/subagent_manager/monitoring.py:751 — when cost_log_identity() returns ("unknown",) (any non-FileNotFoundError stat failure, e.g. the crew subagents/ dir losing read permission on a running gateway), line 740 forces complete = False but line 765's if len(identity) == 3: never records a source, so previous is None stays permanently true and if replaced or previous is None or complete: replaces _learned_costs_gb with the degraded read's result — {} once open() also fails — discarding a published p90 and repricing every warming start at configured_cost, the exact under-price this change exists to remove; the merge branch is unreachable in the one case both the docstring ("or its identity could not be inspected -- is MERGED") and the inline comment ("nothing this read says is proven, so it is additive at most") name for it → Fix: gate the replace branch on complete and len(identity) == 3 (plus replaced), and track "nothing published yet" with a separate flag instead of inferring it from _learned_costs_source is None.

[OPUS-REVIEWED] da66f84

Verdict parsed from the review's SHA-scoped output markers for commit da66f842e1f3cca4dad08b41d4a56e3e4b464d3e.

False positive or not applicable? A repository writer can comment:
/ai-review override fable da66f842e1f3cca4dad08b41d4a56e3e4b464d3e: <one-sentence reason>

@bolichen97
bolichen97 force-pushed the fix/spawn-reserve-learned-cost branch from 28d1f96 to 6ff2bdd Compare September 24, 2026 09:00
@github-actions github-actions Bot added readiness: checking Automated validation is still running and removed readiness: action required A blocking check or review needs attention labels Sep 24, 2026
@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • Learned-cost refresh performs whole-file I/O on the gateway event loop — span=abaf2747bd41 — fixed in 6ff2bdd

Disposition: fixed. The gate no longer opens the cost log at all: the reaper sweep reads it on the maintenance executor (_refresh_learned_cost_impl, once at reaper start and then every sweep) and publishes one float onto the manager (_learned_cost_gb); the gate does arithmetic over that attribute. The stat-keyed cache this finding was about is removed.
self-added: yes
mechanism: off-loop refresh onto a manager attribute, replacing the on-loop cached reader

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • learned cost becomes the per-live-worker RSS target, so an observed worker's reservation never retires — span=e20707424d7b — fixed in 6ff2bdd

Disposition: fixed. _startup_memory_reserve_gb now takes next_start_gb separately: the next start, unregistered claims and dedicated workers not yet sampled are priced at the learned figure; a SAMPLED worker owes only max(configured, own peak) - observed RSS, as before, so the pinned property ([{last_rss_gb: 0.5}], 1, 0.5) still holds and no phantom reserve survives a sample. New pin: test_startup_reserve_prices_unmeasured_starts_at_the_learned_cost.
self-added: yes

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • Watch: a stale-high learned p90 can defer every dedicated spawn indefinitely — span=dd2276e71016 — fixed in 6ff2bdd (the named escape hatch)

Disposition: fixed, by the second clearing condition: the deferral now names the per-start price, the learned p90 and the configured cost in the log line, the SEL record (startup_cost_gb, learned_cost_gb) and the durable row's deferral reason, together with the reset path (delete subagents/cost_samples.jsonl under the data home, or lower spawn_min_memory_gb).
Not taken: admit-one-when-idle or sample decay. Both weaken the reserve exactly where it protects against an OOM, and a first start that cannot fit floor + p90 on a host is the reserve doing its job, not a wedge; the wedge case is a p90 that outlived its roster, which is now visible from the deferral itself.

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • Suggestion: fold compute_max_subagents and the reserve through one cost function — span=d430066d4e6a — rebutted

Disposition: rebutted. The two guards deliberately read the learned cost differently: the cap is a COUNT and a learned cost below the configured one should raise it (learned-over-configured), while the reserve is a safety floor against an unrecoverable OOM, so an operator's higher pin must never be lowered by a learned figure (max). Folding them would force one semantics onto the other. Both semantics and the reason are now stated in _startup_cost_gb's docstring.
Class: same-quantity-different-policy is intentional here, not drift.

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • Item 3 — one consumer, generalized: kwarg-keyed cache dict for one caller — span=eca02346c330 — fixed in 6ff2bdd (subtraction)

Disposition: fixed. read_learned_cost_cached and _LEARNED_COST_CACHE are removed entirely; the learned figure is published off-loop by the reaper sweep onto one manager attribute, which is the single slot the item asked for and also removes the on-loop read.
self-added: yes

@github-actions github-actions Bot added readiness: action required A blocking check or review needs attention and removed readiness: checking Automated validation is still running labels Sep 24, 2026
@bolichen97
bolichen97 force-pushed the fix/spawn-reserve-learned-cost branch from 6ff2bdd to b5306db Compare September 24, 2026 09:49
@github-actions github-actions Bot added readiness: checking Automated validation is still running and removed readiness: action required A blocking check or review needs attention labels Sep 24, 2026
@bolichen97

Copy link
Copy Markdown
Contributor Author
  • Degraded reads erase the learned reserve — span=e9fb8da2b635 — fixed in b5306db

Disposition: fixed. read_learned_cost degrades an unreadable log to None instead of raising, so the previous 'keep on exception' did not cover it. _refresh_learned_cost_impl now drops a held figure ONLY when the log is absent (cost_log_present() -- first boot or the operator's documented reset); a present log that yields None keeps the previous value. Pinned in test_reaper_sweep_publishes_the_learned_cost_off_loop (raising read, degraded read, deleted log).
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • sampled workers charged against another worker's peak — span=6a3a78cb8f4d — fixed in b5306db

Disposition: fixed. Each sampled worker's gap is now max(0, max(cost_gb, ITS OWN peak_rss_gb) - last_rss_gb); the global max peak still raises the next start's price only. New reserve case: a 7.5 GB sibling beside a 1 GB worker reserves 8.0, not 14.5.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • 'from the learned per-run cost' misstates the source — span=8ba58aa868ef — fixed in b5306db

Disposition: fixed. The refusal text and the deferral reason name the figure that actually set the price: 'learned per-run p90' only when it is the larger one, otherwise 'configured agent.subagent_cost_gb'. Pinned by test_low_memory_deferral_names_the_configured_price_when_nothing_is_learned.
self-added: yes

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • sampled_gap uses the max peak over all dedicated workers, not its own — span=88a3825dcf3c — fixed in b5306db

Disposition: fixed, both instances of this span: the gap is per-worker (own peak floored at the configured cost), and the docstring sentence 'Until sharing is known, reserve the configured process cost' now reads 'a row is priced as an unmeasured dedicated start (next_start_gb), since it may yet become one' -- which is what the code does for a row whose sharing is not yet known.
self-added: yes

@bolichen97

bolichen97 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor Author
  • docstring promises an unreadable log keeps the previous value, but read_learned_cost never raises — span=5d3768fe7110 — fixed in b5306db

Disposition: fixed. The refresh keeps the held figure whenever the log is present but yields None, and drops it only when the log is absent; the docstring now states exactly that rule.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • stale "heaviest known" comments and read_learned_cost cross-reference — span=e20707424d7b — fixed in 5c598a3

Disposition: fixed. Both comments now state "configured cost plus live dedicated peaks"; _startup_cost_gb's docstring cites read_learned_costs with dedicated_only.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • Watch: the sweep's merge defeats the 30-day expiry on a long-lived gateway — span=66ab6f0d4206 — fixed in 5c598a3

Disposition: fixed, by the first clearing condition: a clean, fully-parsed read is authoritative and drops the held buckets it did not yield; merge is reserved for the incomplete read it was built for. An expired bucket now retires within one sweep of its last sample crossing the horizon, on the running process, with no restart or log deletion.
self-added: yes

@github-actions github-actions Bot added readiness: checking Automated validation is still running readiness: action required A blocking check or review needs attention and removed readiness: action required A blocking check or review needs attention readiness: checking Automated validation is still running labels Sep 24, 2026
The spawn guard's `_startup_memory_reserve_gb` is the only thing that prices a
dedicated start between admission and the reaper's first RSS sample (60 s), and
a runtime takes tens of seconds to reach its resident size. It read the
first-boot fallback `subagent_cost_gb` (0.5 GB) even when the cost store already
held a learned p90 of ~6 GB -- the figure `compute_max_subagents` sizes the cap
from -- so a burst of starts each cleared the raw free-memory check and then grew
into the same headroom together, taking a 64 GB host from 25 GB free to 0.4 GB
in under two minutes.

Two prices now apply. A WARMING start (the next one, a claim awaiting
registration, a dedicated worker fewer than two sweeps have measured -- one
reading can land mid-growth) is priced by `_effective_next_start_gb` at the
larger of the configured fallback, the learned p90 for the agent being spawned
(its own history when it has one, else the heaviest known: `learned_cost_for`)
and any live dedicated peak, less the RSS it already holds. A SETTLED worker owes
only the gap between the larger of the configured cost and its OWN peak and its
observed RSS, so a learned p90 above what that worker needed never becomes a
reserve no later sample can retire.

The learned figures reach the gate as `SubagentManager._learned_costs_gb` (per
agent) and their max, published by the reaper sweep's off-loop
`_refresh_learned_cost` (once at reaper start, then every sweep); the gate does
arithmetic only and never opens the cost log on the event loop. A held figure is
dropped only when the log is ABSENT (`cost_log_present`: FileNotFoundError, the
operator's documented reset); a present log that yields nothing keeps it. A
low-memory deferral names the effective per-start price, which figure set it
(live peak / learned p90 / configured pin), and the reset path, in the log
line, the SEL record and the deferral reason.
@bolichen97
bolichen97 force-pushed the fix/spawn-reserve-learned-cost branch from 5c598a3 to da66f84 Compare September 24, 2026 16:15
@github-actions github-actions Bot added readiness: checking Automated validation is still running and removed readiness: action required A blocking check or review needs attention labels Sep 24, 2026
@bolichen97

Copy link
Copy Markdown
Contributor Author
  • Learned-cost parsing retains the entire agent-writable log — span=b201b7e862d6 — fixed in da66f84

Disposition: fixed. _iter_samples streams the log one record at a time; read_learned_costs_checked folds each record into bounded per-bucket deques (window values) under a parse ceiling on distinct buckets, and compact_cost_log streams the same way, holding only the records it will keep. Nothing materialises the file; _read_samples remains for tests only.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • docstring "dropped ONLY when the log is absent" contradicts complete-read replacement — span=51047c440b4e — fixed in da66f84

Disposition: fixed. _refresh_learned_cost_impl's docstring now states the three outcomes: absent clears, a complete read of an inspectable log replaces, an incomplete read merges.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • "only the heaviest _MAX_BUCKETS" contradicts first-seen retention — span=92c76ae60709 — fixed in da66f84

Disposition: fixed. The parse-time count bound is now a memory CEILING (_PARSE_BUCKET_CEILING=1024, first seen) that no real log reaches, and the returned/held map is the heaviest _MAX_BUCKETS via cap_buckets; the comment states both.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • parse-time first-seen bound can drop the heaviest bucket — span=8c09a750b6c3 — fixed in da66f84

Disposition: fixed. The heaviest-wins selection is cap_buckets on the returned map; the parse keeps a much higher first-seen ceiling purely as a memory bound. Pinned by test_the_held_map_is_bounded_against_an_agent_writable_log (74 buckets → the heaviest 64 returned).
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • a present-but-unopenable log reads as complete and wipes the map — span=5d3768fe7110 — fixed in da66f84

Disposition: fixed. _iter_samples marks the read incomplete on any OSError other than FileNotFoundError, and the refresh treats an uninspectable identity as incomplete too, so both take the merge arm; only a complete read of an inspectable log replaces. Docstring and subagent.md say so.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • generation bumped after the resets leaves a window for a straddling sweep — span=24940bf6b055 — fixed in da66f84

Disposition: fixed. The respawn increments _rss_generation BEFORE clearing _rss_samples and last_rss_gb, so any sweep that has not yet re-checked observes the new generation and discards its reading.
self-added: yes

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • Item 5 — compat claim about read_learned_cost contradicted; horizon changes the cap's input — span=3ed144a49fcd — fixed in da66f84 (the subtraction)

Disposition: fixed, by the named subtraction: the age horizon is a max_age_secs argument the reserve's refresh passes and read_learned_cost does not, so the cap's input carries no horizon; the bucket bound returns the heaviest buckets, so the cap's max is unchanged for any real log. Pinned in test_expired_samples_do_not_price_a_start (read_learned_cost still 6.0 on expired samples). The compat sentence now states exactly this.
self-added: yes

@github-actions github-actions Bot added readiness: action required A blocking check or review needs attention and removed readiness: checking Automated validation is still running labels Sep 24, 2026
@bolichen97

Copy link
Copy Markdown
Contributor Author
  • Watch: self-locking deferral for up to 30 days on a memory-tight host — span=82b6e208c4a1 — accepted-and-deferred

Disposition: accepted-and-deferred to #13323 (deferred-finding, owner @bolichen97, Due 2026-10-24), which names exactly the bounded re-probe this Watch asks for (one admission at the configured price after sustained all-starts deferral). Residual accepted for this PR: a stale price now expires within 30 days on its own, prices only its own bucket, and the deferral names the figure and the reset path; a re-probe is the very admission the reserve exists to defer, so it needs its own design and test.

@bolichen97

Copy link
Copy Markdown
Contributor Author
  • an uninspectable identity on first publication takes the replace branch — span=5d3768fe7110 — rebutted

Disposition: rebutted as not-a-defect. previous is None only while nothing has ever been published from an inspectable log, and in that state the held map is empty, so the replace branch writes {} over {}: no learned price exists to wipe. Once a source has been recorded, an uninspectable identity forces complete=False and replaced=False, so the merge arm runs and held prices are kept, exactly as the comment states. A log that is never inspectable leaves the map empty, which is the pre-existing "nothing learned" behaviour, not a regression.

@bolichen97

Copy link
Copy Markdown
Contributor Author

/ai-review override gpt da66f84: The merge path is reachable only via a record the log's writer cannot emit (>128 MiB or non-UTF-8 line) or an EIO mid-read, and must then clear the configured floor and every live dedicated peak before it can under-reserve; the proposed max(previous, parsed) would freeze buckets at their historical maximum on a permanently-unreadable log.

@github-actions

Copy link
Copy Markdown
Contributor

Human judgment recorded

@bolichen97 marked the gpt AI finding as false positive, not applicable, or explicitly accepted for da66f842e1f3cca4dad08b41d4a56e3e4b464d3e.

The merge path is reachable only via a record the log's writer cannot emit (>128 MiB or non-UTF-8 line) or an EIO mid-read, and must then clear the configured floor and every live dedicated peak before it can under-reserve; the proposed max(previous, parsed) would freeze buckets at their historical maximum on a permanently-unreadable log.

This decision applies only to this commit. A new push requires a new judgment.

@github-actions github-actions Bot added readiness: passed Eligible automated validation passed for the current revision and removed readiness: action required A blocking check or review needs attention labels Sep 24, 2026
@bolichen97
bolichen97 merged commit 6b4556e into main Sep 24, 2026
96 of 104 checks passed
@bolichen97
bolichen97 deleted the fix/spawn-reserve-learned-cost branch September 24, 2026 18:59
@github-actions github-actions Bot removed the readiness: passed Eligible automated validation passed for the current revision label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants