Deferred from the Design Review Watch on #13294 (span=b2211c56f0c9).
Owner: @bolichen97
Due: 2026-10-24
Task
The spawn guard now prices every unmeasured dedicated start at the learned per-run p90 (never below agent.subagent_cost_gb). The p90 is only ever refreshed from runs that actually execute, so on a host whose real per-run cost DROPPED (a heavy MCP roster removed, a backend switched) and whose free memory never clears spawn_min_memory_gb + old p90, deferred starts produce no new samples and the stale figure never self-corrects. #13294 makes that state visible (the deferral names the price, its source and the reset: delete subagents/cost_samples.jsonl under the data home, or lower spawn_min_memory_gb) but leaves the remedy manual.
Pick one and pin it with a test:
- Sample decay: age out samples older than N days (or beyond a wall-clock window) in
compact_cost_log / read_learned_cost, so a roster change is forgotten without operator action.
- Probe admission: when zero dedicated workers are live and the ONLY thing blocking a start is the learned reserve (the configured-cost reserve would admit it), admit exactly one start so a fresh sample lands; its observed RSS then re-prices the next one.
Not a goal: weakening the reserve while dedicated workers are live -- that is the window #13294 exists to close.
Pointers: _startup_memory_reserve_gb / _startup_cost_gb in src/kiro_crew/subagent.py, _refresh_learned_cost_impl in src/kiro_crew/subagent_manager/monitoring.py, read_learned_cost in src/kiro_crew/subagent_cost.py.
Deferred from the Design Review Watch on #13294 (span=b2211c56f0c9).
Owner: @bolichen97
Due: 2026-10-24
Task
The spawn guard now prices every unmeasured dedicated start at the learned per-run p90 (never below
agent.subagent_cost_gb). The p90 is only ever refreshed from runs that actually execute, so on a host whose real per-run cost DROPPED (a heavy MCP roster removed, a backend switched) and whose free memory never clearsspawn_min_memory_gb + old p90, deferred starts produce no new samples and the stale figure never self-corrects. #13294 makes that state visible (the deferral names the price, its source and the reset: deletesubagents/cost_samples.jsonlunder the data home, or lowerspawn_min_memory_gb) but leaves the remedy manual.Pick one and pin it with a test:
compact_cost_log/read_learned_cost, so a roster change is forgotten without operator action.Not a goal: weakening the reserve while dedicated workers are live -- that is the window #13294 exists to close.
Pointers:
_startup_memory_reserve_gb/_startup_cost_gbinsrc/kiro_crew/subagent.py,_refresh_learned_cost_implinsrc/kiro_crew/subagent_manager/monitoring.py,read_learned_costinsrc/kiro_crew/subagent_cost.py.