Skip to content

subagent: startup memory reserve prices a new spawn at earlier workers' whole-process-tree peaks (test/build children included) #13489

Description

@Premshay

Problem

The spawn memory guard prices a new dedicated subagent from memory figures that include everything a worker ran, not just the agent process itself. On a host where subagents run test suites or builds, a new spawn is priced at the size of the biggest workload any worker has run, and it waits for that much free memory.

Three places use the whole-tree figure:

  1. Live peaks set the next start's price. _effective_next_start_gb (src/kiro_crew/subagent.py:1379) takes the maximum over every live dedicated worker's peak_rss_gb. That peak comes from _proc_subtree_sample(info._pid) (src/kiro_crew/subagent_manager/monitoring.py:663, src/kiro_crew/subagent.py:915), i.e. the whole process tree.
  2. Settled workers keep their peak reserved. In _startup_memory_reserve_gb (src/kiro_crew/subagent.py:1395), a settled worker still owes max(cost_gb, peak_rss_gb) - last_rss_gb. A worker that reached 20 GB during a test run and is now idle at 1 GB reserves 19 GB until it exits.
  3. Learned costs persist it (new in fix(subagent): price the startup memory reserve at the learned cost #13294). _record_cost_impl (src/kiro_crew/subagent_manager/monitoring.py:768) stores the whole-tree peak_rss_gb, and _startup_cost_gb (src/kiro_crew/subagent.py:1454) prices every start at that agent's 90th percentile. Workload peaks now set a lasting price even when nothing heavy is running.

Line references are to main at 6b4556e.

Observed

On a 94 GB host with several agents running tests in parallel:

  • One live worker reached a 20.8 GB tree peak while running a test suite. The admission floor rose to 24.8 GB (4.0 GB spawn_min_memory_gb plus 20.8 GB). With about 24 GB free, every spawn on the host was deferred; one caller's spawn_run waited about 8 minutes, then slot-queued.
  • Learned-cost history (cost_samples.jsonl, last 50 dedicated samples per agent):
Agent Median 90th percentile Max
Agent A 1.15 GB 7.39 GB 11.29 GB
Agent B 1.88 GB 6.84 GB 12.97 GB
Agent C 1.83 GB 5.70 GB 12.55 GB

After #13294, a typical ~1.2 GB start for Agent A is priced at ~7.4 GB each time, because the 90th percentile reflects what those workers ran, not what the agent needs to start.

Why it matters

The reserve exists to cover the time between admitting a start and the new runtime reaching its resident size (#13294's burst case). A worker's test or build children say nothing about that. Counting them over-reserves in proportion to the heaviest workload anyone runs. On hosts with heavy local work, parallel spawning collapses without anything visibly failing: the result is long deferred_low_memory waits.

Suggested direction

Open to maintainers' preference: measure two things separately. Price the next start from the agent's own processes (the runtime plus its MCP children, which the subtree walk already recognises), and record that figure for the learned cost. Keep the whole-tree measurement for what it's good for: live host-pressure decisions such as the raw free-memory check and the reaper. A worker's future growth could still be covered by a separate, capped per-worker allowance rather than by its historical peak.

Relation to #13294

Complementary. #13294 correctly stops pricing starts at the 0.5 GB fallback. This issue is about what the learned and live figures measure.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: agentsACP runtime, sub-agents, session lifecyclebugSomething is not workingneeds-investigationTriage: requires deep analysis before a fix

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions