Skip to content

disk-hygiene: add --auto-split, or name the largest children in the 250k cap error #5517

Description

@kyle-sexton

Split from #5214.

Problem

Whole-home scans hit the entry cap and there is no automatic split. Repro: run a whole-home inventory with --confirmed-large-scan. It exits 2 with:

snapshot exceeds 250000 entries; rerun with --max-depth or split the audit into bounded subtrees

The cap is MAX_SNAPSHOT_ENTRIES = 250_000 at skills/clean/scripts/hygiene.py:34, raised at hygiene.py:2009. Selecting children with --root-children hits the same cap when the selected children are large. There is no multi-target form and no automatic split, so the worker divided home into four runs by hand and merged the results. --sizes-only returns exact totals but writes no per-entry inventory (entries: 0, no findings), so every audit needs a second pass over each subtree. Each scan took minutes.

Expected

One command that audits a large tree, either by accepting several targets or by splitting on the depth-1 frontier itself and returning a documented set of snapshots.

Suggested direction

Add an --auto-split option (or a documented orchestration helper) that uses the --sizes-only frontier to plan bounded subtree scans, runs them, and reports coverage. At a minimum, make the cap error name the largest children so the caller can choose the split.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority: mediumReal value, no hard deadline; normal backlog flow.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions