Repository navigation
feat(cloud-sandbox): Phase 20.15 M1 Cloud Sandbox Plane foundation - #2
paopaonyapi-creator wants to merge 8 commits into
Conversation
|
✅ Deterministic PR hygiene checks passed. |
|
Don't read this PR's red CI as M1 breakage — the baseline is red independently of these 21 files, Proof it is not this branch: the failed-job set of Causes, filed separately from this feature PR so neither blocks the other:
Consequence for review: |
|
Amendment to Verification. The Zero of those errors originate in Also adds the M2 lifecycle slice: One honest loose end: the first combined run of those 9 files reported a single 5s test timeout and |
…ndation Lay the provider-neutral contracts, sidecar persistence, policy capabilities and optional-subsystem guard before any emulator is reachable. Reuse rather than rewrite where the repo already has it: agent-os/policy.ts and gateway.ts drive approval, and generation/cloud/lifecycle-manager.ts shapes the sandbox lifecycle. Two corrections to the source spec are load-bearing. The schema is SQLite, not PostgreSQL, because this repository has no Postgres; and the Docker port stays null-backed because Lambda and RDS are Docker-backed and this host has no daemon, so reporting UNAVAILABLE is honest where faking a pass would violate the fidelity gate. Isolation testing surfaced two real defects in the host-protection rules: the forbidden bind list was POSIX-only, so a container could mount the Windows profile, and bind parsing split `C:\Users\me:/x` on the first colon, yielding a source that matched no rule. The boundary guard was driven red by importing the subsystem from management-api.ts and reverted, and the lossy debt-ledger round trip it exposed is filed upstream as lidge-jun#5650.
…olicy gate Add SandboxManager: the governed lifecycle the source spec asks for in sections 8, 11, 12, 46 and 60. Every create runs flag check, then policy, then service screening, then the idempotency fence, and only then an adapter call, so anything that can refuse without touching a runtime refuses while there is still nothing to tear down. Extending a TTL is bounded by the same ceiling as creating one, because an uncapped extension is the easiest way for an agent to pin a sandbox open forever. Two decisions that only look like preferences: - authorize is injected and DENIES when absent. Defaulting to allow would make a missed wiring the one bug that silently disables every gate below it. - A Docker-backed service is admitted and reported UNAVAILABLE, while a service whose capability requires approval is refused. The first keeps the fidelity marker honest per section 42; the second stops an agent obtaining a gated service by bundling it with an ungated one. The lifecycle tests caught a real ordering fault, not a test fault: cloud_operations carries a foreign key to cloud_sandboxes, so writing the audit row before the sandbox row made every create die on SQLITE_CONSTRAINT_FOREIGNKEY. A provisioning row is now recorded first, which also leaves a visible sandbox behind if the process dies mid-create. Also records the upstream distribution reality in docs section 10.1: Floci is a Java (JAX-RS/Vert.x) CLI plus an image, published on neither npm nor PyPI, and its Docker Hub tags are only latest and dated nightlies with no stable semver -- so the pinning rule in section 66 has to be satisfied by digest, and the daemon-presence assumption in section 10 is corrected from "absent here" to "never assume either way".
…ied launch plan Add FlociAwsAdapter and the launch-plan layer that configures it, against the upstream configuration reference rather than LocalStack habits. The decisive fact is that FLOCI_PORT exists, which is what makes one emulator instance per sandbox possible and therefore makes the isolation guarantee in section 11.2 real instead of aspirational. Pin by digest, not tag: the published channel is only latest and dated nightlies with no stable semver, so the section 66 rule can only be satisfied immutably. Record that, the Java and Vert.x runtime, the absence of npm and PyPI packages, and the per-service image variables, in docs section 10.2 -- those variables are also the answer to the section 74 question about where the image allowlist lives. Two refusals that matter more than the features: - listResources throws RESOURCE_DISCOVERY_FAILED instead of returning an empty list. An empty list would tell the observatory, the diff engine and the promotion gate that the sandbox holds nothing, and a "no leaks" verdict built on that is worse than no verdict. Inventory comes from Terraform state in M3, and upstream documents no resource-listing endpoint or core health route. - The launch plan has no path to process.env and emits no AWS_ variables. Upstream's own quickstart is `eval $(floci env)`, which would put fake credentials and a local endpoint into every later process; configuring the emulator and configuring a client of it are kept as separate surfaces for exactly that reason. Also corrects three fidelity markers against upstream's own service table -- EventBridge and API Gateway REST are in-process, so the old PARTIAL would have forced a section 43 production retest the emulator does not need. Step Functions stays PARTIAL because upstream says nothing about it, and a guessed fidelity is precisely what section 42 forbids. Verified: 175 pass / 0 fail across the subsystem, policy and hygiene sets; typecheck reports zero errors from src/agent-os/cloud-sandbox (the three remaining TS errors are issue #3, removed by PR #4); privacy:scan clean.
e2ccdee to
17f0078
Compare
|
Stacked on #4 now — and this supersedes my earlier typecheck amendment. Base retargeted from Consequence, and it is the proof the fix actually works: Branch was rewritten by the rebase (4 commits replayed, |
…ontainer Running the image that M2 pins replaced three guesses with observations, and one of the guesses was wrong in a way worth naming: docs section 10.2 claimed upstream publishes no core health route. It does -- /_floci/health, which the container's own healthcheck.sh probes over a raw TCP socket and accepts only HTTP 200. /health is an identical alias and the trailing-slash form 404s, so the adapter now targets the path the image itself trusts and requires a parseable document instead of "any HTTP response". A captive proxy that answers 200-with-html would have satisfied the old rule and been treated as a live emulator. The health document registers 121 services and reports every one of them "running", including lambda, rds and eks, in a container started with no Docker socket mounted. That is the strongest possible argument for keeping fidelity in the capability registry: if health had been trusted, an agent would have been told a Lambda sandbox works. So HealthReport carries registeredServices beside readyServices rather than merging them. An unsigned GET /?list-type=2 returns a real ListAllMyBucketsResult, because S3 auth enforcement is off by default. Loopback binding is therefore a control and not a style preference: published on another interface the emulator is an open object store. Both findings are now assertions in a live test that skips unless PAO_CLOUD_FLOCI_REAL=1, so the default suite stays hermetic and Docker-less CI is unaffected. Verified: typecheck clean, 145 pass and 4 skip in the unit run, 4 pass against the live container, privacy:scan clean, and the debt ledger still round-trips with zero lines lost.
…it on a live daemon
M7 landed early because the Floci adapter could not be verified honestly any other way.
This is not the request-filtering socket proxy of source spec 15.1; the transport is the
operator's own docker CLI, invoked as an argv array with no shell. What 15 actually requires
still holds: the agent has no transport to the daemon, only five verbs through a broker that
validates every spec before it spawns, and there is one run() to audit rather than a call site
per adapter.
The guard this repository already has caught me shipping the exact bug in issue 3: index.ts
imported the new port while that file was still untracked, and the tracked-import test failed
as designed.
Two real defects found before committing, both by reasoning about the daemon rather than
trusting my own assertions:
- --publish was assembled as `127.0.0.1:4570/tcp`, which is not Docker's syntax at all.
The mapping is now rebuilt as hostPort:containerPort from a numeric container port, and a
non-numeric port is refused outright.
- mapState understood `docker inspect` vocabulary ("running") but not `docker ps` vocabulary
("Up 2 seconds"), so every live container read as unknown and leak detection could not tell
a running orphan from an exited one.
The live block also stopped assuming somebody had started an emulator: it now provisions its
own through the broker in beforeAll and sweeps by label in afterAll, because the previous
version reported four failures the moment a manually-run container was cleaned up -- a test
reading the state of the machine instead of the behaviour of the code.
Verified with PAO_CLOUD_FLOCI_REAL=1: 7 pass / 0 fail, including a privileged spec refused
before the daemon saw it and destroy leaving no labelled orphan, confirmed independently with
docker ps --filter label=pao.sandbox.id rather than trusted to the test's own assertion.
Default run stays at 1432 pass / 0 fail with the live block skipped, typecheck clean and
privacy:scan clean.
…olicy gate Adds CleanupController, which is the part of source spec 46 and 47 that can be honest about its own coverage: it checks containers, registry rows and sandboxes stranded mid-lifecycle, and says nothing about networks, volumes, host processes and ports, because enumerating those would mean widening DockerControlPort past the five verbs it brokers -- a broker that lists everything the daemon knows is a slow path back to the raw socket access section 15 forbids. Two rules that decide behaviour: - A leak is reported, never silently repaired, except for orphans whose sandbox row is already destroyed or absent. Over-reporting is annoying; a detector that calls a running sandbox's container an orphan hands the sweeper permission to destroy live work. - A container that refuses to die stays in outstandingLeaks rather than being counted as removed, and the cycle re-measures after sweeping instead of trusting its own bookkeeping. That is what catches a daemon that accepted a stop and never released the container. destroyWithVerification exists because the adapter cannot see its own emulator's side containers: Floci starts a labelled worker for a Docker-backed service and its own teardown leaves it. The test models exactly that and asserts the controller is the only layer that notices. The import-graph guard from the previous commit caught this change before it was staged: index.ts imported cleanup.ts while it was still untracked, the same failure as issue 3.
The version committed in 4534014 read an emulator that a human had started on 4566 behind an opt-in env flag, and its commit message quoted the 7-pass result of this rewrite -- so the pushed commit did not contain the file it described. Correcting that here rather than amending, per the repository's no-amend rule. Functionally this replaces an environment assumption with setup: beforeAll starts Floci through the brokered port, afterAll destroys it and sweeps by label so a crashed run cannot leave a container bound to a port. The previous arrangement produced four failures on a correctly behaving host the moment the manual container was cleaned up, which is the machine's state reporting, not the code's.
Summary
CloudEmulatorAdapterSPI, a sidecar SQLite store with aPRAGMA user_versionmigration, anenv-var feature-flag layer where every switch defaults off, a brokered
DockerControlPortwith adeny-everything default implementation, a capability/service-fidelity registry, the error
taxonomy with its retry verdicts, and an in-memory adapter that is the only implementation today.
four new capabilities added to
src/agent-os/policy.ts(cloud.sandbox,cloud.iac.local,cloud.staging.apply,cloud.production.apply) are denied bydefault_denyuntil someone writesa policy row — asserted by tests, not by intent.
because this repository has no Postgres client at all; and the Docker port stays null-backed
because Lambda and RDS are Docker-backed services and the development host has no reachable
daemon. A service with no daemon reports fidelity
UNAVAILABLErather than passing quietly,which is what keeps a local run from being mistaken for production proof.
human permit flow (
agent-os/policy.ts,agent-os/gateway.ts) gate infrastructure actions,generation/cloud/lifecycle-manager.tsshapes the sandbox lifecycle, and the optional-subsystemboundary follows the
AGENTS.mdrule that an inactive subsystem must not be imported at all.the forbidden bind list was POSIX-only, so a container spec could mount the Windows user profile
(
C:\Users\...) while matching no rule, and bind parsing splitC:\Users\me:/workspaceon thefirst colon, yielding the source
"C". Both are covered by regression tests.docs/, and files the lossydebt-ledger round trip this work exposed as upstream #5650.
Verification
Ran on the committed tree (
346c224ce), Windows 11 build 26200, Bun 1.4.2:subsystem from
src/server/management-api.tsfails both the import-graph walk and the direct-namescan.
src/server/management-api.tsis byte-identical todevafterwards.reported itself broken. A per-file specifier cache brings it to 446ms.
recordDebt()against a copy of the ledger loses 0 lines and parses 6 entries with no empty fields.
Not run, stated plainly: the full
bun run testsuite.bun run test:changedis not usable as agate for this branch — it resolves its comparison ref to
origin/dev, whose merge base with HEAD is28c69b03a, so it selects tests for 236 unrelated commits and exceeded its own 900s budget.Coverage for the touched subsystem is therefore the focused sets above. Repository CI
(typecheck + full suite on Linux, Windows and macOS) is the remaining signal, which is why this
opens as a draft.
No UI is added or changed in this milestone, so no screenshot applies.
Checklist
Note for review:
src/agent-os/policy.tsis a security-boundary file underMAINTAINERS.md, sothis still needs the explicit security review that file requires. What was checked here: the new
capabilities are additive and fail closed with no policy row, the two out-of-local capabilities are
approval-required, no credential value is ever logged or persisted (only digests and redacted
summaries, via the existing
recordWebMcpCallpath), andprivacy:scanis green.Deliberately deferred to later milestones, and recorded in
docs/Phase-20.15-*.md§12: the realFloci adapter (M2), the Terraform/OpenTofu runner (M3), any Docker-backed service (M7, needs a
daemon), and all promotion beyond local, which stays flag-off.