Repository navigation
fix(exec,api): disabled spill means no spill; durable projects spill into capped project scratch (#1595) - #1609
Conversation
…into capped project scratch (#1595) A disabled spill policy used to leave DataFusion's default disk manager in place, which spills into the OS temporary directory with no cap, so queries spilled when the policy said they must not. Now: - spill disabled disables the disk manager: a query over its budget fails with a resource error; - the default policy (enabled, no directory) spills a durable project's queries into <project>/.graphforge-query-spill/, one locked subdirectory per open instance, capped at 8 GiB per query unless the caller sets max_bytes; an in-memory instance does not spill; - a configured directory works as before; - scratch left by an exited process is reclaimed when an instance next acquires scratch; live instances' scratch is never touched; - search-index adjacency builds keep their own stage spill root unless a directory is configured. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Repository: CurateLabs/graphforge/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Warning Billing warning: we have not been able to collect payment for this subscription for more than 72 hours. Please update the payment method or pay any pending invoices in Billing to avoid service interruption. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
This comment has been minimized.
This comment has been minimized.
- A lock file is created under a temporary name, locked, then renamed into place, so a reclaimer never sees an unheld live lock; a claim a reclaimer raced is retried with a fresh token. A subdirectory counts as lockless only if its lock is absent when examined. - Reclaiming other owners' leftovers is best effort: an unremovable entry is skipped and never blocks acquiring this instance's scratch. - Building a spilling session is fallible: an unusable spill directory or runtime error is a storage error for that query, not a panic. - Read-only views of a durable project (checkpoints, inspection) spill into project scratch like any query. - EXPLAIN renders its plan with a no-spill configuration and never acquires scratch. - Docs: DataFusion's 100 GiB default cap for a configured directory without max_bytes; the read-only and EXPLAIN rules. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
CI's shared-directory test hit the race the review described: a reader's reclamation locked another reader's freshly created lock file. The claim now publishes its lock only once held and retries a raced claim; two tests drive each interleaving deterministically through a test-only hook, and fail when the retry is removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closes #1595.
What changes
This implements the maintainer decision recorded on #1595.
SpillPolicy { enabled: false }now disables DataFusion's disk manager. Before, the session left DataFusion's default in place, which spills into the OS temporary directory with no cap. Now a query over its memory budget fails with a resource error.enabled: truewith nodirectory. Queries spill into<project>/.graphforge-query-spill/, capped at 8 GiB per query unless the caller setsmax_bytes..lockfile. Dropping the instance removes both.enabled: falseand now fail closed instead of silently spilling.Docs:
docs/development/execution-resource-policy.mdgains a "Query spill" section.Evidence
Tests, with each defect-targeting test proven by mutation:
graphforge-execResources exhausted. Mutating back to OS-temp spill fails this test.exceeded.graphforge-storagequery_spillgraphforge-apiruntime_ownership::spill_tests, checking what each policy and instance resolves to:End to end on the release
gf(SHA-256eb9840cf…6b334d), with the 9M-leaf star (#1585, #1591) at default resources:MATCH (a)-[r]->(b) RETURN a.node_uuid AS s, b.node_uuid AS d ORDER BY s, dreturned 9,000,000 rows, with peak RSS 830 MB.<project>/.graphforge-query-spill/, observed by polling during the query.TMPDIR: pointed at an empty directory, it stayed empty.gf verifyandgf storage-attributionsucceed with the scratch root present.Suites and gates:
cargo test -p graphforge-execcargo test -p graphforge-storage --libcargo test -p graphforge-apicargo test -p graphforge-clinode --testand checkpoints tests passmulti_ontology.py,checkpoints.pyandconcurrency_parity.pypasscargo fmt --all -- --check,cargo clippy --workspace -- -D warningsmake pre-push-fastReview and CI
An independent review found no defects in the policy semantics, and five issues, all verified and fixed in
f45df116and812c1076:shared_directory_semantics_testshit exactly this on the first head: "a new scratch lock is already held". Now the lock is created under a temporary name, locked, then renamed into place, so a visible<token>.lockis always held. A raced claim retries with a fresh token. Two tests drive each interleaving deterministically through a test-only hook, and both fail when the retry is removed.GfError::Storage, with a test.max_bytesis unset). Both are fixed, EXPLAIN with a test.Not in this change
tests/with pytest (beyond CI's selection) showstests/release_workflows/atomic-recoveryfailing: thegenerator.yamlfingerprint recorded in itsscenario.yamldoesn't match the file. It is unchanged since chore: remove opaque mXX milestone shorthand #689 and unrelated to this change.🤖 Generated with Claude Code
Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.