Skip to content

feat(storage): recover interrupted project transactions during open #750

Description

@DecisionNerd

Problem

GraphForge resolves CURRENT during GraphForge::new, but recover_project_transactions is invoked explicitly in tests and selected workflows rather than as a complete open lifecycle. Abandoned journals, attempts, and trash can therefore persist until a separate recovery operation, while callers lack a stable report of recovered or deferred work.

Objective

Make deterministic recovery-on-open a first-class Rust-owned lifecycle step without blocking ordinary snapshot reads behind a live writer or deleting ambiguous evidence.

Requirements

  • Resolve and validate CURRENT before cleanup authority is considered.
  • Detect whether recoverable journals/attempts/trash exist using bounded, link-safe inspection.
  • Acquire the existing writer/transaction/checkpoint locks in the established order.
  • Recover idempotently, re-resolve CURRENT, and open the final pinned generation.
  • If a live writer owns recovery locks, allow a safe read of valid CURRENT and report recovery as deferred; do not steal locks or fail an otherwise valid snapshot open.
  • Never choose a generation by UUID/time/directory order.
  • Preserve unknown or corrupt entries for explicit inspection and return typed integrity errors when authority is ambiguous.
  • Expose a safe recovery summary through the Rust facade and applicable open evidence.
  • Keep initialization and read-only checkpoint opens semantically distinct.

Acceptance Criteria

  • A fresh open cleans or quarantines provably abandoned state and returns the same selected generation on repeated open.
  • A concurrent live writer does not cause false abandonment or prevent a safe reader from pinning valid CURRENT.
  • A killed pre-publication writer reopens to the parent and cleans only its proven private attempt.
  • A killed post-publication writer reopens to the committed child and repairs advisory journal state.
  • Corrupt CURRENT/manifest state fails closed without fallback election.
  • Recovery work and deferral are bounded, observable, and tested across supported platforms.

BDD Completion Scenarios

  • Given a writer died before CURRENT replacement, when a user opens the project, then the parent opens and abandoned private state is handled idempotently.
  • Given a writer died after the durable commit boundary, when the project opens, then the committed generation opens and the journal is reconciled.
  • Given a different writer is live, when a reader opens valid CURRENT, then the reader succeeds and recovery cleanup is safely deferred.

Implementation Notes

Likely surfaces: GraphForge::open_dir_with_options, open_or_initialize_project, project_recovery.rs, open evidence, and concurrency/recovery matrices.

Observability

Report selected generation class, repaired/aborted journal counts, removed/quarantined entries, deferral reason, and elapsed phase without payload data or sensitive paths.

Security And Privacy

Retain existing path, symlink, hard-link, and bounded-enumeration protections. Never delete caller-controlled or unclassified entries.

Testing

Use the deterministic fault model plus native multi-process barriers/failpoints. Cover initialization, live writers, read-only views, checkpoints, repeat open, and corrupt authority.

Non-Goals

Repairing a corrupt CURRENT by guessing, blocking indefinitely, or remote recovery.

Related Issues

Canonical tracker: #747. Direct prerequisite: #749. Existing recovery foundation: #212.

Metadata

Metadata

Assignees

No one assigned

    Labels

    coreCore source code changesenhancementNew feature or requesttestingTest coverage and testing infrastructure

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions