Skip to content

self-host: keep evaluated tool listings across restarts - #2235

Open
0oAstro wants to merge 1 commit into
UsefulSoftwareCo:v2from
0oAstro:self-host-durable-tool-listings
Open

0oAstro wants to merge 1 commit into
UsefulSoftwareCo:v2from
0oAstro:self-host-durable-tool-listings

Conversation

@0oAstro

@0oAstro 0oAstro commented Oct 9, 2026 •

Copy link
Copy Markdown

createExecutor accepts cache.durable, and Cloud uses it to keep evaluated tool listings in each app's data supervisor. Self-host only passes cache.memory. When a self-host server restarts, it loses every listing it has evaluated, and the first search loads every app Worker to evaluate them again. With about 20 apps, that first search took 12 s and about 10 s of CPU, and peak memory roughly doubled. Freed memory went back to workerd's allocator, so RSS stayed high for the life of the process. Over a few days of restarts and updates, this looked like a memory leak.

This PR gives self-host a DurableDeclarations store in its product database:

  • Table. Migration step 7_evaluated_declarations adds hosted_evaluated (app, key, at, until, json). The step only adds a table. The running server never reads it.
  • get. Returns the row while until is in the future. A failed read counts as a miss.
  • set. Upserts the row and keeps the newer at. Results over 4 MiB are not stored, matching Cloud's size limit. A failed write logs a warning.
  • Invalidation. forgetting wraps the memory cache. When changed(app, at) runs, it also deletes that app's durable rows with at <= at. Without this, a removed or changed app would come back with a stale listing after a restart. The DurableDeclarations contract requires this behavior.

Results can include text derived from credentials. They stay in the product database, which already holds the operator's encrypted account state. This follows the rule in the DurableDeclarations doc comment.

Measurements

I measured on executor@2.0.0-beta.8 with this change backported. The host is Docker on an ARM64 machine (8 cores) running a copy of a live install with about 20 apps. Values are container RSS and the duration of the first search call after a restart.

Both runs use the same image, so the store's contents are the only difference between them. That image also ran PGlite with shared_buffers=32MB, which lowers both memory figures by about the same amount. Run 1 starts with an empty hosted_evaluated table, which matches today's behavior after any restart. Run 2 restarts the same container after run 1 filled the table.

After a restart Empty store (today) Populated store
First search 9,427 ms 508 ms
Memory after first search 930 MB 551 MB
  • The first search after a deploy fills the store once (12.4 s on my install, which had more apps connected). Later restarts read from it.
  • On the test copy, a removed app did not reappear after a restart.
  • Real calls to several apps (Moodle, Grafana, Swiggy, GitHub and others) returned data after a restart.

Verification on v2

  • I ported the change from beta.8 to v2. The changes were cache.durable, the effect/sql import, and a migration step instead of a runtime create table.
  • tsc --noEmit -p . reports no errors. I could not run bun run typecheck because tsc-rs has no linux-arm64 build.
  • oxlint passes.
  • The self-host image builds. A fresh container applies the migration, passes its healthcheck, and restarts cleanly with the migration already recorded.

I did not run the e2e suites. I did not repeat the restart benchmark on v2.

Open questions

  • hostedProductMigrations also runs on Cloud, so Cloud gets an empty hosted_evaluated table it never uses. If you prefer, I can move the step into a migration only self-host runs.
  • forgetting deletes rows in a fire-and-forget fiber, because DeclarationCache.changed is synchronous. If the server stops before that delete runs, a stale row can survive until its until. The same window exists in memory today. The row's at guard keeps it from overwriting newer results.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant