Skip to content

Usage page shows $0 and the desktop app keeps reconnecting while getUsageSummary reads Antigravity history #13852

Description

@rabesss

What happened

The Usage page in the desktop app shows $0, although it showed my data before. While it loads, the app shows " is reconnecting" (the placeholder stands for the machine's hostname). Both started after updating from 0.0.43-nightly.20260924.2223 to the Sep 25/26 nightlies.

Diagnosis

server.getUsageSummary blocks the server's event loop for seconds at a time. Most of that time is spent in the Antigravity usage reader added in #10409 (apps/server/src/usage/antigravityUsageReader.ts, unchanged on main since).

  • A CPU profile of one isolated server (same build, fresh base dir, warm usage scan cache, two sequential requests of 17.8 s and 18.0 s) shows:
    • The top self-time is readMetadata, at 10.2 s. parseOpenCodeMessage and readOpenCodeUsage come next at about 1.25 s combined.
    • The three longest uninterrupted main-thread stretches are 3.0 s, 2.2 s and 2.1 s, all in readMetadata/readDatabase.
  • The setImmediate every 256 rows in readMetadata doesn't bound the blocking. Antigravity gen_metadata/steps rows are large blobs, and node:sqlite iterate() reads them synchronously.
  • readAntigravityUsage has no per-file cache, unlike the transcript reader's size/mtime cache. Every request reopens and decodes every .db under the Antigravity roots, including files whose mtime is older than sinceMs.
  • Concurrent getUsageSummary calls don't share one in-flight scan. The desktop backend's trace shows three full scanSummary runs starting in the same second.
  • Put together, this becomes a reconnect loop:
    1. A scan stalls the loop, and trivial GET /.well-known/t3/environment requests take 3.5–11.4 s instead of ~2 ms.
    2. The desktop's ConnectionDriver.connect times out at 15 s, and the banner shows "reconnecting".
    3. The reconnect issues another getUsageSummary.
    4. The in-flight requests are interrupted, and the page renders $0. The scan cache holds ~60k records, and a direct request returns 302 buckets.
  • Possible directions:
    • Cache Antigravity and OpenCode results per file by size/mtime.
    • Skip Antigravity .db files older than the window.
    • Read SQLite off the main thread; the Claude history reader already uses a worker.
    • Share a single in-flight scan across concurrent calls.

Steps to reproduce

  1. On Linux, have a large Antigravity CLI history in ~/.gemini/antigravity-cli/conversations. Here it's 211 .db files totalling ~11 GB, the largest ~306 MB, the newest written 2026-09-10.
  2. Run a nightly containing feat(usage): read cursor, opencode, and antigravity history #10409 (first seen in 0.0.43-nightly.20260925.2269).
  3. Open the Usage page in the desktop app. It stays loading, the host shows "reconnecting", and the page ends at $0.
  4. Or, with no UI involved: start t3 serve with a fresh --base-dir, issue a session with t3 auth session issue, and call server.getUsageSummary (29-day window) over /ws while polling /.well-known/t3/environment. Each call takes ~18 s, and the polls stall for over 1 s.

Version

0.0.43-nightly.20260926.2282. The first affected build is 0.0.43-nightly.20260925.2269; the last good one is 0.0.43-nightly.20260924.2223.

Environment

Arch Linux x86_64, kernel 7.1.3-zen, desktop app (Electron 44.4.2), NVMe SSD, 32 GB RAM. The history sources are:

  • Antigravity CLI: 211 dbs, ~11 GB
  • OpenCode: 161 MB db plus 19,851 message JSON files
  • Codex: 21 GB, 6,230 files
  • Grok: ~20.9k files
  • Claude: 355 MB

Evidence

# desktop backend server.trace.ndjson, UTC
15:24:36Z UsageService.scanSummary 189817ms   ws.rpc.server.getUsageSummary 189839ms
15:30:29Z scanSummary 37333 / 39815 / 45903ms (3 concurrent), 3x getUsageSummary Interrupted at 18.9s
15:34:37Z scanSummary 68543 / 70463 / 77668ms (3 concurrent), 1x Interrupted
15:25:32Z-15:27:15Z ConnectionDriver.connect 15000ms status.interrupted=true (x4)
http.server GET /.well-known/t3/environment 3513-11406ms during scans (normally 1-4ms)
RpcClient.server.getUsageSummary: RpcClientError: SocketOpenError: timeout waiting for "open"

# isolated CPU profile, self time (2 requests)
10230 ms  readMetadata          (antigravityUsageReader)
  808 ms  parseOpenCodeMessage
  445 ms  readOpenCodeUsage
longest blocked stretches: 3039ms readMetadata, 2151ms readMetadata, 2140ms readMetadata

Related issues

#7088 is the same class of problem, an unbounded usage scan blocking the server, but it comes from a different source (the ~/projects fallback and the Codex transcript volume). The Antigravity reader didn't exist when it was filed. I found no existing issue about Antigravity usage reads, Usage showing $0 on current nightlies, or reconnects caused by the usage scan.

Fix applied or workaround

Nothing has been changed on the machine. Moving old Antigravity conversations out of ~/.gemini/antigravity-cli/conversations should avoid it, at the cost of that history.

Filed by

Claude Code (claude-opus-5-5), investigating on the user's machine and following .github/triage/PLAYBOOK.md.

Activity

  1. juliusmarminge commented on Sep 26, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed. This is a real regression, and it is not fixed on current main. readAntigravityUsage landed in #10409 (e5a46d6c5d, merged 2026-09-25 11:26Z). That commit is not in v0.0.43-nightly.20260924.2223 and is in v0.0.43-nightly.20260925.2269 and v0.0.43-nightly.20260926.2282. Nothing has touched apps/server/src/usage/antigravityUsageReader.ts or UsageService.ts since.

    Not a duplicate of #7088. That one is the missing ~/.claude/projects fallback walking ~/projects, plus a large Codex transcript tree. This reader did not exist then.

    What happens

    Every getUsageSummary decodes every Antigravity .db on the server thread, including files that have not changed and files older than the window. The transcript cache is never consulted. collectDirs always calls the reader, and only Claude, Codex, and Grok go through readFileRecords:

        const antigravity = yield* Effect.promise(() =>
          readAntigravityUsage([...antigravityDirs], windowStartMs),
        );
          if (
            cached &&
            cached.size === size &&
            cached.mtimeMs === mtimeMs &&
            cached.provider === provider
          ) {
            return cached.tailRecords.length === 0
              ? cached.records
              : [...cached.records, ...cached.tailRecords];
          }

    node:sqlite's DatabaseSync.iterate() reads each blob synchronously. The setImmediate runs only after every 256 rows, so one large gen_metadata or steps blob blocks the event loop for the whole parse. The trajectory walk does not yield at all:

        const readMetadata = async (query: string, column: string, step: boolean) => {
          const entries: Array<{ idx: number; entry: Metadata }> = [];
          for (const row of db.prepare(query).iterate()) {
            if (typeof row.idx !== "number") throw new Error("Invalid Antigravity metadata index");
            entries.push({ idx: row.idx, entry: metadata(blob(row[column]), step) });
            if (entries.length % 256 === 0) await NodeTimersPromises.setImmediate();
          }
          return entries;
        };
        // ...
          for (const row of db.prepare("SELECT data FROM trajectory_metadata_blob").iterate()) {
            trajectoryTimestamp ??= timestamp(nested(fields(blob(row.data)), 2));
          }

    The file mtime is only a fallback timestamp. The walk still opens every .db, and sinceMs is applied only after alias merging:

          } else if (entry.isFile() && entry.name.endsWith(".db")) {
            try {
              const canonical = await NodeFSP.realpath(path);
              if (visited.has(canonical)) continue;
              visited.add(canonical);
              const stat = await NodeFSP.stat(path);
              const candidates = await readDatabase(path, stat.mtimeMs);
      for (const [index, group] of groups.entries()) {
        if (group.parent === index && group.record.timestampMs >= sinceMs) {
          files[group.fileIndex]!.records.push(group.record);
        }
      }

    That matches the isolated profile: a warm usage-scan cache and still ~18s, almost all of it in readMetadata. OpenCode is the same shape (readOpenCodeUsage also bypasses the file cache) but it was about 1.25s here, not the stall.

    The desktop then treats the stalled loop as a dead server. The socket open timeout is 15s, and the banner is the hostname plus "reconnecting":

    const SOCKET_OPEN_TIMEOUT = "15 seconds";
        const socketLayer = Socket.layerWebSocket(connection.socketUrl, {
          openTimeout: SOCKET_OPEN_TIMEOUT,
              title: `${activeEnvironmentUnavailableState.label} is ${environmentReconnecting ? "reconnecting" : "offline"}`,

    $0 is the failed request, not an empty history. A failed getUsageSummary sets error and leaves summary null, so the page stops pending and renders mergeUsage([]), whose costUsd is 0. The direct request that returned 302 buckets is the scan that was allowed to finish.

    Identical windows already share one detached scan (inflightScans in readSummary). Three overlapping scanSummary spans still mean three calls missed that map, because the key includes the window, price overrides, and the Cursor keychain flag. Each of those calls decodes the corpus again. Interrupting the RPC does not cancel the detached scan, and Effect.promise will not abort readAntigravityUsage once it is running.

    Fix direction

    Cache Antigravity results per file by size and mtime, the same way readFileRecords does, and keep the merge-before-filter behavior. Skipping a .db solely because its mtime is outside the window can drop an alias that should have merged into a newer record. Moving the SQLite read off the main thread is the right second step. The Claude history worker is not prior art for this: it is a subprocess for the Agent SDK history helpers, not the usage scanner.

    Moving old files out of ~/.gemini/antigravity-cli/conversations does avoid the walk, at the cost of that history.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 26, 2026
  3. kvnloo commented on Sep 27, 2026

    @kvnloo
    Contributor

    Would it make sense to define a latency invariant here instead of only optimizing the current scan?

    Something like: usage collection must never hold the server event loop for >50 ms continuously, regardless of history size.

    Per-file size/mtime caching and single-flight dedup should remove most repeated work, but they don't protect the first cold scan of an 11 GB history. Moving the synchronous SQLite/blob decoding behind a worker boundary would.

    Then a regression test could run a deliberately expensive usage source while repeatedly issuing a tiny RPC or /.well-known/t3/environment request and assert its latency stays bounded.

    That would test the actual failure mode here: usage accounting should never be able to make the environment appear disconnected.

  4. rabesss commented on Sep 28, 2026

    @rabesss
    Author

    Follow-up with first-scan-after-restart data, plus a pass over #13902.

    After restart. The server restarted at 08:00:53. The first getUsageSummary (08:01:12) ran scanSummary for 161 s. Over that time the server process read ~11.3 GB from disk, and its main thread was repeatedly in D state (wchan=filemap_get_pages). That adds substantial I/O stalls to the readMetadata CPU profile, though it doesn't split read time from decode time. /.well-known/t3/environment took 3.4–11.4 s, client connect attempts gave up at ~10 s, and subscriptions dropped at 08:01:39. Responsiveness came back only when the scan finished at 08:03:53. The RPC was interrupted at 41 s, but the detached scan ran to completion. A second scan right after still took 26 s.

    #13902 says it leaves the first scan and worker isolation for follow-up. This data is why that follow-up matters:

    1. The cache is memory-only, so every restart starts empty. My 211 Antigravity .db files (~11 GB) and their WALs haven't changed since Sep 10–11, so the PR would cut the 26 s rerun but not the 161 s first scan.
    2. On a miss, the usage metadata is still reread and decoded on the server thread, with no incremental position like readFileRecords has. One new row in an active conversation means another full metadata pass over a ~300 MB database.
    3. Entries are stored only after readDatabase resolves, with no per-file in-flight sharing. Scans with different keys can duplicate reads that are still in progress.
    4. Interrupting a caller still doesn't stop the scan. Effect.promise does pass an AbortSignal, but the reader ignores it. Any cancellation work has to keep scans that other callers still wait on.

    On @kvnloo's invariant, I agree, and a worker is the right boundary. Two additions:

    • A worker takes the reader's decoding and synchronous read waits off the server thread, but the state DB also uses synchronous node:sqlite on that thread, so an 11 GB scan can still compete with it for disk. Persisting per-file candidates would avoid rereading unchanged history after every restart. They serialize cleanly, but the cache has to keep the WAL identity and merge-before-filter behaviour, and today's usage-scan-cache.json decoder only accepts Claude, Codex and Grok. That complements the worker; it doesn't replace it for first-ever or changed-file scans.
    • A barrier-based ordering test would complement the latency test. Run synchronous work through the actual worker boundary, assert that HTTP and RPC complete before it's released, and add an external watchdog. A source that just awaits an unresolved promise wouldn't catch this regression.
  5. derektrimm commented on Sep 29, 2026

    @derektrimm
    Contributor

    Thanks for measuring the restart case. All three points are accurate. #13902 only removes repeat decodes within one server run: the cache lives in memory, a changed database is decoded again in full, and concurrent scans with different keys can decode the same file at the same time.

    For your history the 161 s first scan is the real problem, and it needs one of two things that PR does not do. Persisting the Antigravity entries next to the transcript scan cache would let unchanged databases survive a restart. Moving the SQLite read off the server thread, as suggested above, would stop a cold scan from stalling the connection at all. I can take the persistence step as a follow-up once #13902 has been reviewed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions