Repository navigation
Usage page shows $0 and the desktop app keeps reconnecting while getUsageSummary reads Antigravity history #13852
Description
Activity
Triage
Confirmed. This is a real regression, and it is not fixed on current
main.readAntigravityUsagelanded in #10409 (e5a46d6c5d, merged 2026-09-25 11:26Z). That commit is not inv0.0.43-nightly.20260924.2223and is inv0.0.43-nightly.20260925.2269andv0.0.43-nightly.20260926.2282. Nothing has touchedapps/server/src/usage/antigravityUsageReader.tsorUsageService.tssince.Not a duplicate of #7088. That one is the missing
~/.claude/projectsfallback walking~/projects, plus a large Codex transcript tree. This reader did not exist then.What happens
Every
getUsageSummarydecodes every Antigravity.dbon the server thread, including files that have not changed and files older than the window. The transcript cache is never consulted.collectDirsalways calls the reader, and only Claude, Codex, and Grok go throughreadFileRecords:const antigravity = yield* Effect.promise(() => readAntigravityUsage([...antigravityDirs], windowStartMs), );
if ( cached && cached.size === size && cached.mtimeMs === mtimeMs && cached.provider === provider ) { return cached.tailRecords.length === 0 ? cached.records : [...cached.records, ...cached.tailRecords]; }
node:sqlite'sDatabaseSync.iterate()reads each blob synchronously. ThesetImmediateruns only after every 256 rows, so one largegen_metadataorstepsblob blocks the event loop for the whole parse. The trajectory walk does not yield at all:const readMetadata = async (query: string, column: string, step: boolean) => { const entries: Array<{ idx: number; entry: Metadata }> = []; for (const row of db.prepare(query).iterate()) { if (typeof row.idx !== "number") throw new Error("Invalid Antigravity metadata index"); entries.push({ idx: row.idx, entry: metadata(blob(row[column]), step) }); if (entries.length % 256 === 0) await NodeTimersPromises.setImmediate(); } return entries; }; // ... for (const row of db.prepare("SELECT data FROM trajectory_metadata_blob").iterate()) { trajectoryTimestamp ??= timestamp(nested(fields(blob(row.data)), 2)); }
The file
mtimeis only a fallback timestamp. The walk still opens every.db, andsinceMsis applied only after alias merging:} else if (entry.isFile() && entry.name.endsWith(".db")) { try { const canonical = await NodeFSP.realpath(path); if (visited.has(canonical)) continue; visited.add(canonical); const stat = await NodeFSP.stat(path); const candidates = await readDatabase(path, stat.mtimeMs);
for (const [index, group] of groups.entries()) { if (group.parent === index && group.record.timestampMs >= sinceMs) { files[group.fileIndex]!.records.push(group.record); } }
That matches the isolated profile: a warm usage-scan cache and still ~18s, almost all of it in
readMetadata. OpenCode is the same shape (readOpenCodeUsagealso bypasses the file cache) but it was about 1.25s here, not the stall.The desktop then treats the stalled loop as a dead server. The socket open timeout is 15s, and the banner is the hostname plus "reconnecting":
const SOCKET_OPEN_TIMEOUT = "15 seconds";
const socketLayer = Socket.layerWebSocket(connection.socketUrl, { openTimeout: SOCKET_OPEN_TIMEOUT,
title: `${activeEnvironmentUnavailableState.label} is ${environmentReconnecting ? "reconnecting" : "offline"}`,
$0is the failed request, not an empty history. A failedgetUsageSummarysetserrorand leavessummarynull, so the page stops pending and rendersmergeUsage([]), whosecostUsdis 0. The direct request that returned 302 buckets is the scan that was allowed to finish.Identical windows already share one detached scan (
inflightScansinreadSummary). Three overlappingscanSummaryspans still mean three calls missed that map, because the key includes the window, price overrides, and the Cursor keychain flag. Each of those calls decodes the corpus again. Interrupting the RPC does not cancel the detached scan, andEffect.promisewill not abortreadAntigravityUsageonce it is running.Fix direction
Cache Antigravity results per file by size and mtime, the same way
readFileRecordsdoes, and keep the merge-before-filter behavior. Skipping a.dbsolely because its mtime is outside the window can drop an alias that should have merged into a newer record. Moving the SQLite read off the main thread is the right second step. The Claude history worker is not prior art for this: it is a subprocess for the Agent SDK history helpers, not the usage scanner.Moving old files out of
~/.gemini/antigravity-cli/conversationsdoes avoid the walk, at the cost of that history.- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 26, 2026 Would it make sense to define a latency invariant here instead of only optimizing the current scan?
Something like: usage collection must never hold the server event loop for >50 ms continuously, regardless of history size.
Per-file
size/mtimecaching and single-flight dedup should remove most repeated work, but they don't protect the first cold scan of an 11 GB history. Moving the synchronous SQLite/blob decoding behind a worker boundary would.Then a regression test could run a deliberately expensive usage source while repeatedly issuing a tiny RPC or
/.well-known/t3/environmentrequest and assert its latency stays bounded.That would test the actual failure mode here: usage accounting should never be able to make the environment appear disconnected.
Follow-up with first-scan-after-restart data, plus a pass over #13902.
After restart. The server restarted at 08:00:53. The first
getUsageSummary(08:01:12) ranscanSummaryfor 161 s. Over that time the server process read ~11.3 GB from disk, and its main thread was repeatedly inDstate (wchan=filemap_get_pages). That adds substantial I/O stalls to thereadMetadataCPU profile, though it doesn't split read time from decode time./.well-known/t3/environmenttook 3.4–11.4 s, client connect attempts gave up at ~10 s, and subscriptions dropped at 08:01:39. Responsiveness came back only when the scan finished at 08:03:53. The RPC was interrupted at 41 s, but the detached scan ran to completion. A second scan right after still took 26 s.#13902 says it leaves the first scan and worker isolation for follow-up. This data is why that follow-up matters:
- The cache is memory-only, so every restart starts empty. My 211 Antigravity
.dbfiles (~11 GB) and their WALs haven't changed since Sep 10–11, so the PR would cut the 26 s rerun but not the 161 s first scan. - On a miss, the usage metadata is still reread and decoded on the server thread, with no incremental position like
readFileRecordshas. One new row in an active conversation means another full metadata pass over a ~300 MB database. - Entries are stored only after
readDatabaseresolves, with no per-file in-flight sharing. Scans with different keys can duplicate reads that are still in progress. - Interrupting a caller still doesn't stop the scan.
Effect.promisedoes pass anAbortSignal, but the reader ignores it. Any cancellation work has to keep scans that other callers still wait on.
On @kvnloo's invariant, I agree, and a worker is the right boundary. Two additions:
- A worker takes the reader's decoding and synchronous read waits off the server thread, but the state DB also uses synchronous
node:sqliteon that thread, so an 11 GB scan can still compete with it for disk. Persisting per-file candidates would avoid rereading unchanged history after every restart. They serialize cleanly, but the cache has to keep the WAL identity and merge-before-filter behaviour, and today'susage-scan-cache.jsondecoder only accepts Claude, Codex and Grok. That complements the worker; it doesn't replace it for first-ever or changed-file scans. - A barrier-based ordering test would complement the latency test. Run synchronous work through the actual worker boundary, assert that HTTP and RPC complete before it's released, and add an external watchdog. A source that just awaits an unresolved promise wouldn't catch this regression.
- The cache is memory-only, so every restart starts empty. My 211 Antigravity
Thanks for measuring the restart case. All three points are accurate. #13902 only removes repeat decodes within one server run: the cache lives in memory, a changed database is decoded again in full, and concurrent scans with different keys can decode the same file at the same time.
For your history the 161 s first scan is the real problem, and it needs one of two things that PR does not do. Persisting the Antigravity entries next to the transcript scan cache would let unchanged databases survive a restart. Moving the SQLite read off the server thread, as suggested above, would stop a cold scan from stalling the connection at all. I can take the persistence step as a follow-up once #13902 has been reviewed.
What happened
The Usage page in the desktop app shows $0, although it showed my data before. While it loads, the app shows " is reconnecting" (the placeholder stands for the machine's hostname). Both started after updating from 0.0.43-nightly.20260924.2223 to the Sep 25/26 nightlies.
Diagnosis
server.getUsageSummaryblocks the server's event loop for seconds at a time. Most of that time is spent in the Antigravity usage reader added in #10409 (apps/server/src/usage/antigravityUsageReader.ts, unchanged onmainsince).readMetadata, at 10.2 s.parseOpenCodeMessageandreadOpenCodeUsagecome next at about 1.25 s combined.readMetadata/readDatabase.setImmediateevery 256 rows inreadMetadatadoesn't bound the blocking. Antigravitygen_metadata/stepsrows are large blobs, andnode:sqliteiterate()reads them synchronously.readAntigravityUsagehas no per-file cache, unlike the transcript reader's size/mtime cache. Every request reopens and decodes every.dbunder the Antigravity roots, including files whose mtime is older thansinceMs.getUsageSummarycalls don't share one in-flight scan. The desktop backend's trace shows three fullscanSummaryruns starting in the same second.GET /.well-known/t3/environmentrequests take 3.5–11.4 s instead of ~2 ms.ConnectionDriver.connecttimes out at 15 s, and the banner shows "reconnecting".getUsageSummary..dbfiles older than the window.Steps to reproduce
~/.gemini/antigravity-cli/conversations. Here it's 211.dbfiles totalling ~11 GB, the largest ~306 MB, the newest written 2026-09-10.t3 servewith a fresh--base-dir, issue a session witht3 auth session issue, and callserver.getUsageSummary(29-day window) over/wswhile polling/.well-known/t3/environment. Each call takes ~18 s, and the polls stall for over 1 s.Version
0.0.43-nightly.20260926.2282. The first affected build is 0.0.43-nightly.20260925.2269; the last good one is 0.0.43-nightly.20260924.2223.
Environment
Arch Linux x86_64, kernel 7.1.3-zen, desktop app (Electron 44.4.2), NVMe SSD, 32 GB RAM. The history sources are:
Evidence
Related issues
#7088 is the same class of problem, an unbounded usage scan blocking the server, but it comes from a different source (the
~/projectsfallback and the Codex transcript volume). The Antigravity reader didn't exist when it was filed. I found no existing issue about Antigravity usage reads, Usage showing $0 on current nightlies, or reconnects caused by the usage scan.Fix applied or workaround
Nothing has been changed on the machine. Moving old Antigravity conversations out of
~/.gemini/antigravity-cli/conversationsshould avoid it, at the cost of that history.Filed by
Claude Code (claude-opus-5-5), investigating on the user's machine and following
.github/triage/PLAYBOOK.md.