Skip to content

[Bug]: 30d and Available history totals remain bounded by the moving 64 MiB usage tail #1497

Description

@alanxchen85

Client or integration

OpenCodex Dashboard

Area

Dashboard / Usage

Current status

This issue is partially mitigated but not fixed.

PR #1532 landed on dev and made the bounded read visible to users by exposing the observed snapshot window and warning when history is truncated. The product should therefore no longer describe the returned tail as if it were silently known to cover the selected range.

The core aggregation defect remains: GET /api/usage still reads the bounded management snapshot first and applies 7d, 30d, or all filtering afterward. On a high-volume installation, the loaded tail can cover much less than the requested period.

The issue stays open until date-range and all-time aggregates are complete across the retained ledger rather than merely truthful about truncation.

Problem

managementUsageMaxReadBytes is 64 MiB by default. If usage.jsonl is larger than that, the management reader loads only the newest tail.

The range filter is applied after that read. As a result:

  • 30d can omit persisted requests that are less than 30 days old;
  • Available history / all can return essentially the same bounded tail;
  • cumulative totals can decrease as older rows fall out of the moving read window;
  • detail rows may be bounded intentionally, but aggregate totals become incomplete too.

PR #1532 improves disclosure only. It does not change that aggregation boundary.

Original measured example

On OpenCodex 2.11.0 with a ledger larger than 64 MiB:

Source Requests Tokens Time covered
Complete usage.jsonl 175,818 23.74B Aug 4-11
Bounded Dashboard/API snapshot 46,417 6.60B roughly 39 hours

The API reported truncation, and 129,401 requests plus 17.14B tokens inside the nominal 30-day range were not included in the aggregate.

Root cause

The management route performs the operations in this order:

readUsageSnapshotForManagement(effectiveReadLimit)
        ↓
newest bounded byte tail
        ↓
summarizeUsage(..., range, ...)

The desired contract requires the date/range aggregation boundary to precede, or be independent of, the bounded recent-detail read.

Required behavior

  • 7d aggregates every retained valid request in the selected 7-day range.
  • 30d aggregates every retained valid request in the selected 30-day range.
  • All-time totals do not decrease merely because the raw tail advances.
  • Detail rows may remain bounded if necessary.
  • Aggregate coverage is explicit when an index/projection is incomplete or rebuilding.
  • The existing snapshotWindowStart / snapshotWindowEnd disclosure from fix(usage): disclose the window a truncated read actually covers #1532 remains truthful and must not be reinterpreted as a proof of complete range coverage.

Important correctness constraint

Do not derive a rangeFullyCovered assertion only from the oldest loaded row timestamp.

usage.jsonl is appended when a request completes while a row can carry the request start time. Completion order and start-time order can differ, so the oldest timestamp in the loaded tail does not prove what timestamps may exist in the dropped prefix.

Preferred implementation direction

Use a rebuildable indexed or incremental aggregate projection for summaries, while keeping usage.jsonl as canonical evidence.

Possible implementations include:

Raising managementUsageMaxReadBytes is only a temporary mitigation.

Close condition

Close when 7d/30d aggregates cover all retained in-range requests and all-time totals remain monotonic across raw-tail rotation/truncation and restart.

Related

Activity

  1. self-assigned this
    on Aug 11, 2026
  2. assigned and unassigned on Aug 11, 2026
  3. Ingwannu commented on Aug 12, 2026

    @Ingwannu
    Owner

    Confirmed on current dev: this is a real aggregation-boundary defect, not only a misleading label. GET /api/usage reads managementUsageMaxReadBytes through readUsageSnapshotForManagement() first and applies the 7d/30d/all filter afterward, so a busy installation can return the same moving tail for 30d and all while explicitly reporting historyTruncated: true.

    The durable direction already exists in draft PR #1008 (daily rollup sidecar plus raw-tail merge), so I will not create a competing implementation. That PR is currently stale/conflicting and still has unresolved correctness findings around crash-safe append/commit validation, truncated-ledger invalidation, partial-day range overlap, request deduplication, and disabled-rollup behavior. Those are material because an incorrect derived aggregate would be worse than an explicitly truncated view.

    Keeping this issue open as the user-visible bug contract. The acceptance criteria should remain:

    • 7d/30d totals cover every retained valid row in the selected timestamp range;
    • all-time totals are monotonic across raw-tail rotation/truncation and restart;
    • detail rows may remain bounded, but the response must expose exact aggregate coverage/cutline state;
    • the legacy path must not label a truncated tail as complete history;
    • the derived projection remains rebuildable from usage.jsonl and fails closed on stale/corrupt sidecars.

    Until #1008 is rebased and those boundaries are fixed with exact-head CI, increasing the byte limit is only a temporary mitigation, not a resolution.

  4. lidge-jun commented on Aug 12, 2026

    @lidge-jun
    Owner

    Partial mitigation landed on dev as #1532, merged as 9e777fbaa57331da50e1375217ce8fbf6a53d50d. This does not close the issue — it stops the product from presenting a truncated read as the range you selected, but aggregation is still bounded by the byte limit.

    What changed: GET /api/usage now reports snapshotWindowStart / snapshotWindowEnd, the timestamp bounds of the rows the reader actually loaded, computed before the range and surface filters. The dashboard names that window and the notice moved from an informational tone to a warning, since a total that omits in-range rows is a caveat rather than a status line.

    One thing worth recording, because it constrains any future fix here. An earlier draft of this change carried a rangeFullyCovered boolean derived from the oldest retained entry. That is unsound: usage.jsonl is appended when a request completes while each row carries the request start time, so a long-running request can be appended after shorter ones that started later. The oldest loaded timestamp therefore does not bound what the dropped prefix contains, and any flag asserting "this range is complete" could be wrong. For the same reason the new fields are documented as observed extrema of loaded rows, not as a covered interval — an initial wording that implied coverage was corrected before merge.

    Still open, and what the issue actually asks for: 30d aggregating every persisted in-range request, and all-time totals that do not decrease. That needs a durable index or the daily rollup sidecar in #1008, which remains the right home for it. Raising managementUsageMaxReadBytes only postpones recurrence and makes each summary parse a growing file.

    Verification at the merged head: tsc exit 0; usage suites 75 pass / 0 fail; GUI 772 pass / 0 fail with lint, i18n lint, and build clean; 23 CI checks green including all Linux shards and macOS. All eight locales carry the new string.

  5. changed the title [-][Bug]: 30d and Available history usage stats silently use the same moving 64 MiB tail[/-] [+][Bug]: 30d and Available history totals remain bounded by the moving 64 MiB usage tail[/+] on Aug 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions