Skip to content

fix: sanitize inflated relay usage reports + suppress empty-range growth nudges (#476) - #477

Open
ranxianglei wants to merge 1 commit into
masterfrom
2026-10-06_usage-inflation-nudge-fix
Open

ranxianglei wants to merge 1 commit into
masterfrom
2026-10-06_usage-inflation-nudge-fix

Conversation

@ranxianglei

Copy link
Copy Markdown
Owner

Closes #476

Problem

A third-party OpenAI-compatible gateway reported prompt usage at 1.76–2.05× reality (measured per-request payload comparison in #476). ACP inherited that number verbatim as prePruneTokens / currentTokens, which:

  1. fabricated growthSinceBaseline → growth nudges injected with zero actionable ranges (candidateCount=0) in sessions whose visible context was entirely compressed-block summaries;
  2. sent the model to guess IDs → endId mNNNNN is not available — likely consumed by an existing block. → retry loop (single step 15m02s in the incident).

Both defects verified in code before implementing (see triage comments on #476).

Changes

  • AC1+AC2 — new lib/usage-sanity.ts: resolveCurrentUsage(state, messages, logger?) cross-turn inflation detector. Fires only when the reported usage jumps >1.5× while the local chars/4 payload estimate jumps ≤1.5× (both estimates ≥ 5000-token floor); on detection it returns the re-estimated size and logs WARN Relay usage inflation detected — re-estimating context size. Honest reports pass through unchanged; baseline pair {reported, estimated} persisted as new optional SessionState.usageSanity (old state files load fine; detection skips one turn until seeded).
  • All three call sites use the same sanitized channel: isContextOverLimits (lib/messages/inject/utils.ts), prePruneTokens + postTokens (lib/hooks.ts).
  • AC3 — lib/messages/inject/inject.ts: new noRangesAtAll flag (zero compressible AND zero protected messages) OR-ed into nothingToCompress; new audit reason no_compressible_ranges. This closes the one remaining hole: all other zero-ranges combinations were already covered by allProtected / allInProtectedZone / allBelowMin.
  • AC4 — lib/compress/search.ts: buildBoundaryRecoveryHint now enumerates the actual contiguous free spans instead of a first–last window.

Behavior changes (old → new)

What Old New Why
prePruneTokens / postTokens / currentTokens raw reported usage, inherited verbatim identical unless relay inflates (>1.5× report jump with ≤1.5× payload jump, both ≥ 5000 est.) inflated relays must not reach growth heuristics
nudge suppression audit reasons all_protected, in_protected_zone, below_effective_floor, no_executable_candidates + no_compressible_ranges turns that previously injected phantom nudges now suppress; emergency path falls to the cadence-gated /compact notice
compress boundary-error hint Current visible: first–last (N msgs). … Call acp_status() to see which blocks consumed which IDs, then retry with valid IDs. Free ranges: <contiguous spans, cap 8 (+N more)>. K active compressed blocks. Retry with IDs inside the free ranges. first–last implied continuity across consumed gaps
persisted state no such field optional usageSanity {reported, estimated} appended cross-turn baseline survives restart

Known accepted limitation (documented in module docstring): a one-shot inflation that then holds steady passes the jump gate. An absolute reported-vs-estimated check was deliberately NOT added because chars/4 undercounts CJK up to ~4× and would false-positive on honest CJK-heavy sessions.

Verification

  • New tests (19 total): tests/usage-sanity.test.ts (10), tests/nudge-no-ranges-suppression.test.ts (3, multi-turn with side-effect assertions per §5.7), tests/boundary-recovery-hint.test.ts (6) — all pass.
  • Fail-on-bug matrix (§5.7): each fix temporarily reverted → corresponding test confirmed FAILING → restored (AC3 revert → phantom nudge injected; AC1/AC2 bypass → baseline not corrected from 100K to 7K; AC4 revert → old hint text detected).
  • npm run typecheck clean. Global suite left to CI.

@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown
Contributor

📦 Built Plugin Artifact

Branch: 2026-10-06_usage-inflation-nudge-fix (89b0347)

Option A — Install from npm PR tag (recommended)

opencode plugin opencode-acp@pr-477 --global

Each push to this PR publishes a new version under the pr-477 npm tag.

Option B — Install from GitHub

opencode plugin "github:ranxianglei/opencode-acp#2026-10-06_usage-inflation-nudge-fix" --global

Option C — Download artifact

  1. Download the artifact from the Actions run
  2. Extract the tarball and install:
tar xzf opencode-acp-pr477.tgz
cp -r package/dist ~/.cache/opencode/packages/opencode-acp@latest/node_modules/opencode-acp/dist
  1. Restart opencode to pick up changes.

This comment is automatically updated on each push.

@ranxianglei

Copy link
Copy Markdown
Owner Author

🤖 Powered by ework · qwen3.8-27b

[bot] 🏷

结论先行:不建议合并。 CI 红(Node 24 job:13 个既有测试失败,Node 22 随之取消)。我没有只信 CI——已本地复现全部失败并逐点隔离出根因:13 个回归全部由本 PR 自己的三处改动引入,不是环境抖动。PR 自己新增的 19 个测试是过的;问题在于它改变了既有行为却没有对齐既有回归测试。

方向上先给一个提醒(对应原 bug #476,外部用户实测):#476 测到的是稳态 1.76–2.05× 上游虚高。而本 PR 标题主打的 AC1/AC2「sanitize inflated usage」对这种稳态场景基本无效(详见第 6 问);真正堵住事故症状的是 AC3(零区间抑制)+ AC4(错误提示)。下面按七问展开。


1. 是什么问题
修 #476:第三方 OpenAI 兼容网关把 prompt usage 报成真实的 1.76–2.05×,ACP 原样继承为 prePruneTokens/currentTokens → 造出虚假 growthSinceBaseline → 在「可见上下文全是压缩块摘要、没有任何可压区间」的会话里注入 growth nudge → 模型乱猜已被消费的 ID → endId mNNNNN is not available 重试死循环(事故单步耗时 15m02s)。

2. 什么原因导致的
两个独立缺陷叠加:(a) 上报 token 数无校验直接进增长启发式;(b) 「零可压 + 零受保护」这个组合原本没有任何抑制分支(allProtected / allInProtectedZone / allBelowMin 三者皆假时的真空洞)。

3. 怎么修复的

  • AC1/AC2 新增 lib/usage-sanity.ts:跨轮比值检测(report 跳升 >1.5× 且 payload(chars/4) 跳升 ≤1.5×、两者 ≥5000)时改用重估值;isContextOverLimits、prePruneTokens、postTokens 三个调用点统一走该通道。
  • AC3 lib/messages/inject/inject.ts:488 新增 noRangesAtAll(compressible 与 protected 皆空)并入 nothingToCompress,新增审计原因 no_compressible_ranges。
  • AC4 lib/compress/search.ts 边界报错改为枚举真实连续空闲 span(上限 8,超出显示 +N more),去掉误导性的 “Call acp_status()” 指引。

4. 是否有回归问题 —— 有,且阻断性(13 个既有测试)
本地复现 + 逐点隔离确认,三类根因:

5. 是否破坏缓存
不涉及 prompt 前缀缓存失效路径;改动都在 nudge 决策与 token 计量层。唯一口径变化:prePruneTokens/postTokens/currentTokens 从“原始上报”切到“可能重估的值”(见行为变更披露),不触碰缓存键。

6. 是否改了设计方向 / 方案是否打到核心
需要提醒:AC1/AC2 对 #476 描述的稳态虚高几乎不起作用——恒定倍率下 reportRatio ≈ payloadRatio,检测条件(前者 >1.5 且后者 ≤1.5)不可能同时成立,所以只在“诚实→虚高”的那一轮过渡期生效一次即自停;而 #476 测到的正是稳态。也就是说标题主打的机制恰好覆盖不到主场景,事故的真正止血靠 AC3+AC4。同时 AC1/AC2 又带来上面的误报回归,当前性价比为负。建议三选一,请你拍板:① 砍掉 AC1/AC2,只留 AC3+AC4(最简,已解决严重症状);② 保守化为“持续多轮判据 + 默认关的配置项”,并改用 BPE 而非 chars/4 做对照;③ 保留但把阈值大幅上调并补充分层证据。我倾向 ①或②。

7. 是否引入新配置
新增可选持久化字段 SessionState.usageSanity{reported, estimated}(向后兼容:旧状态文件可读,损坏记录会被丢弃)。若保留 AC1/AC2,RELAY_INFLATION_RATIO=1.5 / PAYLOAD_FLOOR_TOKENS=5000 目前硬编码,建议提为配置项。


收工判定:不可合并。 待作者处理三件事:

  1. 更新 tests/compress-search.test.ts:293 到新文案(或删掉,新语义已由 boundary-recovery-hint.test.ts 覆盖)。
  2. 按 §5.7 把 10 个受影响的 nudge 回归夹具补成带 ref、内容达标的真实多轮场景,并与新增 nudge-no-ranges-suppression 测试对齐语义(那个用纯摘要 = 真无可压 → 抑制正确;其余夹具应有真实可压内容 → 应触发)。
  3. 就 AC1/AC2 的取舍给方向(见第 6 问)。

第 1、2 点方向明确,你确认后我就接手改并走双 agent 评审再交你合并;第 3 点是设计取舍,等你拍板。以上均基于本地复现与逐点隔离验证,未跑全量套件(交由 CI)。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant