Skip to content

Cap adaptive-thinking effort per model's supported levels - #52

Closed
neomnezia wants to merge 2 commits into
ClickHouse:mainfrom
neomnezia:fix/model-aware-effort-capping
Closed

neomnezia wants to merge 2 commits into
ClickHouse:mainfrom
neomnezia:fix/model-aware-effort-capping

Conversation

@neomnezia

@neomnezia neomnezia commented Apr 24, 2026 •

Copy link
Copy Markdown
Contributor

Problem

Global agent.effort (default "max") shared across main agent + cron_model / memory.recall_model / memory.memorize_model. Auxiliary models default to claude-sonnet-4-6 . Sonnet no advertise max.

Forward max to Sonnet → 400. CLIProxyAPI error:

thinking: validation failed | provider=claude model=claude-sonnet-4-6
  error=level "max" not supported, valid levels: low, medium, high

(validator: internal/thinking/validate.go:125; source: internal/registry/models/models.json)

Every cron source job + every memU recall/memorize call on proxy.enabled: true → 400. Stall inbox-processor + background work.

Same shape on Opus 4.6 + xhigh (4.6 supports max not xhigh). Any config where cron_model / memU model tier below main model.

Behavior

effort model result
max claude-opus-4-7 max
max claude-opus-4-6 max
max claude-sonnet-4-6 high ← fix
xhigh claude-opus-4-7 xhigh
xhigh claude-opus-4-6 high
xhigh claude-sonnet-4-6 high
low/medium/high any unchanged
any unknown / None unchanged (backward compatible)
invalid any None (pre-PR behavior)

Why table not proxy query

  1. GET /v1/models at startup → couple to CLIProxyAPI, break direct-API users, need cache/refresh.
  2. Retry-on-400 downgrade → complex, latency cost, error-path only.

Static table: simple, testable, explicit. Bump per new model. Few level-based Claude models today; bump = 1-line PR.

Entries track Anthropic capabilities + CLIProxyAPI internal/registry/models/models.json. Unknown models (Haiku, legacy Sonnet/Opus on budget-tokens) pass through. Fully backwa
rd compatible.

Verification

Unit tests: 17 cases — opus 4.7, opus 4.6, sonnet 4.6, haiku, dated alias, unknown, None, empty, invalid. All pass.

Live proxy mode (proxy.enabled: true), before:

  • cron:inbox-processor → 400 level "max" not supported every tick
  • Telegram Bot admin reply → same 400 on memU recall/memorize

After:

  • Opus 4.7 main agent: effort: max (verified via subprocess env / request inspection).
  • Sonnet 4.6 auxiliary: effort: high. Proxy 200. Requests succeed.

Global `agent.effort` (e.g. "max") is shared across the main agent and
auxiliary models (`cron_model`, memU recall/memorize). Sonnet/Haiku tiers
only advertise {low, medium, high}; forwarding "max" triggers a 400 from
CLIProxyAPI tier validation (`internal/thinking/validate.go:125`):

    thinking: validation failed | provider=claude model=claude-sonnet-4-6
      error=level "max" not supported, valid levels: low, medium, high

Same shape on Opus 4.6 with "xhigh" (supports max but not xhigh), and on
any auxiliary model the user configures below their main Opus 4.7.

Add `_effective_effort(value, model)` — a small table of known Claude
models with their advertised effort levels, and a step-down lookup that
caps the requested effort to the highest level the target model supports.
Unknown models pass through unchanged (backward compatible).

Preserves "max" for Opus 4.7 main agent while letting Sonnet cron/memU
calls succeed with the tier-appropriate level.

Symmetric with existing `_parse_thinking_config(value, model)` which is
already model-aware.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

- `_MODEL_EFFORT_LEVELS` keys changed to substring form (`opus-4-7`,
  `opus-4-6`, `sonnet-4-6`) iterated against the full model name so
  dated Anthropic aliases like `claude-opus-4-7-20260416` resolve.
  Mirrors the MODEL_PRICING pattern in nerve/db/usage.py and
  `_model_supports_legacy_enabled_thinking` in the same file.
- `_effective_effort(value, model=None)` — default matches the sibling
  `_parse_thinking_config(value, model=None)`.
- `logger.debug` when capping happens so users can trace why a
  configured `effort: max` was forwarded as `high` to Sonnet.
- Docstring trimmed to one line to match the sibling method style.
- New `tests/test_engine.py` covering the cap table (incl. dated
  aliases, unknown models, None/empty model, invalid effort string).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@neomnezia
neomnezia marked this pull request as ready for review April 24, 2026 19:08
@neomnezia neomnezia closed this Apr 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants