Skip to content

fix(opencode): cap session retries with jitter - #41939

Merged
rekram1-node merged 1 commit into
devfrom
retry-adjustments
Aug 12, 2026
Merged

fix(opencode): cap session retries with jitter#41939
rekram1-node merged 1 commit into
devfrom
retry-adjustments

Conversation

@rekram1-node

Copy link
Copy Markdown
Collaborator

Issue for this PR

Closes #37076

Type of change

  • Bug fix
  • New feature
  • Refactor / code improvement
  • Documentation

What does this PR do?

Caps the high-level session retry policy at five retries so persistent retryable provider failures terminate instead of running indefinitely. Adds up to 25% jitter to locally computed exponential delays while continuing to honor explicit provider retry delays unchanged.

How did you verify your code works?

  • bun test test/session/retry.test.ts from packages/opencode: 54 passed
  • bun typecheck from packages/opencode: passed
  • Added coverage for jitter bounds and retry schedule exhaustion

Screenshots / recordings

N/A

Checklist

  • I have tested my changes locally
  • I have not included unrelated changes in this PR

@rekram1-node
rekram1-node merged commit c789868 into dev Aug 12, 2026
11 checks passed
@rekram1-node
rekram1-node deleted the retry-adjustments branch August 12, 2026 03:50
eyalatox added a commit to eyalatox/opencode that referenced this pull request Aug 12, 2026
Builds on the fixed RETRY_MAX_RETRIES cap from anomalyco#41939: providers differ
in how aggressively they should be retried, so expose 'retry' (max
attempts, 0 disables) and 'backoffDelay' (initial backoff ms) provider
options in opencode.json. retry-after headers from the provider still
take precedence over the configured backoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
HQ123-BOOP pushed a commit to HQ123-BOOP/LibreCode that referenced this pull request Aug 12, 2026
sepo-eng added a commit to Draugur-AI/opencode that referenced this pull request Aug 13, 2026
…alyco/dev (#56)

* fix(ui): correct OC-2 weak icon color (anomalyco#41504)

(cherry picked from commit 3a90639)

* fix(provider): scope DeepSeek V4 Flash sampling defaults (anomalyco#41620)

Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
Co-authored-by: 侯秀凯 <9627034+hoooou@users.noreply.github.com>
(cherry picked from commit 5d95348)

* fix(opencode): detect Copilot PDF input support (anomalyco#41522)

(cherry picked from commit 561afb4)

* fix(opencode): cap session retries with jitter (anomalyco#41939)

(cherry picked from commit c789868)

* fix(provider): add Merge Gateway reasoning variants (anomalyco#41867)

(cherry picked from commit 8571a92)

* docs: fix broken DigitalOcean and Daytona links (anomalyco#42048)

Co-authored-by: skyzhao1223 <skyzhao1223@users.noreply.github.com>
(cherry picked from commit ca3df21)

* docs: fix provider display name and PAT typos (anomalyco#42034)

Co-authored-by: skyzhao1223 <skyzhao1223@users.noreply.github.com>
(cherry picked from commit 959c8bd)

* fix(xai): pass through reasoning effort (anomalyco#42160)

Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
(cherry picked from commit 502310f)

* fix(mistral): pass through reasoning effort (anomalyco#42164)

Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
(cherry picked from commit beeabe2)

* fix(groq): pass through reasoning effort (anomalyco#42166)

Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
(cherry picked from commit 6fea419)

* fix(opencode): select Kimi prompt by provider (anomalyco#42161)

Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
(cherry picked from commit 91df883)

---------

Co-authored-by: OpeOginni <107570612+OpeOginni@users.noreply.github.com>
Co-authored-by: opencode-agent[bot] <219766164+opencode-agent[bot]@users.noreply.github.com>
Co-authored-by: Aiden Cline <63023139+rekram1-node@users.noreply.github.com>
Co-authored-by: 侯秀凯 <9627034+hoooou@users.noreply.github.com>
Co-authored-by: Steven Ao <stevenao@users.noreply.github.com>
Co-authored-by: Matthew Feroz <136640686+MatthewFeroz@users.noreply.github.com>
Co-authored-by: SKY ZHAO <boyzhaotian@hotmail.com>
Co-authored-by: skyzhao1223 <skyzhao1223@users.noreply.github.com>
Co-authored-by: Aiden Cline <rekram1-node@users.noreply.github.com>
Qiiks added a commit to Qiiks/opencode that referenced this pull request Aug 15, 2026
…hydration-fix

Sync 86 upstream commits (Aug 6-14) including:
- ID-wrap/chronological message ordering fixes (anomalyco#40987-anomalyco#41006) — upstream
  independently fixed the same bug class as our 07889a1/21964da41;
  adopted upstream versions (parentID === user.id for prompt exit, same
  findIndex boundaries for revert/session, isAfter helper in latest())
- v1.18.18 release (4 releases: v1.18.15-18)
- v1 database compatibility preservation (anomalyco#42444)
- ignore unknown config fields (anomalyco#41312)
- session retry cap with jitter (anomalyco#41939)
- compaction instructions for smaller models (anomalyco#42045)
- DeepSeek V4 Flash sampling defaults (anomalyco#41620)
- Copilot PDF input support detection (anomalyco#41522)
- web search for opencode-go (anomalyco#42630)
- compaction plugin hooks (experimental.session.compacting, messages.transform)
- New models: GLM 5.3, Gemini 3.7 Flash, Grok 4.6, Zen updates

Fork-only changes preserved:
- Synthetic web search backend (synthetic flag in webSearchEnabled)
- Config npm propagation to inherited models
- Tool-result media extraction for models without attachment capability
- packageManager pin to bun 1.4.-canary.1 (Rust runtime)

Conflicts resolved: message-v2, prompt, revert, session, server-session
(adopted upstream), compaction (adopted upstream), registry + websearch
test (merged both synthetic + opencode-go).
Spark-Liang pushed a commit to Spark-Liang/opencode that referenced this pull request Aug 18, 2026
@chenxuan520

Copy link
Copy Markdown

This is a behavior regression for users who intentionally wait through provider capacity/rate-limit shortages. Before this PR, retryable provider failures with explicit Retry-After / retry-after-ms could keep retrying until the provider recovered. After this PR, all high-level session retries are forced to stop after 5 attempts, even when the provider is explicitly telling the client to wait and retry.

The hard cap is too blunt, and there is no config/opt-out in opencode.json to restore the previous behavior. That makes this a user-visible breaking change for long-running sessions and enterprise/custom providers where transient capacity shortages can last longer than five seconds or five retry windows.

Please either make the retry limit configurable, including an unlimited option, or scope the cap only to locally computed exponential retries without provider retry headers. Explicit provider retry headers should not be treated the same as blind infinite retry loops.

@chenxuan520

Copy link
Copy Markdown

Concrete fix proposal:

  1. Do not apply the hard 5-attempt cap to provider-directed retries. If the API error has retry-after or retry-after-ms, the provider is explicitly telling the client to wait and retry. That should not be treated the same as a blind infinite exponential loop.

  2. Add a user-visible config knob instead of hard-coding this behavior. For example:

{
  "retry": {
    "max_retries": 5,
    "provider_retry_after_max_retries": false
  }
}

or a simpler first pass:

{
  "retry": { "max_retries": 5 } // number | false, where false means unlimited
}
  1. Thread that config into SessionRetry.policy(...) from packages/opencode/src/session/processor.ts instead of relying on RETRY_MAX_RETRIES as an unconfigurable constant.

  2. Minimal behavior change if you want to keep the current default:

const headers = SessionV1.APIError.isInstance(error) ? error.data.responseHeaders : undefined
const hasProviderRetryAfter = !!headers?.["retry-after"] || !!headers?.["retry-after-ms"]
if (!hasProviderRetryAfter && meta.attempt > maxRetries) return Cause.done(meta.attempt)
  1. Tests that should be added:
  • retryable 503/429 without retry headers stops at the configured default
  • APIError with retry-after-ms continues past the default cap
  • APIError with retry-after continues past the default cap
  • retry.max_retries = false preserves the previous unlimited behavior
  • retry.max_retries = N caps at N for non-provider-directed retries

The important distinction is: local exponential retries need a safety cap; provider-directed retry headers need either unlimited behavior or a separate user-configurable cap. Right now this PR collapses both cases into one hard-coded limit, which is the regression.

@chenxuan520

Copy link
Copy Markdown

@rekram1-node please take a look at the regression described above. The problematic part is not just the default cap; it is that the cap is hard-coded and has no user-visible config or provider retry-header exception.

For provider responses that include retry-after or retry-after-ms, the provider is explicitly instructing the client to wait and retry. Those should either remain unlimited like before, or have a separate configurable cap. Applying the same hard 5-attempt limit to both blind exponential retries and provider-directed retries breaks long-running/internal-provider workflows.

Can you make this configurable or exempt provider-directed retry headers from the hard cap?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

infinite retry loop on Azure 429 with Retry-After: 0

2 participants