Skip to content

⛲ fix: Prevent Agent Model Stream Idle Timeouts - #16201

Merged
danny-avila merged 7 commits into
devfrom
lia/investigate-agent-timeouts
Sep 22, 2026
Merged

danny-avila merged 7 commits into
devfrom
lia/investigate-agent-timeouts

Conversation

@lia-by-librechat

@lia-by-librechat lia-by-librechat Bot commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

OpenAI-compatible Agent model calls can fail with terminated when Undici's five-minute body-idle timeout expires, even when the SDK request timeout is longer. MCP's transport fixes do not configure the model client. The SDK timer ends when response headers arrive; it is not a total-response deadline.

This change supplies an explicit, configurable transport policy for OpenAI-compatible Agent calls and summarization, including built-in cross-provider summaries and directEndpoint custom URLs. Non-Agent callers of getOpenAIConfig retain their existing behavior unless they explicitly supply the policy.

Agent settings -> provider/summarization configuration -> proxy or SSRF-safe dispatcher -> model SDK
endpoints:
  agents:
    modelResponseBodyTimeoutMs: 900000
    modelResponseHeadersTimeoutMs: 300000

The defaults are a finite 15-minute body-idle allowance and a five-minute header allowance. Values are integer milliseconds from 0 through 86400000; 0 explicitly disables that transport timer. Received body chunks reset the idle timer. These settings do not define total run duration and do not replace the SDK's timeout or cancellation. An earlier SDK timeout can still win while waiting for headers. Native non-OpenAI providers and MCP timeouts are unchanged.

Direct Agent endpoints retain their exact URL while using Undici so the same timeout, proxy, and SSRF policy is enforced. Cached dispatchers accept only numeric timeout options; private SSRF connection policies remain isolated. Cancellation and retry policy are unchanged.

Related to #16195. The transport failure is reproduced locally through the real locked Agent model client; the reporter's deployment has not been independently reproduced.

Change Type

  • Bug fix (non-breaking change which fixes an issue)

Testing

  • 9 real-socket Agent SDK tests: standard/direct header and body timeout enforcement, longer allowances, active streams, cancellation, non-streaming bodies, and SSRF redirect rejection. Executed in two short filtered runs.
  • 208 summarization tests passed, including built-in cross-provider timeout defaults/overrides and credential/URL preservation.
  • 255 focused configuration, initialization, compatibility, resolver, and proxy tests passed across 8 suites.
  • npm run static-checks -- --against origin/dev: all affected checks passed across 17 files, including ESLint, Prettier, imports, and circular dependencies.
  • Current-head npx tsc --noEmit in packages/api was stopped at a 20-second execution limit, without a result. Not claimed as passing. No long build, install, full suite, or full static-check run was performed for this revision.

Remaining verification

CI and a new review must cover the latest pushed head. The earlier review does not cover these corrections.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Ready for review at exact head 32a6aa7605ada01ddc674dca16c292b5700cb3dd. This head applies explicit Undici body/header timeout policy to OpenAI-compatible Agent model dispatchers, preserves proxy and SSRF paths, and includes focused configuration and dispatcher-cache coverage.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Updated review head: bbc5cddb551080e94e074d57613acaf46dc5ab38. Added direct dispatcher option/reuse coverage; focused tests, packages/api typecheck, and static checks all pass on this head.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Updated review head: c8d997b1dd0542a01191e76300ef55aeabf9df46. This fixes the failing OpenAI compatibility shards by keeping getOpenAIConfig transport-neutral unless Agent timeout policy is explicitly supplied. Agent calls now use configurable finite body/header timeout defaults. Focused results: 8 API suites, 254 tests passed; data-provider config suite, 221 tests passed; both changed-workspace typechecks and affected-file static checks passed.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-22T12:27:36.031147Z c06db67 Manual request
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c8d997b1d1

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +264 to +265
mergeFetchOptions(configOptions, {
dispatcher: proxyDispatcher ?? getDirectDispatcher(transportTimeouts),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Apply transport timeouts to direct-endpoint fetches

When a custom endpoint sets directEndpoint: true, getOpenAIConfig later replaces the fetch implementation with createFetch, which calls node-fetch; that implementation ignores the Undici dispatcher installed here. Consequently, neither configured timeout is enforced on this supported path—for example, modelResponseHeadersTimeoutMs: 1000 can still wait indefinitely for response headers. Pass the timeout policy into the custom fetch or retain an Undici-based fetch so the documented operator settings work consistently.

AGENTS.md reference: AGENTS.md:L66-L68

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in ae6711f. Direct Agent endpoints now rewrite only the URL and use Undici with the configured dispatcher instead of node-fetch. Real Agent SDK/socket tests verify header and body-idle timeouts on both standard and direct endpoints, exact URL preservation, cancellation, non-streaming response bounds, and SSRF redirect rejection. The legacy no-policy fetch path is unchanged.

tenantId: headerContext.tenantId,
body: headerContext.requestBody,
}),
transportTimeouts: resolveModelTransportTimeouts(appConfig?.endpoints?.agents),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Cover cross-provider built-in summarizers

When a non-OpenAI agent, such as an Anthropic agent, selects summarization.provider: openAI, this timeout-aware resolver is not used because resolveOpenAISummarization is called only for Azure agents. The request instead goes through resolveBuiltInClientOverrides, whose getOpenAIConfig call omits transportTimeouts and whose returned shaping deliberately excludes configuration, so both the new defaults and explicit operator overrides are ignored and the summarization stream retains Undici's five-minute idle timeout. Apply the policy to this built-in cross-provider path as well.

AGENTS.md reference: AGENTS.md:L66-L68

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in ae6711f. The built-in cross-provider OpenAI-family path now supplies the resolved timeout dispatcher without replacing credentials, base URLs, or model settings. Explicit URL overrides retain timeout policy but do not acquire first-party request shaping. Same-provider summaries continue inheriting their Agent client's configuration. The full summarization test file passes (208 tests), including policy defaults, operator overrides, disabled timers, and existing Azure/subagent paths.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Updated review head: ae6711f5801cae7a5509b8ec1e7d13c9b4db7df4.

Both inline findings are addressed: direct endpoints now enforce policy through Undici, and cross-provider built-in summaries receive the configured dispatcher. Also fixed the formatting failure and restricted runtime cache inputs to the options represented by cache keys.

Verification:

  • 9 real-socket Agent SDK regressions passed in short split runs.
  • 208 summarization tests passed.
  • 255 focused compatibility/configuration/proxy/initialization tests passed.
  • Static checks passed against origin/dev for the full committed 17-file diff, including ESLint and Prettier.
  • API TypeScript check hit the 20-second execution limit and was stopped; this head is not claimed typecheck-green. No full suites/builds/installs were run in this revision.

Subsystem self-review covered direct vs SDK fetch paths, cross-provider and inherited summarization, cancellation, retry preservation, SSRF/proxy composition, and cache identity. No persistence or authorization contract is changed. CI is running; a maintainer can trigger review of this exact head.

Correction to the previous handoff: the earlier remote SHA was c8d997b1d19c04147084b2a5ad29f1b92284d428, not the longer SHA typed in that comment. The earlier static run covered the old four-file committed diff, so it was not sufficient evidence for those pending edits. This handoff reports the actual committed diff.

@lia-by-librechat

Copy link
Copy Markdown
Contributor Author

Updated review head: c06db6757daa022aea776d49b4878f4b65faa808.

Addresses both errors from TypeScript job 106739412963:

  • config.ts: narrow the SDK's cross-runtime fetch-options type at the Undici adapter boundary. The options are constructed locally with an Undici dispatcher; their runtime values are unchanged.
  • transport.spec.ts: explicitly omit verbosity as the existing requests.spec.ts adapter does. LibreChat permits a broader nullable string type than the SDK accepts.

All 9 real-socket Agent transport regressions passed in two short runs. Affected-file static checks passed on the full committed 17-file PR diff. Local npx tsc --noEmit hit the 20-second cap again and was stopped, so the fresh CI typecheck remains the validation gate. No install, build, full suite, or long local check was run. No new timeout policy, retry, proxy/SSRF, persistence, or cleanup changes in this commit.

A new exact-head review is pending. The previous head's reviews do not cover this correction.

@danny-avila

Copy link
Copy Markdown
Collaborator

@codex review the latest head

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: c06db6757d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@danny-avila
danny-avila merged commit 45300a6 into dev Sep 22, 2026
36 checks passed
@danny-avila
danny-avila deleted the lia/investigate-agent-timeouts branch September 22, 2026 12:30
KinseyD pushed a commit to KinseyD/LibreChat that referenced this pull request Sep 24, 2026
* 🌊 fix: Prevent Agent Model Stream Idle Timeouts

* fix: Preserve Default Proxy Dispatcher Construction

* test: Assert Agent Transport Timeout Policy

* test: Cover Direct Model Dispatcher Reuse

* fix: Scope Agent Model Transport Timeouts

* fix: Enforce Agent Timeouts Across Direct and Summary Clients

* fix: Align Model Transport Adapter Types

---------

Co-authored-by: Lia <lia@librechat.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants