Skip to content

[Bug]: Deferred tools incorrectly counted in pre-flight context token check, causing empty_messages error on small-context models #12702

Description

@ErikDakoda

What happened?

What happened?

When an agent has a large number of MCP tools enabled but mostly deferred (via defer_loading: true), requests to models with smaller context windows fail immediately with an empty_messages error, even though the actual model payload only contains the non-deferred tools plus a single tool_search stub.

Error:

An error occurred while processing the request: {"type":"empty_messages","info":"Message pruning removed all messages as none fit in the context window. Tool definitions consume 147877 tokens (98% of instructions) across 292 tools, exceeding maxContextTokens (140000). Reduce the number of tools or increase maxContextTokens. Summarization was skipped because the summary would further increase the instruction overhead.\nToken budget breakdown:\n maxContextTokens: 140000\n instructionTokens: 151181 (system: 3304, tools: 147877 [292 tools])\n summaryTokens: 0\n messageTokens: 34 (1 messages)\n availableForMessages: 0"}

The error claims 292 tools consuming 147,877 tokens — but the actual system prompt sent to the model on large-context runs shows only the non-deferred MCP server instructions (~2,000 chars) and no deferred tool schemas at all.

Root cause

The pre-flight context token check in @librechat/agents (Graph.cjs) counts all entries in the toolDefinitions array, including those with defer_loading: true. However, deferred tools are correctly excluded from the actual model payload at runtime.

The call chain:

  1. ToolService.js → loadToolDefinitions() returns a toolDefinitions array containing all tools (deferred + non-deferred), because the full registry is needed for tool_search to function.
  2. initialize.ts → initializeAgent() passes this full toolDefinitions array into AgentInputs unchanged.
  3. run.ts → createRun() passes toolDefinitions (including deferred entries) to the graph.
  4. Inside @librechat/agents, the pre-flight check tokenizes every entry in toolDefinitions to compute instructionTokens, without filtering out defer_loading: true entries.

Result: The pre-flight check sees 292 full tool schemas and fails, even though the actual request to the model would only include a handful of non-deferred tools.

Why it only affects small-context models: Large-context models (e.g. 200k tokens) pass the inflated pre-flight check silently. Small-context models (e.g. 140k tokens) fail it before any messages are even considered.

Expected behavior

The pre-flight token check in @librechat/agents should exclude tool definitions where defer_loading === true when computing instructionTokens. Those tools are not sent to the model and should not count against the context budget.

Version Information

@librechat/agents: ^3.1.65
LibreChat: v0.8.5-rc1 (latest main)

Steps to Reproduce

  1. Configure an agent with a large number of MCP servers (e.g. 10+ servers, 200+ total tools).
  2. Enable deferred_tools capability so most tools are deferred; only tool_search is sent to the model by default.
  3. In Model Parameters set Max Context Tokens to a number smaller than the sum of all tool schema tokens (100000 with 147k tokens worth of tool definitions).
  4. Send any message.

Actual: Immediate empty_messages error — pre-flight check counts all 292 tool definitions including deferred ones.

Expected: Request succeeds — pre-flight check only counts the non-deferred tools actually included in the payload (typically just tool_search + a few non-deferred tools).

Suggested Fix

In the pre-flight token budget calculation inside @librechat/agents, filter toolDefinitions before counting:

// Only count tools that will actually be sent to the model
const activeToolDefinitions = toolDefinitions.filter(t => t.defer_loading !== true);
// use activeToolDefinitions for instructionTokens calculation

Alternatively, pass a separate activeToolDefinitions alongside toolDefinitions in AgentInputs so the graph can distinguish between the full registry (needed for tool_search lookup) and the tools actually included in the model payload.

Additional Context

  • The deferred tools feature itself works correctly when Max Context Tokens is large enough (maybe exceeding the actual context size of the model) — tools are not included in the model payload and tool_search functions as expected.
  • This is a pure token-counting discrepancy in the pre-flight check and context auto-summarization.
  • No existing issue covers this. The only current workarounds are: increase maxContextTokens beyond the inflated count, or reduce the total number of enabled tools.

What browsers are you seeing the problem on?

Chrome

Relevant log output

N/A

Screenshots

Image

Code of Conduct

  • I agree to follow this project's Code of Conduct

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    🐛 bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions