Skip to content

🧊 fix: In-Memory Endpoint Token Config Cache Isolation - #12673

Merged
danny-avila merged 5 commits into
devfrom
fix/endpoint-token-config-caching
Apr 15, 2026
Merged

danny-avila merged 5 commits into
devfrom
fix/endpoint-token-config-caching

Conversation

@danny-avila

@danny-avila danny-avila commented Apr 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Without Redis, standardCache() created a new in-memory Keyv instance on every call. Each instance has its own isolated Map, so data written by one call-site (e.g., fetchModels at startup) was invisible to reads from another call-site (e.g., initializeCustom during agent init). This meant endpointTokenConfig for providers like OpenRouter was never available when initializing agents, causing a redundant fetchModels HTTP call on every request and the token config being absent for usage tracking.

Changes

packages/api/src/cache/cacheFactory.ts

  • standardCache now memoizes in-memory Keyv instances by namespace (via inMemoryCacheMap). Repeated calls with the same namespace return the same instance, so writes and reads share the same store. This only applies to the plain in-memory path (no Redis, no custom fallbackStore). Redis-backed instances are unaffected since they already share a backend.
  • Added tokenConfigCache() convenience accessor for the TOKEN_CONFIG namespace.

packages/api/src/endpoints/models.ts

  • fetchModels now writes token config via tokenConfigCache() instead of creating an ad-hoc standardCache instance.

packages/api/src/endpoints/custom/initialize.ts

  • initializeCustom now reads token config via tokenConfigCache(), ensuring it sees data written at startup by fetchModels.

packages/api/src/agents/initialize.ts

  • Replaced hardcoded 18000 fallback with a named DEFAULT_MAX_CONTEXT_TOKENS constant (32000).

Test plan

  • Existing standardCache namespace isolation and cache integration tests pass (10/10)
  • Verify OpenRouter agent init reads cached endpointTokenConfig on first request (no extra HTTP call to /models)
  • Verify token usage tracking includes correct per-model rates from the endpoint token config
  • Verify behavior is unchanged when Redis is enabled (memoization only affects in-memory path)

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.
- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.
Copilot AI review requested due to automatic review settings April 15, 2026 12:58
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@danny-avila danny-avila changed the title fix/endpoint token config caching fix: endpoint token config cache isolation without Redis Apr 15, 2026

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: ffd3440697

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +91 to +93
const existing = inMemoryCacheMap.get(namespace);
if (existing) {
return existing;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Respect TTL when reusing in-memory caches

Returning an existing in-memory Keyv instance by namespace without considering the requested ttl makes cache expiration order-dependent. TOKEN_CONFIG is now accessed both without TTL (tokenConfigCache()) and with a 30-minute TTL (api/cache/getLogStores.js); if the no-TTL path runs first (e.g., package-level initialization/tests before legacy cache bootstrapping), subsequent TTL-enabled calls silently reuse the no-expiry instance, so token config entries never expire and can remain stale indefinitely.

Useful? React with 👍 / 👎.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Fixes endpoint token configuration caching so that token-config writes and reads share the same cache instance when Redis is disabled/forced in-memory, preventing “empty cache” behavior across call sites.

Changes:

  • Added in-memory Keyv instance memoization to standardCache() and introduced a tokenConfigCache() convenience accessor for CacheKeys.TOKEN_CONFIG.
  • Updated token-config caching call sites to use tokenConfigCache() (custom endpoint initialization and model fetch flow).
  • Changed agent default max-context fallback from 18000 to 32000 tokens via a shared constant.

Reviewed changes

Copilot reviewed 4 out of 4 changed files in this pull request and generated 4 comments.

File Description
packages/api/src/endpoints/models.ts Writes endpoint token-config into the dedicated token-config cache accessor.
packages/api/src/endpoints/custom/initialize.ts Reads token-config from the dedicated token-config cache accessor to avoid cache-instance mismatch.
packages/api/src/cache/cacheFactory.ts Memoizes in-memory caches by namespace and adds tokenConfigCache() helper.
packages/api/src/agents/initialize.ts Replaces hardcoded max-context fallback with a constant and changes its value.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

const endpointTokenConfig = processModelData(input);
const cache = standardCache(CacheKeys.TOKEN_CONFIG);
await cache.set(tokenKey ?? name, endpointTokenConfig);
await tokenConfigCache().set(tokenKey ?? name, endpointTokenConfig);
Comment thread packages/api/src/agents/initialize.ts Outdated
maxToolResultChars?: number;
};

const DEFAULT_MAX_CONTEXT_TOKENS = 32000;
Comment thread packages/api/src/agents/initialize.ts Outdated
Comment on lines +71 to +72
const DEFAULT_MAX_CONTEXT_TOKENS = 32000;

Comment on lines +22 to +26
/**
* Memoized in-memory Keyv instances keyed by namespace.
* Without Redis, each `new Keyv()` gets its own internal Map, so callers that
* write in one call-site and read in another would see an empty store.
* Memoizing ensures a single shared Map per namespace across the entire bundle.
Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.
@danny-avila

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Keep it up!

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@github-actions

Copy link
Copy Markdown
Contributor

GitNexus: ❌ deploy failed

The deploy failed — the previous index (if any) continues to be served.
Deploy run

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.
@github-actions

Copy link
Copy Markdown
Contributor

GitNexus: ❌ deploy failed

The deploy failed — the previous index (if any) continues to be served.
Deploy run

@danny-avila danny-avila changed the title fix: endpoint token config cache isolation without Redis 🧊 fix: In-Memory Endpoint Token Config Cache Isolation Apr 15, 2026
@danny-avila
danny-avila merged commit a613cac into dev Apr 15, 2026
10 checks passed
@danny-avila
danny-avila deleted the fix/endpoint-token-config-caching branch April 15, 2026 13:41
krgokul pushed a commit to syedhabib39/LibreChat that referenced this pull request Apr 20, 2026
…12673)

* fix: endpoint token config not using shared cache in same process (initializing clients)

* refactor: Update default max context tokens for agent initialization

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.

* refactor: shared caching mechanism for token configuration

- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.

* fix: match TOKEN_CONFIG TTL and add memoization tests

Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.

* fix: export DEFAULT_MAX_CONTEXT_TOKENS, address review nits

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.
krgokul pushed a commit to syedhabib39/LibreChat that referenced this pull request Apr 21, 2026
…12673)

* fix: endpoint token config not using shared cache in same process (initializing clients)

* refactor: Update default max context tokens for agent initialization

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.

* refactor: shared caching mechanism for token configuration

- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.

* fix: match TOKEN_CONFIG TTL and add memoization tests

Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.

* fix: export DEFAULT_MAX_CONTEXT_TOKENS, address review nits

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.
yidianyiko pushed a commit to yidianyiko/LibreChat that referenced this pull request May 6, 2026
…12673)

* fix: endpoint token config not using shared cache in same process (initializing clients)

* refactor: Update default max context tokens for agent initialization

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.

* refactor: shared caching mechanism for token configuration

- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.

* fix: match TOKEN_CONFIG TTL and add memoization tests

Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.

* fix: export DEFAULT_MAX_CONTEXT_TOKENS, address review nits

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.

(cherry picked from commit a613cac)
jcbartle pushed a commit to jcbartle/LibreChat that referenced this pull request May 11, 2026
…12673)

* fix: endpoint token config not using shared cache in same process (initializing clients)

* refactor: Update default max context tokens for agent initialization

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.

* refactor: shared caching mechanism for token configuration

- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.

* fix: match TOKEN_CONFIG TTL and add memoization tests

Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.

* fix: export DEFAULT_MAX_CONTEXT_TOKENS, address review nits

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.
ThomasVuNguyen pushed a commit to ThomasVuNguyen/LibreChat that referenced this pull request Jul 15, 2026
…12673)

* fix: endpoint token config not using shared cache in same process (initializing clients)

* refactor: Update default max context tokens for agent initialization

- Introduced a constant `DEFAULT_MAX_CONTEXT_TOKENS` set to 32000.
- Updated the `initializeAgent` function to use this constant instead of hardcoded values for maximum context tokens, improving maintainability and clarity.

* refactor: shared caching mechanism for token configuration

- Introduced a memoized in-memory cache for Keyv instances to ensure shared access across the same namespace, improving cache efficiency.
- Updated the `standardCache` function to utilize the new in-memory cache for the TOKEN_CONFIG namespace.
- Refactored the `initializeCustom` function to use the `tokenConfigCache` for better cache management.
- Removed redundant tokenCache parameter from `fetchModels` to streamline the function signature.

* fix: match TOKEN_CONFIG TTL and add memoization tests

Pass Time.THIRTY_MINUTES to tokenConfigCache() to match the TTL used
by getLogStores.js, preventing load-order-dependent expiry behavior.

Add 7 automated tests covering in-memory memoization: referential
identity, cross-call-site data sharing, namespace isolation,
first-caller TTL semantics, fallbackStore bypass, and tokenConfigCache
parity with direct standardCache access.

* fix: export DEFAULT_MAX_CONTEXT_TOKENS, address review nits

- Export the constant so tests (and future consumers) reference it
  directly instead of hardcoding the numeric value.
- Add independent TTL assertion for tokenConfigCache (R-1 nit).
- Add tokenConfigCache mock to custom/initialize.spec.ts.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants