Skip to content

feat: serve custom endpoints their declared models intersected with the live catalog - #72

Open
sindriii wants to merge 27 commits into
apro-deployfrom
feat/model-availability-filter
Open

sindriii wants to merge 27 commits into
apro-deployfrom
feat/model-availability-filter

Conversation

@sindriii

@sindriii sindriii commented Aug 21, 2026 •

Copy link
Copy Markdown

Why

We front LiteLLM with several custom endpoints — one per provider dialect, because customParams.defaultParamsEndpoint is what selects the Anthropic/Google request builder, and Claude models are measurably worse on the OpenAI builder. All of those endpoints share one baseURL, apiKey and header set.

Today they cannot have different model lists:

  • With fetch: true, a successful fetch replaces models.default wholesale, so every endpoint advertises LiteLLM's entire catalog — including embeddings, rerankers and OCR models, each on the wrong dialect.
  • With fetch: false and curated lists, the lists come apart but nothing tracks reality any more: a model retired in LiteLLM lingers in the picker, and the per-user catalog our OIDC endpoints already return collapses to one static list for everyone.

There is no third option in config, which is what this PR adds.

The same split leaves a second gap. Once the model list comes from the gateway, every row in the picker is a raw model id — claude-opus-4-8, gpt-5.4-nano. Agents and assistants get a name to show instead; no other endpoint does. Commits 5 and 6 give an endpoint somewhere to put a human label.

What

Six commits, reviewable in order.

1. models.filter — declared ∩ fetched

models.default becomes an allowlist rather than a fallback: the endpoint serves the declared names the gateway actually returned, in declared order. Declaring a model a deployment does not have is inert, so one shared list covers a fleet where each deployment holds a subset, and rolling a model out becomes a gateway-side change with no config edit.

Supporting changes:

  • models.default no longer requires one entry — an empty declared list is now meaningful (a template a deployment fills in).
  • Both prefilters deciding whether an endpoint has any model source tested models.default for truthiness, and [] is truthy. They now test length, and share one predicate instead of holding two copies of it.
  • A fetch that never answered is kept distinct from a fetch that answered with nothing.

2. One models-config resolution per request

getModelsConfig is reached from seven places per page (models route, startup-config spec pruning, validateModel, token config, agent initialization, both agent response controllers) and each re-ran the full resolution including a live /models call per gateway. MODEL_QUERIES cannot absorb it — fetchModels skips that cache whenever an endpoint forwards user-bound headers, and must, since the response is identity-scoped. Memoized on the request object instead; a failure evicts so a later caller can retry.

3. Withhold endpoints with nothing to serve

An endpoint's existence is decided by declaration, never content, so an empty endpoint still renders — useEndpoints computes hasModels but drops only agents, and the agent builder offers it as a provider with no model to pick. Already reachable today: an endpoint sending an authorization header whose fetch returns nothing gets [] rather than the declared fallback. Now withheld from the endpoints config, which is the single object the selector, the agent builder and spec pruning all derive from.

Custom endpoints only. Fails open — only an explicit empty list withholds an endpoint.

4. No violation for a request to an endpoint with nothing to offer

validateModel treats any model absent from an endpoint's list as an ILLEGAL_MODEL_REQUEST, which carries a violation score and contributes toward a ban. That reading only holds when the endpoint has models and the requested one is not among them.

An endpoint whose list is empty is unavailable, and the requests that arrive are stored conversations and agents naming an endpoint that has stopped serving them — commit 3 makes that an ordinary state. Their owners would collect violations for a configuration change they had no part in. Those are now rejected as "Endpoint unavailable" without logging. An unlisted model on an endpoint that does have models still logs, unchanged.

5. modelLabels — display labels on the endpoint

An endpoint-level Record<modelId, label>, alongside modelDisplayLabel and following tokenConfig, passed through to the client.

Display-only. The id stays what is declared, fetched, intersected, selected, stored on the conversation and sent upstream, so a label is safe to change at any time and a model with no entry renders its id.

Not routed through models.default: modelItemSchema has no label, and both the declared and the fetched path resolve to bare id strings, so modelsConfig is a Record<string, string[]> and stays that way. Keeping the map on the endpoint also means a label for a model that endpoint does not serve is simply never rendered — one map covers a fleet where each deployment serves a subset, which is the same property commit 1 gives the declared list.

6. Render the label

One helper, getModelName, returns the agent name, the assistant name or the declared label for a model, and undefined when there is none so each caller keeps its own fallback. An empty string counts as no name — that is what agentNames stores for an unnamed agent.

Five sites read it: the endpoint's model list, a search result row, the selector's closed trigger, the announcement made on selecting a model, and the Agent Builder's model combobox. The last matters most on deployments setting modelSpecs.addedEndpoints: [agents], where provider endpoints are kept out of the selector entirely and the Agent Builder is the only place a user meets a model id.

Search covers the label and the id, across all three filters — which endpoints survive a query, the open endpoint's list, and the search results' own predicate. A label is additive there: both strings match. An agent or assistant name still replaces the id, unchanged.

Selection is untouched: handleSelectModel and ControlCombobox both keep the id as the value they hand back.

Behaviour change to review deliberately

A rejected fetch now falls back to the declared list for every endpoint, including those sending an authorization header. Previously a rejected promise was flattened to [] before the header check, so those endpoints went empty during an outage.

This is intentional and paired with commit 3: once an empty list can remove an endpoint, treating an unreachable gateway as an authoritative empty would delete endpoints mid-blip and raise INVALID_AGENT_PROVIDER for stored agents naming them. The gateway, not this list, is the enforcement point — it rejects a call for a model the caller lacks regardless.

An answer of [] still yields nothing. For filtered endpoints that now falls out of the intersection without inspecting headers, so hasAuthorizationHeader should become redundant on the fulfilled path once every endpoint carries filter: true — left in place until that is confirmed on a live deployment.

Blast radius

  • Commits 1–4 need no client changes. useEndpoints and AgentPanel both derive from endpointsConfig / modelsConfig. Commits 5–6 touch the client, but only the string rendered for a model.
  • filterModelSpecsByAvailability untouched — it starts pruning against real availability again as a consequence of getting a better list. Note it fails open on a missing key and closed on [], which is why commit 1 keeps the key present.
  • endpoints.agents.allowedProviders unaffected — initialize.ts rejects a provider missing from the list, so extra entries stay harmless.
  • Existing endpoints without filter behave as before on the fulfilled path, OIDC empty-answer handling included.

Tests

  • packages/api/src/endpoints/config/availability.spec.ts — 16 new: intersection and declared order, two endpoints sharing one coalesced fetch, empty answer, rejected fetch with and without an authorization header, filter with no fetch, the empty-declaration prefilter, and three pinning unfiltered behaviour.

  • packages/api/src/endpoints/config/endpoints.spec.ts — 7 new: withholding, built-ins never withheld, fail-open on error / on a missing key / with no resolver supplied.

  • api/server/services/Config/__tests__/getModelsConfig.spec.js — 4 new: merge, one resolution per request, no sharing between requests, retry after failure.

  • api/server/middleware/__tests__/validateModel.spec.js — 2 new: no violation on an empty list, violation still logged for an unlisted model.

  • packages/data-provider/specs/config-schemas.spec.ts — 4 new: modelLabels survives parsing rather than being stripped as an unknown key, accepts labels for undeclared models, stays optional, rejects a non-string label.

  • packages/api/src/endpoints/custom/config.spec.ts — 2 new: passed through to the client, absent when undeclared.

  • client/.../Endpoints/__tests__/utils.test.ts — 9 new: getModelName precedence and its empty-name and null handling, label-and-id search through filterItems and filterModels, getDisplayValue labelled and unlabelled.

  • client/.../components/__tests__/EndpointModelItem.test.tsx — 3 new: label rendered instead of the id, unlabelled model falls back, selection still keys on the id.

  • client/.../components/__tests__/SearchResults.test.tsx — 3 new: found by label, found by id, selected by id.

Green: packages/api 6576 tests (the 20 remaining failures are the Redis *_integration suites plus one timing-sensitive flow/manager spec that passes in isolation — no local Redis), packages/data-provider 1254, and the client's Endpoints / SidePanel/Agents / hooks/Endpoint suites, 168.

The api suite is flaky on this baseline — three consecutive runs gave 2, 4 and 3 failures out of 2976, all in strategies/openIdJwtStrategy, the OpenID cookie suites, requestPasswordReset and the CloudFront cookie suite. None sits in a path this branch touches. The two stable ones were reproduced at 650e62089 with this branch's changes reverted and both packages rebuilt. The subset covering everything changed here — server/services/Config, server/middleware, server/controllers, server/routes, 1349 tests — passes.

Not yet exercised against a live gateway; that is the next step on apro-sandbox.

One thing this does not fix

RemoteAgents / Remote-style endpoints — fetch: true, no filter, authenticating as a shared key — still replace their declared list with the gateway's whole catalog. That is what filter is for; adopting it there is a config change, not a code one.

Known issue — not fixed here

Commit 9679baa4b moves the custom-endpoint lookup ahead of the case-folded native lookup, and getCustomEndpointConfig throws when appConfig is absent. The old ordering reached the native map first, so a builtin resolved without an app config. Now getProviderConfig({ provider: 'Anthropic' }) with no appConfig throws Config not found for the Anthropic custom endpoint. instead of returning initializeAnthropic — same for Google, Bedrock, VertexAI.

resolveActivityLabelModel (packages/api/src/agents/activityLabels/host.ts) passes a possibly-undefined req.config into an uncaught getProviderConfig call, so it is reachable in principle, though req.config is set on every real request path. Fix when it matters: guard the custom lookup with appConfig &&, and keep a !appConfig arm so the existing Config not found wording survives on the paths that already raised it.

sindriii and others added 3 commits August 21, 2026 12:42
…fetched catalog

`models.default` has only ever been a fallback: with `fetch: true` a
successful fetch replaces it wholesale, so every endpoint pointed at one
gateway advertises that gateway's entire catalog. Deployments that put
several endpoints over a single OpenAI-compatible gateway — one per
provider dialect, which is the only way to send each provider the
parameters it needs — therefore cannot give those endpoints different
model lists. Curating `models.default` with `fetch: false` gets the lists
apart, but then nothing reflects what the gateway actually serves: a model
retired upstream lingers in the picker, and a per-user catalog collapses
to one static list for everyone.

Add `models.filter`. With it, `models.default` is an allowlist rather than
a fallback and the endpoint serves `declared ∩ fetched`, in declared
order. Declaring a model the gateway does not have is inert, so one shared
list can cover a fleet of deployments that each hold a subset, and an
endpoint's contents follow the gateway without a config edit.

An empty declared list is now meaningful — an endpoint template that a
deployment fills in — so `models.default` no longer requires one entry,
and the two prefilters that decide whether an endpoint has any model
source at all now test that list's length. They were testing it for
truthiness, and `[]` is truthy, which would have admitted an endpoint that
can never produce a model. Both prefilters were separate copies of that
predicate; they now share one, next to the intersection it guards.

A fetch that never answered is kept distinct from a fetch that answered
with nothing. Transport failure falls back to the declared list, for every
endpoint: an empty list is about to mean "no endpoint", and treating an
unreachable gateway as an authoritative empty would take endpoints away
mid-outage. This is the one behaviour change for endpoints that send an
`authorization` header — a rejected fetch used to collapse to `[]` for
them, because it was flattened into an empty answer before that check ran.
An answer of `[]` still yields nothing, which is what the header check
already did for a successful empty response, now reached for filtered
endpoints without inspecting headers at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`getModelsConfig` is the accessor for a request's available models and is
already reached from seven places — the models route, the startup-config
route pruning model specs, model validation on submit, token config, agent
initialization and both agent response controllers — each of which
re-runs the whole resolution, including a live `/models` call per gateway.

The shared `MODEL_QUERIES` cache cannot absorb that. `fetchModels` skips
it whenever an endpoint forwards user-bound headers, and it has to: the
response is scoped to the caller's identity, so one user's list must never
be served to another.

Memoize on the request object instead, which is the exact scope that makes
the result reusable — same identity, same moment — and lets the entry be
collected with the request. A failure evicts, so a later caller in the
same request retries rather than inheriting a settled rejection.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An endpoint's existence is decided by its declaration, never by whether it
can serve anything, so an endpoint with an empty model list still renders.
`useEndpoints` computes `hasModels` but drops only `agents`, and the agent
builder offers the endpoint as a provider with no model to pick.

That is already reachable: for an endpoint sending an `authorization`
header, a fetch returning nothing yields an empty list rather than falling
back to declared models, so a user whose token grants no models gets an
endpoint that cannot answer. `models.filter` makes it ordinary — an
endpoint whose declared list intersects the catalog to nothing is an
endpoint this deployment does not have.

Withhold those from the endpoints config, which is the one object the
model selector, the agent builder and model-spec pruning all derive from,
so a single decision covers all three instead of three that can drift.

Custom endpoints only; built-in model lists are not per-request. Fails
open — only an explicit empty list withholds an endpoint, while a models
config that cannot be resolved, or that has no entry for an endpoint,
leaves every declared endpoint in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 21, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Team

Run ID: 31f89130-a33d-4264-a4d6-622a15b2327a

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Comment @coderabbitai help to get the list of available commands.

sindriii and others added 15 commits August 21, 2026 13:05
`validateModel` treats any model absent from an endpoint's list as an
`ILLEGAL_MODEL_REQUEST`, which carries a configurable violation score and
contributes toward a ban. That reading only holds when the endpoint has
models and the requested one is not among them.

An endpoint whose list is empty is unavailable — the gateway serves none of
what it declares, or none of it for this user's grants. Every model named
against it is unserveable rather than illegitimate, and the requests that
arrive are stored conversations and agents pointing at an endpoint that has
since stopped serving them. Their owners would collect violations, and
eventually a ban, for a change in configuration they had no part in.

Reject those without logging a violation, and say the endpoint is
unavailable rather than the request illegal. An unlisted model on an
endpoint that does have models still logs, unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
An endpoint's model list renders raw ids — `claude-opus-4-8`, `gpt-5.4-nano`.
Only agents and assistants carry a name to show instead; every other endpoint
shows the id it was configured with, which is the wrong string to put in front
of a non-technical user.

Adds `modelLabels`, an endpoint-level `Record<modelId, label>` alongside
`modelDisplayLabel` and `tokenConfig`, and passes it through to the client.

It is display-only. The id stays what is declared, fetched, intersected,
selected, stored on the conversation and sent upstream, so a label is safe to
change at any time and a model with no entry renders its id. Keeping the map at
endpoint level rather than in `models.default` also keeps `modelsConfig` a
`Record<string, string[]>`: `modelItemSchema` has no `label`, and both the
declared and the fetched path resolve to bare id strings.

Because a label for a model the endpoint does not serve is simply never
rendered, one map can cover a fleet of deployments that each serve a subset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Resolves the name to show through one helper, `getModelName`, which returns the
agent name, the assistant name or the endpoint's declared label for the model —
and `undefined` when there is none, so each caller keeps its own fallback. An
empty string counts as no name, which is what `agentNames` stores for an unnamed
agent.

Five sites read it: the endpoint's model list, a search result row, the
selector's closed trigger, the announcement made on selecting a model, and the
Agent Builder's model combobox. The last one matters most on deployments that
set `modelSpecs.addedEndpoints: [agents]`, since that keeps provider endpoints
out of the selector entirely and the Agent Builder becomes the only place their
users meet a model id.

Search covers the label and the id, across all three filters — which endpoints
survive a query, the open endpoint's list, and the search results' own
predicate. A label is additive there: both strings match. An agent or assistant
name still replaces the id, unchanged.

Selection is untouched: `handleSelectModel` and `ControlCombobox` both keep the
id as the value they hand back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`withholdEmptyCustomEndpoints` ran on every `getEndpointsConfig` call, and that
function has seven callers — most of which want the endpoint's *configuration*,
not a view of what a user may be offered. Two of them make this actively wrong,
not merely slow:

- `buildEndpointOption` reads `defaultParamsEndpoint` off the merged config to
  pick a request builder. Withhold the endpoint and it reads `undefined`, and the
  conversation falls back to the OpenAI schema — the exact dialect degradation
  the per-provider split exists to prevent, reachable on the message path
  whenever a user's catalog comes back empty.
- `validateModel` reads `userProvide` to let a user-keyed endpoint through. A
  withheld endpoint loses that and takes the wrong branch.

And `checkCapability` shares the same function, so a capability check — three of
which sit on the file-upload path — waited on a model catalog it never reads.

Withholding is now opt-in, and `/api/endpoints` is the only caller that opts in:
it is the one answering "what may this user be offered", and the selector, the
Agent Builder and spec pruning all derive from its response.

Cache the user-scoped catalog instead of skipping the cache. Skipping it was
right about the hazard — keyed by baseURL+apiKey alone, one user's filtered list
is served to the next request over that gateway — and wrong about the remedy: it
made every request pay a live fetch, and a page load issues three. The key now
carries the requesting user and the header templates in play, with a
thirty-second TTL rather than the catalog's two minutes, so a page load costs one
round trip and a revoked grant surfaces within half a minute. The gateway remains
the enforcement point either way. With no identity to key on, nothing is cached.

Cut this path's timeout from 5s to 2s. Every caller of the config loader degrades
to the declared list, so the only thing a longer wait buys is a more accurate
list — paid for with a blank picker in front of a waiting user.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Thirty seconds was picked to cover a page load and nothing more, on the
assumption that a short window was the safe choice. It buys less safety than it
looks: a stale entry cannot grant anything, because the gateway rejects a call
for a model the user lacks regardless. What a long window actually costs is a
list that lags — a revoked grant, or a model just rolled out gateway-side.

Five minutes trades a little of that lag for far fewer round trips, which is the
right way round while the cache has no shared store: without Redis it is an
in-memory map per task, wiped on restart, so entries have to survive longer than
one page load to be worth anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The 5s ceiling was cut to 2s on the reasoning that a caller which degrades
gracefully should not make a user wait — but the premise was wrong. `fetchModels`
catches its own transport errors and returns an empty list, so a timeout is not a
failure this code sees: it is an empty catalog, which intersects to nothing and
withholds the endpoint. A shorter timeout therefore produces *more* empty
pickers, which is the symptom it was meant to relieve.

Removes the parameter as well as the value: nothing else wanted it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An empty model list on a user-provided endpoint reflects the user's stored
key — missing, expired, or granted nothing — and the picker entry is the
only way to set or fix that key. Withholding it would lock the user out,
so it stays visible even when empty, mirroring validateModel's exemption.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Keying MODEL_QUERIES by identity instead of skipping it for user-scoped
responses was a convenience, not something this feature needs. The
per-request memo already collapses a page load's several resolutions into
one fetch per gateway, which is the cost withholding actually adds; the
cross-request cache only reduced it further, and paid for that with a TTL
to reason about, an identity in the cache key to get right, and 43 lines in
a file upstream touches often.

Restores `fetchModels` and its spec to their upstream behaviour: the cache
is skipped whenever a caller resolves headers against a specific user.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Withholding was an option threaded through `getEndpointsConfig`, which meant
an injected models resolver, an options interface, a signature change and a
try/catch inside the loader — and a standing risk that some other caller
would pass `withholdEmpty` and lose the `defaultParamsEndpoint` /
`userProvide` keys the request path reads off those entries.

It is a pure transform of two objects, so it is one now: a
`withholdEmptyEndpoints(endpointsConfig, modelsConfig)` beside the other
availability helpers, applied by `/api/endpoints` — the only caller that
answers "what may this user be offered". `getEndpointsConfig` returns to its
upstream shape and cannot withhold anything, so the hazard is structural
rather than a comment asking callers not to.

Fail-open moves with it: an unresolvable models config yields `null`, which
withholds nothing, and the route logs rather than failing the picker.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…e each

Three small cleanups in the label rendering, no behaviour change:

- `modelSearchNames` states the rule once — a declared label is additive, an
  agent or assistant name replaces the id — and `SearchResults` uses it instead
  of an inline branch that had to restate it. `filterItems` and `filterModels`
  keep their own name lookups: they resolve agents and assistants from
  `agentsMap`/`assistantsMap` rather than off the endpoint, and unifying that is
  a separate change from labelling.
- `getDisplayValue` went through `getModelName` for agents and assistants but
  hand-rolled the same chain first, then appended labels at the end. It calls
  the helper now, keeping only the `agentsMap` fallback the helper has no access
  to.
- `isGlobal` was gated on the model having a name, in both the model list and
  the search results. It is a property of an agent, not of its naming; gating it
  on `isAgentsEndpoint` alone drops a condition from each site.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The reasoning behind each decision belongs in the commit that made it, and
these were carrying whole paragraphs of it inline — an essay per branch in the
fetch loop, a five-line note on a two-line schema field. Cut to the invariants
a reader needs at the line: why an empty declared list is legal, why a failed
fetch is not an empty answer, why an empty endpoint must not log a violation.

No code changes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Nested ternary in `getModelName`, import formatting in `availability.ts`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both are opt-in keys an admin has no other way to discover.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`getProviderConfig` looked up the native provider map case-insensitively
before considering custom endpoints, and the custom lookup sat in the `else`
of that test. So a custom endpoint whose name case-folds onto a builtin was
never consulted at all: `Anthropic` resolved to the native `anthropic`
provider, `initializeAnthropic` read `ANTHROPIC_API_KEY` from the
environment, and a deployment that routes everything through a gateway threw
"Anthropic API key not provided" before making any call. The endpoint's own
`baseURL`, `apiKey` and headers were never read.

Four names are affected — `anthropic`, `google`, `bedrock` and `vertexai`,
case-insensitively. `xai`, `deepseek`, `moonshot` and `openrouter` fold onto
builtins too, but those map to `initializeCustom` and recover their config
through the known-custom-provider branch below, so they behave correctly
today. `openAI` and `azureOpenAI` never collide: the map keys are camelCase,
so no lowercased name equals them.

The custom lookup now runs first when the exact-case lookup misses, and the
case-folded native lookup becomes the fallback. Three properties are kept:

- An exact-case hit still wins outright. The agent flow re-enters this
  function with normalized lowercase enum values, so a genuine native
  provider is unaffected — including when a custom endpoint shares its name.
- A CamelCase known custom provider keeps its own normalized
  `overrideProvider` rather than collapsing to `openai`, since that value
  selects the token and context-window maps.
- `provider: anthropic` on the endpoint still routes to the native
  `/v1/messages` wire format, now reachable for an endpoint named
  `Anthropic` as well.

Deliberate behaviour change: naming a custom endpoint `Google` and expecting
it to inherit the native Google client now gets the custom client instead. An
explicitly declared endpoint should never be silently discarded, and
`provider:` is the supported way to ask for a native wire format.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@sindriii

Copy link
Copy Markdown
Author

⏸️ Deferred: Claude and Gemini render as sender AI on custom endpoints

Status: not caused by this PR, not blocking it. Parked deliberately — recorded here so it isn't rediscovered from scratch.

Observed on apro-sandbox, image vitinn/librechat-custom:v0.8.7-9679baa4b. Talking to three models in one session: GPT models are attributed GPT-5.6, Kimi Kimi, but every Claude response is attributed AI.

Root cause

getResponseSender — packages/data-provider/src/parsers.ts:284-308, the branch for custom endpoints — resolves the message sender by sniffing the model id:

modelLabel → chatGptLabel → omni → mistral/codestral → deepseek
  → kimi → moonshot → gpt- → modelDisplayLabel → 'AI'

There is no claude branch and no gemini branch, so both fall through to return 'AI' at :307.

Verified by running the real function against our declared model lists:

model matches sender
gpt-5.6-terra includes('gpt-') → extractGPTVersion GPT-5.6
kimi-k2.6-azure includes('kimi') Kimi
deepseek-v4-pro includes('deepseek') Deepseek
claude-opus-5 — AI
gemini-3.7-flash — AI

Gemini is affected too — not yet reported, same as the endpoint-name collision.

Why it only bites us

The native branches earlier in the same function do handle these:

:266   endpoint === anthropic  →  modelLabel || 'Claude'
:274   endpoint === google     →  'Gemma' | 'Gemini'

Our Anthropic and Google endpoints are custom endpoints over LiteLLM, so they take the custom branch instead. Every vendor sniff in that branch is one with no native endpoint — added one PR at a time as each was needed:

sniff added PR
codestral 2025-04-07 #6775
deepseek 2025-04-29 #7132
kimi, moonshot 2026-02-04 #11621

Anthropic and Google never needed the patch because they had native endpoints. Serving them through a gateway is what exposes the gap.

Contributing factor on our side: none of our four endpoints set modelDisplayLabel, which is the last fallback before 'AI'. OpenAI and Other survive only because their model ids happen to match a sniff.

Note on this PR's modelLabels

modelLabel in getResponseSender is the per-conversation preset label, not the modelLabels map this PR adds. modelLabels is picker-display only and never reaches the sender, so labelling claude-opus-4-8 as "Opus 4.8" does not change the attribution — it stays AI. Whether it should feed through is a design question, deliberately out of scope here.

Options when we pick this up

  1. Config only, no code. Set modelDisplayLabel: 'Claude' / 'Gemini' on those endpoints in vitinn-infra. One deploy, no image rebuild. Every Claude model then reads "Claude" rather than "Opus 5".
  2. Upstream, ~4 lines. Add claude and gemini to the custom chain, mirroring the native branches. Upstream has accepted this same patch three times for other vendors.
  3. Wire modelLabels into the sender. Largest change; moves this feature from display-only into stored message data.

Current lean: (1) to unblock, (2) as its own upstream PR. Not started.


📌 Upstream status of the commits on this branch

commit upstream status
9679baa4b endpoint-name collision danny-avila/LibreChat#15310 closed — permanently fork-only
modelLabels (2 commits) danny-avila/LibreChat#15311 draft, open
models.filter + withhold — not yet submitted

The collision fix cannot be upstreamed. Discussion danny-avila/LibreChat#6766 has danny-avila ruling the behaviour intentional: "'Google' and 'Anthropic' are reserved names ... you can no longer use these for your custom endpoint names." PR #15310 argued the opposite precedence and was closed. We need the endpoint named Anthropic, so the patch stays here indefinitely.

Practical notes for future rebases:

  • Keep 9679baa4b isolated to providers.ts + its spec so re-applying stays trivial.
  • An upstream-based twin of the same fix exists as c0540d4a5 on the fork branch fix/custom-endpoint-name-collision — that's the version to re-apply when rebasing onto newer upstream, since upstream's resolveCustomEndpointSecrets work (#14510) changed the surrounding block.
  • Upstream promised a reserved-name warning in #6766 (April 2025) that still does not exist anywhere in main as of 2026-08-28. If it ever lands, expect a conflict and possibly a need to suppress it for these names.

sindriii and others added 8 commits August 28, 2026 13:37
`fetchModels` catches every axios error and returns `[]`, so the loader's
fallback for a fetch that never answered was unreachable: a refused
connection, a 5s timeout and a 500 all arrived as a fulfilled empty list.
Under `filter` that intersects to nothing, and an endpoint disappears for
the length of a gateway blip — the opposite of the intended fallback.

Adds an opt-in `throwOnError` so a failed fetch rejects instead. Only the
config loader passes it; every other caller keeps the swallowing default.
For an endpoint that does not filter the outcome is unchanged, since a
rejection and an empty answer both fall back to the declared list.

Also logs the declared models the gateway did not offer. Dropping them is
by design, but a typo and a retired model look identical from the picker.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`filter` is opt-in, but three of the changes around it were not, and a
deployment that never sets it was paying for all three.

`/api/endpoints` resolved the whole models config on every call — live
fetches to OpenAI, Anthropic, Azure, Google, Bedrock and every gateway —
on a route that had always been a cached config read, and on the first
page-load path. It now derives the filter-managed endpoints from the app
config first and returns early when there are none, so the resolution
happens only where an empty list is actually reachable. The route also
gains `configMiddleware`, matching `/token-config` beside it, which
supplies `req.config` and drops a duplicate `getAppConfig`.

`withholdEmptyEndpoints` now takes that set and considers nothing else,
and `validateModel`'s exemption from the illegal-model violation is gated
the same way, so endpoints that do not filter keep the violation they
have always logged.

Dropping `.min(1)` from `models.default` likewise applied to every custom
endpoint. It is now conditional: an empty list is accepted under `filter`,
where it is an endpoint template a deployment fills in, and still rejected
otherwise — where it validates and then silently drops the endpoint.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`validateModel` gained an exemption from the illegal-model violation for
an endpoint with nothing to serve, but it only runs on the assistants
routes — it is commented out on the agents route. The agents path reaches
`validateAgentModel`, which had no such exemption and still called
`logViolation`, and `logViolation` calls `banViolation`.

So the case the exemption was written for was the one path it did not
cover: a stored agent whose filtered endpoint empties out while its
gateway is unreachable, whose owner then accrues violations on every
retry, and is banned where `BAN_VIOLATIONS` is on.

Returns `ENDPOINT_MODELS_NOT_LOADED` instead, which the client already
renders. Gated on the endpoint filtering, so every other endpoint keeps
logging the violation it always did.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The component tests for the label fallback mocked ~/utils and so asserted
a reimplementation of the helper rather than the helper: AgentConfig.test
existed only for that, and ModelPanel.test carried a duplicate. With the
one-liner inlined at its two call sites, getModelDisplayName is gone and
the remaining ModelPanel test exercises the real lookup.

AddedConvo and Favorites labelling moves to a follow-up — neither is one
of the five sites this PR is about.

Also folds the schema/passthrough specs down to the cases that can fail
(parse-and-keep, non-string rejection, config passthrough), merges the
SearchResults and EndpointModelItem scenarios that rendered identical
setups, and reads isGlobal off the model already in hand instead of
re-finding it in endpoint.models.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Mirror the diet applied to the upstream PR: require at least one declared
model again (empty-default endpoint templates are dropped), fold the shared
availability predicate back inline, delete the unused `loadModels` alias,
and remove tests that duplicated coverage or exercised the schema library.
The fork-only behaviour is untouched: model labels, and the loader's
dead-gateway/empty-answer distinction that keeps a per-user empty catalog
authoritative under `filter`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`models.filter: "complement"` serves the endpoint's declared models first,
then every fetched model no endpoint on the same gateway fetch declares.
Lets one endpoint pick up models that exist in the gateway but are curated
nowhere, while its siblings stay curated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@sindriii
sindriii force-pushed the feat/model-availability-filter branch from 2228575 to 8c023da Compare September 7, 2026 17:49
`fc7060159` restored upstream's unconditional `min(1)` on a custom
endpoint's `models.default` while trimming this feature to the shape of the
upstream PR. Upstream can require a declared model: without `filter`,
`default` *is* the served list, so an empty one is an endpoint that serves
nothing and never can. Under `filter` the same empty list is meaningful — it
declares an endpoint that shows only what a deployment actually has, and
nothing until it has something. `withholdEmptyEndpoints` then drops it from
the picker, which is the intended outcome, not a failure.

Restore the conditional rule the trim replaced: an empty list is rejected
unless the endpoint filters. `Boolean(models.filter)` also admits
`'complement'`, matching how `filterManagedEndpoints` tests the same field.

Without this, a base config shipping an empty endpoint template fails
`configSchema.strict()` at load and `loadCustomConfig` exits the process, so
every deployment inheriting the template crash-loops on startup. Verified
against the 22 rendered aprochat tenant configs: 7 declare no models for
their `Other` endpoint and all 7 are rejected before this change, accepted
after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
sindriii added a commit that referenced this pull request Sep 7, 2026
…re model labels

Squash of feat/model-availability-filter (PR #72) at 9137621 onto apro-deploy.

- serve a custom endpoint's declared models intersected with the fetched catalog (models.filter)
- models.filter: complement additionally serves every fetched model no sibling endpoint declares
- let a filtering endpoint declare an empty model list
- withhold custom endpoints with no models available to the request
- let a custom endpoint declare display labels for its models (modelLabels)
- let a declared custom endpoint outrank a case-folded builtin
- do not ban an agent's owner when its endpoint stops serving

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant