Skip to content

[Bug]: OpenCode 2 custom providers / models omitted during boot probe race in makeOpenCode2ModelLoader #15155

Description

@hardik88t

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. Configure custom providers or OpenAI-compatible proxy models in OpenCode v2 configuration (~/.config/opencode/opencode.jsonc), for example:
    {
      "$schema": "https://opencode.ai/config.json",
      "provider": {
        "proxy": {
          "npm": "@opencode/ai/providers/openai-compatible",
          "name": "Proxy Gateway",
          "options": {
            "baseURL": "http://127.0.0.1:8317/v1",
            "apiKey": "..."
          },
          "models": {
            "claude-sonnet-4-6": { "name": "Claude Sonnet 4.6" },
            "gpt-6-luna": { "name": "GPT 6 Luna" }
          }
        }
      }
    }
  2. Run opencode models in the terminal to verify the custom models are listed and active.
  3. Start T3 Code (using the default managed OpenCode 2 server with no serverUrl configured).
  4. Open the composer model picker or navigate to Settings → Providers → OpenCode → Models.

Expected behavior

All models configured in OpenCode (including custom providers and plugins) are discovered and listed in T3 Code's model picker and provider settings without requiring manual entry into customModels.

Actual behavior

Only the 29 built-in opencode-go models are discovered. Custom providers and proxy models defined in OpenCode config are missing.

Root Cause Analysis

In apps/server/src/provider/Layers/OpenCodeProvider.ts (inside makeOpenCode2ModelLoader):

const makeOpenCode2ModelLoader = (list) => sync(() => {
  let lastLoaded = [];
  return list.pipe(
    repeat({
      until: (models) => models.length > 0, // <--- Stops prematurely on first non-empty batch
      schedule: spaced("250 millis")
    }),
    timeoutOption("5 seconds"),
    map((loaded) => {
      if (loaded._tag === "Some" && loaded.value.length > 0) lastLoaded = loaded.value;
      return lastLoaded;
    })
  );
});

In OpenCode v2, provider plugins initialize asynchronously:

  • OpenCode's own OpenAPI specification for /api/model explicitly documents:

    "Retrieve the current snapshot of available models ordered by release date. The snapshot may precede initial plugin settlement."

  • When T3 Code spawns opencode serve, /api/model returns the 29 built-in opencode-go models within ~450ms.
  • Because models.length > 0 evaluates to true (29 > 0), makeOpenCode2ModelLoader immediately terminates polling on the first retry.
  • The external/custom providers finish plugin settlement ~400ms later (~850ms–900ms after boot), but T3 Code has already cached the incomplete 29-model snapshot and does not refresh.

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261003.2610 (commit 8ed276c on main)

Environment

Linux x86_64, T3 Code Desktop 0.0.46-nightly, OpenCode 2.0.22

Logs or stack traces

Timing breakdown of opencode serve startup vs /api/model:
  +0.20s: 0 models
  +0.48s: 29 models (only 'opencode-go' built-ins; repeat condition until: models.length > 0 exits here)
  +0.89s: 57 models (all custom providers and plugins settle)

Workaround

Either:

  1. Manually specify every model slug in settings.json under providerInstances.opencode.config.customModels.
  2. Or start opencode serve independently and configure T3 Code to connect via serverUrl, bypassing the boot probe race.

Activity

  1. hardik88t commented on Oct 3, 2026

    @hardik88t
    Author

    #14962 might be related to this

  2. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @hardik88t for the timing breakdown and for tracking down the OpenCode API note! I confirmed this on current main (18b21325). Custom OpenCode providers get dropped because the startup catalog probe treats the first non-empty model list as complete.

    What happens

    makeOpenCode2ModelLoader in apps/server/src/provider/Layers/OpenCodeProvider.ts polls GET /api/model every 250ms and stops at until: (models) => models.length > 0, with a 5s cap. The comment above it only talks about waiting out an empty catalog. The tests cover "empty, then one model" and "keep the last list if a later read is empty", but not a list that keeps growing after the first models appear. OpenCode's list-models API says the snapshot may come before plugins settle, which matches the 29 built-in opencode-go models arriving before your custom providers.

    That partial list is what the picker keeps. The OpenCode driver (apps/server/src/provider/Drivers/OpenCodeDriver.ts) probes once at startup. refreshOnInterval is false, and changing settings doesn't re-check the provider. Opening a project only refreshes that directory's skills and commands. A ready probe counts as successful discovery, so the shorter list also replaces any fuller catalog already in the provider status cache.

    Settings → refresh runs the same loader again, but it isn't a reliable fix. The owned server stops after 30s idle, so the next check spawns a fresh opencode serve and can stop on the built-ins again.

    Workarounds

    Both of yours hold. customModels gets merged in after the probe, and an external serverUrl that has already settled returns the full list on the first read.

    I didn't find an existing issue for this boot-probe race. The stop condition hasn't changed since 8ed276c (0.0.46-nightly.20261003). A maintainer will decide on the fix direction.

  3. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 3, 2026
  4. p7gg commented on Oct 3, 2026

    @p7gg

    Additional, distinct failure path for the same symptom (custom / plugin providers missing from the picker), found while debugging commandcode + litellm:

    opencode models --verbose was removed in OpenCode v2

    loadInventoryFromCli in apps/server/dist builds the OpenCode inventory by shelling out to:

    const runModelsCli = () => runOpenCodeCommand({
      args: ["models", "--verbose"],   // ← flag no longer exists in OpenCode v2
      ...
    })

    and parses the result with parseModelsCliOutput, which expects each model as a slug line followed by a JSON block (the old --verbose format), deriving connected from providers whose JSON parsed. On OpenCode 2.0.22 that flag is rejected:

    $ opencode models --verbose
    ERROR
      Unrecognized flag: --verbose in command opencode models
    $ echo $?
    1

    So no JSON blocks are produced, the locally-registered providers never enter connected, and flattenOpenCodeModels filters them out (if (!connected.has(provider.id)) continue).

    Evidence

    OpenCode itself is healthy and exposes everything:

    $ opencode models | cut -d/ -f1 | sort | uniq -c
         75 commandcode/
         10 litellm/
         29 opencode-go/
         81 opencode/
    GET /api/provider?directory=<dir>  → [opencode-go, opencode, commandcode, litellm]
    GET /api/model?directory=<dir>     → 195 models, 75 with providerID=commandcode
    

    But the picker shows 110 (= 81 opencode + 29 opencode-go) — i.e. exactly the 85 models from the locally-registered providers are dropped.

    Relationship to the boot-probe race in this issue

    These look like two independent causes producing the same symptom:

    1. The boot-probe race you describe (first non-empty snapshot wins) — affects /api/model timing.
    2. This one — the CLI inventory path depends on a removed --verbose flag, so plugin/config providers are dropped even when settlement is fine.

    Note (2) also survives a full T3 restart, since it re-runs the same broken command every time — consistent with the report that a reboot doesn't help.

    Suggested fix

    Don't depend on the removed --verbose flag. Prefer the HTTP API (/api/provider + /api/model), or use plain opencode models and resolve metadata elsewhere, gated by OpenCode version. Adding a fallback so a CLI-contract change degrades gracefully would also prevent silent provider loss.

    Environment

    • T3 Code 0.0.46-nightly.20261003.2632, OpenCode 2.0.22, orchestrator v2
    • Backend in WSL (wslOnly: true), archlinux
    • Providers: commandcode (OpenCode v2 plugin registered via ~/.config/opencode/opencode.json → "plugins": ["./plugin/commandcode"]), litellm (inline in opencode.json)
  5. yusufameri commented on Oct 4, 2026

    @yusufameri

    Confirming the race on macOS with plain inline config providers (no plugins), against a managed server:

    • T3 Code 0.0.46-nightly.20261003.2638, OpenCode 2.0.22, no serverUrl.
    • ~/.config/opencode/opencode.jsonc defines two custom providers: one @ai-sdk/openai-compatible (75 models) and one @ai-sdk/anthropic (9 models).
    • The managed server's GET /api/model returns 123 models (75 + 9 + 29 opencode-go + 10 opencode).
    • T3's live provider catalog holds 114 models and none from the @ai-sdk/anthropic provider, i.e. exactly the snapshot where the first provider had settled and the second had not.
    • The server has been up for 12+ minutes and reports all 123, but the driver's refreshOnInterval: false means the picker keeps the partial list.

    So this also hits inline config providers, and the stale snapshot outlives the race window on a long-lived settled server.

  6. jerome-marshall-ntx commented on Oct 4, 2026

    @jerome-marshall-ntx

    Independent reproduction confirming the root cause, with a case where the first non-empty read is not even close to complete.

    I hit this on OpenCode 2.0.22 with T3 Code nightly (0.0.46-nightly.20261003.2638 and 0.0.46-nightly.20261004.2644) across several machines, and the symptom was exactly the picker showing the 29 built-in opencode-go models while a custom OpenAI-compatible provider stayed missing even though the OpenCode TUI listed it fine.

    The first non-empty read can be a tiny fraction of the final catalog

    Polling a freshly-spawned opencode serve every 300 ms (same env T3 uses):

       0ms  0  {}
     600ms  29 {'opencode-go': 29}
     900ms  43 {'opencode-go': 29, 'opencode': 10, '<custom>': 4}
    1200ms+ 43 {…}
    

    That machine's final catalog was only 43, so "stop at 29" loses one custom provider. But on another machine the same race is much worse — OpenCode exposes more providers (cloud built-ins plus the custom one), and the catalog kept growing well past the 5 s cap:

       0ms   0
       6s    0
       8s  264   ← first non-empty read: built-ins only
       ...    ...
     settled 207+ across 8 providerIDs, and still climbing
    

    So until: (models) => models.length > 0 is not just "off by one provider" — the first non-empty read can be a small early slice, and the 5 s timeoutOption doesn't help because polling has already stopped. The reporter's ~450 ms → ~900 ms timeline is one instance of a general "first non-empty ≠ settled" problem.

    Direct evidence from T3's own trace

    The trace shows T3 making the exact call and reading the partial value:

    GET http://127.0.0.1:<port>/api/model?location[directory]=/<home>
        -> 200 len=60      (empty)
    GET http://127.0.0.1:<port>/api/model?location[directory]=/<home>
        -> 200 len=2096    (~1 s later, still the built-in-only set)
    

    and checkOpenCode2 exits Success with that shorter list, which then replaces the provider snapshot.

    A second shape of the same bug: persisted empty list

    Because the loader only caches a > 0 read (lastLoaded starts [] and is only written when loaded.value.length > 0), a server whose catalog is still empty at the 5 s timeout leaves lastLoaded empty, and the provider probe can end up with zero models. I saw this as "OpenCode 2 is running, but it did not list any models yet." even though opencode api model.list against the same server returned a full catalog seconds later. Worth covering in the fix: "never got a non-empty read in the window" should be an error/retry state, not a persisted empty snapshot.

    Workarounds that both held for me (matches the triage note)

    1. providerInstances.opencode.config.customModels — merged in after the probe.
    2. External serverUrl pointed at an already-warm opencode serve — returns the full catalog on the first read. This is what I standardized on: a persistent opencode serve on a fixed port (systemd user unit on Linux, LaunchAgent on macOS) plus serverUrl/serverPassword, which makes the race disappear entirely.

    Environment

    • opencode: 2.0.22 (curl v2 install, ~/.opencode/bin/opencode)
    • T3 Code: 0.0.46-nightly.20261003.2638 and 0.0.46-nightly.20261004.2644
    • OS: Linux (Rocky 8.10, Ubuntu 26.04), macOS 26
    • Reproducible consistently on the default managed server; disappears with a warm external serverUrl.
  7. kelchm commented on Oct 4, 2026

    @kelchm

    I reproduced a functional variant of this startup race: discovery caches providers that enabled_providers later disables, which blocks working models and advertises unusable ones.

    Recorded install: T3 Code 0.0.46-nightly.20261003.2638, OpenCode 2.0.22, macOS. Local verification: macOS 26.5.2 arm64; upstream main 4ee6bfd, #15202 diff at 2a963b6, #15510 diff at e555f5f.

    Repro with a managed server (no serverUrl):

    1. Connect OpenCode Console/OpenRouter, configure Ollama, and set enabled_providers to ["opencode-go"]. Use fresh discovery state.
    2. Confirm warm opencode models lists only 29 Go models. On the reported install, CLI turns using opencode-go/kimi-k3 and opencode-go/glm-5.3 both returned subagent-ok.
    3. Delegate to opencode-go/kimi-k3 in T3: rejected before child creation with "Model opencode-go/kimi-k3 is not advertised by provider opencode."
    4. Delegate to T3's advertised opencode/kimi-k3: a child starts, then fails with "Model unavailable: opencode/kimi-k3". The reported install showed the same pair of failures for GLM 5.3; restarting T3's OpenCode servers did not repair its catalog.

    I independently reproduced both Kimi failures on upstream main in an isolated headless backend harness. Its actual loader accepted 473 models (opencode 81, openrouter 390, ollama 2), with zero Go models. With either PR's diff, discovery returned the correct 29 and a T3-owned delegated Kimi task completed with subagent-ok. This exercises the real delegation service, orchestrator and OpenCode 2 adapter; the parent and provider-registry facade are test fixtures, and no UI was tested. Only credentials/config were copied into disposable XDG directories; no live sessions were copied.

    A separate 50 ms catalog trace measured: 0 at 179 ms → 471 at 342 ms → 473 at 410 ms → 502 at 692 ms → 29 at 761 ms. integration.list completed at 738 ms; the directory's model.updated arrived at 757 ms. Reads immediately after either returned 29. Both PRs also retained three custom models with a controlled 800 ms plugin-activation delay, which main omitted.

  8. nkoynov commented on Oct 5, 2026

    @nkoynov
    Contributor

    Another way this shows up: delegating to an OpenCode plugin model from a Cursor parent. I reproduced it on main @ e22c880 with OpenCode 2.0.22 and cursor-opencode-provider@0.8.0, using a managed server in a fresh HOME.

    Model: Claude Opus 5.5 (1M). Harness: Claude Code in T3 Code.

  9. maxvandongenTikkie commented on Oct 7, 2026

    @maxvandongenTikkie

    I'm having a similar problem in makeOpenCode2ModelLoader, from a different cause, with one extra twist that (perhaps) matters for the fix.

    My setup: macOS, T3 Code Nightly, OpenCode 2.0.20. The only enabled provider is github-copilot (Copilot Enterprise), and my organisation limits which models are allowed. A corporate proxy blocks models.opencode.ai:

    level=ERROR message="Failed to fetch models.dev" cause="... 403 GET https://models.opencode.ai/api.json"
    

    So ~/.cache/opencode/models.json hasn't been updated in months. Once OpenCode is fully running, it gets the correct list from the Copilot Enterprise API, which only includes my organisation's allowed models. Running OpenCode directly shows that list. While the server is still starting, /api/model answers from the outdated saved file instead.

    What T3 saves (~/.t3/caches/opencode.json):

    message: OpenCode 2.0.20 lists 35 models.
    contains github-copilot/gpt-6.1-sol: false   ← allowed and working
    contains github-copilot/gpt-6-astra: true    ← no longer offered, fails with "Model unavailable"
    

    The same managed server, asked a few seconds later:

    GET /api/model?directory=<project>  →  12 models, includes gpt-6.1-sol
    

    How this differs from the original report: here the first answer has more models than the final one (35 → 12), and some of them don't work. So a fix like "keep polling until the model count stops growing" wouldn't help in my case. The early list has to be replaced, not added to. T3's merge step (mergeProviderModels keeping previous models) could also keep the dead models around.

    What doesn't help: restarting T3 and deleting ~/.t3/caches/opencode.json. Every startup runs into the same timing and saves the outdated list again.

    What would fix both cases:

    • After the server starts, check the model list again: on OpenCode's provider/plugin-settled event if it has one, otherwise after a short delay. Replace the saved list with the result instead of merging it.
    • Make sure the refresh/re-check action in provider settings asks the running server again instead of reusing T3's saved list.

    Workaround for now: add the missing slug by hand under customModels.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions