Skip to content

[Bug]: Devin catalog drops supportsImages, leaving vision-capable models without inputModalities #4530

Description

@thisisjun786

Client or integration

OpenCodex dashboard

Area

Catalog / models

Summary

Devin's GetCascadeModelConfigs response includes ClientModelConfig.supportsImages (protobuf field 5), but OpenCodex does not parse or propagate this field. As a result, ocx models live --provider devin --json and /api/models omit inputModalities even for models that the signed-in account's upstream catalog explicitly marks as image-capable.

In a live catalog read on 2026-09-14, 190 available model variants collapsed to 42 base IDs. All variants of 36 base IDs advertised image support, including kimi-k3, swe-2, deepseek-v4-1-flash, and glm-5-3-flash. Before the local workaround, all 42 OpenCodex rows omitted inputModalities.

Expected: preserve upstream image capability during discovery and expose inputModalities: ["text", "image"] for confirmed image-capable models. Keep false/unknown handling explicit and handle effort variants conservatively. Do not require users to maintain a manual vision allowlist for a capability already supplied by the provider.

Reproduction

  1. Use OpenCodex 2.52.0 with an authenticated devin provider and liveModels: true, without a providers.devin.modelInputModalities override.
  2. Run ocx models live --provider devin --json and inspect kimi-k3 or swe-2. inputModalities is absent.
  3. Compare with the same account's GetCascadeModelConfigs response: field 5 is true for these models, including every available effort variant observed.
  4. The dropped-field behavior can also be reproduced without credentials or inference, from a checkout with Bun:
bun --eval '
import { parseCatalogBuffer } from "./src/adapters/devin/cloud-direct/catalog.ts";
import { encodeMessage, encodeString, encodeVarintField } from "./src/adapters/devin/cloud-direct/wire.ts";
const row = Buffer.concat([
  encodeString(1, "Vision fixture"),
  encodeVarintField(5, 1),
  encodeVarintField(18, 1048576),
  encodeString(22, "kimi-k3"),
]);
const entry = parseCatalogBuffer(encodeMessage(1, row), "unused-fixture", "https://server.codeium.com").byUid.get("kimi-k3");
console.log(JSON.stringify(entry));
console.log("retainsSupportsImages=" + Object.hasOwn(entry, "supportsImages"));
'

The fixture contains supportsImages=true, but the parsed entry loses it. This fixture tests parsing only; the real provider capability values above came from a separate live catalog read.

Version

2.52.0 (@bitkyc08/opencodex). The same missing field was also confirmed in upstream dev at df7dc1be530141c55bd10e70ebe79cb80914d98e.

Operating system

Ubuntu 26.04.1 LTS, Linux x86_64

Provider and model

devin; examples: kimi-k3, swe-2, gemini-3-8-flash, deepseek-v4-1-flash, glm-5-3-flash.

Logs or error output

Before workaround:
  Devin base models: 42
  Rows with image in inputModalities: 0

Upstream GetCascadeModelConfigs:
  Available variants: 190
  Base models with supportsImages=true for every variant: 36
  Other base models: 6

Examples (upstream -> OpenCodex before workaround):
  kimi-k3:             true on 3/3 variants -> inputModalities absent
  swe-2:               true on 3/3 variants -> inputModalities absent
  gemini-3-8-flash:     true on 3/3 variants -> inputModalities absent
  deepseek-v4-1-flash:  true on 2/2 variants -> inputModalities absent
  glm-5-3-flash:        true on 3/3 variants -> inputModalities absent

Screenshots and supporting files

Source trace at the inspected upstream revision:

  • catalog.ts:116: parseCatalogBuffer reads fields 1, 4, 18, and 22, but omits field 5. ModelCatalogEntry has no image capability property.
  • live-models.ts: fetchDevinUsableModels returns model IDs and context windows, with no modality map.
  • provider-fetch.ts: the Devin discovery branch consequently has no upstream image metadata to propagate.

Devin's adapter already encodes image parts as ChatMessagePrompt.images (field 10); this report concerns capability discovery, not a missing image transport implementation.

Local workaround verified: add ["text", "image"] overrides for the 36 upstream-confirmed base IDs, then run ocx sync. Afterward the live Devin catalog exposes image input on 36/42 rows, and the generated Codex catalog exposes text,image for the selected Devin models. The six other models were left unchanged. No image inference requests were made, so this is catalog/configuration evidence, not an end-to-end vision benchmark.

Redacted configuration

Minimal provider shape before the workaround:

{
  "providers": {
    "devin": {
      "adapter": "devin",
      "authMode": "oauth",
      "baseUrl": "https://server.codeium.com",
      "liveModels": true,
      "defaultModel": "swe-2"
    }
  }
}

Example of the supported local workaround, applied only to confirmed model IDs:

{
  "providers": {
    "devin": {
      "modelInputModalities": {
        "kimi-k3": ["text", "image"],
        "swe-2": ["text", "image"]
      }
    }
  }
}

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Activity

  1. added
    bugSomething isn't working
    catalogModel catalog, slugs, visibility, routed entries
    on Sep 13, 2026
  2. lidge-jun commented on Sep 13, 2026

    @lidge-jun
    Owner

    리뷰 · 우선순위 70 / 80

    설명
    이 이슈는 Devin(Cognition) 라이브 카탈로그가 이미 ClientModelConfig.supportsImages(protobuf 필드 5)로 이미지 가능 여부를 알려 주는데, OpenCodex가 그 필드를 읽지도 옮기지도 않아서 ocx models live --provider devin와 /api/models 행에 inputModalities가 비는 버그입니다. 현재 dev HEAD(df7dc1be5, #4520 문서 팁)에서도 그대로입니다.

    초등학생도 따라갈 수 있게 풀면 이렇습니다. Devin 계정으로 로그인한 뒤 라이브 모델을 물어보면, 클라우드가 GetCascadeModelConfigs로 “이 계정에 쓸 수 있는 모델 목록”을 줍니다. 그 목록 한 줄에는 이름(필드 1), 사용 불가 여부(필드 4), 컨텍스트 크기(필드 18), 모델 UID(필드 22)뿐 아니라 “이미지를 받을 수 있나”(필드 5)도 들어 있습니다. 그런데 OpenCodex 파서는 1·4·18·22만 읽고 5는 버립니다. 그래서 업스트림이 kimi-k3, swe-2 같은 모델에 이미지를 켠다고 표시해도, OpenCodex 카탈로그에는 글자만 받는 모델처럼 보입니다.

    코드로 보면 구멍이 세 층입니다. src/adapters/devin/cloud-direct/catalog.ts의 ModelCatalogEntry에 이미지 필드가 없고, parseCatalogBuffer(대략 116–130줄)도 필드 5를 무시합니다. src/adapters/devin/live-models.ts의 fetchDevinUsableModels는 쓸 수 있는 베이스 ID·컨텍스트·reasoning ladder만 돌려주고 모달리티 맵은 없습니다. src/codex/catalog/provider-fetch.ts Devin 분기(대략 1706–1756줄)도 라이브 결과에 inputModalities를 붙일 재료가 없어서, 설정 힌트(modelInputModalities/modelCapabilities)가 없으면 행이 비어 나갑니다. 이슈에 적힌 Bun 픽스처로도 supportsImages=true가 파싱 후 사라지는 걸 재현할 수 있습니다.

    왜 지금 더 아프냐면, 와이어 쪽 이미지 전송은 이미 #4513(그리고 문서 마감 #4518)으로 고쳤기 때문입니다. src/adapters/devin.ts는 data URL 이미지를 ChatMessagePrompt 필드 10으로 넣을 수 있습니다. 그런데 카탈로그가 “이미지 가능”을 광고하지 않으면 대시보드·비전 후보(src/vision/eligibility.ts)·콤보 경로가 그 모델을 텍스트만으로 취급할 수 있습니다. 전송 파이프는 있는데 문패가 꺼져 있는 상태입니다. 보고자가 말한 대로 수동 modelInputModalities로 36개 베이스 ID를 메우면 카탈로그는 고쳐 보이지만, 업스트림이 이미 주는 값을 사람이 다시 적는 우회입니다.

    관련 이슈와의 경계도 분명합니다. #4505는 OpenCode Go / CommandCode 쪽 DeepSeek·GLM 플래시 모달리티 구멍이고, #4501은 네이티브 행에서 운영자 modelCapabilities를 존중하는 이야기입니다. 이 이슈는 Devin 라이브 카탈로그 파서·전파만의 문제입니다. effort 변형이 베이스로 접힐 때 “모든 변형이 true일 때만 이미지로 올릴지, 하나라도 true면 올릴지”만 정책으로 정하면 됩니다. 보고 관측으로는 36개 베이스가 변형 전부 true였고, 나머지 6개는 그대로 두는 보수적 처리가 맞습니다.

    src/adapters/devin/cloud-direct/catalog.ts의 parseCatalogBuffer - ClientModelConfig 필드 5(supportsImages)를 읽지 않아 파싱 단계에서 능력이 사라진다
    src/adapters/devin/cloud-direct/catalog.ts의 ModelCatalogEntry - 이미지 능력 속성이 없어 캐시/조회 계층에 옮길 자리가 없다
    src/adapters/devin/live-models.ts의 fetchDevinUsableModels - 반환이 models/contextWindows/efforts뿐이라 베이스별 모달리티를 만들 수 없다
    src/codex/catalog/provider-fetch.ts Devin 분기 - 라이브 행에 inputModalities: ["text","image"]를 심을 업스트림 입력이 없다
    설정 providers.devin.modelInputModalities / modelCapabilities - 우회는 되지만 업스트림 필드 5를 대체하는 수동 목록이라 유지비가 생긴다

    메인테이너의 판단이 필요한 지점

    • effort 변형을 베이스로 접을 때: 모든 변형이 supportsImages=true일 때만 ["text","image"]로 올릴지, 하나라도 true면 올릴지 (보고 데이터는 전부 true)
    • 필드 5가 없거나 false인 모델을 ["text"]로 명시할지, 필드를 아예 생략해 기존 “모름” 동작과 맞출지
    • 운영자 modelInputModalities 오버라이드가 라이브 값보다 항상 우선인지 (지금은 힌트 경로가 후순위/병합이라 확인 필요)
    • feat(devin): pass user and tool-result images to the wire #4513 와이어 수정과 같은 릴리즈 노트에 묶을지, 카탈로그 전용 작은 PR로 따로 닫을지

    너의 추천
    ModelCatalogEntry에 supportsImages?: boolean을 추가하고 parseCatalogBuffer에서 필드 5를 읽으세요. fetchDevinUsableModels가 베이스별로 보수적으로 모달리티 맵을 만든 뒤 provider-fetch Devin 분기에 inputModalities를 붙이세요. 이슈 본문 Bun 픽스처를 단위 테스트로 넣고, 라이브 없이도 필드 5 보존을 잠그세요. 수동 allowlist PR은 열지 말고 이 이슈만 닫히게 하세요.

    이 댓글은 grok-bot이 작성했습니다

  3. lidge-jun commented on Sep 13, 2026

    @lidge-jun
    Owner

    Fixed on dev at 72335fc6ac, end to end rather than at the parser alone.

    parseCatalogBuffer now has a field-5 varint arm and carries supportsImages on ModelCatalogEntry as a genuine tri-state (src/adapters/devin/cloud-direct/catalog.ts:93, 139, 156). Absence stays undefined rather than collapsing to false, which is the #1796 precedent — a model the vendor simply did not describe must not be asserted text-only.

    fetchDevinUsableModels then votes per base over only the rows that actually asserted the field (src/adapters/devin/live-models.ts:169): unanimous true advertises ["text","image"], unanimous false advertises ["text"], and mixed or absent advertises nothing. That asymmetry is deliberate. One unsuffixed row with no flag must not poison a base whose measured variants agree, and one measured false must not be outvoted by its siblings.

    provider-fetch.ts spreads the live modalities before catalogHintsFromProviderConfig, so exact modelCapabilities, the legacy modelInputModalities record, and the sidecar-consumer rewrite all keep precedence over discovery. A model that asserts field 5 true on every variant now reaches the client advertised with image input.

    Landed as #4547 (6329f30389e0653af0e592669f508407f3bd7691, CI run 34775751844) and #4556 (72335fc6ac5fc69e150c7ceddfd1fc8f6ea03b0e, CI run 34782050873), each verified at its exact head.

    One routing detail is documented rather than silently assumed: resolveWireModelUid prefers the plain UID when it is listed, so a base whose image support was measured on its effort variants can still route a no-effort request to a row that never asserted the field. That is recorded in the code and the structure docs as a product decision, not an oversight.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingcatalogModel catalog, slugs, visibility, routed entries

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions