Repository navigation
[Bug]: Devin catalog drops supportsImages, leaving vision-capable models without inputModalities #4530
Description
Activity
- addedbugSomething isn't workingSomething isn't workingcatalogModel catalog, slugs, visibility, routed entriesModel catalog, slugs, visibility, routed entries
on Sep 13, 2026 리뷰 · 우선순위 70 / 80
설명
이 이슈는 Devin(Cognition) 라이브 카탈로그가 이미ClientModelConfig.supportsImages(protobuf 필드 5)로 이미지 가능 여부를 알려 주는데, OpenCodex가 그 필드를 읽지도 옮기지도 않아서ocx models live --provider devin와/api/models행에inputModalities가 비는 버그입니다. 현재devHEAD(df7dc1be5, #4520 문서 팁)에서도 그대로입니다.초등학생도 따라갈 수 있게 풀면 이렇습니다. Devin 계정으로 로그인한 뒤 라이브 모델을 물어보면, 클라우드가
GetCascadeModelConfigs로 “이 계정에 쓸 수 있는 모델 목록”을 줍니다. 그 목록 한 줄에는 이름(필드 1), 사용 불가 여부(필드 4), 컨텍스트 크기(필드 18), 모델 UID(필드 22)뿐 아니라 “이미지를 받을 수 있나”(필드 5)도 들어 있습니다. 그런데 OpenCodex 파서는 1·4·18·22만 읽고 5는 버립니다. 그래서 업스트림이kimi-k3,swe-2같은 모델에 이미지를 켠다고 표시해도, OpenCodex 카탈로그에는 글자만 받는 모델처럼 보입니다.코드로 보면 구멍이 세 층입니다.
src/adapters/devin/cloud-direct/catalog.ts의ModelCatalogEntry에 이미지 필드가 없고,parseCatalogBuffer(대략 116–130줄)도 필드 5를 무시합니다.src/adapters/devin/live-models.ts의fetchDevinUsableModels는 쓸 수 있는 베이스 ID·컨텍스트·reasoning ladder만 돌려주고 모달리티 맵은 없습니다.src/codex/catalog/provider-fetch.tsDevin 분기(대략 1706–1756줄)도 라이브 결과에inputModalities를 붙일 재료가 없어서, 설정 힌트(modelInputModalities/modelCapabilities)가 없으면 행이 비어 나갑니다. 이슈에 적힌 Bun 픽스처로도supportsImages=true가 파싱 후 사라지는 걸 재현할 수 있습니다.왜 지금 더 아프냐면, 와이어 쪽 이미지 전송은 이미 #4513(그리고 문서 마감 #4518)으로 고쳤기 때문입니다.
src/adapters/devin.ts는 data URL 이미지를ChatMessagePrompt필드 10으로 넣을 수 있습니다. 그런데 카탈로그가 “이미지 가능”을 광고하지 않으면 대시보드·비전 후보(src/vision/eligibility.ts)·콤보 경로가 그 모델을 텍스트만으로 취급할 수 있습니다. 전송 파이프는 있는데 문패가 꺼져 있는 상태입니다. 보고자가 말한 대로 수동modelInputModalities로 36개 베이스 ID를 메우면 카탈로그는 고쳐 보이지만, 업스트림이 이미 주는 값을 사람이 다시 적는 우회입니다.관련 이슈와의 경계도 분명합니다. #4505는 OpenCode Go / CommandCode 쪽 DeepSeek·GLM 플래시 모달리티 구멍이고, #4501은 네이티브 행에서 운영자
modelCapabilities를 존중하는 이야기입니다. 이 이슈는 Devin 라이브 카탈로그 파서·전파만의 문제입니다. effort 변형이 베이스로 접힐 때 “모든 변형이 true일 때만 이미지로 올릴지, 하나라도 true면 올릴지”만 정책으로 정하면 됩니다. 보고 관측으로는 36개 베이스가 변형 전부 true였고, 나머지 6개는 그대로 두는 보수적 처리가 맞습니다.src/adapters/devin/cloud-direct/catalog.ts의parseCatalogBuffer- ClientModelConfig 필드 5(supportsImages)를 읽지 않아 파싱 단계에서 능력이 사라진다
src/adapters/devin/cloud-direct/catalog.ts의ModelCatalogEntry- 이미지 능력 속성이 없어 캐시/조회 계층에 옮길 자리가 없다
src/adapters/devin/live-models.ts의fetchDevinUsableModels- 반환이 models/contextWindows/efforts뿐이라 베이스별 모달리티를 만들 수 없다
src/codex/catalog/provider-fetch.tsDevin 분기 - 라이브 행에inputModalities: ["text","image"]를 심을 업스트림 입력이 없다
설정providers.devin.modelInputModalities/modelCapabilities- 우회는 되지만 업스트림 필드 5를 대체하는 수동 목록이라 유지비가 생긴다메인테이너의 판단이 필요한 지점
- effort 변형을 베이스로 접을 때: 모든 변형이
supportsImages=true일 때만["text","image"]로 올릴지, 하나라도 true면 올릴지 (보고 데이터는 전부 true) - 필드 5가 없거나 false인 모델을
["text"]로 명시할지, 필드를 아예 생략해 기존 “모름” 동작과 맞출지 - 운영자
modelInputModalities오버라이드가 라이브 값보다 항상 우선인지 (지금은 힌트 경로가 후순위/병합이라 확인 필요) - feat(devin): pass user and tool-result images to the wire #4513 와이어 수정과 같은 릴리즈 노트에 묶을지, 카탈로그 전용 작은 PR로 따로 닫을지
너의 추천
ModelCatalogEntry에supportsImages?: boolean을 추가하고parseCatalogBuffer에서 필드 5를 읽으세요.fetchDevinUsableModels가 베이스별로 보수적으로 모달리티 맵을 만든 뒤provider-fetchDevin 분기에inputModalities를 붙이세요. 이슈 본문 Bun 픽스처를 단위 테스트로 넣고, 라이브 없이도 필드 5 보존을 잠그세요. 수동 allowlist PR은 열지 말고 이 이슈만 닫히게 하세요.이 댓글은 grok-bot이 작성했습니다
- effort 변형을 베이스로 접을 때: 모든 변형이
Fixed on
devat72335fc6ac, end to end rather than at the parser alone.parseCatalogBuffernow has a field-5 varint arm and carriessupportsImagesonModelCatalogEntryas a genuine tri-state (src/adapters/devin/cloud-direct/catalog.ts:93, 139, 156). Absence staysundefinedrather than collapsing tofalse, which is the #1796 precedent — a model the vendor simply did not describe must not be asserted text-only.fetchDevinUsableModelsthen votes per base over only the rows that actually asserted the field (src/adapters/devin/live-models.ts:169): unanimous true advertises["text","image"], unanimous false advertises["text"], and mixed or absent advertises nothing. That asymmetry is deliberate. One unsuffixed row with no flag must not poison a base whose measured variants agree, and one measuredfalsemust not be outvoted by its siblings.provider-fetch.tsspreads the live modalities beforecatalogHintsFromProviderConfig, so exactmodelCapabilities, the legacymodelInputModalitiesrecord, and the sidecar-consumer rewrite all keep precedence over discovery. A model that asserts field 5 true on every variant now reaches the client advertised with image input.Landed as #4547 (
6329f30389e0653af0e592669f508407f3bd7691, CI run 34775751844) and #4556 (72335fc6ac5fc69e150c7ceddfd1fc8f6ea03b0e, CI run 34782050873), each verified at its exact head.One routing detail is documented rather than silently assumed:
resolveWireModelUidprefers the plain UID when it is listed, so a base whose image support was measured on its effort variants can still route a no-effort request to a row that never asserted the field. That is recorded in the code and the structure docs as a product decision, not an oversight.- added a commit that references this issue
on Sep 13, 2026
Client or integration
OpenCodex dashboard
Area
Catalog / models
Summary
Devin's
GetCascadeModelConfigsresponse includesClientModelConfig.supportsImages(protobuf field 5), but OpenCodex does not parse or propagate this field. As a result,ocx models live --provider devin --jsonand/api/modelsomitinputModalitieseven for models that the signed-in account's upstream catalog explicitly marks as image-capable.In a live catalog read on 2026-09-14, 190 available model variants collapsed to 42 base IDs. All variants of 36 base IDs advertised image support, including
kimi-k3,swe-2,deepseek-v4-1-flash, andglm-5-3-flash. Before the local workaround, all 42 OpenCodex rows omittedinputModalities.Expected: preserve upstream image capability during discovery and expose
inputModalities: ["text", "image"]for confirmed image-capable models. Keep false/unknown handling explicit and handle effort variants conservatively. Do not require users to maintain a manual vision allowlist for a capability already supplied by the provider.Reproduction
devinprovider andliveModels: true, without aproviders.devin.modelInputModalitiesoverride.ocx models live --provider devin --jsonand inspectkimi-k3orswe-2.inputModalitiesis absent.GetCascadeModelConfigsresponse: field 5 istruefor these models, including every available effort variant observed.The fixture contains
supportsImages=true, but the parsed entry loses it. This fixture tests parsing only; the real provider capability values above came from a separate live catalog read.Version
2.52.0 (
@bitkyc08/opencodex). The same missing field was also confirmed in upstreamdevatdf7dc1be530141c55bd10e70ebe79cb80914d98e.Operating system
Ubuntu 26.04.1 LTS, Linux x86_64
Provider and model
devin; examples:kimi-k3,swe-2,gemini-3-8-flash,deepseek-v4-1-flash,glm-5-3-flash.Logs or error output
Screenshots and supporting files
Source trace at the inspected upstream revision:
parseCatalogBufferreads fields 1, 4, 18, and 22, but omits field 5.ModelCatalogEntryhas no image capability property.fetchDevinUsableModelsreturns model IDs and context windows, with no modality map.Devin's adapter already encodes image parts as
ChatMessagePrompt.images(field 10); this report concerns capability discovery, not a missing image transport implementation.Local workaround verified: add
["text", "image"]overrides for the 36 upstream-confirmed base IDs, then runocx sync. Afterward the live Devin catalog exposes image input on 36/42 rows, and the generated Codex catalog exposestext,imagefor the selected Devin models. The six other models were left unchanged. No image inference requests were made, so this is catalog/configuration evidence, not an end-to-end vision benchmark.Redacted configuration
Minimal provider shape before the workaround:
{ "providers": { "devin": { "adapter": "devin", "authMode": "oauth", "baseUrl": "https://server.codeium.com", "liveModels": true, "defaultModel": "swe-2" } } }Example of the supported local workaround, applied only to confirmed model IDs:
{ "providers": { "devin": { "modelInputModalities": { "kimi-k3": ["text", "image"], "swe-2": ["text", "image"] } } } }Checks