Repository navigation
bug: Anthropic image input capabilities missing from integration configurations #4667
Description
Activity
- addedbugSomething isn't workingSomething isn't workingaccount-poolOAuth, credentials, Codex pool, quota, failover, plansOAuth, credentials, Codex pool, quota, failover, planscatalogModel catalog, slugs, visibility, routed entriesModel catalog, slugs, visibility, routed entries
on Sep 14, 2026 - changed the title
[-]bug: Anthropic models export text-only input to Aside, Pi and GJC[/-][+]bug: Anthropic image input capabilities missing from integration configurations[/+]on Sep 15, 2026 - added a commit that references this issue
on Sep 15, 2026 리뷰 · 우선순위 60 / 80
이 이슈는 내장
anthropic/anthropic-apikey정적 카탈로그에modelInputModalities시드가 없어서, 클라이언트로 내보내는 설정이 Claude 를 텍스트 전용(input: ["text"])으로 만든다는 메타데이터 버그다. 지금dev(HEAD9b711073a, 패키지 2.56.0)의src/providers/registry/entries-core.ts에서 두 Anthropic 엔트리는models/modelContextWindows/modelReasoningEfforts는 있지만modelInputModalities키가 없다.model-seeds.ts에도ANTHROPIC_MODEL_INPUT_MODALITIES심볼이 없다.configuredInputModalities가 비면 카탈로그·익스포터가 text-only/unknown 바닥으로 떨어지는 기존 계약과 맞다. Aside 설정에"image"를 수동으로 넣자 첨부가 됐다는 관찰과도 일치한다.이슈 범위는 Aside/Pi/GJC 만이 아니다. capability 를 선언하는 익스포터 전반과, 선언이 있어도 버리는 OpenClaw·Kimi 구멍을 같이 본다. 현재
dev의buildOpenclawClientConfig는 모델에id/name/contextWindow만 넣고input배열을 내지 않는다(OpenClaw 기본이 text-only).buildKimiClientConfig는capabilities: ["image_in"]를 내지 않는다. Codexinput_modalities·Claude discoverycapabilities.image_input.supported도 같은 시드에 의존하므로 뿌리는 레지스트리 한곳이다.이미 고치는 초안 PR #4668 이 열려 있다(작성자 TykanN, Closes #4667). Anthropic 시드에 image 를 심고 OpenClaw/Kimi 익스포터를 고치며 capability-aware 익스포터 회귀를 넓힌다. grok-bot 은 #4668 에 별도 리뷰를 이미 남겼다(우선순위 67). 이 이슈는 재현 절차·감사 표·관련 이슈(#3332/#3474/#3454/#4497/#4534) 구분이 잘 되어 있고, 와이어 변환(#4497 계열)과 “내보내기 메타데이터”를 섞지 않았다. types.ts/config.ts 스플릿과 무관하다.
라인 / 경로 수준:
경로
src/providers/registry/entries-core.ts(anthropic/anthropic-apikey) - 현재dev에modelInputModalities부재. 이슈 재현의 뿌리이다
경로src/providers/registry/model-seeds.ts- Anthropic 전용 input modality 시드 심볼 없음
경로src/clients/config-export.tsbuildOpenclawClientConfig/buildKimiClientConfig- 선언된 image 도 클라이언트 문서로 안 나간다(이슈의 “추가 구멍”)
경로src/clients/config-export/model-metadata.ts- text|image 허용·누락 시 text-only 폴백은 의도적. Anthropic 시드만 채우면 Aside/Pi/GJC 쪽은 따라온다
경로 PR #4668 - draft. ready·포커스 스위트 초록 후 머지하면 이 이슈를 닫을 수 있다메인테이너의 판단이 필요한 지점
- fix: preserve image input capabilities across integrations #4668 의 OpenClaw/Kimi 확장을 이번 이슈 범위에 같이 넣을지(이슈 본문은 같이 넣으라고 함)
- unknown 모델에 image 를 추론하지 않는 보수 폴백을 유지할지(권장: 유지)
- draft CI/환경 이슈(bubblewrap·SIGSEGV)를 메인테이너가 어느 선까지 재확인할지
너의 추천
새 PR을 또 열지 말고 #4668 초안을 ready 로 밀어서 이 이슈를 닫는 쪽을 권한다. 레지스트리 시드 + 익스포터 두 구멍이 이슈·PR·현재dev코드에서 같은 이야기다. 중복 이슈로 닫을 대상이 아니라 PR 대기 상태로 두면 된다. types/config 스플릿 때문에 닫을 이슈가 아니다.이 댓글은 grok-bot이 작성했습니다
- added 2 commits that reference this issue
on Sep 15, 2026 - added 2 commits that reference this issue
on Sep 17, 2026
Client or integration
Other
Area
Catalog / models
Summary
This is a cross-integration capability-metadata defect, not an Aside/Pi/GJC-only bug. The built-in
anthropicandanthropic-apikeyproviders omitmodelInputModalitiesfor their static Claude model catalog. All integration configuration paths consuming that metadata need auditing for omitted image-input declarations or their schema-equivalent capability fields.Aside, Pi and GJC (internal client id
gajae) were the initial reproduction examples, not the scope boundary. The audit covers all 15 registered config exporters plus Codex and Claude integration surfaces.Two independent exporter gaps also need fixing: OpenClaw drops even declared input modalities (its omitted
inputdefaults to["text"]); Kimi CLI drops the documentedcapabilities: ["image_in"]flag. These two fixes must apply to all catalog-backed image models, not only Anthropic.I observed that adding
"image"to the input arrays for Fable, Opus, Sonnet and Haiku in my Aside configuration enabled image attachments in Aside chat. This manual observation is separate from the deterministic configuration-generation reproduction below; Pi/GJC live image requests have not been tested here.Expected: every integration that declares model input capabilities must advertise image support for the known image-capable Claude models, using
"input": ["text", "image"]or the integration's supported equivalent (such asmodalities.inputor a vision boolean), without manual client-file edits. Configurations with no per-model capability field must not receive invented unsupported fields. Unknown models should retain the existing conservative fallback, and explicit operator overrides should remain authoritative.Reproduction
liveModels: falseto reproduce offline:loadExportModels(config)fromsrc/server/management/model-rows.ts, thenbuildClientConfig(client, { config, models, baseUrl: "http://127.0.0.1:10100/v1" })for each ofaside,pi,gajae.document.providers.opencodex.models: Claude entries haveinput: ["text"]instead of["text", "image"].A regression added beside the existing bare-Anthropic reasoning test in
tests/server/management-client-config-route.test.tsfails for all three clients. A registry-level test also fails becauseproviderConfigSeed(entry).modelInputModalities?.[model]isundefined.Version
Reproduced on upstream
devcommitaa91958e3b050084e1edc07dcd66b05ef6eac604. Releasemainexamined at1cc89cf88c39160e4bb1ee21f0fd37404dd6930b(2.55.0). The installed version used for my manual Aside observation was not recorded.Operating system
macOS / Darwin 25.6.0; Bun 1.4.2 (automated reproduction).
Provider and model
Built-in
anthropic(OAuth) andanthropic-apikey(API key); Claude Fable 5.1/5, Sonnet 5/4.6, Opus 5/4.8/4.7/4.6, Haiku 4.5 static catalog.Logs or error output
Screenshots and supporting files
Audit matrix (capability declarations, not a claim of live image-request testing in every application):
inputarrays; consume shared catalog metadatamodalities.input; OpenCode alsoattachmentmodalities.inputandsupportsVisionsupports_visionabilities.vision.supportedinput; defaults to text-only when omittedcapabilities: ["image_in"]input_modalitiescapabilities.image_input.supportedOpenClaw input schema, attachment gate, text-only default.
Kimi CLI capability schema and image input gate.
Official model documentation: https://platform.claude.com/docs/en/models/overview — "All current models support text and image input".
Vision API documentation: https://platform.claude.com/docs/en/build-with-claude/vision
Current registry:
src/providers/registry/entries-core.tshas context windows and reasoning efforts for both Anthropic providers, but no input modalities.src/providers/derive.tsalready fills modality declarations per model while preserving explicit overrides.src/clients/config-export/model-metadata.tsalready allows bothtextandimage; the text-only fallback for missing metadata is intentional, not an Anthropic-specific restriction.Related but different: fix(catalog,anthropic): keep Claude combo image/effort capabilities and honor provider output budget #3332 / feat(catalog,codex): carry Claude combo capabilities and give reset-credit redeems a stable identity #3474 addressed combo metadata fallback; fix(providers): advertise the reasoning-effort ladder for native Anthropic models #3454 added Anthropic reasoning metadata; fix(chat): accept Pi and Anthropic-shaped image parts on the chat wire #4497 / [agent] fix: normalize inbound Chat images before route selection and preserve an explicit reasoning disable #4534 addressed image wire conversion rather than exported capabilities. I found no exact existing report for this direct-provider export gap in all-state issue/PR searches.
Redacted configuration
{"id":"anthropic/claude-fable-5-1","input":["text"]}Proposed fix: declare image input for the existing shared Anthropic model seed on both auth variants, audit every integration configuration writer, fix any remaining dropped image declaration, and cover every supported output shape through the registry-to-client-document chain. No blanket image claim for unknown models, and no unsupported fields added to client schemas.
Checks