Skip to content

[WRONG BRANCH] fix(responses): refuse a combo failover that cannot replay mandatory reasoning (#4696) - #4769

Merged
lidge-jun merged 14 commits into
codex/pw1-vision-sidecar-truncationfrom
codex/pw2-deepseek-reasoning-replay
Sep 16, 2026
Merged

lidge-jun merged 14 commits into
codex/pw1-vision-sidecar-truncationfrom
codex/pw2-deepseek-reasoning-replay

Conversation

@lidge-jun

@lidge-jun lidge-jun commented Sep 16, 2026 •

Copy link
Copy Markdown
Owner

Summary

A combo failover to official DeepSeek returned 400 invalid_request_error: "The reasoning_text in the thinking mode must be passed back to the API." The first target failed with upstream_server_error, the ladder moved to DeepSeek, and the replayed history was missing the reasoning DeepSeek requires.

Reproduction status first, because the report was filed against 2.54.0: it partially reproduces on current dev, and the surviving half is narrow.

Not reproducing: a history that already carries plaintext reasoning_text. preserveResponsesReasoningContent keeps it, and combo failover never reuses a previous target's converted body — previous_response_id is expanded once before the ladder and each target structuredClones the same original request. Neither commit since 2.54.0 changes this (b42d573331 is an ancestor of the reported build, and 369be813c4 only backfills summary: []), so there is no basis to close this as already-fixed.

Still reproducing: when the previous successful turn's reasoning exists only as opaque encrypted_content minted by a different provider, account or model. The replay identity (provider + destination + adapter + model + credential) no longer matches, so the sanitizer strips the foreign blob. Stripping is right — that blob is not decodable by the new route — but it left a reasoning item carrying neither encrypted_content nor reasoning_text, and forwarded that hollow shell to a provider whose thinking mode requires the real thing. That is exactly the shape of the reported 400.

The fix fails closed rather than guessing. A target that requires plaintext reasoning replay is ineligible for a history proven to have lost it across a route change; failover continues to the next target; an exhausted combo returns 400 target_incompatible. Reasoning is never fabricated, and an opaque blob belonging to another provider or account is never forwarded — the serializer additionally strips any encrypted blob for such a provider regardless of provenance.

Two deliberate details:

  • An unroutable target stays eligible here. Treating a routing exception as a reasoning incompatibility would answer a missing provider row with a 400 about plaintext reasoning, sending operators down the wrong path. Only a positively proven incompatibility triggers the new refusal; anything else still surfaces as 503 combo_unavailable.
  • requiresPlaintextReasoningReplay() names the contract in one place. It derives from preserveResponsesReasoningContent and requiresAdjacentResponsesToolResults because official DeepSeek is the only registry entry setting both. That derivation is incidental rather than meaningful, so it is documented as such and should become an explicit capability the moment a second provider needs it. (Verified against this stack: the sibling layer adding requiresAdjacentResponsesToolResults to kimi/kimi-code does not trip it, because neither Kimi entry sets preserveResponsesReasoningContent.)

The refusal uses the existing core-errors.ts factory convention rather than a second error shape — same form as unreadable_encrypted_agent_task, which is the same class of failure: a history the selected provider cannot consume. It fires only before a response has started, so it is a clean refusal and never needs a mid-stream terminal variant.

Upstream Codex has no multi-provider inference failover at all, so this is not reinventing an upstream mechanism; its closest analogue is comp_hash-guarded compaction, which fails closed on incompatible opaque state the same way.

Closes #4696

Verification

Static source review only, plus hosted CI. No local test suite, typecheck, test:changed, or build was run — the repository owner prohibits local suite execution in this lane after a past local run deleted real ~/.opencodex data.

Static checks performed:

  • Traced the reproducing path end to end: src/responses/reasoning-replay-cache.ts (identity includes credential), src/server/responses/core-replay.ts (_stripReasoningEncryptedContent on mismatch), src/adapters/openai-responses/reasoning.ts (strips the blob, never synthesizes plaintext).
  • Confirmed the non-reproducing path against the existing pins in tests/providers/deepseek-reasoning-replay.test.ts and the body-cloning in src/combos/request.ts / core-combo.ts.
  • Verified the registry conjunction claim directly: kimi (entries-core.ts:426) and kimi-code (entries-extended.ts:894) set only the adjacency flag; preserveResponsesReasoningContent appears on neither.
  • Confirmed no existing test was deleted or weakened anywhere in this layer (git diff --numstat on tests/ shows additions only).
  • git diff --check clean.

Regression coverage:

  • upstream_server_error failover to strict Responses preserves existing reasoning_text
  • a foreign opaque-only replay skips the strict target without forwarding or fabrication
  • an exhausted combo reports target_incompatible when mandatory plaintext is unavailable
  • a route switch never forwards foreign opaque reasoning or invents plaintext
  • an exact-response regression proving a missing provider row still returns 503 combo_unavailable, not the new refusal.

structure/transports/responses.md documents the new target-eligibility rule, since this changes observable routing behaviour.

Hosted CI: this is a non-tip layer of a stacked lane and carries [skip ci] under the maintainer-approved DEV-STACK-08 tip-only policy. The lane's CI gate runs on the tip branch, which contains this commit.

Checklist

  • Scope stays focused and avoids unrelated cleanup.
  • Docs or release notes were updated when needed. (structure/transports/responses.md; no user-facing configuration changed)
  • Security-sensitive changes were reviewed for secrets, auth, and unsafe defaults. (the change makes credential-scoped opaque reasoning less likely to cross an account boundary; no token or account identifier is logged or serialized)

Update after the first hosted CI run

The first push of this layer failed the file-size ratchet: tests/server/server-combo-failover-e2e.test.ts reached 4,322 lines against its committed cap of 4,166 (verdict: GREW).

The cap was not raised — that guard exists precisely to stop this, and the tooling only ever lowers caps. The tests this layer adds now live in a new file, tests/server/server-combo-reasoning-replay-eligibility.test.ts (300 lines), registered in both scripts/test-layout/layout.json and tests/fixtures/test-layout-expected.json. The original file is back to 4,156 lines with content byte-identical to its pre-change state, and the test-name multiset across the two files is unchanged at 101.

The regression coverage listed above is unchanged in substance; it simply resides in the new file.

@lidge-jun
lidge-jun requested a review from Ingwannu as a code owner September 16, 2026 02:30
@coderabbitai

coderabbitai Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

🗂️ Base branches to auto review (2)
  • ^dev$
  • ^preview$

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 6141b96f-6e69-4695-82ca-b2f3ce5302d1

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-16T02:33:21.051874Z 862eef4 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@lidge-jun

Copy link
Copy Markdown
Owner Author

리뷰 · 우선순위 76 / 80

이 PR은 콤보 페일오버가 공식 DeepSeek로 넘어갈 때 나오던 400 invalid_request_error — thinking mode에서는 reasoning_text를 다시 넘겨야 한다 — 를 고칩니다. 이슈 #4696(조합에서 DeepSeek 페일오버 후 필드 유실)입니다. 지금 dev HEAD(3070d64d8, 패키지 2.57.0)에서는 이렇게 됩니다. 이전 턴의 reasoning이 다른 프로바이더/계정/모델이 만든 opaque encrypted_content만 있을 때, 리플레이 identity(provider+destination+adapter+model+credential)가 안 맞으면 새니타이저가 그 blob을 지웁니다. 지우는 것 자체는 맞습니다. 새 라우트가 그 blob을 못 읽으니까. 그런데 지운 뒤에 encrypted_content도 없고 reasoning_text도 없는 빈 reasoning 껍데기를 DeepSeek thinking 경로로 그대로 보내서 400이 납니다. 반대로 이미 plaintext reasoning_text가 있는 히스토리는 preserveResponsesReasoningContent가 지키므로 재현되지 않습니다. 그래서 “이미 고쳐졌다”고 닫을 근거는 없고, 살아 남은 절반만 좁게 막는 PR입니다.

고치는 방식은 추측해서 plaintext를 만들어 넣지 않고 fail-closed입니다. plaintext reasoning 리플레이가 필수인 타깃은, 라우트 변경으로 그 reasoning이 이미 사라졌다고 증명되면 후보에서 빠집니다. 페일오버는 다음 타깃으로 가고, 후보가 다 떨어지면 400 target_incompatible을 돌려줍니다. 에러 모양은 기존 core-errors.ts의 unreadable_encrypted_agent_task과 같은 급 — 선택된 프로바이더가 소화 못 하는 히스토리 — 이고, 응답이 시작되기 전에만 막아서 mid-stream terminal 변형이 필요 없습니다. 라우팅 자체가 안 되는 타깃은 여기서 부적격으로 만들지 않습니다. 프로바이더 행이 없는데 갑자기 reasoning 400을 주면 운영자가 엉뚱한 길을 봅니다. 그 경우는 예전처럼 503 combo_unavailable로 남깁니다.

코드로 보면 세 층입니다. (1) requiresPlaintextReasoningReplay()을 src/adapters/openai-responses/passthrough.ts에 두고, preserveResponsesReasoningContent AND requiresAdjacentResponsesToolResults로 현재 DeepSeek 계약을 한곳에 이름 붙입니다. 직렬화 쪽에서는 stripEncryptedContent를 identity 변경뿐 아니라 이 계약에도 켜서, 출처를 몰라도 foreign opaque blob이 DeepSeek로 나가지 않게 마지막 가드를 겁니다. (2) src/server/responses/core-replay.ts는 routeReasoningReplayIdentity를 빼고, mandatoryResponsesReasoningReplayUnavailable()을 추가합니다. openai-responses + plaintext 계약 + tool-bearing body + opaque-only reasoning item + serving identity 변경이 모두 맞을 때만 true입니다. (3) src/server/responses/core-combo.ts의 executeComboResponses는 기존 payloadEligible 위에 reasoningReplayEligible을 얹어 targetEligible로 고르고, 남은 후보가 전부 리플레이 부적격이면 targetIncompatibleResponse()(core-errors.ts)를 반환합니다. structure/transports/responses.md에 라우팅 관측 규칙도 한 절 추가됐습니다.

테스트는 추가만 했습니다. tests/providers/deepseek-reasoning-replay.test.ts는 라우트 스위치 시 foreign opaque를 전달·조작하지 않음을 고정하고, tests/server/server-combo-failover-e2e.test.ts는 (a) strict 타깃으로 페일오버해도 기존 reasoning_text는 유지, (b) opaque-only면 strict 타깃을 건너뛰고 전달·조작 없음, (c) 후보 소진 시 target_incompatible, (d) 프로바이더 행 부재는 여전히 503 combo_unavailable 같은 회귀를 넣었습니다. 로컬 스위트/타입체크는 이 레인 금지이고, 커밋에 [skip ci]가 있습니다. 베이스는 dev가 아니라 codex/pw1-vision-sidecar-truncation(부모 #4752, vision sidecar 잘린 캡션 fail-closed / #4727)입니다. mergeable: MERGEABLE, mergeStateStatus: UNSTABLE이며 hygiene/label/enforce가 큐에 있습니다. types.ts/config.ts 스플릿에 무효화되는 모놀리스 편집은 아닙니다.

현재 dev 대비 방향은 맞고, #4696의 살아 남은 재현 절반을 정확히 조준합니다. 다만 부모 #4752가 BLOCKED이고 이 레이어는 tip-only CI 정책상 스킵이라, 스택 안착 전에 단독 머지하면 안 됩니다. 또한 requiresPlaintextReasoningReplay의 두 플래그 교집합 추론은 PR 본문이 스스로 “두 번째 프로바이더가 생기면 명시 capability로 바꿔라”고 적어 둔 기술 부채입니다. 레지스트리에서 kimi/kimi-code는 adjacency만 켜고 preserve는 안 켜서 지금 트리거는 안 탄다고 본문에서 검증했습니다.

라인 ~41+ (passthrough.ts requiresPlaintextReasoningReplay) - DeepSeek 전용 계약을 한 함수로 묶은 건 읽기 좋습니다. 다만 두 플래그 AND는 우연의 증거이지 의미 동치가 아닙니다. 두 번째 프로바이더가 생기면 레지스트리 명시 필드로 바꿔야 합니다.
라인 ~388+ (passthrough.ts stripEncryptedContent) - identity 변경 OR plaintext 계약으로 켜서 serializer 최종 가드가 됩니다. 콤보 거절과 역할이 겹치지만, 콤보를 안 타는 단일 라우트 경로까지 막아주므로 중복이 아니라 방어층입니다.
경로 mandatoryResponsesReasoningReplayUnavailable (core-replay.ts) - tools 있음 + opaque-only + identity 변경일 때만 true. 일반 plaintext 보존 게이트웨이나 provenance 불명은 적격으로 남겨 과잉 거절을 피합니다.
경로 reasoningReplayEligible catch (core-combo.ts) - routeConcreteModel 예외를 리플레이 부적격으로 취급하지 않습니다. 503 vs 400 표면을 섞지 않으려는 의도가 분명합니다.
경로 base codex/pw1-vision-sidecar-truncation / #4752 - dev 직접 머지 대상이 아닙니다. 부모 안착 후 리타겟하거나 스택 순서 머지가 필요합니다.
경로 [skip ci] - 이 커밋만으로는 호스티드 스위트 증거가 없습니다. 부모/팁 레인 CI 그린을 머지 게이트로 봐야 합니다.

메인테이너의 판단이 필요한 지점

너의 추천
열어 두고 부모 #4752가 dev에 안착한 뒤 스택 순서 머지(또는 리타겟)하세요. plaintext를 추측 생성하거나 foreign opaque를 DeepSeek로 다시 흘리지 마세요. types/config 스플릿 close-dont-rebase 대상 아닙니다. #4696은 머지 시 함께 닫되, opaque-only cross-route 절반만 막았다는 한 줄을 이슈에 남기면 이후 회귀 추적이 쉽습니다.

이 댓글은 grok-bot이 작성했습니다

@github-actions

Copy link
Copy Markdown
Contributor

✅ Deterministic PR hygiene checks passed.

@github-actions github-actions Bot added the bug Something isn't working label Sep 16, 2026
…4726) [skip ci]

Kimi's Code Plan Responses endpoint requires a tool result to follow its call
immediately. When the desktop LSP hook injects a developer message between a
code-mode exec call and its output, Kimi rejects the whole request with HTTP 400
naming the unanswered tool_call_id. Because the row replays full history every
turn, the session then fails permanently rather than once.

opencodex already implements exactly this repair in
normalizeResponsesToolResultAdjacency, but passthrough gates it on
requiresAdjacentResponsesToolResults, which only DeepSeek's entry seeded. Kimi
inherited no normalization, so a user who configures Kimi onto the Responses
wire hits the 400 on every affected turn.

Seed the flag on both kimi and kimi-code. The existing fill-only derivation in
providerConfigSeed, enrichProviderFromRegistry and routedProviderConfig carries
it into new and already-persisted rows without overriding an explicit user
value, and the flag is inert while these presets use the Chat wire.

The repair reorders; it does not delete. The intervening developer message is
preserved and moves after the batch, so the fix cannot be mistaken for silencing
the 400 by dropping hook context. Coverage pins that, plus call_id pairing with
two outstanding calls and interleaved noise, and an already-adjacent input being
left untouched.

No upstream specification documents the requirement; the evidence is the
reported 400 and DeepSeek's identical failure shape under #1292. Upstream Codex
deliberately leaves an intervening developer message where it is, so this stays
a per-provider capability rather than a wire-wide default.
…reasoning (#4696) [skip ci]

Official DeepSeek rejects a tool-bearing continuation whose prior reasoning is
not replayed, with 400 invalid_request_error: the reasoning_text in the thinking
mode must be passed back to the API. A combo failover reached that state.

The reported shape is narrow and still reproduces. When the previous successful
turn's reasoning exists only as opaque encrypted_content minted by a different
provider, account or model, the replay identity no longer matches, so the
sanitizer strips the foreign blob. Stripping is correct — that blob is not
decodable by the new route — but it left a reasoning item carrying neither
encrypted_content nor reasoning_text, and forwarded that hollow shell to a
provider whose thinking mode requires the real thing.

A history that already carries plaintext reasoning_text was never affected and
is unchanged: preserveResponsesReasoningContent keeps it, and each combo target
clones the original request body rather than reusing a previous target's
converted one.

Fail closed instead of guessing. A target that requires plaintext reasoning
replay is ineligible for a history proven to have lost it across a route change;
failover continues to the next target, and an exhausted combo returns 400
target_incompatible through the same core-errors factory convention as
unreadable_encrypted_agent_task. Reasoning is never fabricated, and an opaque
blob belonging to another provider or account is never forwarded.

A target that cannot be routed at all stays eligible here, so an unrelated
routing failure still surfaces as 503 combo_unavailable rather than being
misreported as a reasoning incompatibility.

requiresPlaintextReasoningReplay() names the provider contract in one place. It
currently derives from preserveResponsesReasoningContent plus
requiresAdjacentResponsesToolResults because official DeepSeek is the only entry
setting both; that derivation is documented and should become an explicit
capability as soon as a second provider needs it.
@lidge-jun
lidge-jun force-pushed the codex/pw2-deepseek-reasoning-replay branch from 862eef4 to 1eccd0f Compare September 16, 2026 03:11
lidge-jun and others added 12 commits September 16, 2026 14:01
…) [skip ci]

A custom provider whose baseUrl ends in a slash produced a doubled discovery
path: https://gateway.example.com/v1//models. Gateways that route the doubled
path as a distinct route reject it — the report observed HTTP 403 — so discovery
failed and the catalog silently fell back to the configured models.

The send paths were normalized already, by openaiChatCompletionsUrl and
openaiResponsesUrl, but discovery was not. buildModelsRequest appended the
endpoint verbatim, and resolveProviderModelDiscoveryUrl returns that default
unchanged for a provider with no registry spec, which is exactly the custom
case. Registry providers escaped it because new URL(spec.path, base) collapses
the doubled slash.

providerModelsUrl mirrors openaiChatCompletionsUrl rather than inventing a
second policy: trim outer whitespace and trailing slashes, drop an already
pasted /models, then append exactly one. Both production callers of the default
discovery URL use it — catalog discovery and API-key validation.

An existing path prefix is preserved, so /api/openai/v1 is not collapsed to the
origin. Registry spec.path, absolute endpoint overrides and relative endpoint
overrides all resolve exactly as before; a baseUrl already written without a
trailing slash is byte-identical to its previous output.
The CodeBuddy route launches the vendor CLI with --tools "" and
--strict-mcp-config, so the routed model has no native tool channel and writes
its call as prose. The shared coding-agent projection forwards text_delta
unrepaired, so that markup reached the client as an ordinary assistant answer.

Qoder's guard does not match it. The leaked tags are wrapped in FULLWIDTH
VERTICAL LINE (U+FF5C), which none of the shipped UNREPAIRABLE_MARKERS cover, so
this needed a signature of its own rather than a port.

Refusal requires the observed two-line grammar: a calls control line at column
zero, outside a Markdown fence, immediately followed by an invoke line naming a
functions.* tool. A lone tag, a quoted or inline-code literal, a fenced example,
a blockquote, indented source, or prose discussing the markup all carry extra
syntax before the tag and are forwarded untouched. Matching the marker alone
would refuse a legitimate answer that merely explains this protocol, which is
why the detector is narrower than the marker spelling.

A detected leak preserves the answer text already proven safe, emits one
non-retryable vendor_scaffold_detected error, and suppresses the vendor's later
success terminal so the client never sees a completed turn. Markers split across
streamed deltas are caught by holding only a bounded suffix that could still
complete a control sequence or a fence; unrelated pending text is released at
the next mismatch or terminal. The reasoning channel is guarded independently.

Leaked prose is never promoted into a real tool call. The text channel carries
no authenticated call envelope and no validated arguments, so converting it
would manufacture execution authority out of model output.

Kept CodeBuddy-owned rather than lifted into the shared coding-agent path, the
same containment #4234 chose for Qoder: the contract observed here is this
vendor's, and #4190's lane packet asked for a report rather than symmetry.

Co-authored-by: Ingwannu <ingwannu@users.noreply.github.com>
…#4679) [skip ci]

Command Code's gateway rejects a request outright with 400 name must be at most
64 characters, got 66. Codex Desktop built-in app tools flatten to
<namespace>__<name> past that bound, a user cannot exclude them, and
Responses-Lite catalogs bundle every declared tool, so the surface cannot be
shrunk from configuration.

The bound belongs to the adapter, not to the shared name helper. Three adapters
already solve this for themselves: Kiro normalizes to its own charset with a
deterministic 8-hex suffix, Google compiles and restores names in its wire
compiler, and Meta Muse aliases names on api.meta.ai. The translated
openai-chat path is the only one with no answer, and it is the path Command Code
uses.

A request-scoped registry now owns one collision domain per translated Chat
Completions request, following Kiro's shape. A namespaced name whose flattened
spelling exceeds 64 characters becomes a charset-safe alias derived purely from
the native identity, so it is stable across processes, catalog order and catalog
membership. Declarations, replayed assistant tool calls and tool_choice all pass
through the same registry, and both the streaming and buffered parsers restore
the echoed alias before tool_call_start, so the existing bridge map still hands
the client its native {namespace, name}.

The registry is seeded from the union of the current catalog and the structured
tool calls still present in replay history, because a historical call can keep
its namespace without being redeclared; seeding from the catalog alone would let
exactly the reported over-limit name reach the gateway again on a later turn.

Nothing else changes. Names at or under 64 characters and bare names are
byte-identical on the wire, and Kiro, Google and Muse still receive the raw
flattened name and run their own normalization.

64 is the Chat Completions function-name limit and a strict-gateway
compatibility concern, not OpenAI Responses parity: upstream Codex raised its own
MCP ceiling to 128 bytes in openai/codex#39594 because native Responses accepts
128. Applying it on this wire is correct for that wire alone.

Carried from #4715. That PR placed the bound in the shared namespacedToolName
helper and was provisionally accepted there. Hosted CI then showed twice that the
shared point intercepts adapters which already had an answer: it broke Google's
wire-compiler restore, and after that was narrowed it broke Kiro's normalizer.
The problem statement and issue analysis are the original author's; only the
placement changed.

Co-authored-by: Hulian Buligon <205309211+HulianBuligon@users.noreply.github.com>
# Conflicts:
#	structure/transports/responses.md
…guard

fix(codebuddy): refuse leaked vendor tool-call scaffolding (#4596)
…ames

fix(openai-chat): bound flattened tool wire names for strict gateways (#4679)
fix(catalog): normalize the custom-provider model-discovery join (#4724)
fix(providers): give Kimi the Responses tool-result adjacency repair (#4726)
@lidge-jun

Copy link
Copy Markdown
Owner Author

Cascading downward. A DeepSeek failover whose history carries only another route's opaque reasoning is refused as target-incompatible instead of dispatched with an empty placeholder that the provider rejects anyway.

Evidence at the verified tip 49f815d (tree cea66da56f10b8d1aee2290fa185d321442eab98), from run 35061163092:

  • test 1-4/4 and macos 1-2/2 all completed with conclusion success, confirmed through the check-runs API rather than the check rollup, so the heavy jobs actually executed and were not path-filtered. gates, changes, storage policy, api usage, docker smoke, keyring and npm-global on three platforms, the three service-lifecycle jobs, and the aggregate ci check all succeeded.
  • The ci failure at this commit belongs to run 35061161660, which this push superseded; run 35061163092 is the live one and it concluded success.
  • The lane absorbed dev at cf6e939 from the bottom layer upward, so each pull request keeps its own layer diff (4 / 9 / 6 / 7 / 7 / 5 files) and no dev commit appears in any layer's diff.
  • The single conflict was structure/transports/responses.md, where both sides appended a new section to the same empty base. It was resolved by keeping both: the section count goes from 14 on dev to 15 here, and every dev section name is still present. That loss is the kind CI cannot detect, so the names were compared directly rather than trusting the count.
  • The file-size ratchet reports no offenders after the absorption; openai-chat.ts sits exactly at its 822 cap.
  • git merge-tree --write-tree origin/dev <tip> reports a clean merge, and origin/dev is itself an ancestor of this tip.
  • Ancestry verified so each layer closes as MERGED: pw1 through pw5 are all ancestors of this tip.

Maintainer integration decision under MAINTAINERS.md / AGENTS.md: a maintainer with maintain or admin access may integrate into dev without a second maintainer approval, recording the decision and exact-head CI evidence.

@lidge-jun
lidge-jun merged commit 53455e0 into codex/pw1-vision-sidecar-truncation Sep 16, 2026
7 checks passed
@lidge-jun
lidge-jun deleted the codex/pw2-deepseek-reasoning-replay branch September 16, 2026 06:08
@github-actions github-actions Bot changed the title fix(responses): refuse a combo failover that cannot replay mandatory reasoning (#4696) [WRONG BRANCH] fix(responses): refuse a combo failover that cannot replay mandatory reasoning (#4696) Sep 16, 2026
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

⏳ DRAFT

  • wrong target branch (codex/pw1-vision-sidecar-truncation); retarget to dev.

What to do

  • Retarget this PR to dev — all contributions go to dev.

Its title has been prefixed with [WRONG BRANCH].
Automatic draft conversion failed (token cannot change draft status). Please convert this pull request to a draft manually. The required enforce-target check will keep failing until every issue above is resolved.

agentHits pushed a commit to agentHits/opencodex that referenced this pull request Sep 17, 2026
…easoning-replay

fix(responses): refuse a combo failover that cannot replay mandatory reasoning (lidge-jun#4696)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant