Skip to content

DeepSeek V4 web-search continuation drops reasoning_raw_delta and returns 502 #688

Description

@Michael-Han0608

Client or integration

VS Code Codex extension

Area

Provider adapters / hosted Web Search loop

Summary

DeepSeek V4 thinking-mode requests can still fail after a synthetic hosted web_search call on opencodex v2.7.42, even though ordinary tool-call history replay is covered by #61 / #58.

The direct DeepSeek provider is correctly configured with preserveReasoningContentModels for deepseek-v4-pro and deepseek-v4-flash. However, the internal Web Search loop appears to lose DeepSeek's raw reasoning when it reconstructs the assistant tool-call message for the next model iteration.

The upstream rejection is:

Provider error 400: The `reasoning_content` in the thinking mode must be passed back to the API.

opencodex surfaces this to Codex as a 502/in-stream failure, after which Codex retries the sampling request.

Reproduction

  1. Configure the built-in direct deepseek provider with deepseek-v4-pro.
  2. Start a fresh Codex task through the Responses endpoint.
  3. Ask for a research/current-information task that causes the synthetic hosted web_search tool to be invoked.
  4. Let the model emit thinking/reasoning followed by the web_search tool call.
  5. After the search result is injected, observe the next internal DeepSeek iteration.

This was observed after several normal shell/plan tool continuations had succeeded. The failure occurred specifically during the Web Search continuation, not on the first prompt.

Expected behavior

The assistant message replayed by the Web Search loop should include the original DeepSeek reasoning_content together with the synthetic tool_calls, allowing the next thinking-mode request to continue.

Actual behavior

The follow-up DeepSeek request is rejected with HTTP 400 because the replayed assistant tool-call message does not contain reasoning_content. opencodex records the request as:

status=502
errorCode=upstream_server_error
terminalStatus=failed
upstreamError=Provider error 400: The `reasoning_content` in the thinking mode must be passed back to the API.

The Codex client then reports:

codex_core::responses_retry: stream disconnected - retrying sampling request

Suspected root cause

The OpenAI Chat adapter emits DeepSeek reasoning chunks as reasoning_raw_delta:

https://github.com/lidge-jun/opencodex/blob/v2.7.42/src/adapters/openai-chat.ts#L742-L744

But extractIterationThinking() in the Web Search loop only collects thinking_delta, thinking_signature, and redacted_thinking:

https://github.com/lidge-jun/opencodex/blob/v2.7.42/src/web-search/loop.ts#L107-L116

The returned value is then passed as precedingThinking when reconstructing the assistant web_search tool-call message:

https://github.com/lidge-jun/opencodex/blob/v2.7.42/src/web-search/loop.ts#L465-L470

For DeepSeek/OpenAI-compatible streams, extractIterationThinking() therefore returns null, so the assistant replay has the tool call but no thinking part. The normal OpenAI Chat request builder can only restore reasoning_content when an internal thinking part is present.

This looks like a path missed by the earlier #61 fix: ordinary Codex tool-call history replay works, but opencodex's internal hosted-Web-Search replay has its own event collection logic.

Suggested fix

Collect raw reasoning deltas in extractIterationThinking() as well:

if (e.type === "thinking_delta") thinking += e.thinking;
else if (e.type === "reasoning_raw_delta") thinking += e.text;
else if (e.type === "thinking_signature") signature = e.signature;
else if (e.type === "redacted_thinking") redacted.push(e.data);

A regression test could feed reasoning_raw_delta followed by a synthetic web_search tool call into the loop and assert that the next OpenAI Chat request contains an assistant message with both reasoning_content and tool_calls. Existing thinking_delta/Anthropic coverage should remain unchanged.

Version

2.7.42

Operating system

macOS arm64

Provider and model

Built-in direct deepseek provider, deepseek/deepseek-v4-pro, thinking mode enabled

Logs or error output

See the redacted error excerpts above. No credentials, request bodies, local paths, account identifiers, or conversation IDs are included.

Redacted configuration

{
  "adapter": "openai-chat",
  "baseUrl": "https://api.deepseek.com",
  "defaultModel": "deepseek-v4-pro",
  "preserveReasoningContentModels": [
    "deepseek-v4-pro",
    "deepseek-v4-flash"
  ]
}

Checks

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions