Skip to content

[Bug] google-antigravity/gemini-3.8-flash: intermittent first-send 400 "function call turn comes immediately after a user turn" in long agentic sessions #5008

Description

@GoldenLoaf24h

Client or integration

Direct HTTP/API client (agentic coding client driving /v1/chat/completions against the local proxy; tool-heavy long-running sessions)

Area

Provider adapter

Summary

On google-antigravity/gemini-3.8-flash, a long-running agentic conversation (~440 messages, 37 tools, heavy interleaved tool traffic) intermittently fails with HTTP 400 on the first upstream send (sendCount: 1):

Provider error 400: Antigravity invalid request: Please ensure that function call turn comes
immediately after a user turn or after a function response turn.

The same conversation continues fine when re-pointed at another provider, and fresh short conversations on gemini-3.8-flash work. I expected the google adapter to translate the replayed OpenAI history into a wire-valid contents array (or fail with an actionable client-side error), as it already does for many malformed-history cases.

What I ruled out on 2.58.0 (local proxy, port 10100). A static repro matrix covering every adjacency hazard I could think of all returns 200:

  • (A) assistant text turn immediately followed by a separate assistant tool_calls turn → serializes as model(text) → model(functionCall)
  • (B) tool_call turn directly after user turn (control)
  • (C) text + tool_calls mixed inside one assistant message
  • (D) assistant-tail history (model-tail)
  • (E) orphan + duplicate tool results pointing at one call id
  • (F) tool result placed before its tool call (order flipped)

So messagesToGeminiFormat's repairs (adjacent-response batching, orphan handling, (continue) nudge) cover all the static shapes I could construct. The failure appears to require the real long history and/or upstream session state.

Additional observations (timeline correlation)

  • ~/.opencodex/thought-signature-replay.json: the failing request (requestId ocx-a278d15d1374d24e79cbd7ee52dcf1ad, 2026-09-18 10:24:11 UTC+8) landed 606 s after the last signature save for the same session key; that session bucket had accumulated 6,201 signature entries over ~17 h.
  • Hypothesis: with requestType: "agent" + a long-lived Antigravity session, the upstream appears to hold server-side conversation state. A first-send 400 about turn adjacency may indicate the upstream session history diverged from what the proxy replays (concurrent in-flight turns, retry, or session-key reuse), and/or a signature-cache miss pushing the turn into a shape the adapter cannot repair.

Reproduction

Not deterministic yet; recurs under real long agentic sessions. Observed sequence:

  1. ocx start (2.58.0, default provider google-antigravity)
  2. Drive a long tool-heavy conversation (~440 messages, 37 tools) over /v1/chat/completions on google-antigravity/gemini-3.8-flash
  3. At some point the next turn 400s on first send with the adjacency error; switching the conversation to another provider recovers immediately
  4. Minimal static repros of the same error class (A–F above) all return 200

Happy to capture ocx debug provider logs around the next occurrence and attach the sanitized contents tail if that helps.

Version

2.58.0

Operating system

Windows 11 24H2

Provider and model

google-antigravity / gemini-3.8-flash

Logs or error output

requestId: ocx-a278d15d1374d24e79cbd7ee52dcf1ad
status 400, errorCode invalid_request_error, sendCount 1, closeReason non_stream
upstreamError: Provider error 400: Antigravity invalid request: Please ensure that function
call turn comes immediately after a user turn or after a function response turn.
routeDecision: explicit-provider → google-antigravity/gemini-3.8-flash (single eligible candidate)

Redacted configuration

{
  "google-antigravity": {
    "adapter": "google",
    "googleMode": "cloud-code-assist",
    "authMode": "oauth",
    "models": ["gemini-3.8-flash"]
  }
}

Checks

  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.

Related prior art: #1429 (thought-signature replay cache miss → 400), #2065 (assistant-tail prefill 400), #4726 (requiresAdjacentResponsesToolResults splitting tool call from result).

Activity

  1. added
    bugSomething isn't working
    providerProvider adapters, OpenAI-compat presets, upstream API quirks
    toolstool_calls, MCP, web-search / sidecar tools
    on Sep 18, 2026
  2. Ingwannu commented on Sep 18, 2026

    @Ingwannu
    Owner

    This remains a valid report, but the evidence does not yet justify changing messagesToGeminiFormat: your A-F matrix shows the static history repairs already accept the obvious adjacency shapes, while the real failure depends on long-lived upstream state.

    One concrete isolation defect was found during this investigation and is now tracked separately as #5033: Antigravity derives its upstream session id from the child thread id alone, while the general lane identity includes parent plus child. Parallel children can therefore collapse onto one upstream session. That is a real bug worth fixing, but it is only a plausible amplifier here, not proven as the cause of this 400.

    Please keep this issue for the long-session failure. On the next occurrence, attach only a sanitized comparison from the same minute: whether sibling turns were concurrent, the parent/child relationship (hashed or boolean, not raw ids), request and attempt counts, input/history byte estimate, tool count and aggregate schema bytes, whether a thought-signature lookup hit, and the last few translated role/part kinds without content. Do not attach prompts, tool arguments, signatures, OAuth material, account ids, cookies, or raw request bodies. The next fix should target the first demonstrated divergence rather than adding another static adjacency rewrite.

  3. lidge-jun commented on Sep 18, 2026

    @lidge-jun
    Owner

    One candidate explanation is now off the table, which narrows this rather than closing it.

    #5033 landed on dev: antigravitySessionId was anchored on x-codex-parent-thread-id, which every parallel child of one parent presents identically, so concurrent children collapsed onto a single upstream Cloud Code Assist session. That is exactly the shape your hypothesis describes — upstream session history diverging from what the proxy replays, through concurrent in-flight turns sharing one session lane. It now anchors on the request's own thread-id.

    It is worth being precise about what that does and does not mean for your report. It was filed as a plausible amplifier, never as the cause, and it was never shown to produce your 400. If this recurs on a build containing #5033, collapsed session lanes are no longer available as the explanation, and that is genuinely useful information.

    Your static matrix is the reason this is hard to take further from here: A through F cover every adjacency hazard the adapter can repair, and they all return 200. So the remaining candidates all involve state the proxy does not hold — upstream session history, or a signature-cache miss reshaping the turn. The 606-second gap after the last signature save for a bucket holding 6,201 entries over ~17 hours is the most concrete lead in the report.

    The ocx debug provider logs capture you offered would help most, specifically the sanitized contents tail of the failing request alongside the immediately preceding successful one on the same session. The adjacency error names a relationship between two turns, so the pair is what distinguishes "the proxy replayed an invalid shape" from "the upstream believed a different history".

  4. lidge-jun commented on Sep 19, 2026

    @lidge-jun
    Owner

    The observation half of this landed on dev as 4f3c22353b (#5079). This does not close the issue — nothing here explains your 400 yet, and no behaviour changed.

    What it adds is a content-free structural description of the request that is actually about to leave, taken after compilation and immediately before the send. It records per-role turn counts, function-call and function-response counts, where an ordering violation sits and of what class, request-local ordinals, booleans for signature presence and scope binding, and which of four anchor classes the Antigravity session id was derived from. It records none of your content: no prompts, tool arguments or results, original call ids, signature text or hashes, account ids, or session and thread ids.

    It is off unless provider debug is on. With it on, the projection runs synchronously before the send and its cost is linear in history length, which is stated plainly in the module rather than hidden — an operator turning it on for a long session pays that on every turn.

    Three things were fixed during review before this landed, and they are worth naming because they were the kind of defect a diagnostic must not have. The projection had been evaluated in argument position, outside the logger's own try/catch, so a throw inside it would have turned a request that was about to be sent into a rejected one; it is now passed as a builder and evaluated inside the fail-safe. The detail ceilings did not agree with the debug buffer's 16 KiB per-line cap, so a worst case could be cut mid-JSON while still reporting truncated: false; the summary now budgets below the buffer and trims from the tail, keeping the opening turns where a first-send violation would be. And the send-count test proved nothing because it only called buildRequest, which never sends; it now drives fetchResponse with a real send budget and onPhysicalSend, and there is a case where the projection itself throws.

    What would help next: if you can reproduce the 400 with provider debug enabled, the [ocx:google:antigravity-wire-shape] line for the failing first send, paired with one from a preceding success, would show whether the outbound structure differs at all. If the structures match, the cause is not in the wire we build and the remaining hypotheses are the signature scope and the upstream's own session state.

  5. GoldenLoaf24h commented on Sep 19, 2026

    @GoldenLoaf24h
    Author

    Thanks for the swift follow-up and the clean diagnostic probe in #5079!

    The observation approach makes total sense — having a content-free [ocx:google:antigravity-wire-shape] diff between the failed turn and the preceding success will decisively isolate whether this is an outbound serialization regression or upstream session state drift.

    I'll track dev / the next build with provider debug enabled on our long-running agentic workloads. As soon as this hits again, I'll extract and share the exact pair of wire-shape log lines.

  6. GoldenLoaf24h commented on Sep 19, 2026

    @GoldenLoaf24h
    Author

    400 Error Reproduced — Local Evidence Captured

    Core Discovery

    The 400 error was reproduced again in a v2.59.0 release long-running session on Google Antigravity.

    Upstream Error:

    Antigravity invalid request: Please ensure that function call turn comes immediately after a user turn or after a function response turn.
    

    Failed Request ID: ocx-6f75f787e5425efb139d1236ff9505d5
    Timestamp: 2026/9/19 13:59:05
    Model: gemini-3.8-flash / max


    Adjacent Request Comparison

    Timestamp Request ID Status Tokens
    13:56:38 ocx-546272b6bfc016058bb05d037a7b411a ✅ 200 27k
    13:56:50 ocx-4f8aa6422dea3757fde99cf6914331bc ❌ 400 —
    13:57:56 ocx-f4ecf847e92e393c999d8fe64e0c40f1 ❌ 400 —
    13:59:05 ocx-6f75f787e5425efb139d1236ff9505d5 ❌ 400 —

    Same session: one 200 followed by three consecutive 400s, all reporting function call turn sequence violation.


    Wire-Shape Comparison Missing

    The [ocx:google:antigravity-wire-shape] diagnostic probe (PR #5079) was merged into the dev branch but is not included in the current v2.59.0 release.

    Therefore I cannot provide the wire-structure comparison log between the failed request and its preceding successful request.

    Request:

    1. Could you publish the next release (or nightly/pre-release) containing PR diag(google): describe the Antigravity wire request without its contents (#5008) #5079?
    2. Or is there a dev-branch build artifact available for testing?

    With wire-shape comparison data, we can definitively determine whether this is an OpenCodeX local message sequence assembly error or Google upstream Session state inconsistency.

  7. GoldenLoaf24h commented on Sep 23, 2026

    @GoldenLoaf24h
    Author

    Wire-Shape Evidence Captured: Root Cause Confirmed (call-turn-opens-request)

    We captured the diagnostic logs using the [ocx:google:antigravity-wire-shape] probe (PR #5079) during a first-send 400 failure on a long agentic session.

    This decisively rules out upstream session drift and confirms outbound request assembly as the direct root cause.

    Request Metadata

    • Request ID: ocx-554a8a2ff47908ae04dfad134cd3dbb0
    • Timestamp: 2026/9/23 14:22:53
    • Model: google-antigravity/gemini-3.8-flash
    • Status: 400 (invalid_request_error, sendCount: 1)
    • Upstream Error: Provider error 400: Antigravity invalid request: Please ensure that function call turn comes immediately after a user turn or after a function response turn.

    Diagnostic Wire-Shape Output

    The probe explicitly flagged the ordering violation locally right before physical transmission:

    [14:22:53] [ocx:google:antigravity-wire-shape] {
      "version": 1,
      "truncated": true,
      "turns": 213,
      "roles": { "user": 107, "model": 106, "other": 0 },
      "functionCalls": 103,
      "functionResponses": 103,
      "distinctCallIds": 103,
      "toolDeclarations": 37,
      "orderingViolations": 1,
      "firstOrderingViolation": {
        "index": 0,
        "kind": "call-turn-opens-request"
      },
      "turnShapes": [
        {
          "index": 0,
          "role": "model",
          "parts": 1,
          "kinds": ["functionCall"],
          "calls": [1],
          "responses": [],
          "signedCalls": 1,
          "sentinelCalls": 1,
          "violation": "call-turn-opens-request"
        },
        {
          "index": 1,
          "role": "user",
          "parts": 1,
          "kinds": ["functionResponse"],
          "calls": [],
          "responses": [1]
        }
      ]
    }

    Analysis & Proposed Fix

    1. Root Cause: When long-session context compaction or pruning truncates history on the client side, the conversation slice presented to /v1/chat/completions can start with an assistant tool call. messagesToGeminiFormat compiles this into index: 0 as role: "model" with a functionCall.
    2. Gemini Protocol Constraint: Gemini wire rules require that any functionCall turn is immediately preceded by a user turn or a functionResponse turn. At index: 0, there is no preceding turn, triggering an immediate 400 on the first send.
    3. The Fix: In messagesToGeminiFormat, check if the compiled contents starts with role: "model" containing functionCall (or when firstOrderingViolation.kind === "call-turn-opens-request"), and either prepend a synthetic user nudge (e.g. (continue)) or drop/heal the opening orphan model turn before sending.
  8. added
    priority: P1High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,
    on Sep 24, 2026
  9. devin-ai-integration commented on Sep 25, 2026

    @devin-ai-integration
    Contributor

    New matched pull request: #5825 (would resolve this issue) — fix: P0 #5369 spill growth + P1 bug batch [priority: P0]

  10. added a commit that references this issue on Sep 25, 2026
  11. lidge-jun commented on Sep 25, 2026

    @lidge-jun
    Owner

    Fixed by #5825 (22b22ae): if compaction leaves a request that opens on a functionCall model turn, a (continue) user turn is now added in front of it so the upstream doesn't return a 400.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingpriority: P1High: reproducible failure in a core path (routing, failover, account pool, streaming, usage, auth,providerProvider adapters, OpenAI-compat presets, upstream API quirkstoolstool_calls, MCP, web-search / sidecar tools

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions