Skip to content

[safeoutputs] Safe-output tool descriptions: submit_pull_request_review lacks anti-probing guidance (empty-args probe pattern, 4 occurrences/2 [Content truncated due to length] #51467

Description

@github-actions

Summary

Analysis of agentic workflow runs in github/gh-aw from the last 24h (2026-08-08) found a reproducible pattern: agents call submit_pull_request_review with completely empty arguments ({}) as an apparent schema-discovery probe, get an ERR_VALIDATION error, then retry ~9s later with the correct fields and succeed. This happened 4 times across 6 invoking runs (67%), in 2 unrelated workflows, ruling out a workflow-prompt cause. Two secondary, lower-priority patterns on push_to_pull_request_branch and create_issue are also included since they reinforce the same root cause category (missing/weak anti-probing tool descriptions).

This is a tool-description issue: submit_pull_request_review's MCP-served description lacks the WRITE-ONCE: do NOT call this tool with empty or placeholder arguments to probe or discover its schema preamble that sibling tools (add_comment, create_issue) already carry.

Error Analysis Details

submit_pull_request_review — empty-args probe (4 occurrences)

Observed identically in two different workflows on two different PRs:

Exact JSON-RPC exchange (from mcp-logs/safeoutputs.log, run 31268864783):

recv: {"jsonrpc":"2.0","id":N,"method":"tools/call","params":{"name":"submit_pull_request_review","arguments":{}}}
send: {"jsonrpc":"2.0","id":N,"error":{"code":-32602,"message":"ERR_VALIDATION: ..."}}

... ~9 seconds later ...

recv: {"jsonrpc":"2.0","id":N+1,"method":"tools/call","params":{"name":"submit_pull_request_review","arguments":{"body":"...","event":"APPROVE","pull_request_number":51425}}}
send: {"jsonrpc":"2.0","id":N+1,"result":{...,"isError":false}}

Of 6 runs across the window that invoked this tool at all, 4 began with the empty-args probe (67%); the 2 clean runs (§31273655356, run 31278288729) called it correctly on the first try. Every occurrence self-corrected within seconds, so no task ultimately failed — but each probe wastes an LLM turn and an MCP round-trip, and reduces confidence in structured-output reliability at scale.

push_to_pull_request_branch — invalid/missing required params (2 occurrences, lower priority)

Both in the Design Decision Gate 🏗️ workflow (a batch, multi-PR gate), on different PRs and different engines:

  • §31268864790: agent first called with an unsupported expected_head_sha param, then retried missing message, then succeeded on the 3rd attempt.
  • §31278288713 (claude engine): agent called with only pull_request_number (missing required message), got ERR_VALIDATION: ... missing or empty 'message', retried immediately with correct fields, succeeded.

This tool's description already explicitly warns against passing expected_head_sha/head_sha/base_sha, and already states the accepted-fields list — so the existing mitigation is present but imperfect; agents still sometimes omit the required message on the first call. Lower priority than finding #1 since a mitigation already exists here (unlike submit_pull_request_review, which has none).

create_issue — placeholder body/title probe (1 occurrence, lowest priority)

§31278189210: agent called create_issue with literal placeholder {"body":"test","title":"test"}, got 'body' is too short (minimum 20 characters, got 4), retried ~86s later with real content, succeeded. Notable because create_issue already has the strongest anti-probing language of any tool observed, yet a slip still occurred once — this suggests diminishing (but nonzero) returns from the preamble, and is included mainly as a baseline/control data point, not a strong signal on its own.

Current Tool Description

submit_pull_request_review (current, in pkg/workflow/js/safe_outputs_tools.json)
Submit a pull request review with a status decision. By default this tool targets the pull request that triggered the workflow. When the workflow is configured with `target: "*"`, you must specify `pull_request_number` to indicate which PR to target. REQUIRED: every call must include either a non-empty body or be preceded by at least one create_pull_request_review_comment call; calling with no body and no prior comments is rejected with ERR_VALIDATION. All preceding create_pull_request_review_comment outputs are automatically attached as inline comments. If this tool is not called, buffered review comments are submitted as a COMMENT review at workflow end. Use COMMENT for non-blocking feedback; use REQUEST_CHANGES only for merge-blocking. Example (inline-only review): call create_pull_request_review_comment one or more times, then call this tool with event: COMMENT and no body. CONSTRAINTS: Maximum 1 review(s) can be submitted.

For comparison, add_comment's description already has this preamble prepended:

WRITE-ONCE: do NOT call this tool with empty or placeholder arguments to probe or discover its schema — the required `body` field is listed in this schema; if you are not ready to post a real comment, call `noop` instead. ...

and this suffix appended:

... This tool records a real comment intent. Do not use it for placeholder comments, auth checks, or probing. Call it only when the final comment body is ready; otherwise use noop or report_incomplete.

Root Cause Analysis

This is a tool-description issue, not a workflow-prompt issue:

  • The 4 submit_pull_request_review occurrences span 2 independent workflows ("Impeccable Skills Reviewer", "Test Quality Sentinel") with different prompts, both on the copilot engine. Neither workflow's prompt instructs empty-arg probing (both instruct calling submit-pull-request-review — note the hyphenated form used in prose, a minor doc/tool-name mismatch, but this alone doesn't explain calling with zero arguments).
  • The tool already documents its validation rule ("REQUIRED: every call must include either a non-empty body or...") but does not carry the imperative anti-probing instruction ("do NOT call with empty/placeholder arguments to probe") that measurably correlates with lower probe rates on add_comment and create_issue.
  • push_to_pull_request_branch already has partial anti-probing language and a still-nonzero (but much lower, and now with a concrete required-field omission rather than an unsupported param) occurrence rate — consistent with the preamble being only partially, not perfectly, effective, and reinforcing that adding it to submit_pull_request_review should meaningfully reduce (not necessarily eliminate) its 67% probe rate.

Recommended Improvements

  1. Add the same WRITE-ONCE anti-probing preamble used by add_comment/create_issue to submit_pull_request_review's description in pkg/workflow/js/safe_outputs_tools.json (~line 473-474), e.g.:

    "WRITE-ONCE: do NOT call this tool with empty or placeholder arguments to probe or discover its schema — this call must include a non-empty body, or be preceded by at least one create_pull_request_review_comment call; if you are not ready to submit the real review, call noop instead."

  2. Add the matching closing sentence used elsewhere: "This tool records a real review-submission intent. Do not use it for placeholder reviews, auth checks, or probing. Call it only when the final review is ready; otherwise use noop or report_incomplete."
  3. Optionally, restate the required-field example inline (as push_to_pull_request_branch does) so the very first call has a concrete non-empty template to copy, e.g. showing a minimal {"body": "...", "event": "COMMENT"} example directly in the description.

Affected Workflows

  • Impeccable Skills Reviewer (copilot engine)
  • Test Quality Sentinel (copilot engine)
  • Design Decision Gate 🏗️ (secondary finding, claude + unspecified engine)

Testing Plan

  • After updating the description in pkg/workflow/js/safe_outputs_tools.json, recompile affected workflow lock files and re-run the two named workflows against a PR to confirm submit_pull_request_review is called with valid arguments on the first attempt.
  • Monitor the next few days of run logs (via this same optimizer's daily scan) for a drop in the empty-args-probe occurrence rate for submit_pull_request_review.

Implementation Checklist

  • Add anti-probing preamble + closing sentence to submit_pull_request_review description in pkg/workflow/js/safe_outputs_tools.json
  • Recompile .lock.yml files for affected workflows
  • Re-run affected workflows and confirm no more empty-args probe calls
  • Track probe rate for push_to_pull_request_branch and create_issue in subsequent daily scans to see if the imperfect-but-present mitigation there offers a useful comparison baseline

References

Generated by ⚡ Daily Safe Output Tool Optimizer · agent · 280.7 AIC · ⌖ 39.2 AIC · ⊞ 9.5K ·

  • expires on Aug 10, 2026, 1:22 PM UTC-08:00

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions