Skip to content

review: three sites disagree on the mandatory Confidence: label, and the rank order punishes honesty #4267

Description

@kyle-sexton

Observed in .claude/plugins/data/plugin-quality-melodic-software/evidence/c7ff503b-19e2-40d2-b99e-38c747f7fc0f/review-code-reviewer/20260919T205610Z/audit-notes.md, finding F6 (auditor severity: IMPORTANT, confidence high).

Three hand-maintained copies of one rule, disagreeing:

  • agents/code-reviewer.md:65: "Give every finding an explicit Confidence: high|medium|low line (per the severity baseline's confidence axis)", that is, always label, including low.
  • context/severity.md:30, the file the agent cites as its authority and which declares "This file owns the order": "Emit high or omit the field." That is, do not emit medium or low, omit instead.
  • skills/fanout/context/run-everything-mode.md, the COVERAGE_CLAUSE that the skill's dispatch contract says every finding-producing leaf prompt carries verbatim: "For each finding, include your confidence level (high / medium / low) and an estimated severity." Always label, including low.

severity.md's rank order is high > medium > unscored > low. A reviewer following the agent body or the fanout prompt honestly emits low for an uncertain finding, and the ranking sorts it below a finding nobody scored. The agent body warns about this inversion in the sentence after the one that causes it.

From the run: the session's review returned "3 should-fix and 7 nits" under a caller-supplied path:line: severity: problem. fix. contract, carrying neither the plugin's CRITICAL/IMPORTANT/SUGGESTION tiers nor any Confidence: line. So the mandatory field the normalization pipeline depends on was not emitted at all, and the agent has no rule for reconciling a caller-supplied output contract with its own non-negotiable fields.

Cheapest fix: pick one rule in context/severity.md, the file that claims ownership, and have the agent body and COVERAGE_CLAUSE cite it rather than restate it. That removes two copies and the drift risk.

Better, and favoured by the evidence: fix the rank order instead. Putting unscored below low removes the incentive to hide uncertainty and makes "always label" safe, which is what two of the three sites already say. Then add one line to the agent's Output format: when a caller supplies its own finding shape, keep the Confidence: field inside it.

Verification: a run under a caller-supplied output contract must still carry a confidence value per finding, and sorting a set containing one low and one unscored finding must place the unscored one last.

🤖 Generated with Claude Code

https://claude.ai/code/session_01F7GFS5autRMSaea6kSoXzx

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-readyFully specified and briefed; eligible for autonomous pickup from the frontier.priority: mediumReal value, no hard deadline; normal backlog flow.work-class: scopedA briefed fix or small feature; blast radius bounded by the brief, tests exist.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions