Skip to content

[Safe Outputs Conformance] WTD-001/WTD-002: Conformance checker false-positives on centralized threat-detection label string #49723

Description

@github-actions

Conformance Check Failure (false positive in checker)

Check ID: WTD-001 (CRITICAL) / WTD-002 (HIGH)
Severity: CRITICAL / HIGH as reported by the script — see analysis below for actual risk assessment
Category: Implementation / Test Tooling

Problem Description

scripts/check-safe-outputs-conformance.sh reports two failures on this run:

[CRITICAL] WTD-001: Footer generator missing 'agentic threat detected' label string (WTD1 requirement 2)
[HIGH] WTD-002: push_to_pull_request_branch fallback missing 'agentic threat detected' text (WTD2 / WTD1)

Both checks grep -q "agentic threat detected" directly against actions/setup/js/generate_footer.cjs and actions/setup/js/push_to_pull_request_branch.cjs. After investigating the actual implementation, this looks like a false positive in the checker itself, not a spec violation in the handlers:

  • The WTD1 "agentic threat detected" label is centralized in actions/setup/js/threat_detection_warning.cjs, in getThreatWarningPresentation() (title: "agentic threat detected"), and in the rendered template actions/setup/md/threat_detection_caution.md (literal line > agentic threat detected).
  • generate_footer.cjs::getExpiredEntityCautionAlert() renders that template via renderTemplateFromFile(getPromptPath("threat_detection_caution.md"), context), so the literal string only exists in the .md template, not in generate_footer.cjs source — the grep on generate_footer.cjs alone will never find it.
  • push_to_pull_request_branch.cjs builds its fallback review-PR body from warning.title (const warning = getThreatWarningPresentation(detectionReasonEnv); ... \> ${warning.title}``), so the literal string is constructed at runtime from the centralized helper rather than appearing verbatim in the handler source.

Notably, the XML marker requirement in the same WTD-001 check (requirement 3) does account for this delegation pattern — it falls back to checking threat_detection_warning.cjs via getThreatDetectedMarker when the marker isn't found directly in generate_footer.cjs. The label string check (requirement 2) has no equivalent fallback, so it flags a false CRITICAL failure even though the rendered output correctly includes the label text.

Affected Components

  • Checker script: scripts/check-safe-outputs-conformance.sh (check_wtd_reviewable_annotation, check_wtd_convertible_fallback)
  • Implementation (verified correct, no change needed): actions/setup/js/generate_footer.cjs, actions/setup/js/push_to_pull_request_branch.cjs, actions/setup/js/threat_detection_warning.cjs, actions/setup/md/threat_detection_caution.md
🔍 Current vs Expected Behavior

Current Behavior

The checker greps only the handler .cjs files for the literal string "agentic threat detected" and fails when it isn't found verbatim, even when the string is correctly produced at runtime via a centralized helper/template.

Expected Behavior

The checker should recognize the same delegation pattern it already accepts for the XML marker (requirement 3): if "agentic threat detected" isn't found directly in the handler/footer file, it should also check actions/setup/js/threat_detection_warning.cjs and actions/setup/md/threat_detection_caution.md (and any other centralized rendering paths) before failing.

Remediation Steps

This task can be assigned to a Copilot coding agent with the following steps:

  1. In scripts/check-safe-outputs-conformance.sh, update check_wtd_reviewable_annotation (WTD-001 requirement 2) to also pass if "agentic threat detected" is found in actions/setup/js/threat_detection_warning.cjs or actions/setup/md/threat_detection_caution.md, mirroring the existing fallback logic already used for the requirement 3 marker check.
  2. Update check_wtd_convertible_fallback (WTD-002) similarly: if "agentic threat detected" isn't literal in push_to_pull_request_branch.cjs, also accept the centralized helper/template as satisfying the requirement (e.g. detect the delegation via getThreatWarningPresentation usage plus the string present in threat_detection_warning.cjs/the template).
  3. Do not add a redundant literal "agentic threat detected" string directly into generate_footer.cjs or push_to_pull_request_branch.cjs — the current centralized-template approach is intentional (per code comments in generate_footer.cjs and threat_detection_warning.cjs) and duplicating the string would just be dead/misleading code.
  4. Re-run the checker locally to confirm WTD-001 and WTD-002 pass without a source-level change to the handlers.

Verification

After remediation, verify the fix by running:

bash scripts/check-safe-outputs-conformance.sh

Checks WTD-001 and WTD-002 should pass without errors, and no functional change should be needed in generate_footer.cjs or push_to_pull_request_branch.cjs.

References

  • Safe Outputs Specification: docs/src/content/docs/specs/safe-outputs-specification.md (Section 10.5, Requirements WTD1/WTD2)
  • Conformance Checker: scripts/check-safe-outputs-conformance.sh (lines ~1320-1421)
  • Run ID: 30735966273
  • Date: 2026-08-02

Generated by ✅ Daily Safe Outputs Conformance Checker · agent · 42 AIC · ⌖ 12.8 AIC · ⊞ 6.8K · ◷

  • expires on Aug 2, 2026, 10:33 PM UTC-08:00

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions