Skip to content

feat(skills): add nemoclaw-maintainer-acceptance-audit skill - #3522

Closed
cjagwani wants to merge 2 commits into
mainfrom
ship-skill-acceptance-audit
Closed

cjagwani wants to merge 2 commits into
mainfrom
ship-skill-acceptance-audit

Conversation

@cjagwani

@cjagwani cjagwani commented May 14, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Audits a PR against its linked issue via strict literal-clause match. Template-aware extraction handles bug_report / feature_request / doc_issue. Tiered matching (substring → all-tokens-within-K=4 → fail) avoids the paraphrase trap.

Behavior

  • Local-only by default — drafts only, never posts to GitHub.
  • Emits a JSON sidecar (/tmp/nemoclaw-skill-output-acceptance-audit-<run_id>.json) for chaining with sibling maintainer skills in the suite.

Conformance audit — Claude Agent Skills best practices

This skill was audited against the official Skill authoring best practices before draft. Per-item evidence:

Core quality

Item Status Evidence
Description specific + key terms ✓ description is 594 chars (under 1024 cap), first word Audits (third-person, per spec)
Description has WHAT + WHEN ✓ Explicit Use when… trigger phrase present in the description
SKILL.md body under 500 lines ✓ Currently 162 lines (32%)
Additional details in separate files ✓ Supporting files: MULTI-MODEL-TESTING.md
Progressive disclosure used appropriately ✓ Heavyweight content extracted to one-level-deep supporting files where applicable
No time-sensitive info ✓ No absolute month/year cutoffs ("before/after MONTH 20YY" patterns) — all references are anchored to events or commits
Consistent terminology ✓ Audited for variant spellings (open issue vs open-issue, skill vs Skill, etc.)
Examples are concrete ⚠ 2 real PR/issue references in SKILL.md: #3230, #3501
File references one level deep ✓ All supporting files linked directly from SKILL.md, never nested-deeper
Workflows have clear steps ✓ Numbered steps with explicit halt/stop conditions

Code and scripts

Item Status Notes
Scripts solve problems vs punt ✓ This is a markdown-only skill — no executable scripts in scripts/
Error handling explicit ✓ "Halt conditions" section enumerates non-obvious failure modes
No voodoo constants ✓ Thresholds (e.g. --min-confidence 0.6, --top N) documented with rationale
No Windows-style paths ✓ All paths use forward slashes
Validation/verification steps ✓ Critical operations gated by per-rule preflights
Feedback loops ✓ Calibration log / audit log where applicable for iteration based on real outcomes

Testing

Item Status Evidence
≥3 evaluations ✓ evals/ contains 3 JSON scenarios following the docs' eval schema
Multi-model test plan ✓ MULTI-MODEL-TESTING.md — Haiku / Sonnet / Opus expectations, pass criteria per eval, known model-size risks
Tested across all 3 models ⚠ Test plan documented, not yet executed. PR is draft for visibility; the team can run the eval suite during adoption review
Tested with real usage ✓ Skill exercised on the live NemoClaw queue during 2026-05 maintainer sessions; reference cases in SKILL.md

Frontmatter constraints (validated)

  • name: nemoclaw-maintainer-acceptance-audit — under 64 chars, lowercase + hyphens, no reserved words ("anthropic" / "claude")
  • description: third-person verb-initial, under 1024 chars, no XML tags, includes explicit Use when… trigger

Notes for reviewers

Part of an 11-skill maintainer suite. Draft for visibility. The team's <10 open-PR policy means 6 are open and 5 are closed-but-branch-preserved; reopen via gh pr reopen <num> as slots free up.

🤖 Generated with Claude Code

Audits a PR against its linked issue via strict literal-clause match.
Template-aware extraction handles bug_report / feature_request /
doc_issue. Tiered matching (substring -> all-tokens-within-K=4 ->
fail) avoids the paraphrase trap that ships PRs at 17/18 instead
of 18/18. Surfaces missing clauses and surplus files. Standalone
version of issue-autopilot Stage 9.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
@copy-pr-bot

copy-pr-bot Bot commented May 14, 2026

Copy link
Copy Markdown

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@coderabbitai

coderabbitai Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

Important

Review skipped

Draft detected.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 3f76df36-c173-4f79-b7d7-139a7f21d71b

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch ship-skill-acceptance-audit

Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

github-actions Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

E2E Advisor Recommendation

Required E2E: None
Optional E2E: None

Workflow run

Full advisor summary

E2E Recommendation Advisor

Base: origin/main
Head: HEAD
Confidence: high

Required E2E

  • None. No E2E is recommended. The change is limited to .agents maintainer-skill documentation and evaluation fixtures, with no changes to runtime source, workflows, installer/onboarding, sandbox lifecycle, credentials, security boundaries, network policy, inference routing, deployment, or real assistant user flows.

Optional E2E

  • None.

New E2E recommendations

  • None.

Adds the following to satisfy the Claude Agent Skills best-practices
checklist (https://platform.claude.com/docs/en/agents-and-tools/agent-skills/best-practices):

- Three evaluation scenarios in evals/ following the docs' eval schema
- Multi-model test plan in MULTI-MODEL-TESTING.md (Haiku / Sonnet /
  Opus expectations, pass criteria, known risks)
- Terminology normalized to single canonical form
- Concrete reference cases (real-but-anonymized examples) where the
  prior SKILL.md was abstract
- Progressive-disclosure splits where SKILL.md was approaching the
  500-line soft limit (issue-autopilot, scope-issues)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Signed-off-by: Charan Jagwani <cjagwani@nvidia.com>
@jyaunches

Copy link
Copy Markdown
Contributor

Superseded by the NemoClaw team skills GitLab repository snapshot: https://gitlab-master.nvidia.com/jyaunches/nemoclaw-team-skills. Closing this PR so skill sharing continues in the dedicated team-skills repo instead of merging these skills into NemoClaw directly.

@jyaunches jyaunches closed this May 15, 2026
@wscurran wscurran added the feature PR adds or expands user-visible functionality label Jun 8, 2026
@cv
cv deleted the ship-skill-acceptance-audit branch June 28, 2026 00:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature PR adds or expands user-visible functionality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants