feat: implement issue #1651 — [qa-lead S8] Fleet-wide: 11 of 12 eval sets fail case.schema.json and CI never notices - #1815
Conversation
…sets fail case.schema.json and CI never notices
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
ⓘ Qodo reviews are paused because your trial has ended. Ask your workspace admin to add credits to resume reviews. Manage billing |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (16)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📝 WalkthroughWalkthroughThe change adds per-skill evaluation schemas, resolves those schemas during fleet-wide validation, removes the schema allowlist, updates example cases to use ChangesEvaluation schema validation
Priority: ➖ Normal Estimated code review effort: 3 (Moderate) | ~25 minutes Change: Feature · Severity of issue fixed: Medium Sequence Diagram(s)sequenceDiagram
participant CI as validate-eval-cases job
participant Validator as validate-cases.py
participant Schema as Resolved case.schema.json
participant Cases as Skill cases
CI->>Validator: Run --schema-tree evals
Validator->>Schema: Resolve per-skill schema or root fallback
Validator->>Cases: Validate every case and split
Validator-->>CI: Report validation result and schema usage
Merge Risk: ⚪ Minimal · up to Evaluation cases are now validated against their applicable schema across the fleet, with coverage for schema precedence and invalid cases. No merge-blocking risk remains. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 40.00% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 2 files. (14 skipped: 14 unsupported.) ✨ Finishing Touches 💡 2📝 Generate docstrings 💡
🛠️ Fix failing CI checks 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request implements per-skill schema validation for eval cases, discharging the migration allowlist so that all skills are now schema-validated. It introduces individual case.schema.json files for various skills to govern their specific decision shapes while maintaining the shared root envelope. The validator script evals/validate-cases.py is updated to resolve and apply these per-skill schemas, and the test suite in tests/test_validate_cases.bats is expanded to cover these changes. Feedback on the tests suggests asserting specific non-zero exit codes (e.g., [ "$status" -eq 1 ]) instead of generic non-zero checks ([ "$status" -ne 0 ]) to avoid passing on unexpected errors.
|
This PR touches only Making it green from this PR would require either editing another repo or relaxing what the check asserts (e.g. allowlisting the drifted files) — the "optimize the check, not the intent" failure mode this repo guards against. So I'm leaving it red and preserving the invariant it protects. No action in #1815. |
Dev-Lead — review-changes (applied)Changes committed and pushed. |
Dev-Lead — fix-bot-comment (no-changes)Agent reasoning |
Superseded by automated re-review at
|
Superseded by automated re-review at
|
|
pr-review approved on PARTIAL advisory evidence: 3/6 required advisory bots reported before the gate's quiescence-timeout fallback proceeded. Recorded for the miss-rate metric (#1596). |
Dev-Lead — fix-bot-comment (no-changes)Agent reasoning |
Review — fix requested (cycle 3/3)The automated review identified the following issues. Please address each one: Findings to fixAutomated review — NEEDS HUMAN REVIEWRisk: MEDIUM SummaryCorrectness-focused review of the fleet-wide eval-case schema gate (#1651): removes SCHEMA_TREE_ALLOWLIST, adds per-skill case.schema.json resolution, and validates every skill. The logic is sound and verified — validate-cases.py --schema-tree passes all 13 skills / 113 cases on the PR head, per-skill schemas keep a consistent root envelope, and the change strengthens (not weakens) the gate. Escalating (not approving) on two gate failures: the failing template-drift check and a description missing all 5 required sections. No security-audit tier needed. No downstream consumers impacted (DOWNSTREAM_IMPACT: none). Findings
Reviewed by the PR-review cascade (triage: haiku 4.5 [sonnet 5] → deep: opus 4.8 [sonnet 5] + duck: o4-mini → audit: fable 5). Reply if you need a human review. Additional tasks
The review cascade will automatically re-review after new commits are pushed. |
🤖 CodeAnt AI — Review Status
|
Thanks for using CodeAnt! 🎉We're free for open-source projects. if you're enjoying it, help us grow by sharing. Share on X · |
| "properties": { | ||
| "decomposition": { | ||
| "type": "string", | ||
| "enum": ["good", "too-coarse"], |
There was a problem hiding this comment.
Suggestion: The judge recognizes too-fine, but this enum rejects such cases during schema validation, preventing valid scrum-master evaluations from running. [api mismatch]
Assessment: 🟠 Major · 🔁 Occurrence: Sometimes
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** evals/scrum-master/case.schema.json
**Line:** 39:39
**Comment:**
*Api Mismatch: The judge recognizes `too-fine`, but this enum rejects such cases during schema validation, preventing valid scrum-master evaluations from running.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| except (OSError, json.JSONDecodeError) as exc: | ||
| fail(f"could not read/parse schema {path}: {exc}") | ||
| v = jsonschema.Draft202012Validator(schema) | ||
| validators[path] = v |
There was a problem hiding this comment.
Suggestion: A per-skill schema can be empty or permissive, so Draft202012Validator accepts invalid cases and the fleet-wide schema gate reports success. [security]
Assessment: 🟠 Major · 🔁 Occurrence: Sometimes
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** evals/validate-cases.py
**Line:** 273:273
**Comment:**
*Security: A per-skill schema can be empty or permissive, so `Draft202012Validator` accepts invalid cases and the fleet-wide schema gate reports success.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix|
pr-review approved on PARTIAL advisory evidence: 4/6 required advisory bots reported before the gate's quiescence-timeout fallback proceeded. Recorded for the miss-rate metric (#1596). |
|



User description
Closes #1651
Implemented by dev-lead agent. Please review.
CodeAnt-AI Description
Enforce schema validation across all evaluation cases
What Changed
Impact
✅ Fewer invalid evaluation sets reaching CI✅ Clearer failures for malformed case data✅ Supported validation for skill-specific evaluation outputs💡 Usage Guide
Checking Your Pull Request
Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.
Talking to CodeAnt AI
Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:
This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.
Example
Preserve Org Learnings with CodeAnt
You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:
This helps CodeAnt AI learn and adapt to your team's coding style and standards.
Example
Retrigger review
Ask CodeAnt AI to review the PR again, by typing:
Check Your Repository Health
To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.
Summary by CodeRabbit
promptfields toinput.Problem
[qa-lead S8] Fleet-wide: 11 of 12 eval sets fail case.schema.json and CI never notices
From the issue: As a maintainer who believed the eval sets were validated, I want every skill's cases to actually conform to
evals/case.schema.json, so that the held-out eval infrastructure measures something instead of passing vacuously.Risk
Medium — changes GitHub Actions workflow behavior, which is exercised only post-merge; verify via the affected workflow runs.
Test plan
Tests added/updated:
tests/test_validate_cases.bats. Verification:bash scripts/dev-lead-lint.sh(shellcheck --severity=warning) ran pre-commit; the bats suite runs in CI.Rollback
Revert this PR. No non-revertible side effects (no tags, migrations, or external state).
Monitoring
Watch the affected workflow run(s) in the Actions tab and this PR's Lint check for regressions.