Repository navigation
Emit eval outcomes and post V4 results - #1016
Conversation
Judges report assessed/insufficient_evidence/not_applicable instead of always scoring, and a failed judge invocation carries a coded reason (judge_timeout/chat_error/verdict_parse_failed) instead of vanishing. The daemon negotiates eval protocol 2 (RunQuestionV2/FinalizeEvalV2) on DaemonConnect and posts the aggregate to /evals/v4; a protocol-1 server still gets the legacy verdict-only RPCs and /evals/v3. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A complete coverage record on an insufficient_evidence question self-contradicts the server's validator and 400s the whole run's POST, discarding every correctly-scored question alongside it — null it in that one case, mirroring the server producer's reconciliation. ParseVerdict now tells an absent outcome (defaults to assessed) apart from an explicit JSON null (a parse failure); both bound to the same C# null before. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
PR Summary by QodoAdd outcome-aware V4 evaluations and daemon protocol 2
AI Description
Diagram
High-Level Assessment
Files changed (35)
|
Code Review by Qodo
1. All failed runs lose results
|
| if (assessments.Count == 0) { | ||
| observer.OnFailed("all judge invocations failed"); | ||
|
|
||
| return null; | ||
| } |
There was a problem hiding this comment.
1. All failed runs lose results 🐞 Bug ☼ Reliability
FinalizeAsync returns before calling Aggregate or PersistAggregateV4Async when assessments.Count is zero, despite receiving coded entries in failures and supporting a failure-only aggregate with a nullable score. When every judge times out, encounters a chat error, or emits an invalid verdict, both the CLI and protocol-2 daemon discard the failure taxonomy without posting a V4 payload, and the daemon maps the null return to a failed RPC result.
Agent Prompt
## Issue description
V4 finalization discards an eval when every question produces a coded failure because it rejects runs with no assessments. The V4 aggregate already represents failure-only runs with a null overall score and populated `FailedQuestions`, so these runs must be aggregated and posted rather than treated as unpersistable failures.
## Fix Focus Areas
- src/Capacitor.Cli.Core/Eval/EvalService.cs[739-783]
- src/Capacitor.Cli.Core/Eval/EvalService.cs[1450-1513]
## Recommended Fix
Remove or narrow the zero-assessment early return so finalization proceeds whenever `failures` is non-empty. Build and POST a V4 aggregate with no categories, a null overall score, zero assessed and judged questions, total questions derived from the failures, and every coded failure copied into `FailedQuestions`; skip the retrospective when no assessed question exists, and fail only when both assessments and failures are empty.
ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools
A per-question deadline that fires while still streaming the prompt fell into the generic stdin-error catch and reported chat_error instead of judge_timeout. Separately, an unassessed outcome on the protocol-1 relay sent an older server both a null-score completion and a failed RunQuestion result for the same question; the observer now relays it as a single failure, matching what the RPC already reports. Also uses the shared JsonElement helper for outcome-null detection and refreshes the README/help text for outcomes, nullable scores, and the v4/v3 posting split. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
|
Thanks Qodo — dispositions (9c7c490): Fixed:
Leaving as-is:
|
AI-2744 — no kcap-cli GitHub issue exists for this (driven from the Linear issue); the server half is kurrent-io/kcap-server#1939.
What & why
The server now records an honest per-question evaluation outcome (
assessed/insufficient_evidence/not_applicable) with a nullable score and a versioned evidence-coverage record, accepted atPOST /api/sessions/{id}/evals/v4. This teaches the CLI judge and the dashboard-triggered daemon to emit the same:ParseVerdictyields an outcome-bearing assessment (an absent outcome reads asassessed, a present malformed one is a parse failure), a failed judge invocation becomes a coded failure instead of a silent null,Aggregatebuilds a V4 payload whose overall is the mean of assessed scores only (null when none scored), and the run POSTs V4. The daemon advertises eval protocol 2 at connect and answers the newRunQuestionV2/FinalizeEvalV2RPCs; the V1 RPCs and V3 posting stay for an older server.Where to look
ReconcileEvidenceCoverage— a complete coverage record cannot accompany aninsufficient_evidenceoutcome (the server validator rejects that self-contradiction and 400s the whole run), so it is nulled to "unknown". The failure taxonomy inClaudeCliRunner.RunDetailedAsync.DaemonConnectsendingeval_protocol_version: 2(snake_case, pinned by a golden fixture). This PR merges only after the server route is deployed; until the submodule pin bump and npm release, the CLI still posts V3 and reads as fully assessed.Verification
Core and integration suites green, including
EvalRunnerV4PostTests(exactly one POST to/evals/v4, none to/evals/v3), the D13 parser cases (absent vs present-nullvs unknown outcome), theReconcileEvidenceCoveragecases, and theeval_protocol_versiongolden fixture;dotnet publish -c Releaseis clean for bothkcapandkcap-daemon(no AOT/trim warnings).🤖 Generated with Claude Code