Skip to content

Give the legacy eval judge the declared tasks the server sends - #1301

Merged
realtonyyoung merged 3 commits into
mainfrom
claude-tyoung/ai-3459-limitations
Oct 2, 2026
Merged

realtonyyoung merged 3 commits into
mainfrom
claude-tyoung/ai-3459-limitations

Conversation

@realtonyyoung

@realtonyyoung realtonyyoung commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

AI-3459 — no GitHub issue exists for this change.

What & why

The legacy daemon judge left {TASKS} unfilled on the text path and had no task section on the tools path, so a single-scope judge scored plan completion against the session's prose alone. The eval context now carries the server judge's declared-task block (kurrent-io/kcap-server#2231); the text prompt fills {TASKS} with it in one pass, and the tools prompt gains a declared-tasks section before the question.

Where to look

With no block (an older server) both prompts are byte-identical to what they render without it, which the existing legacy byte pins check. The evidence route still sends its fixed line: there the plan reaches the judge as evidence.

Verification

New unit tests pin the one-pass fill (a task title holding {TRACE_JSON} stays literal), the tools section placement, and parsing the field from a newer and an older server; the legacy byte-identity and prompt tests still hold; AOT publish shows no IL warnings.

Part of AI-3459.

🤖 Generated with Claude Code

A server that sends no block leaves both prompts byte-identical. The evidence
route keeps its fixed line: there the plan reaches the judge as evidence.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@linear-code

linear-code Bot commented Oct 2, 2026

Copy link
Copy Markdown

AI-3459

@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Pass server-declared tasks to legacy eval judges

🐞 Bug fix 🧪 Tests 🕐 20-40 Minutes

Grey Divider

AI Description

• Pass server-declared tasks to legacy judges so plan completion is assessed against the task list.
• Preserve existing prompt output when an older server sends no task block.
• Test task rendering, section placement, and context compatibility.
Diagram

graph TD
  Server["Eval context API"] --> Context["EvalContextResult"] --> Runner["RunQuestionAsync"] --> Text["Text prompt"] --> Judge["Legacy judge"]
  Runner --> Tools["Tools prompt"] --> Judge
Loading
High-Level Assessment

Passing the optional server field through the existing context and prompt builders is the narrowest compatible approach. Reusing the existing one-pass renderer avoids rescanning task text, while retaining the old rendering path preserves prompts from older servers.

Files changed (5) +84 / -18

Bug fix (3) +41 / -15
EvalService.csInclude declared tasks in both legacy question prompts +35/-14

Include declared tasks in both legacy question prompts

• Passes the context's optional task block to the text and tools prompt builders. The text path fills placeholders in one pass when tasks exist; the tools path inserts a declared-tasks section before the question. Both retain their previous rendering behavior without a block.

src/Capacitor.Cli.Core/Eval/EvalService.cs

Models.csDeserialize optional declared tasks from eval context +5/-0

Deserialize optional declared tasks from eval context

• Adds a nullable tasks field to the eval-context response model so older servers can omit it.

src/Capacitor.Cli.Core/Models.cs

prompt-eval-question-tools.txtReserve a declared-tasks section before the tools question +1/-1

Reserve a declared-tasks section before the tools question

• Adds a placeholder immediately before the question heading for the optional task block.

src/Capacitor.Cli.Core/Resources/prompt-eval-question-tools.txt

Tests (2) +43 / -3
EvalServiceDeclaredTasksTests.csTest task rendering and older-server compatibility +40/-0

Test task rendering and older-server compatibility

• Tests one-pass text substitution, tools-section placement, and deserialization with and without the new server field.

test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceDeclaredTasksTests.cs

EvalServiceLegacyByteIdentityTests.csClarify the scope of legacy prompt byte pins +3/-3

Clarify the scope of legacy prompt byte pins

• Updates the test description to specify that existing prompt bytes are pinned when the server sends no declared-task block.

test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceLegacyByteIdentityTests.cs

@qodo-code-review

qodo-code-review Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Remediation recommended

1. Tools judges expect absent attribution ✓ Resolved
Description
DeclaredTasksSection tells the tools judge that each task line names who set it and may say
“status partial,” but kcap-server supplies a root-lane block whose lines contain neither field.
Whenever that block is present, the newly added tools prompt describes metadata the judge cannot
actually inspect.
Code

src/Capacitor.Cli.Core/Eval/EvalService.cs[R1019-1020]

+      + "The agent's own declared task list for the plan this session worked on, when one exists. Each line carries the status and who "
+      + "set it; \"status partial\" means part of that task's history is withheld from this view. When a task list is declared, judge "
Relevance

●●● Strong

Prompt guidance must match the server-rendered task format so the judge is not instructed to inspect
absent metadata.

PR-#200

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new CLI instruction claims fields that the server’s selected root-lane formatter does not
render.

kcap-cli -> kcap-server
src/Capacitor.Cli.Core/Eval/EvalService.cs[1017-1021]
External repo: kurrent-io/kcap-server, src/Capacitor.Server.Services/Sessions/Plans/EvalContextTasks.cs [31-36]
External repo: kurrent-io/kcap-server, src/Capacitor.Server.Core/Evals/PlanTasksPromptBlock.cs [75-86]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The tools prompt describes source and partial-status metadata absent from the server’s root-lane task block.
## Fix Focus Areas
- src/Capacitor.Cli.Core/Eval/EvalService.cs[1017-1021]
- /cross_repos/kcap-server/src/Capacitor.Server.Services/Sessions/Plans/EvalContextTasks.cs[31-36]
- /cross_repos/kcap-server/src/Capacitor.Server.Core/Evals/PlanTasksPromptBlock.cs[78-86]
## Recommended Fix
Remove the unsupported metadata claims from the CLI tools section, or coordinate a server format change that supplies those fields.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Slug evaluations omit declared tasks ✗ Dismissed
Description
RunQuestionAsync passes the server’s Tasks value into both legacy prompts, but kcap-server
returns null for a meta-session slug because it checks the slug against a visible chain containing
linked session IDs. When the CLI evaluates a slug, it receives the session context but still judges
without the declared-task block.
Code

src/Capacitor.Cli.Core/Eval/EvalService.cs[535]

+            var prompt = BuildTextQuestionPrompt(question, ctx.SessionId, ctx.EvalRunId, ctx.TraceJson, ctx.ContextResult.Tasks);
Relevance

●●● Strong

Directly contradicts the PR’s stated intent; slug sessions can lose declared tasks through
server-side ID resolution.

PR-#1297

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The CLI newly consumes Tasks. The server resolves slugs into linked IDs but passes the original slug
to a task reader that requires its input to be one of those IDs.

kcap-cli -> kcap-server
src/Capacitor.Cli.Core/Eval/EvalService.cs[532-535]
src/Capacitor.Cli.Core/Eval/EvalService.cs[349-357]
External repo: kurrent-io/kcap-server, src/Capacitor.Server.Services/Sessions/VisibleSessionChain.cs [41-48]
External repo: kurrent-io/kcap-server, src/Capacitor.Api.Public/Sessions/SessionEvalContextHandler.cs [153-160]
External repo: kurrent-io/kcap-server, src/Capacitor.Server.Services/Sessions/Plans/EvalContextTasks.cs [22-27]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Meta-session slugs resolve to visible session IDs, but the server checks the slug itself for task eligibility and returns no block.
## Fix Focus Areas
- /cross_repos/kcap-server/src/Capacitor.Api.Public/Sessions/SessionEvalContextHandler.cs[153-160]
- /cross_repos/kcap-server/src/Capacitor.Server.Services/Sessions/Plans/EvalContextTasks.cs[22-36]
- src/Capacitor.Cli.Core/Eval/EvalService.cs[532-535]
## Recommended Fix
Coordinate a kcap-server change that selects an authorized linked session as the task root for slug requests, and test a CLI evaluation using that server response.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. A test comment repeats the test names ✓ Resolved
Description
The new summary on EvalServiceDeclaredTasksTests lists the behaviors already named and asserted by
its three tests. A reader gains no additional constraint from it and must keep the summary
synchronized when those tests change.
Code

test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceDeclaredTasksTests.cs[R6-8]

+/// <summary>The legacy prompts carry the declared-task block the eval context sends: the text prompt fills its
+/// <c>{TASKS}</c> in one pass, the tools prompt gains a declared-task section before the question, and a context from a
+/// server that sends no block leaves both prompts as they are without it.</summary>
Relevance

●● Moderate

The repository both accepts redundant test-comment cleanup and rejects similar comments when
rationale is arguably useful.

PR-#591
PR-#672

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The added summary enumerates the text-prompt, tools-prompt, and absent-block cases, which the three
test methods immediately below already name and assert. Rule 2762993 restricts comments that merely
restate what the code clearly does.

Rule 2762993: Restrict comments to documenting non-obvious, behavior‑critical constraints
test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceDeclaredTasksTests.cs[6-38]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The class summary repeats the behaviors covered by the test methods without documenting a separate constraint.

## Fix Focus Areas
- test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceDeclaredTasksTests.cs[6-8]

## Recommended Fix
Remove the class summary; retain the descriptive test names and assertions.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
✅ Compliance rules (platform): 64 rules
✅ Cross-repo context — repo relationships
  Explored: repo: kurrent-io/kcap-server (branch: claude-tyoung/ai-3459-limitations, sha: 78a05c11) — View relationship
Review mode: ⚖️ Balanced: This changes runtime prompt rendering and backward-compatible context deserialization across text and tools evaluation paths, creating meaningful behavioral and compatibility risk that warrants a complete single-pass review.

Grey Divider

Tip of the day
💡 Did you know, you can show, collapse, or hide each part of a finding: code, evidence, and all

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Qodo Logo

Comment thread test/Capacitor.Cli.Core.Tests.Unit/Eval/EvalServiceDeclaredTasksTests.cs Outdated
Comment thread src/Capacitor.Cli.Core/Eval/EvalService.cs
Comment thread src/Capacitor.Cli.Core/Eval/EvalService.cs Outdated
realtonyyoung and others added 2 commits October 2, 2026 15:28
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@realtonyyoung
realtonyyoung merged commit 6eb86e0 into main Oct 2, 2026
8 checks passed
@realtonyyoung
realtonyyoung deleted the claude-tyoung/ai-3459-limitations branch October 2, 2026 23:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant