Skip to content

Add work-item evaluation tools to the work-items MCP server - #1257

Merged
realtonyyoung merged 6 commits into
mainfrom
tonyyoung/ai-3267-work-item-eval-mcp-tools
Sep 30, 2026
Merged

realtonyyoung merged 6 commits into
mainfrom
tonyyoung/ai-3267-work-item-eval-mcp-tools

Conversation

@realtonyyoung

Copy link
Copy Markdown
Collaborator

AI-3267 (no GitHub issue: the work is tracked in Linear, and a GitHub issue would import as a duplicate)

What & why

Adds list_work_item_evals, get_work_item_eval, request_work_item_eval and cancel_work_item_eval to kcap-workitems, over the server's /api/work-items/{id}/evals/runs routes (kurrent-io/kcap-server, AI-3267). An agent can queue an evaluation of the item it works on, cancel its own queued run, and read results, judged requirements and the retrospective.

Where to look

WorkItemEvalToolResults: findings, requirements and retrospectives are model output over transcripts, so they are rendered only inside a <work-item-eval-data> block with every field sanitised. Run ids and cursors must match the shapes the server mints before they go into a route, and an error keeps only a well-formed code. A 404 with no JSON body means a server that predates the routes, and reads the same as evaluations being switched off.

Verification

  • --treenode-filter "/*/*/McpWorkItemEvalToolsTests/*": 19/19 pass. McpToolAnnotationsTests: 5/5 pass.
  • dotnet publish -r osx-arm64 -c Release: no AOT or trim warnings.

🤖 Generated with Claude Code

Findings and retrospectives are model output over transcripts, so they reach
the agent only inside a sanitised data block; run ids and cursors are checked
against the shapes the server mints before they enter a route.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@linear-code

linear-code Bot commented Sep 30, 2026

Copy link
Copy Markdown

AI-3267

@qodo-code-review

Copy link
Copy Markdown

PR Summary by Qodo

Add work-item evaluation tools to the MCP server

✨ Enhancement 🧪 Tests 📝 Documentation 🕐 40+ Minutes

Grey Divider

AI Description

• Add MCP tools to list, read, request, and cancel work-item evaluations.
• Validate route inputs and isolate sanitized evaluation output from agent instructions.
• Document the tools and test routing, responses, and tool annotations.
Diagram

sequenceDiagram
    actor Agent
    participant MCP as Work-items MCP
    participant Helper as Eval results helper
    participant API as Work-items API
    Agent->>MCP: Call evaluation tool
    MCP->>Helper: Validate route inputs
    Helper-->>MCP: Safe route and body
    MCP->>API: GET or POST eval runs
    API-->>MCP: Evaluation response
    MCP->>Helper: Render status and data
    Helper-->>MCP: Sanitized tool text
    MCP-->>Agent: MCP result
Loading
High-Level Assessment

Keep the domain-specific response helper and reuse the existing untrusted-text sanitizer. Passing API JSON directly to the agent would lose the explicit boundary around transcript-derived output; a generic renderer would obscure the evaluation-specific fields and response rules.

Files changed (8) +538 / -2

Enhancement (2) +313 / -1
McpWorkItemsServer.csRegister and dispatch evaluation MCP tools +48/-1

Register and dispatch evaluation MCP tools

• Adds tool schemas and annotations, routes calls to the work-items evaluation API, and sends responses through the dedicated renderer.

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs

WorkItemEvalToolResults.csValidate evaluation inputs and render safe results +265/-0

Validate evaluation inputs and render safe results

• Validates run IDs, cursors, and modes before constructing routes. Formats evaluation details inside a sanitized data block, limits exposed error text, and treats legacy-route 404 responses as unavailable.

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs

Tests (3) +210 / -1
McpToolAnnotationsTests.csCheck evaluation tool annotations +5/-0

Check evaluation tool annotations

• Asserts read-only annotations for list and get, additive for request, and destructive for cancel.

test/Capacitor.Cli.Tests.Unit/Commands/McpToolAnnotationsTests.cs

McpWorkItemEvalToolsTests.csTest evaluation routes and trust boundaries +203/-0

Test evaluation routes and trust boundaries

• Covers HTTP routing, input rejection, request outcomes, unavailable servers, error-code filtering, and sanitized result rendering.

test/Capacitor.Cli.Tests.Unit/Commands/McpWorkItemEvalToolsTests.cs

McpWorkItemsServerTests.csAssert evaluation tools appear in the tool catalog +2/-1

Assert evaluation tools appear in the tool catalog

• Extends the expected work-items MCP tool list with all four evaluation tools.

test/Capacitor.Cli.Tests.Unit/Commands/McpWorkItemsServerTests.cs

Documentation (3) +15 / -0
README.mdDocument the four evaluation tools +2/-0

Document the four evaluation tools

• Adds the evaluation read, request, and cancellation tools to the work-items MCP tool list.

README.md

SKILL.mdAdd evaluation guidance to the work-items skill +4/-0

Add evaluation guidance to the work-items skill

• Describes each tool's inputs, pagination, evaluation modes, active-run behavior, and cancellation constraint.

kcap/skills/work-items/SKILL.md

help-mcp.txtExpose evaluation tools in MCP help +9/-0

Expose evaluation tools in MCP help

• Lists the four tools and their principal arguments in CLI help.

src/Capacitor.Cli.Core/Resources/help-mcp.txt

@qodo-code-review

qodo-code-review Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Code Review by Qodo

🐞 Bugs (0) 📘 Rule violations (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)

Grey Divider


Action required

1. Evaluation tools cannot reach server ✗ Dismissed
Description
McpWorkItemsServer sends all four evaluation tools to /api/work-items/{id}/evals/runs routes,
but the pinned kcap-server registers no evaluation HTTP routes. Against that server, the requests
receive 404 responses, which WorkItemEvalToolResults.Render can report as a non-error “not
enabled” result when the response is not JSON.
Code

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[R228-229]

+                "list_work_item_evals"   => await client.GetAsync(ItemUrl(baseUrl, arguments, "work_item_id",
+                    WorkItemEvalToolResults.ListSuffix(McpToolArguments.OptionalString(arguments, "cursor")))),
Relevance

●● Moderate

Potential cross-repository route incompatibility is serious, but no close rejection or acceptance
precedent confirms expected server versioning.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The PR issues requests to the new paths and treats a non-JSON 404 as unavailable. The server's
public work-item route registration contains no evaluation routes, and its existing evaluation
service is a web-service interface rather than an HTTP endpoint.

kcap-cli -> kcap-server
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-235]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[55-64]
External repo: kurrent-io/kcap-server, src/Capacitor.Api.Public/WorkItems/WorkItemEndpoints.cs [18-78]
External repo: kurrent-io/kcap-server, src/Capacitor.Api.Web.Abstractions/WorkItems/Evals/IWorkItemEvalDataService.cs [1-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new CLI evaluation tools target HTTP routes absent from the pinned kcap-server.

## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-235]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[55-64]
- /cross_repos/kcap-server/src/Capacitor.Api.Public/WorkItems/WorkItemEndpoints.cs[18-78]

## Recommended Fix
Implement and deploy the four authenticated evaluation HTTP operations in kcap-server with request and response contracts matching the CLI, including queued-run cancellation. Coordinate the server release with these CLI tools and test against the deployed server routes.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended

2. Incomplete scope reasons disappear ✓ Resolved
Description
Joined renders only the first 20 scope_incomplete_reasons and does not indicate that it omitted
any. When a run contains more than 20 reasons, the agent receives an incomplete explanation of its
evaluation scope without knowing that the explanation was truncated.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[265]

+            ? string.Join(", ", arr.EnumerateArray().Take(MaxListItems).Select(v => v.IsString ? Text(v.GetString(), cap) : "").Where(s => s.Length > 0))
Relevance

●●● Strong

Accepted history favors explicit omitted-count markers for capped agent-visible lists, matching
existing question and requirement rendering.

PR-#1211
PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
RenderRun puts the value returned by Joined directly on the scope line. Joined takes only
MaxListItems values, which is 20, and returns no omitted count; the renderer uses an explicit
omitted-count message for capped questions and requirements.

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[26-29]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[140-150]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[263-269]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Runs with more than 20 scope-incomplete reasons silently lose the remaining reasons in the rendered result.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[140-143]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[263-266]
## Recommended Fix
Keep the display cap, but append a count of omitted reasons to the scope line when the input array exceeds it. Add a test with more than 20 reasons.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. Retrospective entries disappear silently ✓ Resolved
Description
AppendRetrospective and Items take only the first 20 suggestions, strengths, and issues without
reporting how many remain. When a returned retrospective exceeds that cap, an agent reading the
result cannot tell that some entries were omitted, unlike the capped questions and requirements.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[R225-227]

+        foreach (var s in suggestions.EnumerateArray().Take(MaxListItems))
+            if (s.IsObject && Text(s.Str("text"), LongTextCap) is { Length: > 0 } text)
+                Line(sb, $"  suggestion ({Text(s.Str("audience"), CodeCap)}): {text}");
Relevance

●●● Strong

Existing capped-list behavior reports omissions; adding retrospective omission markers is a
consistent deterministic correctness fix.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The two retrospective rendering loops stop at MaxListItems without an omission marker. The
adjacent capped question and requirement loops call More to disclose truncated results.

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-29]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[146-149]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[201-212]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[217-236]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new retrospective list caps silently discard entries after the twentieth.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[217-236]
## Recommended Fix
After rendering each capped retrospective list, emit an omission count using the existing `More` helper when additional entries are present.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


4. Slow evaluations block other tools ✓ Resolved
Description
The evaluation GETs return after headers, but BoundedHttpContent.ReadAsync then reads their bodies
with CancellationToken.None. If a server sends headers and stalls or trickles the body without
reaching the byte limit, the single-call stdio loop remains blocked.
Code

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[243]

+                var bytes = await BoundedHttpContent.ReadAsync(httpResponse.Content, WorkItemEvalToolResults.MaxResponseBytes, CancellationToken.None);
Relevance

●●● Strong

Nearly identical timeout finding was accepted recently for the same bounded reader and sequential
MCP loop.

PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The GET calls opt into ResponseHeadersRead, the bounded reader waits on stream reads using the
supplied token, and the existing next-work path demonstrates a deadline applied to both the request
and body read because the stdio loop serves one call at a time.

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-245]
src/Capacitor.Cli/BoundedHttpContent.cs[8-16]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[282-310]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation GET body reads have no cancellation deadline after response headers arrive.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-245]
## Recommended Fix
Use one deadline token for each evaluation HTTP request and its bounded body read, and return a tool error when that deadline expires.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


View medium (3)
5. The work-items tool count is outdated ✓ Resolved
Description
The README adds four evaluation tools but leaves the introduction saying `It provides seventeen
tools`. The work-items section now lists 21 tools, so readers get a count that contradicts the
documented commands.
Code

README.md[R742-743]

+- **`list_work_item_evals`** / **`get_work_item_eval`** — list a work item's evaluation runs you can read and read one run's results, judged requirements and retrospective.
+- **`request_work_item_eval`** / **`cancel_work_item_eval`** — queue an evaluation of a work item (`mode` `process` or `root_cause`), or cancel your own run while it is still queued.
Relevance

●●● Strong

Recent README synchronization findings were accepted, including documentation mismatches after
user-facing changes.

PR-#1097
PR-#1015

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The added lines document four new tools, while the same README section still states that the server
provides seventeen tools.

Rule 2270057: Keep CLI documentation in README.md in sync with user-facing CLI changes
README.md[728-743]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The README lists four new work-items tools but still says there are seventeen.

## Fix Focus Areas
- README.md[728-743]

## Recommended Fix
Change the introduction to say the work-items server provides 21 tools.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


6. Large evaluations stall agent tools ✓ Resolved
Description
RenderRun walks every question and nested requirement, while Citations builds a list of every
citation before displaying only ten; the evaluation response is also read without a size limit. A
sufficiently large evaluation response can consume substantial memory and processing time, blocking
other calls because the MCP server handles requests one at a time.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[R227-229]

+        var rendered = arr.EnumerateArray()
+            .Where(c => c.IsObject)
+            .Select(c => {
Relevance

●●● Strong

Recent work-items precedents accepted bounded reads and sequential-loop protection for oversized
responses.

PR-#1178
PR-#291

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new evaluation path uses the unrestricted response read, and the renderer materializes all
citations before applying its ten-entry display cap. The existing next-work path demonstrates a
bounded HTTP-content read, while the server's stdio loop processes requests sequentially.

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-255]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[120-142]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[224-245]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[272-308]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[47-52]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[103-114]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation responses have no size limit, and rendering traverses unbounded collections, including citations that are fully materialized before the display cap applies.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-255]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[120-142]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[224-245]
## Recommended Fix
Apply a response-size bound to evaluation calls before parsing, return a tool error when exceeded, and limit collection traversal and output. Count omitted citations without materializing all rendered citation strings.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


7. A size-limit comment overstates coverage ✓ Resolved
Description
The comment on MaxResponseBytes claims that larger responses are never read, but the evaluation
and cancellation calls use PostAsync, which buffers their responses before
BoundedHttpContent.ReadAsync runs. When either endpoint returns an oversized body, the full reply
can be read into memory before the one-megabyte limit is checked.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[22]

+    /// call at a time, so nothing larger is read.</summary>
Relevance

●● Moderate

Response-size safety findings are accepted, but recent comment-accuracy findings have also been
rejected.

PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Both evaluation POST routes use PostAsync, and the bounded read occurs only after each HTTP call
returns. The bounded-reader documentation states that its limit applies only with
ResponseHeadersRead, so these calls do not support the unconditional read-bound claim in the
MaxResponseBytes comment.

Rule 2762993: Restrict comments to documenting non-obvious, behavior‑critical constraints
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-23]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[232-245]
src/Capacitor.Cli/BoundedHttpContent.cs[3-5]
src/Capacitor.Cli/BoundedHttpContent.cs[1-16]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation and cancellation POST replies are buffered before the bounded reader sees them, making the response-size comment inaccurate for those calls.

## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[232-245]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-23]

## Recommended Fix
Send both evaluation POST requests with `HttpCompletionOption.ResponseHeadersRead`, then apply the existing bounded reader to their response streams. Keep the `MaxResponseBytes` comment aligned with the resulting behavior.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Context sources
✅ Compliance rules (platform): 64 rules
✅ Cross-repo context — repo relationships
  Explored: repo: kurrent-io/kcap-server (sha: dfec14ab) — View relationship
  Explored: repo: kurrent-io/kcap-deployments (sha: de9ba18f) — View relationship
Review mode: ⚖️ Balanced: The push contains multiple behavioral runtime changes across HTTP response handling and evaluation rendering, with enough independent edge cases to warrant a careful single-pass review.

Grey Divider

Tip of the day
💡 Did you know, you can route each severity your way: inline, summary, both, or drop

More tips ↗ | Customize Qodo ↗ | Qodo docs ↗

Grey Divider

Previous reviews

Review updated until commit 0057022

Results up to commit 02cde97 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Action required
1. Evaluation tools cannot reach server ✗ Dismissed
Description
McpWorkItemsServer sends all four evaluation tools to /api/work-items/{id}/evals/runs routes,
but the pinned kcap-server registers no evaluation HTTP routes. Against that server, the requests
receive 404 responses, which WorkItemEvalToolResults.Render can report as a non-error “not
enabled” result when the response is not JSON.
Code

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[R228-229]

+                "list_work_item_evals"   => await client.GetAsync(ItemUrl(baseUrl, arguments, "work_item_id",
+                    WorkItemEvalToolResults.ListSuffix(McpToolArguments.OptionalString(arguments, "cursor")))),
Relevance

●● Moderate

Potential cross-repository route incompatibility is serious, but no close rejection or acceptance
precedent confirms expected server versioning.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The PR issues requests to the new paths and treats a non-JSON 404 as unavailable. The server's
public work-item route registration contains no evaluation routes, and its existing evaluation
service is a web-service interface rather than an HTTP endpoint.

kcap-cli -> kcap-server
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-235]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[55-64]
External repo: kurrent-io/kcap-server, src/Capacitor.Api.Public/WorkItems/WorkItemEndpoints.cs [18-78]
External repo: kurrent-io/kcap-server, src/Capacitor.Api.Web.Abstractions/WorkItems/Evals/IWorkItemEvalDataService.cs [1-10]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new CLI evaluation tools target HTTP routes absent from the pinned kcap-server.

## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-235]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[55-64]
- /cross_repos/kcap-server/src/Capacitor.Api.Public/WorkItems/WorkItemEndpoints.cs[18-78]

## Recommended Fix
Implement and deploy the four authenticated evaluation HTTP operations in kcap-server with request and response contracts matching the CLI, including queued-run cancellation. Coordinate the server release with these CLI tools and test against the deployed server routes.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools



Remediation recommended
2. Large evaluations stall agent tools ✓ Resolved
Description
RenderRun walks every question and nested requirement, while Citations builds a list of every
citation before displaying only ten; the evaluation response is also read without a size limit. A
sufficiently large evaluation response can consume substantial memory and processing time, blocking
other calls because the MCP server handles requests one at a time.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[R227-229]

+        var rendered = arr.EnumerateArray()
+            .Where(c => c.IsObject)
+            .Select(c => {
Relevance

●●● Strong

Recent work-items precedents accepted bounded reads and sequential-loop protection for oversized
responses.

PR-#1178
PR-#291

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The new evaluation path uses the unrestricted response read, and the renderer materializes all
citations before applying its ten-entry display cap. The existing next-work path demonstrates a
bounded HTTP-content read, while the server's stdio loop processes requests sequentially.

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-255]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[120-142]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[224-245]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[272-308]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[47-52]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[103-114]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation responses have no size limit, and rendering traverses unbounded collections, including citations that are fully materialized before the display cap applies.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-255]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[120-142]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[224-245]
## Recommended Fix
Apply a response-size bound to evaluation calls before parsing, return a tool error when exceeded, and limit collection traversal and output. Count omitted citations without materializing all rendered citation strings.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. The work-items tool count is outdated ✓ Resolved
Description
The README adds four evaluation tools but leaves the introduction saying `It provides seventeen
tools`. The work-items section now lists 21 tools, so readers get a count that contradicts the
documented commands.
Code

README.md[R742-743]

+- **`list_work_item_evals`** / **`get_work_item_eval`** — list a work item's evaluation runs you can read and read one run's results, judged requirements and retrospective.
+- **`request_work_item_eval`** / **`cancel_work_item_eval`** — queue an evaluation of a work item (`mode` `process` or `root_cause`), or cancel your own run while it is still queued.
Relevance

●●● Strong

Recent README synchronization findings were accepted, including documentation mismatches after
user-facing changes.

PR-#1097
PR-#1015

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The added lines document four new tools, while the same README section still states that the server
provides seventeen tools.

Rule 2270057: Keep CLI documentation in README.md in sync with user-facing CLI changes
README.md[728-743]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The README lists four new work-items tools but still says there are seventeen.

## Fix Focus Areas
- README.md[728-743]

## Recommended Fix
Change the introduction to say the work-items server provides 21 tools.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit 804c74d ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Remediation recommended
1. Retrospective entries disappear silently ✓ Resolved
Description
AppendRetrospective and Items take only the first 20 suggestions, strengths, and issues without
reporting how many remain. When a returned retrospective exceeds that cap, an agent reading the
result cannot tell that some entries were omitted, unlike the capped questions and requirements.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[R225-227]

+        foreach (var s in suggestions.EnumerateArray().Take(MaxListItems))
+            if (s.IsObject && Text(s.Str("text"), LongTextCap) is { Length: > 0 } text)
+                Line(sb, $"  suggestion ({Text(s.Str("audience"), CodeCap)}): {text}");
Relevance

●●● Strong

Existing capped-list behavior reports omissions; adding retrospective omission markers is a
consistent deterministic correctness fix.

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The two retrospective rendering loops stop at MaxListItems without an omission marker. The
adjacent capped question and requirement loops call More to disclose truncated results.

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-29]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[146-149]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[201-212]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[217-236]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
The new retrospective list caps silently discard entries after the twentieth.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[217-236]
## Recommended Fix
After rendering each capped retrospective list, emit an omission count using the existing `More` helper when additional entries are present.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


2. Slow evaluations block other tools ✓ Resolved
Description
The evaluation GETs return after headers, but BoundedHttpContent.ReadAsync then reads their bodies
with CancellationToken.None. If a server sends headers and stalls or trickles the body without
reaching the byte limit, the single-call stdio loop remains blocked.
Code

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[243]

+                var bytes = await BoundedHttpContent.ReadAsync(httpResponse.Content, WorkItemEvalToolResults.MaxResponseBytes, CancellationToken.None);
Relevance

●●● Strong

Nearly identical timeout finding was accepted recently for the same bounded reader and sequential
MCP loop.

PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
The GET calls opt into ResponseHeadersRead, the bounded reader waits on stream reads using the
supplied token, and the existing next-work path demonstrates a deadline applied to both the request
and body read because the stdio loop serves one call at a time.

src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-245]
src/Capacitor.Cli/BoundedHttpContent.cs[8-16]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[282-310]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation GET body reads have no cancellation deadline after response headers arrive.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[228-245]
## Recommended Fix
Use one deadline token for each evaluation HTTP request and its bounded body read, and return a tool error when that deadline expires.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


3. A size-limit comment overstates coverage ✓ Resolved
Description
The comment on MaxResponseBytes claims that larger responses are never read, but the evaluation
and cancellation calls use PostAsync, which buffers their responses before
BoundedHttpContent.ReadAsync runs. When either endpoint returns an oversized body, the full reply
can be read into memory before the one-megabyte limit is checked.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[22]

+    /// call at a time, so nothing larger is read.</summary>
Relevance

●● Moderate

Response-size safety findings are accepted, but recent comment-accuracy findings have also been
rejected.

PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
Both evaluation POST routes use PostAsync, and the bounded read occurs only after each HTTP call
returns. The bounded-reader documentation states that its limit applies only with
ResponseHeadersRead, so these calls do not support the unconditional read-bound claim in the
MaxResponseBytes comment.

Rule 2762993: Restrict comments to documenting non-obvious, behavior‑critical constraints
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-23]
src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[232-245]
src/Capacitor.Cli/BoundedHttpContent.cs[3-5]
src/Capacitor.Cli/BoundedHttpContent.cs[1-16]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Evaluation and cancellation POST replies are buffered before the bounded reader sees them, making the response-size comment inaccurate for those calls.

## Fix Focus Areas
- src/Capacitor.Cli/Commands/McpWorkItemsServer.cs[232-245]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[21-23]

## Recommended Fix
Send both evaluation POST requests with `HttpCompletionOption.ResponseHeadersRead`, then apply the existing bounded reader to their response streams. Keep the `MaxResponseBytes` comment aligned with the resulting behavior.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Results up to commit e4721f7 ⚖️ Balanced


🐞 Bugs (0) 📘 Rule violations (0) 📎 Requirement gaps (0) 🎨 UX issues (0) 🔗 Cross-repo conflicts (0) 📜 Skill insights (0)


Remediation recommended
1. Incomplete scope reasons disappear ✓ Resolved
Description
Joined renders only the first 20 scope_incomplete_reasons and does not indicate that it omitted
any. When a run contains more than 20 reasons, the agent receives an incomplete explanation of its
evaluation scope without knowing that the explanation was truncated.
Code

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[265]

+            ? string.Join(", ", arr.EnumerateArray().Take(MaxListItems).Select(v => v.IsString ? Text(v.GetString(), cap) : "").Where(s => s.Length > 0))
Relevance

●●● Strong

Accepted history favors explicit omitted-count markers for capped agent-visible lists, matching
existing question and requirement rendering.

PR-#1211
PR-#1178

ⓘ Recommendations generated based on similar findings in past PRs

Evidence
RenderRun puts the value returned by Joined directly on the scope line. Joined takes only
MaxListItems values, which is 20, and returns no omitted count; the renderer uses an explicit
omitted-count message for capped questions and requirements.

src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[26-29]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[140-150]
src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[263-269]

Agent prompt
The issue below was found during a code review. Follow the provided context and guidance below and implement a solution

## Issue description
Runs with more than 20 scope-incomplete reasons silently lose the remaining reasons in the rendered result.
## Fix Focus Areas
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[140-143]
- src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs[263-266]
## Recommended Fix
Keep the display cap, but append a count of omitted reasons to the scope line when the input array exceeds it. Add a test with more than 20 reasons.

ⓘ Copy this prompt and use it to remediate the issue with your preferred AI generation tools


Grey Divider

Qodo Logo

Comment thread README.md
Comment thread src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs Outdated
Comment thread src/Capacitor.Cli/Commands/McpWorkItemsServer.cs Outdated
The stdio loop serves one call at a time, so an oversized run is refused before
it is parsed rather than rendered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@realtonyyoung

Copy link
Copy Markdown
Collaborator Author

/agentic_review

Comment thread src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs
Comment thread src/Capacitor.Cli/Commands/McpWorkItemsServer.cs Outdated
Comment thread src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs
@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit 804c74d

realtonyyoung and others added 2 commits September 30, 2026 16:33
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A headers-first read leaves the body outside HttpClient's timeout, and the stdio
loop serves one call at a time.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@realtonyyoung

Copy link
Copy Markdown
Collaborator Author

/agentic_review

Comment thread src/Capacitor.Cli/Commands/WorkItemEvalToolResults.cs Outdated
@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit e4721f7

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@realtonyyoung

Copy link
Copy Markdown
Collaborator Author

/agentic_review

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@qodo-code-review

Copy link
Copy Markdown

Code review by qodo was updated up to the latest commit f5ade4a

@realtonyyoung
realtonyyoung merged commit 6798456 into main Sep 30, 2026
8 checks passed
@realtonyyoung
realtonyyoung deleted the tonyyoung/ai-3267-work-item-eval-mcp-tools branch September 30, 2026 21:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant