Story
As a person working with the Sidekick in xmd repl, I want it to actively improve how we conduct work, protect my cognitive load and investigate failures, so collaboration becomes easier and more reliable as we use it.
For example, several Agents return long, overlapping reports and a workflow repeatedly fails at the same handoff. The Sidekick notices the reading burden, presents the decision I need to make with supporting detail available, and investigates the failed handoff. It can create and test a reusable reporting component and suggest a targeted workflow or XMD improvement grounded in what actually happened.
Current gap
#883 establishes a conversational Sidekick with operations for working in the REPL. #896 adds component authoring and tests; #897 adds middleware that invokes the Sidekick at selected execution events. Those capabilities do not by themselves establish responsibility for improving its own working methods, the person's experience or flow reliability.
A Sidekick that waits for a direct instruction to summarize, reflect or diagnose misses this outcome even if it can perform all those operations on request.
Accepted roles
Optimizer of its own execution
The Sidekick reflects on its experience conducting work and identifies specific improvements to XMD, the REPL, collaboration with the person and its management of other Agents.
Examples include repeated difficulty expressing a useful operation in XMD, unnecessary manual handoffs, duplicated Agent work, or coordination that repeatedly requires the person to repair context. It connects a suggestion to the observed friction and explains the practical improvement. Optimizing its execution includes working methods and coordination, rather than only speed or token use.
Protector of the user's cognitive load
The Sidekick looks for opportunities to reduce the reading and mental bookkeeping needed to remain productive. It organizes communication around relevant outcomes, decisions and actions, with supporting evidence accessible when needed.
It creates reusable components where they make communication more effective, and suggests better collaboration patterns when the problem extends beyond formatting. For example, it can turn several Agent reports into a concise comparison of disagreements and the decision they require. Reducing reading preserves consequential failures, uncertainty and tradeoffs; a shorter message that hides the decision is insufficient.
Troubleshooter
When a flow fails, the Sidekick investigates why and suggests improvements that make the flow more reliable or robust. It uses available execution history, component source, test results and Agent outcomes to distinguish the observed failure from a supported explanation and an unresolved hypothesis.
It proposes a useful next diagnostic step when the evidence does not establish a cause. Where appropriate, it can create a regression test, revise a component and run its Markdown tests through the capabilities in #896. Repeating the error message or retrying until something passes does not establish a diagnosis or a reliability improvement.
Initiative and communication
These are ongoing responsibilities: the Sidekick identifies useful opportunities without requiring the person to assign each role on every turn. Its observations can arise during work, after an outcome or while reviewing its experience.
An improvement proposal explains what happened, what change could help and why it matters. The Sidekick uses its existing authorized operations to act where appropriate and distinguishes an applied change from a suggestion. This Story does not make every suggestion a required approval step or grant unrestricted authority to change XMD, the REPL or other Agents.
Proactivity must serve the cognitive-load role too. Repeated generic retrospectives, unsolicited report dumps and recurring suggestions the person has already dismissed do not satisfy it.
Decisions before implementation
Settle the following with the Product Owner before freezing a Planner handoff:
- When to intervene. Choose useful triggers and review moments, how urgent failures differ from optional improvements, and how the Sidekick avoids interrupting productive work or repeatedly raising the same point.
- What it can observe. Identify the conversation, execution, test and Agent-coordination evidence available to each role. Show how the Sidekick references that evidence and acknowledges missing information.
- How learning carries forward. Decide where observations, proposed improvements, dismissed suggestions and adopted working methods live, including their lifetime across entries and cold reopen. A retrospective is not evidence that a proposed improvement has already helped.
- What it may change. Define when it can create or revise communication components and tests using existing authority, and how proposals affecting the product, workflow or Agent coordination become concrete work. Do not infer new publishing or repository-change permissions from these roles.
- What the person sees. Show a concise finding, its supporting evidence, an actionable recommendation and any applied result. Define how the person requests detail, corrects a diagnosis, dismisses an idea or adjusts the Sidekick's initiative.
The role responsibilities above are accepted direction. Their activation, persistence and presentation remain design work rather than implementation choices delegated by this Story.
Product verification proposal
Use representative sessions and inspect the resulting behavior, rather than requiring one exact model-generated wording:
- Recurring execution friction: a session contains repeated workarounds or inefficient Agent handoffs. Without an explicit retrospective request, the Sidekick identifies a concrete pattern, cites its evidence and proposes a relevant XMD, REPL or collaboration improvement. Generic advice that fits any session fails the comparison.
- Reading burden: several Agents produce overlapping reports with a meaningful disagreement. The Sidekick reduces the material the person must read to make the decision while keeping the disagreement and evidence accessible. Where a reusable communication component helps, show its creation, tests and use in a later entry.
- Failed flow: a representative failure has an inspectable cause. The Sidekick traces it, proposes a targeted correction and a discriminating regression check. In a second case with insufficient evidence, it states what remains uncertain and chooses a useful diagnostic step instead of claiming a cause.
- Attention and feedback: a straightforward successful flow produces no obligatory improvement essay. Correct or dismiss a suggestion and show the approved adaptation without losing important failure information or repeating the same advice.
The Architect and Product Owner refine these journeys before the Planner freezes acceptance. An implementation consisting only of role labels in a prompt is insufficient evidence; demonstrate the roles in actual REPL collaboration and connect claimed improvements to observed results.
Relationships and boundaries
Follow-up to #883, related to REPL Quest #827. #896 supplies component creation, sharing and Markdown-test execution. #897 supplies event-triggered invocation and suspension where the accepted design needs them. This Story defines why and how the Sidekick uses its capabilities to improve the work; it does not require every retrospective or communication improvement to install middleware or pause execution.
Preserve REPL-owned admission, execution lifetimes, provider feedback and the distinction between live work and historical inspection. Reading a historical projection does not itself execute an Agent or apply a repair. Internal Sidekick instructions and user-facing presentation require Product Owner review under the existing interface rules. Keep the accepted behavior in specs/repl-spec.md and any affected architecture description as part of delivery.
Exclude a new Agent orchestration framework, a requirement to run a separate Agent for each role, automatic modification of XMD itself, and a general product analytics or benchmarking platform.
Story
As a person working with the Sidekick in
xmd repl, I want it to actively improve how we conduct work, protect my cognitive load and investigate failures, so collaboration becomes easier and more reliable as we use it.For example, several Agents return long, overlapping reports and a workflow repeatedly fails at the same handoff. The Sidekick notices the reading burden, presents the decision I need to make with supporting detail available, and investigates the failed handoff. It can create and test a reusable reporting component and suggest a targeted workflow or XMD improvement grounded in what actually happened.
Current gap
#883 establishes a conversational Sidekick with operations for working in the REPL. #896 adds component authoring and tests; #897 adds middleware that invokes the Sidekick at selected execution events. Those capabilities do not by themselves establish responsibility for improving its own working methods, the person's experience or flow reliability.
A Sidekick that waits for a direct instruction to summarize, reflect or diagnose misses this outcome even if it can perform all those operations on request.
Accepted roles
Optimizer of its own execution
The Sidekick reflects on its experience conducting work and identifies specific improvements to XMD, the REPL, collaboration with the person and its management of other Agents.
Examples include repeated difficulty expressing a useful operation in XMD, unnecessary manual handoffs, duplicated Agent work, or coordination that repeatedly requires the person to repair context. It connects a suggestion to the observed friction and explains the practical improvement. Optimizing its execution includes working methods and coordination, rather than only speed or token use.
Protector of the user's cognitive load
The Sidekick looks for opportunities to reduce the reading and mental bookkeeping needed to remain productive. It organizes communication around relevant outcomes, decisions and actions, with supporting evidence accessible when needed.
It creates reusable components where they make communication more effective, and suggests better collaboration patterns when the problem extends beyond formatting. For example, it can turn several Agent reports into a concise comparison of disagreements and the decision they require. Reducing reading preserves consequential failures, uncertainty and tradeoffs; a shorter message that hides the decision is insufficient.
Troubleshooter
When a flow fails, the Sidekick investigates why and suggests improvements that make the flow more reliable or robust. It uses available execution history, component source, test results and Agent outcomes to distinguish the observed failure from a supported explanation and an unresolved hypothesis.
It proposes a useful next diagnostic step when the evidence does not establish a cause. Where appropriate, it can create a regression test, revise a component and run its Markdown tests through the capabilities in #896. Repeating the error message or retrying until something passes does not establish a diagnosis or a reliability improvement.
Initiative and communication
These are ongoing responsibilities: the Sidekick identifies useful opportunities without requiring the person to assign each role on every turn. Its observations can arise during work, after an outcome or while reviewing its experience.
An improvement proposal explains what happened, what change could help and why it matters. The Sidekick uses its existing authorized operations to act where appropriate and distinguishes an applied change from a suggestion. This Story does not make every suggestion a required approval step or grant unrestricted authority to change XMD, the REPL or other Agents.
Proactivity must serve the cognitive-load role too. Repeated generic retrospectives, unsolicited report dumps and recurring suggestions the person has already dismissed do not satisfy it.
Decisions before implementation
Settle the following with the Product Owner before freezing a Planner handoff:
The role responsibilities above are accepted direction. Their activation, persistence and presentation remain design work rather than implementation choices delegated by this Story.
Product verification proposal
Use representative sessions and inspect the resulting behavior, rather than requiring one exact model-generated wording:
The Architect and Product Owner refine these journeys before the Planner freezes acceptance. An implementation consisting only of role labels in a prompt is insufficient evidence; demonstrate the roles in actual REPL collaboration and connect claimed improvements to observed results.
Relationships and boundaries
Follow-up to #883, related to REPL Quest #827. #896 supplies component creation, sharing and Markdown-test execution. #897 supplies event-triggered invocation and suspension where the accepted design needs them. This Story defines why and how the Sidekick uses its capabilities to improve the work; it does not require every retrospective or communication improvement to install middleware or pause execution.
Preserve REPL-owned admission, execution lifetimes, provider feedback and the distinction between live work and historical inspection. Reading a historical projection does not itself execute an Agent or apply a repair. Internal Sidekick instructions and user-facing presentation require Product Owner review under the existing interface rules. Keep the accepted behavior in
specs/repl-spec.mdand any affected architecture description as part of delivery.Exclude a new Agent orchestration framework, a requirement to run a separate Agent for each role, automatic modification of XMD itself, and a general product analytics or benchmarking platform.