Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion plugins/planning/.claude-plugin/plugin.json
Original file line number Diff line number Diff line change
@@ -1,7 +1,7 @@
{
"$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json",
"name": "planning",
"version": "0.22.3",
"version": "0.23.0",
"userConfig": {
"use_ask_user_question": {
"type": "boolean",
Expand Down
22 changes: 22 additions & 0 deletions plugins/planning/CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,28 @@
All notable changes to the `planning` plugin are documented here. Format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning.

## [0.23.0]

### Added

- **`interview` recommends the downstream session's model, effort, and advisor.**
The interview already reads task complexity and ambiguity to drive its rounds; at
the stop/handoff boundary it now turns that read into a recommendation for how the
execution session should be configured — a **model tier** (capability: raise when
the assistant would be confidently wrong despite full context) and an **effort
level** (thoroughness: raise when it would under-explore or under-verify) picked per
the official distinction, plus the **advisor** pairing when the main model is a
faster tier (a faster main without a stronger advisor is not the recommended config
for non-trivial work). The current model names, tiers, and accepted pairings are
read **live** from the official docs each run and never pinned in the skill (the
durable distinction is stable; the names drift) — mirroring `draft-goal-condition`'s
live-doc discipline. A doc-fetch failure **degrades, never halts**: it falls back to
the durable distinction with a visible note rather than guessing a model name. The
recommendation is advisory (applied via `/model`, `/advisor`, the effort setting),
fires for engineering and general sessions alike, and carries an inverse mid-task
direction (surface "too complex for the current model/effort" when execution
warrants). Detail in the new `skills/interview/context/session-config.md` (#231).

## [0.22.3]

### Changed
Expand Down
30 changes: 30 additions & 0 deletions plugins/planning/skills/interview/SKILL.md
Original file line number Diff line number Diff line change
Expand Up @@ -197,6 +197,36 @@ Route the handoff by what the session produced. **A general (non-engineering) se

Do NOT auto-clear or auto-invoke. Recommend; let the user pull the trigger.

## Session-config recommendation (model, effort, advisor)

The interview already reads task complexity and ambiguity to drive its rounds — so at
the stop/handoff boundary, turn that read into a recommendation for how the
**downstream execution session** should be configured. Two orthogonal knobs, picked
per the official distinction:

- **Model tier (capability)** — raise the model when the assistant would be
*confidently wrong despite full context* (a reasoning ceiling, not missing input).
- **Effort level (thoroughness)** — raise effort when the assistant would
*under-explore or under-verify* (right answer reachable, but it stops short).

When the recommendation keeps a faster main model, pair it with the **advisor**: a
faster main without a stronger advisor is not the recommended config for non-trivial
work — the documented efficiency pairing escalates planning, ambiguous failures, and
completion checks to a stronger advisor instead of paying for the top model every
turn.

**Source the current names live, never pin them.** Model names, tiers, effort levels,
and accepted advisor pairings drift between versions; the durable *distinction* above
is stable, the *names* are not. Fetch them once when you form the recommendation from
the official docs (mirror `draft-goal-condition`'s never-pin discipline). A doc-fetch
failure **degrades, never halts** — fall back to the durable distinction and tell the
user the current names could not be verified live so they confirm against `/model` /
`/advisor`. Frame the whole thing as advisory (the skill cannot read the current
effort/advisor state) and applicable to engineering and general sessions alike. The
Comment thread
kyle-sexton marked this conversation as resolved.
same signals run **mid-task** in the inverse direction — surface "too complex for the
current model/effort" when execution warrants. Full detail, sources, and the
knob-picking signals in [`context/session-config.md`](context/session-config.md).

## What this skill does NOT do

- `context/gotchas.md` — failure patterns from real sessions
Expand Down
89 changes: 89 additions & 0 deletions plugins/planning/skills/interview/context/session-config.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
# Session-config recommendation — model, effort, advisor

Reference detail for the `## Session-config recommendation (model, effort, advisor)`
section of `SKILL.md`.
Read on demand when forming the recommendation at the interview's stop/handoff
boundary. The interview already reads task complexity and ambiguity to drive its
rounds; this turns that read into a recommendation for how the **downstream
execution session** should be configured.

## Two orthogonal knobs

The official guidance separates two levers. Recommend against the right one — they
are not interchangeable:

- **Model tier (capability).** Raise the model when the assistant would be
**confidently wrong despite full context** — the failure is a reasoning ceiling,
not missing information. Signals from the interview: the task turned on subtle
correctness, dense cross-module invariants, or tradeoffs the user themselves found
hard to adjudicate.
- **Effort level (thoroughness).** Raise effort when the assistant would
**under-explore or under-verify** — it can reach the right answer but tends to stop
short. Signals: broad surface area, many files, a verification-heavy acceptance
criteria list, or a task where the risk is a missed case rather than a wrong model.

A task can want both, one, or neither. State which knob each recommendation turns and
why, in the interview's own evidence terms.

## Advisor pairing

A faster main model running **without** a stronger advisor is not the recommended
configuration for non-trivial work: the documented efficiency pairing is a faster
main model that escalates planning, ambiguous failures, and completion checks to a
stronger advisor, rather than paying for the stronger model on every routine turn.
The concrete tier names that fill this **faster-main + stronger-advisor** shape are
exactly the values that drift between versions — and which specific pairings are
accepted drifts with them. Source them live (below), never pin them here: the durable
fact is the *shape* of the pairing, not the names that fill it.

When the recommendation is "keep the faster main model," pair it with the advisor
recommendation. When it is "raise the main model to the top tier," the advisor adds
less — note that and let the user decide.

## Read the live contract — never pin

Current model names, tiers, effort levels, and accepted advisor pairings change
between Claude Code versions. Source them at recommendation time from the official
docs; do not bake them into this skill (the durable *distinction* above is stable —
the *names and tiers* are not). This mirrors `draft-goal-condition`'s never-pin,
live-doc discipline — its fetch-**failure** handling differs (below): there the
fetched value is the deliverable so it halts, here the recommendation is auxiliary so
it degrades.

Primary sources, fetched once when you form the recommendation (not per round):

- `https://code.claude.com/docs/en/model-config` — model aliases and the effort setting
- `https://claude.com/blog/claude-model-and-effort-level-in-claude-code` — which model and effort fit which work
- `https://code.claude.com/docs/en/advisor` — advisor enablement and accepted main+advisor pairings
- `https://claude.com/blog/the-advisor-strategy` — why a faster main + stronger advisor works

**Fetch failure degrades, never halts.** The recommendation is an auxiliary output —
a doc-fetch failure must not block the interview or the Brief. Fall back to the
durable distinction above and tell the user, in the same breath, that the current
model names and pairings could not be verified live (cite the URL) so they confirm
against `/model` and `/advisor` themselves. This is a visible degrade, not a silent
one, and never a guessed-from-memory model name.

## Advisory framing — you cannot read the current config

The skill knows its own main model (stated in the system prompt) but cannot reliably
read the current effort level or whether an advisor is already set. Frame the
recommendation as a delta the user applies, not a fact about their current state:
"if you are not already on X, consider it," plus how to apply it — `/model` for the
model, the effort setting for effort, `/advisor` for the advisor. Do not instruct a
capability (reading the live effort/advisor state) that does not exist.

## Both domains

Complexity and ambiguity apply to engineering and general sessions alike — a hard
general decision can warrant the top model just as a subtle refactor can. Surface the
recommendation for both; it is orthogonal to the engineering/general domain split and
to the `me`/`auto`/`lock` action.

## Inverse direction — mid-task

The same two signals run mid-task, not only at the interview boundary. If execution
starts showing **confidently-wrong-despite-context** (raise the model) or
**under-exploration / under-verification** (raise effort), surface "this may be too
complex for the current model/effort" and recommend the upgrade with the same
knob-picking logic — rather than grinding on under a config the task has outgrown.
Comment thread
kyle-sexton marked this conversation as resolved.
14 changes: 14 additions & 0 deletions plugins/planning/skills/interview/evals/evals.json
Original file line number Diff line number Diff line change
Expand Up @@ -99,6 +99,20 @@
"The override-precedence choice is asked as a question with a recommendation, not silently decided",
"Output does not ask the user to confirm a fact the code already answers"
]
},
{
"id": 9,
"name": "recommends-session-config-from-live-docs",
"prompt": "/planning:interview me — I need to re-architect our authorization layer to support per-resource policies across three services, and I'm unsure which invariants can change safely.",
"expected_output": "At the stop/handoff boundary the skill recommends how to configure the downstream execution session — a model tier and effort level chosen per the capability-vs-thoroughness distinction, plus the advisor pairing when the main model is a faster tier — deriving the current model names and accepted pairings from the live official docs rather than pinned values, framed as an advisory delta the user applies, and degrading gracefully (durable distinction + a visible note) if the docs cannot be fetched rather than halting or guessing a model name.",
"files": [],
"expectations": [
"Output recommends a model tier and effort level for the downstream session, distinguishing capability (model) from thoroughness (effort)",
"Output recommends advisor pairing when the main model is a faster tier, rather than a faster main with no advisor",
"Current model names / tiers / pairings are sourced from the live official docs, not pinned in the skill",
"On a doc-fetch failure the skill degrades to the durable distinction with a visible note and does not halt the interview or guess a model name",
"The recommendation is framed as advisory (applied via /model, /advisor, effort setting), not as a read of the user's current config"
]
}
]
}
2 changes: 1 addition & 1 deletion plugins/planning/skills/interview/templates/checklist.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,7 @@ Copy into `<memory_dir>/<topic-slug>/interview-checklist.md` (default `.work/`;
- [ ] Step 2: Drive the frontier-rounds loop — each round asks every settled-prerequisite question as one numbered set in **inline prose** (`AskUserQuestion` only via the `use_ask_user_question` opt-in; `lock` synthesizes without Q&A); order rounds by blast radius; restate decided/open after each round
- [ ] Step 3: Recognize the stop condition — frontier is empty (every load-bearing unknown resolved or captured as a named assumption) AND user has confirmed the restated shared understanding (`me`/`auto`; `lock` is exempt — invoking it IS the confirmation)
- [ ] Step 4: Persist the contract — engineering: write the PLAN.md Brief section with goal + constraints + acceptance criteria + captured assumptions; general: write the shared-understanding summary, never a Brief (`me` mode: persist each answer incrementally as it locks in; flush before context overflows)
- [ ] Step 5: Hand off — engineering: recommend the next skill (exploration/research for engineering-internal; chain after `/prd` for product-driven); general: deliver the summary and stop, no pipeline handoff
- [ ] Step 5: Hand off — engineering: recommend the next skill (exploration/research for engineering-internal; chain after `/prd` for product-driven); general: deliver the summary and stop, no pipeline handoff. Both: recommend the downstream session's model / effort / advisor per the live-doc-sourced session-config guidance (never a pinned model name)

## Decision tree (`me` mode only)

Expand Down
Loading