From 2203a9e5492f014f3388387e057b628b282def61 Mon Sep 17 00:00:00 2001 From: Kyle Sexton <153232337+kyle-sexton@users.noreply.github.com> Date: Mon, 20 Jul 2026 15:42:36 -0400 Subject: [PATCH 1/2] feat(planning): interview recommends downstream model, effort, and advisor MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The interview already reads task complexity and ambiguity to drive its rounds. At the stop/handoff boundary it now turns that read into a recommendation for the downstream execution session: a model tier (capability), an effort level (thoroughness), and the advisor pairing when the main model is a faster tier — picked per the official capability-vs-thoroughness distinction. Current model names and accepted pairings are read live from the official docs each run and never pinned (durable distinction stable, names drift), mirroring draft-goal-condition's live-doc discipline. A doc-fetch failure degrades to the durable distinction with a visible note rather than halting or guessing a name. Advisory only, fires for engineering and general sessions, and carries an inverse mid-task direction. Closes #231 Co-Authored-By: Claude Opus --- plugins/planning/.claude-plugin/plugin.json | 2 +- plugins/planning/CHANGELOG.md | 22 +++++ plugins/planning/skills/interview/SKILL.md | 30 +++++++ .../interview/context/session-config.md | 88 +++++++++++++++++++ .../skills/interview/evals/evals.json | 14 +++ .../skills/interview/templates/checklist.md | 2 +- 6 files changed, 156 insertions(+), 2 deletions(-) create mode 100644 plugins/planning/skills/interview/context/session-config.md diff --git a/plugins/planning/.claude-plugin/plugin.json b/plugins/planning/.claude-plugin/plugin.json index 46a728174f..dea4d62986 100644 --- a/plugins/planning/.claude-plugin/plugin.json +++ b/plugins/planning/.claude-plugin/plugin.json @@ -1,7 +1,7 @@ { "$schema": "https://json.schemastore.org/claude-code-plugin-manifest.json", "name": "planning", - "version": "0.22.2", + "version": "0.23.0", "userConfig": { "use_ask_user_question": { "type": "boolean", diff --git a/plugins/planning/CHANGELOG.md b/plugins/planning/CHANGELOG.md index 4497212b13..ba508e59df 100644 --- a/plugins/planning/CHANGELOG.md +++ b/plugins/planning/CHANGELOG.md @@ -3,6 +3,28 @@ All notable changes to the `planning` plugin are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); this plugin uses semantic versioning. +## [0.23.0] + +### Added + +- **`interview` recommends the downstream session's model, effort, and advisor.** + The interview already reads task complexity and ambiguity to drive its rounds; at + the stop/handoff boundary it now turns that read into a recommendation for how the + execution session should be configured — a **model tier** (capability: raise when + the assistant would be confidently wrong despite full context) and an **effort + level** (thoroughness: raise when it would under-explore or under-verify) picked per + the official distinction, plus the **advisor** pairing when the main model is a + faster tier (a faster main without a stronger advisor is not the recommended config + for non-trivial work). The current model names, tiers, and accepted pairings are + read **live** from the official docs each run and never pinned in the skill (the + durable distinction is stable; the names drift) — mirroring `draft-goal-condition`'s + live-doc discipline. A doc-fetch failure **degrades, never halts**: it falls back to + the durable distinction with a visible note rather than guessing a model name. The + recommendation is advisory (applied via `/model`, `/advisor`, the effort setting), + fires for engineering and general sessions alike, and carries an inverse mid-task + direction (surface "too complex for the current model/effort" when execution + warrants). Detail in the new `skills/interview/context/session-config.md` (#231). + ## [0.22.2] ### Changed diff --git a/plugins/planning/skills/interview/SKILL.md b/plugins/planning/skills/interview/SKILL.md index 805a94085c..c0d487040d 100644 --- a/plugins/planning/skills/interview/SKILL.md +++ b/plugins/planning/skills/interview/SKILL.md @@ -197,6 +197,36 @@ Route the handoff by what the session produced. **A general (non-engineering) se Do NOT auto-clear or auto-invoke. Recommend; let the user pull the trigger. +## Session-config recommendation (model, effort, advisor) + +The interview already reads task complexity and ambiguity to drive its rounds — so at +the stop/handoff boundary, turn that read into a recommendation for how the +**downstream execution session** should be configured. Two orthogonal knobs, picked +per the official distinction: + +- **Model tier (capability)** — raise the model when the assistant would be + *confidently wrong despite full context* (a reasoning ceiling, not missing input). +- **Effort level (thoroughness)** — raise effort when the assistant would + *under-explore or under-verify* (right answer reachable, but it stops short). + +When the recommendation keeps a faster main model, pair it with the **advisor**: a +faster main without a stronger advisor is not the recommended config for non-trivial +work — the documented efficiency pairing escalates planning, ambiguous failures, and +completion checks to a stronger advisor instead of paying for the top model every +turn. + +**Source the current names live, never pin them.** Model names, tiers, effort levels, +and accepted advisor pairings drift between versions; the durable *distinction* above +is stable, the *names* are not. Fetch them once when you form the recommendation from +the official docs (mirror `draft-goal-condition`'s never-pin discipline). A doc-fetch +failure **degrades, never halts** — fall back to the durable distinction and tell the +user the current names could not be verified live so they confirm against `/model` / +`/advisor`. Frame the whole thing as advisory (the skill cannot read the current +effort/advisor state) and applicable to engineering and general sessions alike. The +same signals run **mid-task** in the inverse direction — surface "too complex for the +current model/effort" when execution warrants. Full detail, sources, and the +knob-picking signals in [`context/session-config.md`](context/session-config.md). + ## What this skill does NOT do - `context/gotchas.md` — failure patterns from real sessions diff --git a/plugins/planning/skills/interview/context/session-config.md b/plugins/planning/skills/interview/context/session-config.md new file mode 100644 index 0000000000..16c2e1e09a --- /dev/null +++ b/plugins/planning/skills/interview/context/session-config.md @@ -0,0 +1,88 @@ +# Session-config recommendation — model, effort, advisor + +Reference detail for the `## Session-config recommendation (model, effort, advisor)` +section of `SKILL.md`. +Read on demand when forming the recommendation at the interview's stop/handoff +boundary. The interview already reads task complexity and ambiguity to drive its +rounds; this turns that read into a recommendation for how the **downstream +execution session** should be configured. + +## Two orthogonal knobs + +The official guidance separates two levers. Recommend against the right one — they +are not interchangeable: + +- **Model tier (capability).** Raise the model when the assistant would be + **confidently wrong despite full context** — the failure is a reasoning ceiling, + not missing information. Signals from the interview: the task turned on subtle + correctness, dense cross-module invariants, or tradeoffs the user themselves found + hard to adjudicate. +- **Effort level (thoroughness).** Raise effort when the assistant would + **under-explore or under-verify** — it can reach the right answer but tends to stop + short. Signals: broad surface area, many files, a verification-heavy acceptance + criteria list, or a task where the risk is a missed case rather than a wrong model. + +A task can want both, one, or neither. State which knob each recommendation turns and +why, in the interview's own evidence terms. + +## Advisor pairing + +A faster main model running **without** a stronger advisor is not the recommended +configuration for non-trivial work: the documented efficiency pairing is a faster +main model that escalates planning, ambiguous failures, and completion checks to a +stronger advisor, rather than paying for the stronger model on every routine turn. +As of the contract this skill targets, that pairing is **Sonnet main + Opus +advisor** — but the current model names and which pairings are accepted are exactly +the values that drift, so source them live (below), never from this sentence. + +When the recommendation is "keep the faster main model," pair it with the advisor +recommendation. When it is "raise the main model to the top tier," the advisor adds +less — note that and let the user decide. + +## Read the live contract — never pin + +Current model names, tiers, effort levels, and accepted advisor pairings change +between Claude Code versions. Source them at recommendation time from the official +docs; do not bake them into this skill (the durable *distinction* above is stable — +the *names and tiers* are not). This mirrors `draft-goal-condition`'s never-pin, +live-doc discipline — its fetch-**failure** handling differs (below): there the +fetched value is the deliverable so it halts, here the recommendation is auxiliary so +it degrades. + +Primary sources, fetched once when you form the recommendation (not per round): + +- `https://code.claude.com/docs/en/model-config.md` — model aliases and the effort setting +- `https://claude.com/blog/claude-model-and-effort-level-in-claude-code` — which model and effort fit which work +- `https://code.claude.com/docs/en/advisor.md` — advisor enablement and accepted main+advisor pairings +- `https://claude.com/blog/the-advisor-strategy` — why a faster main + stronger advisor works + +**Fetch failure degrades, never halts.** The recommendation is an auxiliary output — +a doc-fetch failure must not block the interview or the Brief. Fall back to the +durable distinction above and tell the user, in the same breath, that the current +model names and pairings could not be verified live (cite the URL) so they confirm +against `/model` and `/advisor` themselves. This is a visible degrade, not a silent +one, and never a guessed-from-memory model name. + +## Advisory framing — you cannot read the current config + +The skill knows its own main model (stated in the system prompt) but cannot reliably +read the current effort level or whether an advisor is already set. Frame the +recommendation as a delta the user applies, not a fact about their current state: +"if you are not already on X, consider it," plus how to apply it — `/model` for the +model, the effort setting for effort, `/advisor` for the advisor. Do not instruct a +capability (reading the live effort/advisor state) that does not exist. + +## Both domains + +Complexity and ambiguity apply to engineering and general sessions alike — a hard +general decision can warrant the top model just as a subtle refactor can. Surface the +recommendation for both; it is orthogonal to the engineering/general domain split and +to the `me`/`auto`/`lock` action. + +## Inverse direction — mid-task + +The same two signals run mid-task, not only at the interview boundary. If execution +starts showing **confidently-wrong-despite-context** (raise the model) or +**under-exploration / under-verification** (raise effort), surface "this may be too +complex for the current model/effort" and recommend the upgrade with the same +knob-picking logic — rather than grinding on under a config the task has outgrown. diff --git a/plugins/planning/skills/interview/evals/evals.json b/plugins/planning/skills/interview/evals/evals.json index 86895f4428..4fad3f2331 100644 --- a/plugins/planning/skills/interview/evals/evals.json +++ b/plugins/planning/skills/interview/evals/evals.json @@ -99,6 +99,20 @@ "The override-precedence choice is asked as a question with a recommendation, not silently decided", "Output does not ask the user to confirm a fact the code already answers" ] + }, + { + "id": 9, + "name": "recommends-session-config-from-live-docs", + "prompt": "/planning:interview me — I need to re-architect our authorization layer to support per-resource policies across three services, and I'm unsure which invariants can change safely.", + "expected_output": "At the stop/handoff boundary the skill recommends how to configure the downstream execution session — a model tier and effort level chosen per the capability-vs-thoroughness distinction, plus the advisor pairing when the main model is a faster tier — deriving the current model names and accepted pairings from the live official docs rather than pinned values, framed as an advisory delta the user applies, and degrading gracefully (durable distinction + a visible note) if the docs cannot be fetched rather than halting or guessing a model name.", + "files": [], + "expectations": [ + "Output recommends a model tier and effort level for the downstream session, distinguishing capability (model) from thoroughness (effort)", + "Output recommends advisor pairing when the main model is a faster tier, rather than a faster main with no advisor", + "Current model names / tiers / pairings are sourced from the live official docs, not pinned in the skill", + "On a doc-fetch failure the skill degrades to the durable distinction with a visible note and does not halt the interview or guess a model name", + "The recommendation is framed as advisory (applied via /model, /advisor, effort setting), not as a read of the user's current config" + ] } ] } diff --git a/plugins/planning/skills/interview/templates/checklist.md b/plugins/planning/skills/interview/templates/checklist.md index acaaba8651..78767fa3f1 100644 --- a/plugins/planning/skills/interview/templates/checklist.md +++ b/plugins/planning/skills/interview/templates/checklist.md @@ -9,7 +9,7 @@ Copy into `//interview-checklist.md` (default `.work/`; - [ ] Step 2: Drive the frontier-rounds loop — each round asks every settled-prerequisite question as one numbered set in **inline prose** (`AskUserQuestion` only via the `use_ask_user_question` opt-in; `lock` synthesizes without Q&A); order rounds by blast radius; restate decided/open after each round - [ ] Step 3: Recognize the stop condition — frontier is empty (every load-bearing unknown resolved or captured as a named assumption) AND user has confirmed the restated shared understanding (`me`/`auto`; `lock` is exempt — invoking it IS the confirmation) - [ ] Step 4: Persist the contract — engineering: write the PLAN.md Brief section with goal + constraints + acceptance criteria + captured assumptions; general: write the shared-understanding summary, never a Brief (`me` mode: persist each answer incrementally as it locks in; flush before context overflows) -- [ ] Step 5: Hand off — engineering: recommend the next skill (exploration/research for engineering-internal; chain after `/prd` for product-driven); general: deliver the summary and stop, no pipeline handoff +- [ ] Step 5: Hand off — engineering: recommend the next skill (exploration/research for engineering-internal; chain after `/prd` for product-driven); general: deliver the summary and stop, no pipeline handoff. Both: recommend the downstream session's model / effort / advisor per the live-doc-sourced session-config guidance (never a pinned model name) ## Decision tree (`me` mode only) From be3590c95b70739586ffad5b43a601081b6b2558 Mon Sep 17 00:00:00 2001 From: Kyle Sexton <153232337+kyle-sexton@users.noreply.github.com> Date: Mon, 20 Jul 2026 16:34:28 -0400 Subject: [PATCH 2/2] fix(planning): unpin advisor tier names and align doc URLs to convention Two review findings on the session-config recommendation spoke: - Remove the inline "Sonnet main + Opus advisor" pairing. Eval id 9 asserts names are "not pinned in the skill" and the degrade path must "not guess a model name"; the pinned illustration contradicted both and left a stale fallback anchor under a partial-degrade fetch failure. The durable faster-main + stronger-advisor *shape* stays; the drifting names are sourced live only. - Drop the `.md` suffix from the two code.claude.com/docs URLs to match the suffix-free canonical URL convention in CLAUDE.md (L15-29) and the sibling draft-goal-condition skill. Co-Authored-By: Claude Opus --- .../skills/interview/context/session-config.md | 11 ++++++----- 1 file changed, 6 insertions(+), 5 deletions(-) diff --git a/plugins/planning/skills/interview/context/session-config.md b/plugins/planning/skills/interview/context/session-config.md index 16c2e1e09a..4ddd2c29a7 100644 --- a/plugins/planning/skills/interview/context/session-config.md +++ b/plugins/planning/skills/interview/context/session-config.md @@ -31,9 +31,10 @@ A faster main model running **without** a stronger advisor is not the recommende configuration for non-trivial work: the documented efficiency pairing is a faster main model that escalates planning, ambiguous failures, and completion checks to a stronger advisor, rather than paying for the stronger model on every routine turn. -As of the contract this skill targets, that pairing is **Sonnet main + Opus -advisor** — but the current model names and which pairings are accepted are exactly -the values that drift, so source them live (below), never from this sentence. +The concrete tier names that fill this **faster-main + stronger-advisor** shape are +exactly the values that drift between versions — and which specific pairings are +accepted drifts with them. Source them live (below), never pin them here: the durable +fact is the *shape* of the pairing, not the names that fill it. When the recommendation is "keep the faster main model," pair it with the advisor recommendation. When it is "raise the main model to the top tier," the advisor adds @@ -51,9 +52,9 @@ it degrades. Primary sources, fetched once when you form the recommendation (not per round): -- `https://code.claude.com/docs/en/model-config.md` — model aliases and the effort setting +- `https://code.claude.com/docs/en/model-config` — model aliases and the effort setting - `https://claude.com/blog/claude-model-and-effort-level-in-claude-code` — which model and effort fit which work -- `https://code.claude.com/docs/en/advisor.md` — advisor enablement and accepted main+advisor pairings +- `https://code.claude.com/docs/en/advisor` — advisor enablement and accepted main+advisor pairings - `https://claude.com/blog/the-advisor-strategy` — why a faster main + stronger advisor works **Fetch failure degrades, never halts.** The recommendation is an auxiliary output —