diff --git a/.github/workflows/feature-ideation-reusable.yml b/.github/workflows/feature-ideation-reusable.yml index a582231f..1fe12762 100644 --- a/.github/workflows/feature-ideation-reusable.yml +++ b/.github/workflows/feature-ideation-reusable.yml @@ -23,7 +23,7 @@ # place propagates to every BMAD-enabled repo after the v1 tag is bumped # and their next scheduled run executes. Repos tracking @main pick it up # immediately on their next run. -# - The two critical gotchas (github_token override, ANTHROPIC_MODEL env var) +# - The two critical gotchas (github_token override, the `model` input) # are baked in here so callers cannot accidentally regress them. # # Standard: https://github.com/petry-projects/.github/blob/main/standards/ci-standards.md#8-feature-ideation-feature-ideationyml--bmad-method-repos @@ -50,9 +50,12 @@ on: default: 'standard' type: string model: - description: 'Anthropic model ID. Opus 4.6 is the default and recommended — Sonnet produces shallower adversarial passes.' + description: > + Model family — `opus` (default, recommended), `sonnet`, or `haiku`. + The resolver/CLI maps the family to the current version ID. An + operator MAY override with a concrete model ID for a one-off pin. required: false - default: 'claude-opus-4-6' + default: 'opus' type: string timeout_minutes: description: 'Timeout for the analyst job (the gather-signals job has its own short timeout).' @@ -464,7 +467,7 @@ jobs: avoid missing major developments. This keeps each weekly run focused on genuinely new signal rather than - re-processing the same content at Opus 4.6 cost. + re-processing the same content at Opus cost. Using the **Project Context** above as your starting point and the source list as your seed, research the competitive landscape and diff --git a/standards/agent-standards.md b/standards/agent-standards.md index 0a206901..45904d1b 100644 --- a/standards/agent-standards.md +++ b/standards/agent-standards.md @@ -114,6 +114,44 @@ The AgentShield action adds the agent-specific security layer on top. See [AGENTS.md § Decision Logic Lives in a Pure, Tested Script](../AGENTS.md#decision-logic-lives-in-a-pure-tested-script) for the full standard, exemplars, and rationale. +## Model selection + +Agent code names a **model family**, never a pinned version ID. A caller MAY +suggest the family best suited to the task — **`opus`**, **`sonnet`**, or +**`haiku`** — but MUST NOT hard-code a specific version ID such as +`claude-opus-4-6`. A single resolver/CLI maps a family to the current model ID, +and version changes roll out centrally through the normal release channels — so +moving the whole fleet to a newer model is one change in the resolver, not a +find-and-replace of pinned IDs across every workflow, script, and prompt. + +**Why family, not version.** A hard-coded version ID is drift waiting to happen: +every place that names `claude-opus-4-6` must be found and edited on each model +bump, migrations land unevenly, and a stale ID silently pins a repo to an +outdated model. Naming the family defers the ID to the one resolver, so the +version lives in exactly one place and is promoted like any other release. + +### Allowed exceptions — each needs an audit marker + +A literal version ID is permitted **only** in these four cases, and each +occurrence MUST carry an audit marker so the pin is auditable and intentional: + +| # | Exception | Why a real ID is required | Audit marker | +|---|-----------|---------------------------|--------------| +| a | **The resolver itself** | It is the one place that maps family → current ID, so it must name the IDs. | Inline `# model-pin-ok: ` | +| b | **Price data keyed by real IDs** | Cost is per concrete model, so the table is keyed by the actual version IDs. | Inline `# model-pin-ok: ` for comment-supporting formats; metadata field or sidecar file for formats that don't support comments (e.g. a JSON price table). | +| c | **Recorded data** (fixtures, eval sets, baselines) | A captured artifact records the ID that produced it; rewriting it would falsify the record. | Inline `# model-pin-ok: ` for comment-supporting formats; metadata field or sidecar file for formats that don't support comments. | +| d | **A fixed eval judge** | The judge must stay pinned so A/B results stay comparable across runs. | Inline `# model-pin-ok: ` | + +Anything outside these four names a family and lets the resolver supply the ID. + +### Operator override + +An **operator-supplied full model ID** — a workflow input or an Actions variable +set by a human operator — is an operator choice and stays allowed; it is not a +code default. The rule constrains **defaults in code**: those MUST name a family. +An operator MAY still pass a concrete ID to override for a one-off experiment or +pin, without that ID ever becoming the hard-coded default. + ## BMAD Method Workflows Repositories with BMAD Method installed (presence of `_bmad/`, `_bmad-output/`, @@ -123,6 +161,6 @@ the market and produce evidence-grounded feature proposals as GitHub Discussions See [CI Standards §8 — Feature Ideation](ci-standards.md#8-feature-ideation-feature-ideationyml--bmad-method-repos) for the full standard, including the multi-skill ideation pipeline and the -critical configuration gotchas (Opus 4.6 model selection, GitHub token override, -log-secret hygiene). The template is at +critical configuration gotchas (Opus-family [model selection](#model-selection), +GitHub token override, log-secret hygiene). The template is at [`standards/workflows/feature-ideation.yml`](workflows/feature-ideation.yml). diff --git a/standards/ci-standards.md b/standards/ci-standards.md index 46871952..9e8d1b5a 100644 --- a/standards/ci-standards.md +++ b/standards/ci-standards.md @@ -1242,7 +1242,7 @@ These workflows are required only when a specific ecosystem is detected. (warning) by the audit. The BMAD Method framing below reflects the original pilot; the pipeline itself is not BMAD-specific. -Scheduled weekly workflow that runs the BMAD Analyst (Mary) on **Claude Opus 4.6** +Scheduled weekly workflow that runs the BMAD Analyst (Mary) on the **Claude Opus family** through a 5-phase multi-skill ideation pipeline, producing evidence-grounded feature proposals as GitHub Discussions in the **Ideas** category. Each proposal is a separate Discussion, updated by subsequent runs as the market and project @@ -1331,13 +1331,13 @@ The adversarial pass is the load-bearing part: ideas that survive it are | Setting | Value | |---------|-------| -| **Model** | `claude-opus-4-6` (set via `ANTHROPIC_MODEL` env var on the step) | +| **Model** | `opus` family by default (the reusable workflow's `model` input, passed as `--model`) | | **Schedule** | Weekly (template uses Friday 07:00 UTC) | | **Output** | GitHub Discussions in the Ideas category, one per proposal | | **Inputs** | `focus_area` (optional), `research_depth` (quick/standard/deep) | | **Permissions** | `contents: read`, `discussions: write`, `id-token: write` | | **Required secrets** | `CLAUDE_CODE_OAUTH_TOKEN` (org-level) | -| **Typical cost** | ~$2-3 per run on Opus 4.6, standard depth, 25-40 turns | +| **Typical cost** | ~$2-3 per run on Opus, standard depth, 25-40 turns | **Prerequisite:** Discussions must be enabled with an "Ideas" category (see [Discussions Configuration](github-settings.md#discussions-configuration)). @@ -1407,12 +1407,15 @@ understanding why they exist:** no Discussions. Passing the workflow's `GITHUB_TOKEN` makes the job-level `permissions: discussions: write` grant apply. -2. **`ANTHROPIC_MODEL: claude-opus-4-6` is set as a step env var.** - The action does not expose model selection as an input — it reads the - `ANTHROPIC_MODEL` environment variable. Opus is required for the depth - the multi-skill pipeline expects; Sonnet runs cheaper but produces - noticeably shallower adversarial passes. The reusable workflow exposes - this as the optional `model` input for callers that need an override. +2. **The `model` input defaults to the `opus` family.** + The reusable workflow passes this via `--model ${{ inputs.model }}` to the + Claude Code action, which then resolves the family to the current version ID. + Opus is required for the depth the multi-skill pipeline expects; Sonnet runs + cheaper but produces noticeably shallower adversarial passes. The `model` + input is optional — callers may override it per the + [operator override](agent-standards.md#operator-override) + described in agent-standards.md, which is the canonical family-not-version + rule this gotcha follows. 3. **`show_full_output: true` is NOT enabled.** It echoes raw tool results to public action logs, which can leak secrets. @@ -1430,7 +1433,7 @@ understanding why they exist:** | `project_context` | yes | — | 3-5 sentence project description; the only required input | | `focus_area` | no | `''` | Optional research focus, typically wired to `workflow_dispatch` input | | `research_depth` | no | `'standard'` | `quick` / `standard` / `deep` | -| `model` | no | `'claude-opus-4-6'` | Override only for cost experiments — see gotcha #2 | +| `model` | no | `'opus'` | Model family; override only for cost experiments — see gotcha #2 | | `timeout_minutes` | no | `60` | Analyst job timeout (signal collection has its own short timeout) | | Secret | Required | Notes |