Skip to content
Merged
11 changes: 7 additions & 4 deletions .github/workflows/feature-ideation-reusable.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,7 @@
# place propagates to every BMAD-enabled repo after the v1 tag is bumped
# and their next scheduled run executes. Repos tracking @main pick it up
# immediately on their next run.
# - The two critical gotchas (github_token override, ANTHROPIC_MODEL env var)
# - The two critical gotchas (github_token override, the `model` input)
# are baked in here so callers cannot accidentally regress them.
#
# Standard: https://github.com/petry-projects/.github/blob/main/standards/ci-standards.md#8-feature-ideation-feature-ideationyml--bmad-method-repos
Expand All @@ -50,9 +50,12 @@ on:
default: 'standard'
type: string
model:
description: 'Anthropic model ID. Opus 4.6 is the default and recommended — Sonnet produces shallower adversarial passes.'
description: >
Model family — `opus` (default, recommended), `sonnet`, or `haiku`.
The resolver/CLI maps the family to the current version ID. An
operator MAY override with a concrete model ID for a one-off pin.
required: false
default: 'claude-opus-4-6'
default: 'opus'
type: string
timeout_minutes:
description: 'Timeout for the analyst job (the gather-signals job has its own short timeout).'
Expand Down Expand Up @@ -464,7 +467,7 @@ jobs:
avoid missing major developments.

This keeps each weekly run focused on genuinely new signal rather than
re-processing the same content at Opus 4.6 cost.
re-processing the same content at Opus cost.

Using the **Project Context** above as your starting point and the
source list as your seed, research the competitive landscape and
Expand Down
42 changes: 40 additions & 2 deletions standards/agent-standards.md
Original file line number Diff line number Diff line change
Expand Up @@ -114,6 +114,44 @@ The AgentShield action adds the agent-specific security layer on top.
See [AGENTS.md § Decision Logic Lives in a Pure, Tested Script](../AGENTS.md#decision-logic-lives-in-a-pure-tested-script)
for the full standard, exemplars, and rationale.

## Model selection

Agent code names a **model family**, never a pinned version ID. A caller MAY
suggest the family best suited to the task — **`opus`**, **`sonnet`**, or
**`haiku`** — but MUST NOT hard-code a specific version ID such as
`claude-opus-4-6`. A single resolver/CLI maps a family to the current model ID,
and version changes roll out centrally through the normal release channels — so
moving the whole fleet to a newer model is one change in the resolver, not a
find-and-replace of pinned IDs across every workflow, script, and prompt.

**Why family, not version.** A hard-coded version ID is drift waiting to happen:
every place that names `claude-opus-4-6` must be found and edited on each model
bump, migrations land unevenly, and a stale ID silently pins a repo to an
outdated model. Naming the family defers the ID to the one resolver, so the
version lives in exactly one place and is promoted like any other release.

### Allowed exceptions — each needs an audit marker

A literal version ID is permitted **only** in these four cases, and each
occurrence MUST carry an audit marker so the pin is auditable and intentional:

| # | Exception | Why a real ID is required | Audit marker |
|---|-----------|---------------------------|--------------|
| a | **The resolver itself** | It is the one place that maps family → current ID, so it must name the IDs. | Inline `# model-pin-ok: <reason>` |
| b | **Price data keyed by real IDs** | Cost is per concrete model, so the table is keyed by the actual version IDs. | Inline `# model-pin-ok: <reason>` for comment-supporting formats; metadata field or sidecar file for formats that don't support comments (e.g. a JSON price table). |
| c | **Recorded data** (fixtures, eval sets, baselines) | A captured artifact records the ID that produced it; rewriting it would falsify the record. | Inline `# model-pin-ok: <reason>` for comment-supporting formats; metadata field or sidecar file for formats that don't support comments. |
| d | **A fixed eval judge** | The judge must stay pinned so A/B results stay comparable across runs. | Inline `# model-pin-ok: <reason>` |

Anything outside these four names a family and lets the resolver supply the ID.

### Operator override

An **operator-supplied full model ID** — a workflow input or an Actions variable
set by a human operator — is an operator choice and stays allowed; it is not a
code default. The rule constrains **defaults in code**: those MUST name a family.
Comment thread
cubic-dev-ai[bot] marked this conversation as resolved.
An operator MAY still pass a concrete ID to override for a one-off experiment or
pin, without that ID ever becoming the hard-coded default.

## BMAD Method Workflows

Repositories with BMAD Method installed (presence of `_bmad/`, `_bmad-output/`,
Expand All @@ -123,6 +161,6 @@ the market and produce evidence-grounded feature proposals as GitHub Discussions

See [CI Standards §8 — Feature Ideation](ci-standards.md#8-feature-ideation-feature-ideationyml--bmad-method-repos)
for the full standard, including the multi-skill ideation pipeline and the
critical configuration gotchas (Opus 4.6 model selection, GitHub token override,
log-secret hygiene). The template is at
critical configuration gotchas (Opus-family [model selection](#model-selection),
GitHub token override, log-secret hygiene). The template is at
[`standards/workflows/feature-ideation.yml`](workflows/feature-ideation.yml).
23 changes: 13 additions & 10 deletions standards/ci-standards.md
Original file line number Diff line number Diff line change
Expand Up @@ -1242,7 +1242,7 @@ These workflows are required only when a specific ecosystem is detected.
(warning) by the audit. The BMAD Method framing below reflects the original
pilot; the pipeline itself is not BMAD-specific.

Scheduled weekly workflow that runs the BMAD Analyst (Mary) on **Claude Opus 4.6**
Scheduled weekly workflow that runs the BMAD Analyst (Mary) on the **Claude Opus family**
through a 5-phase multi-skill ideation pipeline, producing evidence-grounded
feature proposals as GitHub Discussions in the **Ideas** category. Each proposal
is a separate Discussion, updated by subsequent runs as the market and project
Expand Down Expand Up @@ -1331,13 +1331,13 @@ The adversarial pass is the load-bearing part: ideas that survive it are

| Setting | Value |
|---------|-------|
| **Model** | `claude-opus-4-6` (set via `ANTHROPIC_MODEL` env var on the step) |
| **Model** | `opus` family by default (the reusable workflow's `model` input, passed as `--model`) |
| **Schedule** | Weekly (template uses Friday 07:00 UTC) |
| **Output** | GitHub Discussions in the Ideas category, one per proposal |
| **Inputs** | `focus_area` (optional), `research_depth` (quick/standard/deep) |
| **Permissions** | `contents: read`, `discussions: write`, `id-token: write` |
| **Required secrets** | `CLAUDE_CODE_OAUTH_TOKEN` (org-level) |
| **Typical cost** | ~$2-3 per run on Opus 4.6, standard depth, 25-40 turns |
| **Typical cost** | ~$2-3 per run on Opus, standard depth, 25-40 turns |

**Prerequisite:** Discussions must be enabled with an "Ideas" category
(see [Discussions Configuration](github-settings.md#discussions-configuration)).
Expand Down Expand Up @@ -1407,12 +1407,15 @@ understanding why they exist:**
no Discussions. Passing the workflow's `GITHUB_TOKEN` makes the job-level
`permissions: discussions: write` grant apply.

2. **`ANTHROPIC_MODEL: claude-opus-4-6` is set as a step env var.**
The action does not expose model selection as an input — it reads the
`ANTHROPIC_MODEL` environment variable. Opus is required for the depth
the multi-skill pipeline expects; Sonnet runs cheaper but produces
noticeably shallower adversarial passes. The reusable workflow exposes
this as the optional `model` input for callers that need an override.
2. **The `model` input defaults to the `opus` family.**
The reusable workflow passes this via `--model ${{ inputs.model }}` to the
Claude Code action, which then resolves the family to the current version ID.
Opus is required for the depth the multi-skill pipeline expects; Sonnet runs
cheaper but produces noticeably shallower adversarial passes. The `model`
input is optional — callers may override it per the
[operator override](agent-standards.md#operator-override)
described in agent-standards.md, which is the canonical family-not-version
rule this gotcha follows.

3. **`show_full_output: true` is NOT enabled.**
It echoes raw tool results to public action logs, which can leak secrets.
Expand All @@ -1430,7 +1433,7 @@ understanding why they exist:**
| `project_context` | yes | — | 3-5 sentence project description; the only required input |
| `focus_area` | no | `''` | Optional research focus, typically wired to `workflow_dispatch` input |
| `research_depth` | no | `'standard'` | `quick` / `standard` / `deep` |
| `model` | no | `'claude-opus-4-6'` | Override only for cost experiments — see gotcha #2 |
| `model` | no | `'opus'` | Model family; override only for cost experiments — see gotcha #2 |
| `timeout_minutes` | no | `60` | Analyst job timeout (signal collection has its own short timeout) |

| Secret | Required | Notes |
Expand Down
Loading