Skip to content

yaml: every extracted step is named run — func_start captures no name, so orphan/duplicate census and per-step identity are all dead #2767

Description

@squid-protocol

Same pass as #2763–#2766 (2026-09-05, engine 3b3568e0, corpus main 6183c4b). raw_state_slop_orphans is the corpus's third-weakest metric (61% consistency) after #2765's two, and yaml sits at −100% on it. The cause is upstream of the orphan census: yaml's thirteen extracted functions all carry the same name.

Measured

galaxyscope --db-only over keyword-rosetta/data/yaml, function_data joined to file_data:

a.yml     run   run   run
b.yml     run   run   run
c.yml     run   run   run
main.yml  run   run   run   run

Thirteen slices, one name. yaml.py's func_start has no capture group —

"func_start": re.compile(r"^[ \t]*(?:-?[ \t]*run:|script:|before_script:|after_script:)[ \t]*[|>]*")

— so _extract_name falls back to the matched keyword. The literal-capture helper that builds _keyword_bucket_names (detector.py:1010) does not recover run from this pattern (the -?[ \t]* prefix defeats it), which is confirmable directly:

>>> _closed_literal_capture(LANGUAGE_DEFINITIONS['yaml']['rules']['func_start'].pattern)
frozenset()          # css → {'@media','@keyframes',...}; dockerfile → {'RUN','CMD',...}

So unlike css's keyframes and dockerfile's RUN — which #2728 deliberately excluded from the orphan and duplicate censuses precisely because a keyword cannot be orphaned — yaml's run reaches both censuses as if it were an identifier, and is then dropped by the len(func_name) > 3 guard on detector.py:1496. Two different mechanisms both silently discarding the same slice.

What is lost

  • Orphan census: raw_state_slop_orphans 0.00 against a 2.50 median, and risk_tech_debt inherits it.
  • Duplicate census: thirteen same-named slices with different bodies. The body_hash guard (duplicate_logic heuristic is scope-blind: same-named shadowed local helpers across different where-blocks flagged as duplicates #1498) is the only thing preventing thirteen false duplicates — the name test alone would have fired on every one.
  • Per-step identity, which is the real cost outside the corpus. In a scanned repo, every step of every workflow in function_data is a row named run. Nothing downstream — the tri-comparison function list, the archetype classifier, the call graph, --llm-only output — can tell one CI step from another.

The name is right there

Both major YAML CI dialects name their steps, one line above the run::

- name: Build the wheel        # GitHub Actions
  run: python -m build

- script: python -m build      # Azure Pipelines
  displayName: Build the wheel

GitLab is stronger still: a job's own key is its name, and script: is a child of it.

Candidate shape

Give func_start a capture group fed by the adjacent name key, falling back to the current behaviour when there is none:

"func_start": re.compile(
    r"^[ \t]*-?[ \t]*(?:name:[ \t]*(?P<n>[^\n#]+?)[ \t]*(?:#.*)?\n[ \t]*)?"
    r"(?:run:|script:|before_script:|after_script:)[ \t]*[|>]*",
    re.M,
),

with the displayName:-after-script: (Azure) and job-key (GitLab) forms as the two follow-ups the audit should price. Where no name exists the slice keeps run, so nothing regresses — but then it should also be added to the keyword-bucket exclusion so it is censused consistently with dockerfile's RUN rather than accidentally, by name length. The cleanest form of that is to fix _closed_literal_capture to see through the -?[ \t]* prefix, which is a one-line change and makes the exclusion derived from the rule rather than coincidental.

Strict tests in tests/extraction/languages/test_yaml_strict.py: a named step, an unnamed step, a step whose name: carries a # comment, a name: belonging to the job rather than the step (must not be captured), and a matrix step where name: interpolates ${{ matrix.os }}.

Corpus pairing

keyword-rosetta/data/yaml's steps carry no name: today, so the plant is one name: line per probe — and it needs screening first: name: is not in yaml's doc rule (that one matches description:), but screen_plant.py should confirm it against the full rule set before the shells change. With names, the thirteen slices become distinguishable, the orphan census runs, and raw_state_slop_orphans moves off −100%. Ledger: the yaml half of orphan-detection-is-name-recurrence (shape 2, "SLICE NAME SYNTHETIC OR SHORT") narrows or retires.

Part of #2669 (rosetta[yaml] is #2606). Adjacent to #2728 (keyword buckets) and #2754 (the span-scoped orphan test).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugUnintended behavior or logic failure in the enginecore-engineModifications to the central physics and parsing enginemetricsHeuristics, risk exposures, and topological math updates

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions