You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
yaml: every extracted step is named run — func_start captures no name, so orphan/duplicate census and per-step identity are all dead #2767
Same pass as #2763–#2766 (2026-09-05, engine 3b3568e0, corpus main 6183c4b). raw_state_slop_orphans is the corpus's third-weakest metric (61% consistency) after #2765's two, and yaml sits at −100% on it. The cause is upstream of the orphan census: yaml's thirteen extracted functions all carry the same name.
Measured
galaxyscope --db-only over keyword-rosetta/data/yaml, function_data joined to file_data:
a.yml run run run
b.yml run run run
c.yml run run run
main.yml run run run run
Thirteen slices, one name. yaml.py's func_start has no capture group —
— so _extract_name falls back to the matched keyword. The literal-capture helper that builds _keyword_bucket_names (detector.py:1010) does not recover run from this pattern (the -?[ \t]* prefix defeats it), which is confirmable directly:
So unlike css's keyframes and dockerfile's RUN — which #2728 deliberately excluded from the orphan and duplicate censuses precisely because a keyword cannot be orphaned — yaml's run reaches both censuses as if it were an identifier, and is then dropped by the len(func_name) > 3 guard on detector.py:1496. Two different mechanisms both silently discarding the same slice.
What is lost
Orphan census: raw_state_slop_orphans 0.00 against a 2.50 median, and risk_tech_debt inherits it.
Per-step identity, which is the real cost outside the corpus. In a scanned repo, every step of every workflow in function_data is a row named run. Nothing downstream — the tri-comparison function list, the archetype classifier, the call graph, --llm-only output — can tell one CI step from another.
The name is right there
Both major YAML CI dialects name their steps, one line above the run::
with the displayName:-after-script: (Azure) and job-key (GitLab) forms as the two follow-ups the audit should price. Where no name exists the slice keeps run, so nothing regresses — but then it should also be added to the keyword-bucket exclusion so it is censused consistently with dockerfile's RUN rather than accidentally, by name length. The cleanest form of that is to fix _closed_literal_capture to see through the -?[ \t]* prefix, which is a one-line change and makes the exclusion derived from the rule rather than coincidental.
Strict tests in tests/extraction/languages/test_yaml_strict.py: a named step, an unnamed step, a step whose name: carries a # comment, a name: belonging to the job rather than the step (must not be captured), and a matrix step where name: interpolates ${{ matrix.os }}.
Corpus pairing
keyword-rosetta/data/yaml's steps carry no name: today, so the plant is one name: line per probe — and it needs screening first: name: is not in yaml's doc rule (that one matches description:), but screen_plant.py should confirm it against the full rule set before the shells change. With names, the thirteen slices become distinguishable, the orphan census runs, and raw_state_slop_orphans moves off −100%. Ledger: the yaml half of orphan-detection-is-name-recurrence (shape 2, "SLICE NAME SYNTHETIC OR SHORT") narrows or retires.
Part of #2669 (rosetta[yaml] is #2606). Adjacent to #2728 (keyword buckets) and #2754 (the span-scoped orphan test).
Same pass as #2763–#2766 (2026-09-05, engine
3b3568e0, corpus main6183c4b).raw_state_slop_orphansis the corpus's third-weakest metric (61% consistency) after #2765's two, and yaml sits at −100% on it. The cause is upstream of the orphan census: yaml's thirteen extracted functions all carry the same name.Measured
galaxyscope --db-onlyoverkeyword-rosetta/data/yaml,function_datajoined tofile_data:Thirteen slices, one name.
yaml.py'sfunc_starthas no capture group —— so
_extract_namefalls back to the matched keyword. The literal-capture helper that builds_keyword_bucket_names(detector.py:1010) does not recoverrunfrom this pattern (the-?[ \t]*prefix defeats it), which is confirmable directly:So unlike css's
keyframesand dockerfile'sRUN— which #2728 deliberately excluded from the orphan and duplicate censuses precisely because a keyword cannot be orphaned — yaml'srunreaches both censuses as if it were an identifier, and is then dropped by thelen(func_name) > 3guard ondetector.py:1496. Two different mechanisms both silently discarding the same slice.What is lost
raw_state_slop_orphans0.00 against a 2.50 median, andrisk_tech_debtinherits it.body_hashguard (duplicate_logic heuristic is scope-blind: same-named shadowed local helpers across different where-blocks flagged as duplicates #1498) is the only thing preventing thirteen false duplicates — the name test alone would have fired on every one.function_datais a row namedrun. Nothing downstream — the tri-comparison function list, the archetype classifier, the call graph,--llm-onlyoutput — can tell one CI step from another.The name is right there
Both major YAML CI dialects name their steps, one line above the
run::GitLab is stronger still: a job's own key is its name, and
script:is a child of it.Candidate shape
Give
func_starta capture group fed by the adjacent name key, falling back to the current behaviour when there is none:with the
displayName:-after-script:(Azure) and job-key (GitLab) forms as the two follow-ups the audit should price. Where no name exists the slice keepsrun, so nothing regresses — but then it should also be added to the keyword-bucket exclusion so it is censused consistently with dockerfile'sRUNrather than accidentally, by name length. The cleanest form of that is to fix_closed_literal_captureto see through the-?[ \t]*prefix, which is a one-line change and makes the exclusion derived from the rule rather than coincidental.Strict tests in
tests/extraction/languages/test_yaml_strict.py: a named step, an unnamed step, a step whosename:carries a#comment, aname:belonging to the job rather than the step (must not be captured), and a matrix step wherename:interpolates${{ matrix.os }}.Corpus pairing
keyword-rosetta/data/yaml's steps carry noname:today, so the plant is onename:line per probe — and it needs screening first:name:is not in yaml'sdocrule (that one matchesdescription:), butscreen_plant.pyshould confirm it against the full rule set before the shells change. With names, the thirteen slices become distinguishable, the orphan census runs, andraw_state_slop_orphansmoves off −100%. Ledger: the yaml half oforphan-detection-is-name-recurrence(shape 2, "SLICE NAME SYNTHETIC OR SHORT") narrows or retires.Part of #2669 (rosetta[yaml] is #2606). Adjacent to #2728 (keyword buckets) and #2754 (the span-scoped orphan test).