Skip to content

Call-graph comparison tooling: reference contract + cache, triage, one-command check, explain, session hook (#3773–#3776, #3778) - #3793

Merged
squid-protocol merged 2 commits into
mainfrom
claude/optimistic-heisenberg-bta91d
Sep 26, 2026
Merged

squid-protocol merged 2 commits into
mainfrom
claude/optimistic-heisenberg-bta91d

Conversation

@squid-protocol

Copy link
Copy Markdown
Owner

Summary

This is the infrastructure half of epic #3772. Closes #3773, closes #3774, closes #3775, closes #3776, closes #3778.

The TypeScript round (#3756–#3767) spent its time in five places:

  • building the reference;
  • triaging misses and wrong links by hand, which was the biggest token cost;
  • environment detours;
  • re-running the same six-command verification chain on every PR;
  • serial fixes.

This PR addresses each of those except the serial fixes. It's tooling and docs only; no engine code changes.

# file what it does
#3773 tests/tools/callgraph_refs.py One JSON contract for every reference (defs, edges with the first call site, optional external). A registry holds python (pyan3 2.8.1, language-crucible) and typescript (tsc 6.0.2, pinned import-graph repos). A cache is keyed on the corpus repo's git tree, the tool version and the adapter's own code. call_graph_resolution.py now scores any registered language through it.
#3774 tests/tools/callgraph_triage.py Puts every missed reference edge in exactly one bucket (caller_unmapped, not_extracted/{in_string,after_nested_def,same_name,other}, ambiguous/<step>, resolved_outside, …) and gives every confident link a precision verdict. Buckets are ranked, with source lines, and must add up to the recall gap (checked at runtime).
#3775 tests/tools/callgraph_check.py The verification chain as one command, one line per step, with logs on disk. A missing tool or version pin is MISSING and fails. It chooses the crucible mode by whether tiktoken's encoding is reachable, and says why. --regenerate rewrites baselines only after a full pass.
#3776 tests/tools/explain_call.py Re-runs the real resolver on a master DB's rehydrated inputs and checks the result against the recorded fcall_data rows. For each callee it shows the step, the qualifiers tried, and every candidate with owner, def_shape and whether a bare call or a receiver can reach it.
#3778 .claude/hooks/session-start.sh + settings.json Web sessions only, synchronous and idempotent: CI's pins (pytest/xdist, mypy + stubs, ruff 0.16.0, pyan3, typescript 6.0.2, ML deps), both corpora as siblings, and the v2.4.7 tag the AST audit needs. It also writes a session PATH that prefers this interpreter's tools over uv shims but keeps the current node.

ts_callgraph.js now reports each edge's first call site, and the workflow's paths: includes callgraph_refs.py.

Measured

  • Scores unchanged through the contract: python 99.8% / 51.6% / 97.5% (decorators 100% / 68.1%, references 99.7% / 77.3%), typescript 99.9% / 57.1% / 73.7%.
  • Cache: zod's reference loads in about 1 s cached against about 20 s built. The python gate takes 39 s with cached pyan graphs, down from several minutes.
  • Session hook: 29 s on first start, 6 s after. It's a no-op outside the web.
  • callgraph_check.py on this branch, under the hook's environment, printed 11 lines:
PASS     gate:python        precision 99.8% (1806) recall 51.6% resolution 97.5%  [baseline 99.8% / 51.6% / 97.5%]  (39s)
PASS     gate:typescript    precision 99.9% (1520) recall 57.1% resolution 73.7%  [baseline 99.9% / 57.1% / 73.7%]  (12s)
PASS     tests              290 passed  (10s)
PASS     crucible           mode zero (tiktoken: openaipublic.blob.core.windows.net unreachable (URLError) -- ...)  (85s)
PASS     tree-sitter        30 language(s) checked  (102s)
PASS     lint:ruff / lint:dead-key / lint:mypy / lint:format (7 files)
callgraph_check: all passed

What the tools found while being built (filed, not fixed here)

Type of change

  • Bug fix
  • New feature / language support
  • Parsing or engine logic (gitgalaxy/core/detector.py, language_standards.py, prism.py, or a per-language rule)
  • Docs, tooling, or CI only
  • Other (describe above)

CI checklist

  • ruff_audit.py, dead_key_audit.py and mypy_audit.py (--ci) pass, via callgraph_check.py. ruff format --check is clean on the 7 touched Python files. Under the hook's PATH, mypy is the interpreter's own with stubs, so the four yaml-stub findings reported in earlier PRs no longer appear.
  • tests/tools/test_callgraph_tools.py (16 new tests, including an end-to-end scan of a tiny git repo against a registered fake reference, with no node or pyan3 needed) and test_call_graph_resolution_gate.py: 21 passed, run as plain pytest from /tmp (CI's invocation). The focused set is 290 passed.
  • crucible_check.py: zero-dependency PASS. Full precision can't run here, because this environment blocks tiktoken's encoding host, so CI's crucible-audit (full-precision) is the check for that half. No engine code changed, so no drift is expected.
  • tree_sitter_accuracy_audit.py --ci --all: 30 languages OK.
  • Parsing/engine logic untouched.

Verification

$ python tests/tools/callgraph_check.py                                    # the block above
$ CLAUDE_CODE_REMOTE=true CLAUDE_ENV_FILE=/tmp/env .claude/hooks/session-start.sh    # exit 0, 29 s / 6 s
$ source /tmp/env; cd /tmp; pytest -q tests/tools/test_callgraph_tools.py ...         # 21 passed
$ python tests/tools/callgraph_triage.py typescript        # buckets reconcile: 1142 = 2661 - 1519
$ python tests/tools/explain_call.py <db> "v3/types.ts:_parse@3824" addIssue   # re-run matches the scan

🤖 Generated with Claude Code

https://claude.ai/code/session_01AbvXUJpva5XxxbssoKG9Jt


Generated by Claude Code

…e-command check, explain, session hook (#3773-#3776, #3778)

Epic #3772's infrastructure, so each new language costs an adapter, not a week.

- callgraph_refs.py (#3773): one JSON contract for every reference (defs,
  edges with their first call site, optional external), a registry (python/
  pyan3 2.8.1 on language-crucible, typescript/tsc 6.0.2 on the pinned
  import-graph repos), and a cache keyed on the corpus repo's git tree, the
  tool version and the adapter's own code. call_graph_resolution.py scores any
  registered language through it; python and typescript numbers are unchanged
  (99.8/51.6/97.5 and 99.9/57.1/73.7). zod's reference loads in ~1 s cached vs
  ~20 s built.
- callgraph_triage.py (#3774): every missed reference edge in exactly one
  bucket (not_extracted/in_string, after_nested_def, same_name, caller_unmapped,
  ambiguous/<step>, resolved_outside, ...) and every confident link's verdict,
  ranked, with source lines; the buckets add up to the recall gap exactly. Its
  first run on zod found #3787 (spread calls), #3788 (namespace-alias imports)
  and #3789 (package self-imports).
- callgraph_check.py (#3775): gates, focused tests, crucible (both modes when
  tiktoken's encoding is reachable, zero otherwise, and it says why),
  tree-sitter audit and lint audits, one line per step, logs on disk. A
  missing tool or pin is MISSING and fails; --regenerate only after a pass.
- explain_call.py (#3776): re-runs the real resolver on a master DB's own
  rehydrated inputs, checks it against the recorded rows, and lists every
  candidate with def_shape and reach. It found #3786 (delta scans lose class
  inheritance), which it works around locally.
- .claude/hooks/session-start.sh (#3778): web sessions get CI's pins (pytest,
  mypy+stubs, ruff 0.16.0, pyan3, typescript 6.0.2, ML deps), both corpora,
  the ast-accuracy tag, and a PATH that prefers this interpreter's tools over
  uv shims while keeping the right node. tiktoken is installed only when its
  encoding host is reachable (#3791). 29 s first start, 6 s after.
- ts_callgraph.js now reports each edge's first call site.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AbvXUJpva5XxxbssoKG9Jt
@squid-protocol squid-protocol added this to the Compiler call-graph comparison milestone Sep 26, 2026 — with Claude
@squid-protocol squid-protocol added enhancement New feature, sensor, or structural signature testing Unit, integration, and E2E pipeline verification labels Sep 26, 2026 — with Claude
Comment thread tests/tools/call_graph_resolution.py Fixed
node_env moved to callgraph_refs; the alias only served the gate test, which
now calls callgraph_refs.node_env directly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AbvXUJpva5XxxbssoKG9Jt
@github-actions

Copy link
Copy Markdown
Contributor

🐦‍⬛ Muninn Security Scan

✅ No security issues found.

🐦‍⬛ Powered by Muninn · Skald Lab

@squid-protocol
squid-protocol merged commit 02ae597 into main Sep 26, 2026
30 checks passed
@squid-protocol
squid-protocol deleted the claude/optimistic-heisenberg-bta91d branch September 26, 2026 19:30

Copy link
Copy Markdown
Owner Author

full-suite-matrix (windows-latest, 3.12) is red, and the failure isn't from this PR.

Failing: tests/cobol_mainframe/test_refractor_engine_channels.py::test_open_sites_survive_a_delta_scan and ::test_a_baseline_without_open_sites_still_rehydrates, both with KeyError: 'src/LEDGER.cbl'. That's 2 failed and 12,017 passed. The 16 tests this PR adds passed on Windows.

Why it isn't this PR's:

Root cause: on Windows the master DB stores file_path with backslashes.

  • galaxy_ir.load_galaxy_ir already normalizes for this (file_path=(rec[1] or "").replace("\\", "/"), commented "so target / file_path agree on every platform"). The other tests in that file go through it and pass.
  • StateRehydrator.load_state() keys ram_cache by the raw path, src\LEDGER.cbl on Windows.
  • The two failing tests look up "src/LEDGER.cbl" directly in that cache.

Proposed patch (test-only): look the file up separator-agnostically in tests/cobol_mainframe/test_refractor_engine_channels.py:

def _cached(cache, rel):
    """ram_cache is keyed by the DB's file_path, which is OS-native (`\\` on Windows)."""
    (node,) = [v for k, v in cache.items() if k.replace("\\", "/") == rel]
    return node

# test_open_sites_survive_a_delta_scan
rows = {d["dd_name"]: d["open_sites"] for d in _cached(cache, "src/LEDGER.cbl")["dataset_bindings"]}
# test_a_baseline_without_open_sites_still_rehydrates
assert all("open_sites" not in d for d in _cached(cache, "src/LEDGER.cbl")["dataset_bindings"])

The alternative is normalizing the keys in the rehydrator itself. That's an engine change, and galaxyscope's delta-scan comparison, which uses native paths on Windows, would need to agree with it. So it's better as its own PR if you want it.

A re-run wouldn't change the result: this is deterministic on Windows. I haven't pushed the patch here, to keep this PR to tooling. I can add it to this PR or open it separately, whichever you prefer.


Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature, sensor, or structural signature testing Unit, integration, and E2E pipeline verification

Projects

None yet

3 participants