Skip to content

Epic: network-graph accuracy — calls and imports for every tree-sitter-comparable language #3641

Description

@squid-protocol

Why

The call graph and import graph feed PageRank, blast radius, popularity, dependency_density and call resolution. Until this week we measured them for only a few languages:

  • Calls: 11 of ~36 applicable languages (call_graph_accuracy.py).
  • Imports: 20 of ~32 (import_graph_accuracy.py).

The first sweeps found real bugs quickly:

Scope: calls and imports only. Structural extraction (functions, classes, args) is out of scope; it has its own method.

Coverage today (engine main cfb9de91)

measured tree-sitter-comparable, not measured
calls c, cpp, csharp, go, java, javascript, php, python, ruby, rust, typescript apex, dart, fortran, haskell, kotlin, lua, matlab, objective-c, perl, powershell, scala, swift, tcl, zig, groovy, solidity, ada, scheme, makefile, …
imports c, cpp, dart, go, haskell, java, javascript, kotlin, lua, objective-c, perl, php, python, ruby, rust, scala, shell, solidity, typescript, zig csharp, swift, fortran, matlab, powershell, tcl, apex, groovy, ada, css, html, dockerfile

Approach: three referees per language

  1. Pinned repos (import_graph_accuracy.py, call_graph_accuracy.py, as today): 2–3 repos per language in different styles. They give the benchmark number, the prevalence of each gap, and resolution effects that only a real tree shows.
  2. Semantic truth, generated offline and committed as JSON: SCIP indexers (scip-python, scip-typescript, scip-java for Java/Kotlin/Scala, scip-go, rust-analyzer, scip-clang, scip-ruby, scip-dotnet). Where cheap, native resolvers too: dependency-cruiser, grimp, go list, Composer PSR-4.
    • This gives the first real Level-2 truth (does a call bind to the right definition?) beyond pyan3 for Python.
  3. Shape catalog (tests/graph_shapes/<lang>/): numbered variants such as IMP-PY-012 from . import a as b or CALL-RB-007 receiver-dot without parens.
    • Each variant is a tiny multi-file fixture with declared expected edges. There are also numbered decoys: the form inside a string, comment, or method name.
    • Each entry has a status: supported, gap #issue, out-of-contract (why) or truth-bug.
    • The per-PR gate ratchets on the pass/fail grid.
    • Seeding: from the grammar, from a shape miner, and from every bug found by the other referees. The miner runs tree-sitter / ast-grep with each grammar's tags.scm over a broad sample (The Stack v2 or Software Heritage), and lists every distinct import/call shape by frequency.

A disagreement between referees is a lead, not a verdict (CLAUDE.md's comparative-correctness rule). Each goes through a verification ledger modelled on tri_comparison_ledger. This week our own ground-truth code was wrong twice (Kotlin importing Java classes, Scala backticks).

Steps

  • 1. Disagreement bucketer. Classify each call/import FP/FN by syntactic shape, not callee name. Today's reports list add, get, then; the bucketer should say "62% of JS FPs are callback attribution".
  • 2. Coverage sweep. Extend call_graph_accuracy and import_graph_accuracy to the unmeasured languages with the existing tools. Output: the full gap map, filed per language.
  • 3. SCIP pilot. scip-typescript against the JS/TS call precision gap (79.5% / 86.2%), and scip-python cross-checked against pyan3. The result decides whether SCIP truth is clean enough to adopt.
  • 4. Catalog runner, and shape miner. Seed Python, JS/TS and Ruby first.
  • 5. Fix by shape ID, highest cost first.
  • 6. recheck static ReDoS analysis on changed patterns in PRs.

Fix checklists

Import graph:

  1. Add a failing catalog entry first.
  2. Decide which layer is wrong: capture, resolution, or the truth side.
  3. Capture fixes: anchored to statement position, one line unless the grammar requires more, bounded quantifiers with a ReDoS test. Flags go at the top level of the definition, not in rules (orphaned_logic reports 100% dead code for five languages whose invocation model never names the callee — the family #2727 closed without covering #2806).
  4. Check the side effects: firewall imports_unknown, typosquat hits, local tokens.
  5. Run import_graph_accuracy --ci, then --regenerate.
  6. Golden rebless with a scoped diff.
  7. rosetta-audit: if it moves, add rosetta:rebless-owed and open the companion PR.
  8. Mainframe ledger.
  9. Label the PR so the OS × Python matrix runs (Python 3.9 and Windows bugs only show there).

Call graph:

  1. Add a failing catalog entry first.
  2. Check the contract (docs/calls_out_rule_contract.md C1–C7): is it a call at all?
  3. Decide which layer is wrong: Level-1 name, Level-2 resolution, or attribution (nested callbacks, class ownership). Fixing one layer often exposes the next, as with Ruby in fix(ruby): calls without parentheses, ?/! names, and methods owned by their class (#3377) #3637.
  4. Pattern changes: group 1 is the callee, register the pattern in QUALIFIED_CALLS_OUT_PATTERNS, and check the qualifier and the C5 header.
  5. Run call_graph_accuracy --ci and regenerate; call_graph_resolution where a reference exists.
  6. Run rosetta check_calls_out_truth.
  7. Goldens: popularity and blast radius ripple, so scope the diff.
  8. Label for the matrix.

Out of scope

  • Structural extraction (functions, classes, args).
  • Languages with no tree-sitter grammar (mainframe, assembly). COBOL is covered by the mainframe answer keys instead.

Activity

  1. added
    enhancementNew feature, sensor, or structural signature
    core-engineModifications to the central physics and parsing engine
    on Sep 25, 2026
  2. squid-protocol commented on Sep 25, 2026

    @squid-protocol
    OwnerAuthor

    Decisions (2026-09-25, from the author)

    1. Reconcile, don't grade.

    • The graph comparisons follow the structural tri-comparison process.
    • A disagreement is grouped into a shape: language, call or import, the bucketer's cause label, and the agreeing vs dissenting readers.
    • It moves no score until someone reads the source and records a verdict.
    • A validated verdict moves the number through credit_tools, using the same geometry as tri_comparison_ledger.py:
      • GitGalaxy-only shape, credited to gitgalaxy: GitGalaxy was right.
      • tree-sitter-only shape, credited to tree_sitter: GitGalaxy genuinely missed it.
      • Validated with no credit: that reader's claim was wrong.
    • There are only two readers today, so there is no consensus. Both validated numbers therefore use the ledger's conservative precision rule: until a verdict exists, an unverified disagreement keeps counting against GitGalaxy, so validated never inflates on its own. When SCIP joins as a third reader, it follows the ctags model.
    • Two numbers per language:
      • raw, which keeps the --ci regression gate;
      • validated, which is the only one quoted in claims.

    2. Separate ledger file.

    • docs/self_scan/graph_comparison_ledger.json.
    • Same module (tri_comparison_ledger.py), same schema and lifecycle, same how_to_investigate_a_discrepancy.md.
    • The structural ledger and chart are untouched.

    3. Contract rulings (to be written into docs/calls_out_rule_contract.md):

    • Anonymous functions: a call inside an anonymous function or callback (lambda, arrow, closure, block) belongs to the enclosing named unit, since an anonymous function can't be called by name. This is a harness fix: call_graph_accuracy drops these calls today.
    • Nested named functions: a call inside a nested named function belongs to that inner unit only, extending C5. The engine counting it in the outer unit too is an engine defect.
    • Go conversions: (*T)(x) and []byte(s) are conversions, not calls. tree-sitter's call nodes for them are tree-sitter-side verdicts.
    • Patterns: constructors stay calls (C3). A pattern is not a call: Rust Data::Struct(x) => in a match, and Ok(t) => / Err(e) => match arms. The engine counting them is an engine defect.

    Step 1 of this epic is now: bucketer (done, pushed) → shapes → ledger → validated numbers, for calls first, then imports.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    core-engineModifications to the central physics and parsing engineenhancementNew feature, sensor, or structural signature

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions