Repository navigation
Epic: network-graph accuracy — calls and imports for every tree-sitter-comparable language #3641
Copy link
Copy link
Open
Labels
core-engineModifications to the central physics and parsing engineModifications to the central physics and parsing engineenhancementNew feature, sensor, or structural signatureNew feature, sensor, or structural signature
Milestone
Description
Activity
- addedenhancementNew feature, sensor, or structural signatureNew feature, sensor, or structural signaturecore-engineModifications to the central physics and parsing engineModifications to the central physics and parsing engine
on Sep 25, 2026 - added a commit that references this issue
on Sep 25, 2026 squid-protocol commented
on Sep 25, 2026 OwnerAuthorMore actionsDecisions (2026-09-25, from the author)
1. Reconcile, don't grade.
- The graph comparisons follow the structural tri-comparison process.
- A disagreement is grouped into a shape: language,
callorimport, the bucketer's cause label, and the agreeing vs dissenting readers. - It moves no score until someone reads the source and records a verdict.
- A validated verdict moves the number through
credit_tools, using the same geometry astri_comparison_ledger.py:- GitGalaxy-only shape, credited to gitgalaxy: GitGalaxy was right.
- tree-sitter-only shape, credited to tree_sitter: GitGalaxy genuinely missed it.
- Validated with no credit: that reader's claim was wrong.
- There are only two readers today, so there is no consensus. Both validated numbers therefore use the ledger's conservative precision rule: until a verdict exists, an unverified disagreement keeps counting against GitGalaxy, so validated never inflates on its own. When SCIP joins as a third reader, it follows the ctags model.
- Two numbers per language:
- raw, which keeps the
--ciregression gate; - validated, which is the only one quoted in claims.
- raw, which keeps the
2. Separate ledger file.
docs/self_scan/graph_comparison_ledger.json.- Same module (
tri_comparison_ledger.py), same schema and lifecycle, samehow_to_investigate_a_discrepancy.md. - The structural ledger and chart are untouched.
3. Contract rulings (to be written into
docs/calls_out_rule_contract.md):- Anonymous functions: a call inside an anonymous function or callback (lambda, arrow, closure, block) belongs to the enclosing named unit, since an anonymous function can't be called by name. This is a harness fix:
call_graph_accuracydrops these calls today. - Nested named functions: a call inside a nested named function belongs to that inner unit only, extending C5. The engine counting it in the outer unit too is an engine defect.
- Go conversions:
(*T)(x)and[]byte(s)are conversions, not calls. tree-sitter's call nodes for them are tree-sitter-side verdicts. - Patterns: constructors stay calls (C3). A pattern is not a call: Rust
Data::Struct(x) =>in a match, andOk(t) =>/Err(e) =>match arms. The engine counting them is an engine defect.
Step 1 of this epic is now: bucketer (done, pushed) → shapes → ledger → validated numbers, for calls first, then imports.
Generated by Claude Code
- added 7 commits that reference this issue
on Sep 25, 2026
Metadata
Metadata
Assignees
Labels
core-engineModifications to the central physics and parsing engineModifications to the central physics and parsing engineenhancementNew feature, sensor, or structural signatureNew feature, sensor, or structural signature
Why
The call graph and import graph feed PageRank, blast radius, popularity,
dependency_densityand call resolution. Until this week we measured them for only a few languages:call_graph_accuracy.py).import_graph_accuracy.py).The first sweeps found real bugs quickly:
a.b.{C, D}→a.b. C,D): Scala import recall 33% #3595–Shellsource "$VAR/path"resolves by bare stem: wrong edges to config.yml / lib/completion.bash, misses on repeated names #3598,_dependency_capturereads comments and docstrings: example imports become real edges (Dart, Perl, TS, Python) #3600, Statistical auditor's Packed Payload Guard drops re-export__init__.pyfiles from the scan (flask's package root vanishes) #3583, Kotlin func_start misses extension functions with a generic receiver (Collection<X>.name,fun <T> List<T>.name) #3608, PHP_dependency_capturecaptures multi-line code blocks as import tokens #3609.Scope: calls and imports only. Structural extraction (functions, classes, args) is out of scope; it has its own method.
Coverage today (engine main
cfb9de91)Approach: three referees per language
import_graph_accuracy.py,call_graph_accuracy.py, as today): 2–3 repos per language in different styles. They give the benchmark number, the prevalence of each gap, and resolution effects that only a real tree shows.scip-python,scip-typescript,scip-javafor Java/Kotlin/Scala,scip-go,rust-analyzer,scip-clang,scip-ruby,scip-dotnet). Where cheap, native resolvers too:dependency-cruiser,grimp,go list, Composer PSR-4.tests/graph_shapes/<lang>/): numbered variants such asIMP-PY-012 from . import a as borCALL-RB-007 receiver-dot without parens.supported,gap #issue,out-of-contract (why)ortruth-bug.tags.scmover a broad sample (The Stack v2 or Software Heritage), and lists every distinct import/call shape by frequency.A disagreement between referees is a lead, not a verdict (CLAUDE.md's comparative-correctness rule). Each goes through a verification ledger modelled on
tri_comparison_ledger. This week our own ground-truth code was wrong twice (Kotlin importing Java classes, Scala backticks).Steps
add,get,then; the bucketer should say "62% of JS FPs are callback attribution".call_graph_accuracyandimport_graph_accuracyto the unmeasured languages with the existing tools. Output: the full gap map, filed per language.scip-typescriptagainst the JS/TS call precision gap (79.5% / 86.2%), andscip-pythoncross-checked against pyan3. The result decides whether SCIP truth is clean enough to adopt.recheckstatic ReDoS analysis on changed patterns in PRs.Fix checklists
Import graph:
rules(orphaned_logic reports 100% dead code for five languages whose invocation model never names the callee — the family #2727 closed without covering #2806).imports_unknown, typosquat hits, local tokens.import_graph_accuracy --ci, then--regenerate.rosetta-audit: if it moves, addrosetta:rebless-owedand open the companion PR.Call graph:
docs/calls_out_rule_contract.mdC1–C7): is it a call at all?QUALIFIED_CALLS_OUT_PATTERNS, and check the qualifier and the C5 header.call_graph_accuracy --ciand regenerate;call_graph_resolutionwhere a reference exists.check_calls_out_truth.Out of scope