Skip to content

Ambiguous symbol names silently resolve to defs[0] — blocks main reports 0 blocks when another main has 45 #64

Description

@r0h1tb

Summary

ast-rag blocks <name> resolves the name, takes the first match with no disambiguation, and reports on that one only. When the name is ambiguous — which is normal for main, run, parse, __init__ — the answer is usually about a symbol the user did not mean, and an empty result reads as "this function has no blocks" rather than "I looked at the wrong function".

Reproduce

$ python -m ast_rag blocks main
No blocks found.

That is wrong. main resolves to 13 symbols in this repo, and several of them do have blocks:

resolved symbol blocks
TestCallResolution.main 0 ← defs[0], the one the CLI reports on
test_phase2.main 45
parallel_parsing_benchmark.main 6
benchmark_hybrid.main 4
generate_ground_truth.main 2
watcher_service.main 2
server.main 1

Passing the id directly works, which confirms the query and the data are fine and the defect is purely in resolution:

$ python -m ast_rag blocks ccc5afacfbc1e4fe573c4083
[ { "id": "f36ee5fd...", "block_type": "if", ... } ]

Root cause

# ast_rag/cli.py:1346-1355
defs = api.find_definition(function, lang=lang)
if not defs:
    # Try as function_id directly
    function_id = function
    function_name = function
else:
    function_id = defs[0].id          # <-- silently discards defs[1:]
    function_name = defs[0].qualified_name

find_definition returns every match; only defs[0] survives. Ordering comes from the query, so which symbol wins is essentially arbitrary from the user's point of view — here it lands on a Java test fixture.

Suggested fix

When len(defs) > 1, list the candidates and ask for a qualified name instead of guessing:

$ ast-rag blocks main
'main' is ambiguous — 13 matches. Re-run with a qualified name:
  test_phase2.main                 (45 blocks)
  parallel_parsing_benchmark.main  (6 blocks)
  ...

That keeps the unambiguous case a single command and makes the ambiguous case honest. An --all flag to report on every match would be a reasonable alternative, and --lang already narrows it in some cases.

The header line already prints the resolved name:

console.print(f"\n[bold]Blocks in[/bold] {function_name} (ID: {function_id[:12]}...)\n")

but it only shows under --humanize, and the default JSON path prints nothing identifying at all — so in the default output there is no clue which symbol was chosen.

It is not just blocks — six call sites share the pattern

ast_rag/cli.py:476   callers
ast_rag/cli.py:565   call-graph
ast_rag/cli.py:608   symbol-impact
ast_rag/cli.py:818   evaluate  (MCP tool dispatch: find_callers(defs[0].id, ...))
ast_rag/cli.py:1353  blocks
ast_rag/cli.py:1564  summarize

Confirmed the same wrong resolution on a second command:

$ python -m ast_rag symbol-impact main
📍 Definition: TestCallResolution.main

Same main, same arbitrary pick. symbol-impact is the one I'd worry about most, since "0 references, 0 callers, 0 callees" looks like a finding rather than a misresolution.

A shared _resolve_one(name, ...) helper that errors on ambiguity would fix all six in one place.

Happy to send a PR once you say which behaviour you want (error-and-list vs --all vs first-match-but-announce).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions