Everything you need to make a first contribution in about fifteen minutes. For the full design see ARCHITECTURE.md; this page is the hands-on version.
src/codebase_index/
├── cli.py Typer commands; every command delegates to service.py
├── service.py shared CLI/MCP layer: db resolution, search/impact/stats payloads
├── discovery/ walker.py (file walk), ignore.py (.gitignore & co), classify.py (language, secrets, binary, generated)
├── parsers/ languages.py (LangSpec per language), treesitter.py (symbols + edges), line_chunker.py
├── indexer/ pipeline.py (build_index / update_index), freshness.py
├── graph/ builder.py (edge resolution), expand.py (impact), analysis.py, navigate.py, export.py
├── retrieval/ pipeline.py (search), searchers.py, fusion.py, rerank.py, priors.py, tuning.py, budget.py, skeleton.py
├── storage/ db.py, schema.sql, repo.py (typed SQL accessors)
├── mcp/server.py stdio MCP server over the same service layer
├── output/ markdown.py, json.py, redact.py
└── skill_template/ canonical skill source; copies under .claude/ .codex/ .opencode/ skill/ skills/ are generated
Query path in one line: retrieval/pipeline.py::search → detect_intent →
_run_retrievers (path, symbol, FTS5, optional vector, optional graph) → fuse
(RRF) → rerank → dedup / diversify → token budget → payload.
Python 3.11+ and git are the only prerequisites.
git clone https://github.com/denfry/codebase-index.git
cd codebase-index
python -m venv .venv
. .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -e ".[dev,mcp]"dev brings pytest, ruff, mypy, pyyaml and the mcp SDK; mcp is listed
separately so the server extra stays installable on its own. Optional extras:
embeddings, embeddings-local, watch, build.
Before running the CLI inside this checkout, export
CBX_NO_SKILL_AUTO_UPDATE=1. Without it the skill auto-updater may rewrite the
committed .skill_version stamps when the installed package metadata is stale.
tests/conftest.py sets this for pytest automatically.
CI (.github/workflows/ci.yml) runs exactly these:
ruff check src tests
mypy src/codebase_index
python scripts/sync_skill_copies.py --check
pytest # coverage gate: --cov-fail-under=80 (pyproject.toml)
pytest tests/test_perf_smoke.py --runslow --no-cov # Linux + 3.12 onlyNotes:
-
ruff formatis not enforced; runningruff format --checkon the current tree reports many files. Format new code if you like, but do not reformat files you are not otherwise touching. -
Tests marked
sloware skipped unless you pass--runslow(tests/conftest.py::pytest_addoption). -
Golden snapshots live in
tests/golden/*.json. When an output contract changes on purpose, regenerate them and review the diff:UPDATE_GOLDEN=1 pytest tests/test_cli_golden.py tests/test_mcp_golden.py
tests/golden_utils.pymasks timestamps, commit SHAs and the package version so snapshots are stable across machines;schema_versionis deliberately not masked because it is the contract under test. -
The skill copies gate: anything under
src/codebase_index/skill_template/or the version insrc/codebase_index/__init__.pymust be propagated withpython scripts/sync_skill_copies.py(no flag) before committing.
Run a single test file quickly without the coverage gate:
pytest tests/test_fusion.py -q --no-covexport CBX_NO_SKILL_AUTO_UPDATE=1
codebase-index --root tests/fixtures/sample_repo index
codebase-index --root tests/fixtures/sample_repo search "refresh token" --json
codebase-index --root tests/fixtures/sample_repo impact Usertests/fixtures/sample_repo/ deliberately contains files that must never be
indexed (.env, secrets.pem, huge.json, logo.png); see
tests/fixtures/README.md.
The project rule is measure improvements, do not assert them. Three surfaces exist; know which one you are using.
| Surface | Command | Use it for |
|---|---|---|
| Retrieval eval (ranking gate) | python tests/eval/run_eval.py |
Any change that can move ranking |
| Public synthetic suite | python tests/benchmark_public.py --workdir .tmp-public-benchmark |
Metric-shape regression check (tests/test_public_benchmark.py wraps it in CI) |
| Older single-repo script | python tests/benchmark_honest.py --repo <path> |
Index vs rg+window token/recall comparison against one repository |
The retrieval eval (tests/eval/):
-
harness.pybuilds one index per corpus into a temp dir (build_corpus_index) and reuses it for every variant, then computes recall@5/10, MRR, nDCG@10, hit@3, P@5, MAP,useful@budget, mean tokens, duplicate rate, and latency percentiles (evaluate,format_table,pool). -
metrics.pyholds the IR metrics pluspaired_bootstrap_ciandpaired_permutation_p(seeded, reproducible). -
gen_queries.pymints leak-free ground truth from git history (commit subject → files that commit changed):python tests/eval/gen_queries.py --repo ../some-repo --out /tmp/some-repo.yml python tests/eval/run_eval.py --corpus ../some-repo:/tmp/some-repo.yml --ablate
-
run_eval.py --ablateruns a one-signal-off sweep over the boolean flags inABLATABLEand prints a paired significance table for each row.
Checked-in query sets: tests/eval/queries/self_repo.yml (hand-written) and
self_repo_git.yml (generated). See tests/eval/README.md.
src/codebase_index/discovery/classify.py: add the extension to_LANG_BY_SUFFIXand the language id to_TREE_SITTER_LANGS.src/codebase_index/parsers/languages.py: add aLangSpec(name, ts_name, defs_query, calls_query, imports_query)and register it inLANGS. Definitions are captured as@def.<kind>with the name node as@name; calls capture@callee; imports/inheritance capture@import.module,@extends.base,@implements.iface(mapped by_EDGE_PREFIXESinparsers/treesitter.py).- If the grammar's node kinds are not covered by
_definition_kind/_name_node/_callee_nodeinparsers/treesitter.py, extend them. - Add a fixture under
tests/fixtures/multilang/and cases intests/test_languages.py(query compiles against the grammar) andtests/test_multilang_symbols.py(test_registry_consistency_every_treesitter_lang_extractsis parametrised over the registry, so a language with zero symbols fails loudly). - Update the tier table in LANGUAGES.md.
pytest tests/test_languages.py tests/test_multilang_symbols.py tests/test_graph.py -q --no-covEdges are rows in the edges table (storage/schema.sql): edge_type,
src_kind/src_id, dst_kind/dst_id/dst_name, line, resolved, and a
confidence of extracted, inferred, or ambiguous.
- Emit the edge from
parsers/treesitter.py::_extract_graph_edges(query captures) or_extract_edges(calls). Today's types arecall,reference,import,extends,implements. - Resolve it in
graph/builder.py::resolve_edges. Symbol-target types are in_SYMBOL_EDGE_TYPESand resolve only on a repo-unique name; imports resolve by path suffix. Anything that cannot be pinned to one target is markedambiguousbyrepo.mark_ambiguous_edges. Never guess. - Make sure
graph/expand.py(impact),graph/navigate.py(path/describe) andgraph/export.py(HTML/GraphML/DOT/Neo4j) do the right thing with the new type, and document it in SCHEMA.md and LANGUAGES.md. - Add tests in
tests/test_graph.py/tests/test_graph_coverage.py; goldens that include edges (impact_*,mcp_impact_of*,path_*) may need regeneration.
- Put every new signal behind a boolean on
RetrievalTuning(src/codebase_index/retrieval/tuning.py). Defaults are the shipped configuration;RetrievalTuning.baseline()must keep reproducing the 1.7.0 behaviour, otherwise the ablation table is meaningless. - Wire it in
retrieval/pipeline.py::search(candidate generation,fuse,rerank, selection) or inretrieval/rerank.py/retrieval/priors.py(source_role_prior, capped byMAX_ABS_PRIORso priors stay tiebreakers). - Add the flag to
ABLATABLEintests/eval/run_eval.py. - Run
python tests/eval/run_eval.py --ablateon at least two corpora (--corpus), and paste the pooled table plus the significance row into the PR. A signal ships only if its row is significant (p < 0.05) on the pooled set; otherwise it stays off by default or is removed. - Unit tests:
tests/test_fusion.py,tests/test_hybrid_ranking.py,tests/test_priors.py,tests/test_diversity.py.
src/codebase_index/mcp/server.py: decorate a function with@_tool()and return_emit("<tool_name>", payload). The_emitenvelope addsschema_versionandtool; error paths must use it too (_no_index_payload()is the standard no-index body).- Do the real work in
service.pyso the CLI command and the MCP tool share one implementation. - Register the tool in
tests/test_mcp_server.py::test_mcp_server_has_expected_toolsand in the envelope parametrisation there; add a golden case intests/test_mcp_golden.pyand generatetests/golden/mcp_<name>.jsonwithUPDATE_GOLDEN=1. - Document it in the tool table of MCP.md. Bump
MCP_SCHEMA_VERSIONonly for a breaking change (field removal or type change).
Add it in cli.py, delegate to service.py, support --json, and if the skill
should be allowed to call it, add it to the ALLOWED whitelist in
src/codebase_index/skill_template/scripts/cbx and cbx.ps1, then run
python scripts/sync_skill_copies.py and update references/commands.md in the
template. Add a golden in tests/test_cli_golden.py for --json output.
The version is single-sourced in src/codebase_index/__init__.py
(__version__); pyproject.toml reads it through hatch dynamic versioning.
Two mirrors are kept in sync by scripts/sync_skill_copies.py:
.claude-plugin/plugin.json (version) and requirements.lock (release tag
tarball). Tagging vX.Y.Z triggers .github/workflows/release.yml: test gate,
build, twine check, scripts/release_smoke.py (clean venv install → init →
index → search), GitHub release, and PyPI publish via Trusted Publishing.
Follow RELEASE_CHECKLIST.md top to bottom. Contributors
never bump the version; they add entries under [Unreleased] in
CHANGELOG.md.
Open a discussion for design questions, an issue for bugs and concrete proposals, and read CONTRIBUTING.md for the PR workflow.