Why
Our dev-lead and pr-review agents navigate code by grep + file-reading, which returns textual not semantic matches and gives a sliced, incomplete view of large repos. Language Server Protocol (LSP) tooling — definition, references, hover, diagnostics — exposed to the agent as MCP tools gives real semantic code intelligence. The discussion's strongest, lowest-risk pilot is pr-review finding-verification on Shell in this repo (a bad verification only downgrades a finding; a bad code write breaks a build).
This is strictly opt-in plumbing on top of the MCP mechanism that already shipped — scripts/engine.sh threads --mcp-config/--strict-mcp-config into the deep + rubber-duck Claude tiers only (REVIEW_MCP_CONFIG / REVIEW_MCP_ALLOWED_TOOLS, epic #676, Context7 landed in #826). The triage tier stays fast/restricted and untouched. LSP is one more MCP server behind that same knob — we are not rebuilding the gating, the graceful-degradation path, or the token observatory.
Scope (confirmed from the discussion)
- Agent surface: pr-review deep/audit tiers first; dev-lead follows only after a positive go/no-go (out of scope here).
- Pilot language/repo: Shell in
.github-private (this repo — the agent's own substrate; bash-language-server is lightweight, no cross-repo coordination).
- Server selection is an experiment, not a pre-decision. Per @don-petry, candidate LSP-MCP servers (agent-lsp, Serena, lsp-mcp) are compared on speed, quality, and cost across a corpus of PRs with and without LSP; selection is decided from that data, not the vendor narrative.
Initiative success metric
Go = on the frozen pilot PR corpus, LSP-on deep-tier review delivers a measurable navigation-token reduction (target >=2x fewer navigation tool-call tokens vs the LSP-off control) AND no regression in review quality (false-positive / precision no worse than the frozen LSP-off baseline), achieved within the cold-start SLA (<=30s P95). A win on tokens that costs precision is a no-go, not a win.
Cost cap (explicit)
- LSP is gated to deep/audit tiers only (triage untouched); cold-start is auto-skipped above the 30s P95 SLA (graceful degradation, never a workflow failure).
- The pilot is run-count bounded: <=20 PRs in the corpus x <=2 candidate servers x <=3 runs each ~= <=120 deep-tier review runs total, under the existing 60-min job cap. No fleet-wide rollout, and no dev-lead (code-writing) LSP, until the go/no-go decision record lands.
Shape
Phase 1 scopes the pilot and freezes the measurement corpus/baseline; Phase 2 wires the candidate server(s) into the runtime and adds the finding-verification step; Phase 3 runs the comparison and records a human go/no-go.
Untracked prerequisites
Planned from idea discussion #578 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds initiative:auto.
Why
Our dev-lead and pr-review agents navigate code by
grep+ file-reading, which returns textual not semantic matches and gives a sliced, incomplete view of large repos. Language Server Protocol (LSP) tooling —definition,references,hover,diagnostics— exposed to the agent as MCP tools gives real semantic code intelligence. The discussion's strongest, lowest-risk pilot is pr-review finding-verification on Shell in this repo (a bad verification only downgrades a finding; a bad code write breaks a build).This is strictly opt-in plumbing on top of the MCP mechanism that already shipped —
scripts/engine.shthreads--mcp-config/--strict-mcp-configinto the deep + rubber-duck Claude tiers only (REVIEW_MCP_CONFIG/REVIEW_MCP_ALLOWED_TOOLS, epic #676, Context7 landed in #826). The triage tier stays fast/restricted and untouched. LSP is one more MCP server behind that same knob — we are not rebuilding the gating, the graceful-degradation path, or the token observatory.Scope (confirmed from the discussion)
.github-private(this repo — the agent's own substrate;bash-language-serveris lightweight, no cross-repo coordination).Initiative success metric
Go = on the frozen pilot PR corpus, LSP-on deep-tier review delivers a measurable navigation-token reduction (target >=2x fewer navigation tool-call tokens vs the LSP-off control) AND no regression in review quality (false-positive / precision no worse than the frozen LSP-off baseline), achieved within the cold-start SLA (<=30s P95). A win on tokens that costs precision is a no-go, not a win.
Cost cap (explicit)
Shape
Phase 1 scopes the pilot and freezes the measurement corpus/baseline; Phase 2 wires the candidate server(s) into the runtime and adds the finding-verification step; Phase 3 runs the comparison and records a human go/no-go.
Untracked prerequisites
bash-language-server(npm) must be installable in the pr-review runner environment for the Shell pilot.Planned from idea discussion #578 by the BMAD Scrum Master initiative-planner. Inert until a maintainer adds
initiative:auto.