Skip to content

feat(eval): recall-quality optimization harness — lab foundation + design - #197

Merged
jack-arturo merged 13 commits into
mainfrom
eval/recall-quality-harness
Jun 17, 2026
Merged

jack-arturo merged 13 commits into
mainfrom
eval/recall-quality-harness

Conversation

@jack-arturo

Copy link
Copy Markdown
Member

Recall-Quality Optimization Harness — lab foundation + design

Optimize AutoMem recall quality on a real corpus, then confirm on public benchmarks.

Design docs (docs/superpowers/): spec (corpus-tuned → benchmark-confirmed funnel; NDCG@10 primary + distractor-precision guardrail + simplicity/latency tiebreakers; labelled consolidation arm), Plan A (lab foundation), Plan B (parallel matrix harness — code lives in automem-evals).

Plan A code (scripts/lab/):

  • lab_metrics.py — NDCG@10, distractor_rate@10, config_complexity, pick_winner decision rule
  • lab_corpus.py — parameterized recall, aged labelled distractor injection, real consolidation pass (dry_run=false)
  • run_recall_test.py — scorecard wired into the runner (distractors, recall params, consolidation)

17/17 unit tests pass; black + flake8 clean. Does not touch the unrelated in-flight WIP in the working tree.

🤖 Generated with Claude Code

jack-arturo and others added 12 commits June 16, 2026 23:40
Corpus-tuned, benchmark-confirmed optimization harness. Primary metric
NDCG@10 + distractor-precision guardrail + simplicity/latency tiebreakers;
three-tier funnel (corpus sweep -> usefulness gate -> AMB/BEAM publish);
isolated parallel matrix stacks; labeled dreaming/consolidation arm.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
NDCG@10 scorecard + distractor-precision guardrail + simplicity/latency
tiebreakers + decision rule; parameterized recall; aged labelled distractor
injection; real consolidation pass. Pure logic extracted to lab_metrics.py /
lab_corpus.py for reuse by the parallel matrix harness (Plan B).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
build_scorecard emitted 'config_name' but pick_winner reads 'name' — a latent
KeyError once Plan B wires them. Standardize on 'name'; add a regression test
exercising the producer->consumer path; raise ValueError on unknown baseline;
drop the proven-unreachable regressed branch. Found by Plan A final review.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Corpus-sweep matrix: per-stack baked config, provenance manifest, idempotent
resume, RAM-capped concurrency, compose lint, winner via pick_winner. Scoring
imported from automem lab (DRY). Runs in an isolated automem-evals worktree off
main; synthetic-corpus smoke validates the live path without prod credentials.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings June 17, 2026 01:13

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR lays the foundation for a “recall-quality lab” harness by extracting scoring and corpus-side helpers into importable modules, updating the existing run_recall_test.py runner to compute the new scorecard axes (incl. distractor guardrail + complexity), and adding unit tests + design/plan docs to support Plan A/Plan B execution.

Changes:

  • Added scripts/lab/lab_metrics.py (NDCG@10, distractor rate, config complexity, pick_winner) and scripts/lab/lab_corpus.py (parameterized recall + distractor injection + consolidation helper).
  • Updated scripts/lab/run_recall_test.py to use the extracted modules and emit distractor-rate + complexity in results/CLI flow.
  • Added a focused pytest suite under tests/lab/ plus detailed design spec + implementation plans under docs/superpowers/.

Reviewed changes

Copilot reviewed 10 out of 10 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
scripts/lab/lab_metrics.py New pure scoring primitives + decision rule (pick_winner) for the lab scorecard.
scripts/lab/lab_corpus.py New HTTP/corpus helpers (recall params, distractor injection, consolidation).
scripts/lab/run_recall_test.py Runner updated to compute/track distractor-rate + config complexity and to call new helpers.
tests/lab/conftest.py Makes scripts/lab importable for the new unit tests.
tests/lab/test_lab_metrics.py Unit coverage for metrics + complexity + pick_winner.
tests/lab/test_lab_corpus.py Unit coverage for parameterized recall, ID extraction, distractor injection, consolidation ordering.
tests/lab/test_run_recall_test.py Validates build_scorecard contract and pick_winner compatibility.
docs/superpowers/specs/2026-06-16-recall-quality-optimization-harness-design.md Design spec for the corpus-tuned → benchmark-confirmed funnel and scorecard rule.
docs/superpowers/plans/2026-06-16-lab-metric-foundation.md Step-by-step Plan A implementation plan for the lab foundation.
docs/superpowers/plans/2026-06-17-matrix-parallel-harness.md Step-by-step Plan B implementation plan for the parallel matrix harness (automem-evals).

Comment thread scripts/lab/lab_metrics.py Outdated
Comment thread scripts/lab/run_recall_test.py
Comment thread docs/superpowers/specs/2026-06-16-recall-quality-optimization-harness-design.md Outdated
Comment thread docs/superpowers/specs/2026-06-16-recall-quality-optimization-harness-design.md Outdated
@jack-arturo
jack-arturo merged commit 431433e into main Jun 17, 2026
7 checks passed
@jack-arturo
jack-arturo deleted the eval/recall-quality-harness branch June 17, 2026 03:15
jack-arturo added a commit that referenced this pull request Jun 26, 2026
🤖 I have created a release *beep* *boop*
---


##
[0.16.0](v0.15.2...v0.16.0)
(2026-06-26)


### Features

* **api:** add admin backup endpoint
([#162](#162))
([8b1f264](8b1f264))
* **api:** support bulk memory associations
([1221e36](1221e36))
* **api:** support bulk memory associations
([#198](#198))
([28eb916](28eb916))
* **benchmarks:** LongMemEval failure-mode diagnosis harness + judge
quota preflight
([#183](#183))
([f99bece](f99bece))
* **consolidation:** expose cluster threshold and min size as env vars
([#163](#163))
([7e731f3](7e731f3))
* **enrichment:** expose classification fallback-rate metrics in
/enrichment/status
([#188](#188))
([0b522a9](0b522a9))
* **entity:** harden identity cleanup and repair tooling
([#176](#176))
([827dfbc](827dfbc))
* **eval:** recall-quality optimization harness — lab foundation +
design ([#197](#197))
([431433e](431433e))
* **graph:** support unbounded visualizer snapshots
([#141](#141))
([c730128](c730128))
* **lab:** add aged labelled distractor injection
([cc5d546](cc5d546))
* **lab:** add config_complexity simplicity metric
([dfb10d9](dfb10d9))
* **lab:** add distractor_rate_at_k precision guardrail metric
([872eab2](872eab2))
* **lab:** add lab_corpus with parameterized recall
([5e1e071](5e1e071))
* **lab:** add pick_winner scorecard decision rule
([3187eac](3187eac))
* **lab:** add real consolidation pass helper
([48a7d4a](48a7d4a))
* **lab:** isolate production clone restores
([#171](#171))
([aef90c0](aef90c0))
* **lab:** wire scorecard, distractors, recall params, consolidation
into runner
([589ec30](589ec30))
* **recall:** add metadata sidecar search
([#177](#177))
([4e7956e](4e7956e))
* **recall:** add state_mode=current|history recall alias
([#173](#173))
([b1df86c](b1df86c))
* **recall:** cap tag-score denominator to fix query-length bias
([#193](#193))
([cefa516](cefa516))
* **recall:** date-aware ranking + latest-fact selection
([#158](#158),
[#159](#159))
([#187](#187))
([a6ed945](a6ed945))
* **recall:** make recency decay window and curve configurable
([#182](#182))
([dbb933f](dbb933f))
* **recall:** ranking release — recency config, tag-score cap, relevance
gate, date-aware ranking
([#182](#182),
[#193](#193),
[#186](#186),
[#187](#187),
[#183](#183),
[#184](#184),
[#188](#188))
([#194](#194))
([337fe98](337fe98))
* **scripts:** safer reclassify_with_llm.py with provider flags +
tighter prompt
([#164](#164))
([a742602](a742602))


### Bug Fixes

* **api:** address copilot review on PR
[#198](#198)
([0466a1e](0466a1e))
* **api:** handle grouped association write failures
([cd93df9](cd93df9))
* **backup:** make backup_automem.py runnable as `python
scripts/backup_automem.py`
([#175](#175))
([edd9742](edd9742))
* **benchmarks:** add publication verification bundle
([#166](#166))
([420d721](420d721))
* **consolidation:** skip eager first tick at startup to avoid FalkorDB
load race
([#165](#165))
([1b812cf](1b812cf))
* **docs:** keep dispatch payload arrays stable
([df6e9e8](df6e9e8))
* **embedding:** fall back to per-item real embeddings before
placeholders in batch path
([#189](#189))
([6e9c62c](6e9c62c))
* **entity:** restore person-shape exemption on the slug validation path
([#179](#179))
([5e29960](5e29960))
* **entity:** stop validator over-rejecting real people, code tools, and
event categories
([#178](#178))
([193b730](193b730))
* **lab:** address copilot review on PR
[#197](#197)
([45f80d6](45f80d6))
* **lab:** align scorecard key contract (build_scorecard -&gt;
pick_winner)
([7d91530](7d91530))
* **mcp-sse:** decouple /health liveness from upstream readiness
([#151](#151))
([5bcfb8b](5bcfb8b))
* **mcp:** cap association failure summary
([ea4e08f](ea4e08f))
* **mcp:** surface stored metadata and updated_at in detailed recall
format ([#184](#184))
([230416e](230416e))
* **recall:** address copilot review on PR
[#194](#194)
([50b1647](50b1647))
* **recall:** canonicalize / and : separators in context_tag matching
([3afd9d3](3afd9d3))
* **recall:** canonicalize / and : separators in context_tag matching
([#203](#203))
([ba5e9ff](ba5e9ff))
* **recall:** gate query-independent scoring on topical evidence within
tag scope
([#130](#130))
([#186](#186))
([c11b594](c11b594))
* **recall:** hydrate semantic recall summaries
([#192](#192))
([76e845d](76e845d))
* **recall:** normalize graph keyword scores into the 0-1 component
range ([#191](#191))
([3653ddf](3653ddf))
* **recall:** respect current memory state
([#170](#170))
([ed36b98](ed36b98)),
closes [#169](#169)
[#158](#158)
[#159](#159)
* **scripts:** add sys.path guard to reembed_embeddings.py
([d333cf0](d333cf0))


### Documentation

* add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations
([f20c664](f20c664))
* **bench:** log full judged 500q LongMemEval ship-config run with churn
attribution
([41bf8d0](41bf8d0))
* **eval:** Plan A — lab metric foundation (TDD, 9 tasks)
([0087dda](0087dda))
* **eval:** Plan B — parallel matrix harness (TDD, 9 tasks)
([c8ddfb2](c8ddfb2))
* **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as
diagnostic-only
([#174](#174))
([e8a3285](e8a3285))
* **eval:** spec for recall-quality optimization harness
([b1a1995](b1a1995))
* fix stale claims and document gated flags for 0.16.0
([b152d64](b152d64))
* note develop-branch contribution policy in README
([ccf02dd](ccf02dd))
* **positioning:** add scout reference
([#168](#168))
([922d23b](922d23b))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([3ff95bd](3ff95bd))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([#204](#204))
([89c30e0](89c30e0))
* refresh README and benchmark guidance
([#157](#157))
([bba31cc](bba31cc))
* **runtime:** align Docker viewer paths and setup guidance
([#155](#155))
([bbda79b](bbda79b))
* scripts catalog, recall quality lab guide, and 0.16.0 migration
runbook ([#199](#199))
([f190ae5](f190ae5))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants