fix(benchmarks): add publication verification bundle - #166
Merged
Conversation
Contributor
There was a problem hiding this comment.
Pull request overview
Note
Copilot was unable to run its full agentic suite in this review.
Adds a reproducibility-focused “publication verification” bundle for the May 2026 arXiv effort, updates published benchmark claims across docs, and fixes FalkorDB’s /memories/<id>/related fallback query by inlining a sanitized variable-length depth (with a regression test).
Changes:
- Updated README/docs benchmark numbers to the latest verified LoCoMo/LongMemEval results.
- Added
benchmarks/publication/2026-05-arxiv/bundle (commands, notes, machine-readable manifest). - Fixed FalkorDB fallback related-memories query to inline sanitized depth; added regression test.
Reviewed changes
Copilot reviewed 14 out of 14 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/test_api_endpoints.py | Adds regression test ensuring fallback query inlines depth and avoids $max_depth in relationship range. |
| automem/api/recall.py | Inlines sanitized max_depth into variable-length range for FalkorDB compatibility. |
| pyproject.toml | Introduces Black/isort configuration and generated-directory excludes. |
| Makefile | Updates bench-health to run using venv Python. |
| .flake8 | Extends excludes for generated/worktree/benchmark directories. |
| README.md | Updates headline benchmark claims and adds publication bundle link. |
| docs/TESTING.md | Updates LoCoMo/LongMemEval baseline tables and example output to new results. |
| docs/COMPARISON.md | Refreshes text to reflect current canonical benchmark posture. |
| benchmarks/EXPERIMENT_LOG.md | Updates headline results; adds publication-bundle pointer and verification entry. |
| benchmarks/publication/2026-05-arxiv/README.md | Adds claim posture + bundle index for the arXiv publication effort. |
| benchmarks/publication/2026-05-arxiv/benchmark-summary.md | Adds paper-ready benchmark summary and limitations. |
| benchmarks/publication/2026-05-arxiv/commands.md | Documents repo + supplemental eval commands for verification. |
| benchmarks/publication/2026-05-arxiv/fresh-verification.md | Captures verification run notes and artifact hashes. |
| benchmarks/publication/2026-05-arxiv/artifact-manifest.json | Adds machine-readable manifest for claims/artifacts/hashes. |
jack-arturo
added a commit
that referenced
this pull request
Jun 26, 2026
🤖 I have created a release *beep* *boop* --- ## [0.16.0](v0.15.2...v0.16.0) (2026-06-26) ### Features * **api:** add admin backup endpoint ([#162](#162)) ([8b1f264](8b1f264)) * **api:** support bulk memory associations ([1221e36](1221e36)) * **api:** support bulk memory associations ([#198](#198)) ([28eb916](28eb916)) * **benchmarks:** LongMemEval failure-mode diagnosis harness + judge quota preflight ([#183](#183)) ([f99bece](f99bece)) * **consolidation:** expose cluster threshold and min size as env vars ([#163](#163)) ([7e731f3](7e731f3)) * **enrichment:** expose classification fallback-rate metrics in /enrichment/status ([#188](#188)) ([0b522a9](0b522a9)) * **entity:** harden identity cleanup and repair tooling ([#176](#176)) ([827dfbc](827dfbc)) * **eval:** recall-quality optimization harness — lab foundation + design ([#197](#197)) ([431433e](431433e)) * **graph:** support unbounded visualizer snapshots ([#141](#141)) ([c730128](c730128)) * **lab:** add aged labelled distractor injection ([cc5d546](cc5d546)) * **lab:** add config_complexity simplicity metric ([dfb10d9](dfb10d9)) * **lab:** add distractor_rate_at_k precision guardrail metric ([872eab2](872eab2)) * **lab:** add lab_corpus with parameterized recall ([5e1e071](5e1e071)) * **lab:** add pick_winner scorecard decision rule ([3187eac](3187eac)) * **lab:** add real consolidation pass helper ([48a7d4a](48a7d4a)) * **lab:** isolate production clone restores ([#171](#171)) ([aef90c0](aef90c0)) * **lab:** wire scorecard, distractors, recall params, consolidation into runner ([589ec30](589ec30)) * **recall:** add metadata sidecar search ([#177](#177)) ([4e7956e](4e7956e)) * **recall:** add state_mode=current|history recall alias ([#173](#173)) ([b1df86c](b1df86c)) * **recall:** cap tag-score denominator to fix query-length bias ([#193](#193)) ([cefa516](cefa516)) * **recall:** date-aware ranking + latest-fact selection ([#158](#158), [#159](#159)) ([#187](#187)) ([a6ed945](a6ed945)) * **recall:** make recency decay window and curve configurable ([#182](#182)) ([dbb933f](dbb933f)) * **recall:** ranking release — recency config, tag-score cap, relevance gate, date-aware ranking ([#182](#182), [#193](#193), [#186](#186), [#187](#187), [#183](#183), [#184](#184), [#188](#188)) ([#194](#194)) ([337fe98](337fe98)) * **scripts:** safer reclassify_with_llm.py with provider flags + tighter prompt ([#164](#164)) ([a742602](a742602)) ### Bug Fixes * **api:** address copilot review on PR [#198](#198) ([0466a1e](0466a1e)) * **api:** handle grouped association write failures ([cd93df9](cd93df9)) * **backup:** make backup_automem.py runnable as `python scripts/backup_automem.py` ([#175](#175)) ([edd9742](edd9742)) * **benchmarks:** add publication verification bundle ([#166](#166)) ([420d721](420d721)) * **consolidation:** skip eager first tick at startup to avoid FalkorDB load race ([#165](#165)) ([1b812cf](1b812cf)) * **docs:** keep dispatch payload arrays stable ([df6e9e8](df6e9e8)) * **embedding:** fall back to per-item real embeddings before placeholders in batch path ([#189](#189)) ([6e9c62c](6e9c62c)) * **entity:** restore person-shape exemption on the slug validation path ([#179](#179)) ([5e29960](5e29960)) * **entity:** stop validator over-rejecting real people, code tools, and event categories ([#178](#178)) ([193b730](193b730)) * **lab:** address copilot review on PR [#197](#197) ([45f80d6](45f80d6)) * **lab:** align scorecard key contract (build_scorecard -> pick_winner) ([7d91530](7d91530)) * **mcp-sse:** decouple /health liveness from upstream readiness ([#151](#151)) ([5bcfb8b](5bcfb8b)) * **mcp:** cap association failure summary ([ea4e08f](ea4e08f)) * **mcp:** surface stored metadata and updated_at in detailed recall format ([#184](#184)) ([230416e](230416e)) * **recall:** address copilot review on PR [#194](#194) ([50b1647](50b1647)) * **recall:** canonicalize / and : separators in context_tag matching ([3afd9d3](3afd9d3)) * **recall:** canonicalize / and : separators in context_tag matching ([#203](#203)) ([ba5e9ff](ba5e9ff)) * **recall:** gate query-independent scoring on topical evidence within tag scope ([#130](#130)) ([#186](#186)) ([c11b594](c11b594)) * **recall:** hydrate semantic recall summaries ([#192](#192)) ([76e845d](76e845d)) * **recall:** normalize graph keyword scores into the 0-1 component range ([#191](#191)) ([3653ddf](3653ddf)) * **recall:** respect current memory state ([#170](#170)) ([ed36b98](ed36b98)), closes [#169](#169) [#158](#158) [#159](#159) * **scripts:** add sys.path guard to reembed_embeddings.py ([d333cf0](d333cf0)) ### Documentation * add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations ([f20c664](f20c664)) * **bench:** log full judged 500q LongMemEval ship-config run with churn attribution ([41bf8d0](41bf8d0)) * **eval:** Plan A — lab metric foundation (TDD, 9 tasks) ([0087dda](0087dda)) * **eval:** Plan B — parallel matrix harness (TDD, 9 tasks) ([c8ddfb2](c8ddfb2)) * **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as diagnostic-only ([#174](#174)) ([e8a3285](e8a3285)) * **eval:** spec for recall-quality optimization harness ([b1a1995](b1a1995)) * fix stale claims and document gated flags for 0.16.0 ([b152d64](b152d64)) * note develop-branch contribution policy in README ([ccf02dd](ccf02dd)) * **positioning:** add scout reference ([#168](#168)) ([922d23b](922d23b)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([3ff95bd](3ff95bd)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([#204](#204)) ([89c30e0](89c30e0)) * refresh README and benchmark guidance ([#157](#157)) ([bba31cc](bba31cc)) * **runtime:** align Docker viewer paths and setup guidance ([#155](#155)) ([bbda79b](bbda79b)) * scripts catalog, recall quality lab guide, and 0.16.0 migration runbook ([#199](#199)) ([f190ae5](f190ae5)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Breaking Changes
None. No runtime API shape changed.
Related Issues
Publication/reproducibility prep for the AutoMem arXiv paper.
Test Plan
Notes
The fresh LongMemEval full console output included transient gpt-5-mini empty-answer warnings and one local recall read timeout, but the harness completed with memory_ingest_failures=0, judge_errors=0, and publishable=true. Actual arXiv/Hugging Face submission still needs author metadata and an arXiv ID.