Skip to content

fix(benchmarks): add publication verification bundle - #166

Merged
jack-arturo merged 2 commits into
mainfrom
feat/automem-arxiv-publication
May 22, 2026
Merged

fix(benchmarks): add publication verification bundle#166
jack-arturo merged 2 commits into
mainfrom
feat/automem-arxiv-publication

Conversation

@jack-arturo

Copy link
Copy Markdown
Member

Summary

  • update README/docs benchmark claims to the fresh publication verification results: LongMemEval full 87.00% / recall@5 97.00%, LoCoMo full 84.74%
  • add benchmarks/publication/2026-05-arxiv with claim posture, commands, artifact hashes, fresh verification notes, and a machine-readable manifest
  • fix the /memories//related fallback query by inlining the sanitized FalkorDB variable-length depth, with a regression test

Breaking Changes

None. No runtime API shape changed.

Related Issues

Publication/reproducibility prep for the AutoMem arXiv paper.

Test Plan

  • make test: 238 passed, 1 skipped, 25 deselected
  • .venv/bin/black --check .: 117 files unchanged
  • .venv/bin/isort --check-only .: skipped 10 files
  • make lint: passed
  • make test-integration: 11 passed, 253 deselected
  • make bench-health: HEALTHY, p50=416ms p95=441ms mean=421ms
  • LoCoMo mini pinned judge: 85.20% (259/304), sha256 ba2b98b0055f92ca17de9bc36207d7f39cf90b6270c2c3d903d69b8044aa7015
  • LoCoMo full pinned judge: 84.74% (1683/1986), sha256 a75816e9a6d3302c22b34852b75ac19a9d9f5cb27d1a109e0af7e49359330716
  • LongMemEval mini: 70.00% (21/30), recall@5 96.67%, sha256 7ea922b77e312a17c313bbf8c0e81f0268b48d1082080cae1db3c38e906577b8
  • LongMemEval full: 87.00% (435/500), recall@5 97.00%, sha256 ed6f7cf69b7be6fa0050536ec2b0f947f5510afd8c2a374b3fafb9cde009da75
  • automem-evals: runner unit tests 95 passed, scripts unit tests 10 passed, Writ npm test 72 passed, Writ npm run build passed
  • paper static check: inputs and BibTeX cite keys resolve; no local LaTeX compiler available

Notes

The fresh LongMemEval full console output included transient gpt-5-mini empty-answer warnings and one local recall read timeout, but the harness completed with memory_ingest_failures=0, judge_errors=0, and publishable=true. Actual arXiv/Hugging Face submission still needs author metadata and an arXiv ID.

Copilot AI review requested due to automatic review settings May 18, 2026 06:25

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Adds a reproducibility-focused “publication verification” bundle for the May 2026 arXiv effort, updates published benchmark claims across docs, and fixes FalkorDB’s /memories/<id>/related fallback query by inlining a sanitized variable-length depth (with a regression test).

Changes:

  • Updated README/docs benchmark numbers to the latest verified LoCoMo/LongMemEval results.
  • Added benchmarks/publication/2026-05-arxiv/ bundle (commands, notes, machine-readable manifest).
  • Fixed FalkorDB fallback related-memories query to inline sanitized depth; added regression test.

Reviewed changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
tests/test_api_endpoints.py Adds regression test ensuring fallback query inlines depth and avoids $max_depth in relationship range.
automem/api/recall.py Inlines sanitized max_depth into variable-length range for FalkorDB compatibility.
pyproject.toml Introduces Black/isort configuration and generated-directory excludes.
Makefile Updates bench-health to run using venv Python.
.flake8 Extends excludes for generated/worktree/benchmark directories.
README.md Updates headline benchmark claims and adds publication bundle link.
docs/TESTING.md Updates LoCoMo/LongMemEval baseline tables and example output to new results.
docs/COMPARISON.md Refreshes text to reflect current canonical benchmark posture.
benchmarks/EXPERIMENT_LOG.md Updates headline results; adds publication-bundle pointer and verification entry.
benchmarks/publication/2026-05-arxiv/README.md Adds claim posture + bundle index for the arXiv publication effort.
benchmarks/publication/2026-05-arxiv/benchmark-summary.md Adds paper-ready benchmark summary and limitations.
benchmarks/publication/2026-05-arxiv/commands.md Documents repo + supplemental eval commands for verification.
benchmarks/publication/2026-05-arxiv/fresh-verification.md Captures verification run notes and artifact hashes.
benchmarks/publication/2026-05-arxiv/artifact-manifest.json Adds machine-readable manifest for claims/artifacts/hashes.

Comment thread Makefile Outdated
Comment thread pyproject.toml
Comment thread benchmarks/publication/2026-05-arxiv/commands.md Outdated
Comment thread benchmarks/publication/2026-05-arxiv/fresh-verification.md Outdated
Comment thread tests/test_api_endpoints.py Outdated
Comment thread docs/TESTING.md Outdated
@jack-arturo
jack-arturo merged commit 420d721 into main May 22, 2026
7 checks passed
@jack-arturo
jack-arturo deleted the feat/automem-arxiv-publication branch May 22, 2026 07:39
jack-arturo added a commit that referenced this pull request Jun 26, 2026
🤖 I have created a release *beep* *boop*
---


##
[0.16.0](v0.15.2...v0.16.0)
(2026-06-26)


### Features

* **api:** add admin backup endpoint
([#162](#162))
([8b1f264](8b1f264))
* **api:** support bulk memory associations
([1221e36](1221e36))
* **api:** support bulk memory associations
([#198](#198))
([28eb916](28eb916))
* **benchmarks:** LongMemEval failure-mode diagnosis harness + judge
quota preflight
([#183](#183))
([f99bece](f99bece))
* **consolidation:** expose cluster threshold and min size as env vars
([#163](#163))
([7e731f3](7e731f3))
* **enrichment:** expose classification fallback-rate metrics in
/enrichment/status
([#188](#188))
([0b522a9](0b522a9))
* **entity:** harden identity cleanup and repair tooling
([#176](#176))
([827dfbc](827dfbc))
* **eval:** recall-quality optimization harness — lab foundation +
design ([#197](#197))
([431433e](431433e))
* **graph:** support unbounded visualizer snapshots
([#141](#141))
([c730128](c730128))
* **lab:** add aged labelled distractor injection
([cc5d546](cc5d546))
* **lab:** add config_complexity simplicity metric
([dfb10d9](dfb10d9))
* **lab:** add distractor_rate_at_k precision guardrail metric
([872eab2](872eab2))
* **lab:** add lab_corpus with parameterized recall
([5e1e071](5e1e071))
* **lab:** add pick_winner scorecard decision rule
([3187eac](3187eac))
* **lab:** add real consolidation pass helper
([48a7d4a](48a7d4a))
* **lab:** isolate production clone restores
([#171](#171))
([aef90c0](aef90c0))
* **lab:** wire scorecard, distractors, recall params, consolidation
into runner
([589ec30](589ec30))
* **recall:** add metadata sidecar search
([#177](#177))
([4e7956e](4e7956e))
* **recall:** add state_mode=current|history recall alias
([#173](#173))
([b1df86c](b1df86c))
* **recall:** cap tag-score denominator to fix query-length bias
([#193](#193))
([cefa516](cefa516))
* **recall:** date-aware ranking + latest-fact selection
([#158](#158),
[#159](#159))
([#187](#187))
([a6ed945](a6ed945))
* **recall:** make recency decay window and curve configurable
([#182](#182))
([dbb933f](dbb933f))
* **recall:** ranking release — recency config, tag-score cap, relevance
gate, date-aware ranking
([#182](#182),
[#193](#193),
[#186](#186),
[#187](#187),
[#183](#183),
[#184](#184),
[#188](#188))
([#194](#194))
([337fe98](337fe98))
* **scripts:** safer reclassify_with_llm.py with provider flags +
tighter prompt
([#164](#164))
([a742602](a742602))


### Bug Fixes

* **api:** address copilot review on PR
[#198](#198)
([0466a1e](0466a1e))
* **api:** handle grouped association write failures
([cd93df9](cd93df9))
* **backup:** make backup_automem.py runnable as `python
scripts/backup_automem.py`
([#175](#175))
([edd9742](edd9742))
* **benchmarks:** add publication verification bundle
([#166](#166))
([420d721](420d721))
* **consolidation:** skip eager first tick at startup to avoid FalkorDB
load race
([#165](#165))
([1b812cf](1b812cf))
* **docs:** keep dispatch payload arrays stable
([df6e9e8](df6e9e8))
* **embedding:** fall back to per-item real embeddings before
placeholders in batch path
([#189](#189))
([6e9c62c](6e9c62c))
* **entity:** restore person-shape exemption on the slug validation path
([#179](#179))
([5e29960](5e29960))
* **entity:** stop validator over-rejecting real people, code tools, and
event categories
([#178](#178))
([193b730](193b730))
* **lab:** address copilot review on PR
[#197](#197)
([45f80d6](45f80d6))
* **lab:** align scorecard key contract (build_scorecard -&gt;
pick_winner)
([7d91530](7d91530))
* **mcp-sse:** decouple /health liveness from upstream readiness
([#151](#151))
([5bcfb8b](5bcfb8b))
* **mcp:** cap association failure summary
([ea4e08f](ea4e08f))
* **mcp:** surface stored metadata and updated_at in detailed recall
format ([#184](#184))
([230416e](230416e))
* **recall:** address copilot review on PR
[#194](#194)
([50b1647](50b1647))
* **recall:** canonicalize / and : separators in context_tag matching
([3afd9d3](3afd9d3))
* **recall:** canonicalize / and : separators in context_tag matching
([#203](#203))
([ba5e9ff](ba5e9ff))
* **recall:** gate query-independent scoring on topical evidence within
tag scope
([#130](#130))
([#186](#186))
([c11b594](c11b594))
* **recall:** hydrate semantic recall summaries
([#192](#192))
([76e845d](76e845d))
* **recall:** normalize graph keyword scores into the 0-1 component
range ([#191](#191))
([3653ddf](3653ddf))
* **recall:** respect current memory state
([#170](#170))
([ed36b98](ed36b98)),
closes [#169](#169)
[#158](#158)
[#159](#159)
* **scripts:** add sys.path guard to reembed_embeddings.py
([d333cf0](d333cf0))


### Documentation

* add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations
([f20c664](f20c664))
* **bench:** log full judged 500q LongMemEval ship-config run with churn
attribution
([41bf8d0](41bf8d0))
* **eval:** Plan A — lab metric foundation (TDD, 9 tasks)
([0087dda](0087dda))
* **eval:** Plan B — parallel matrix harness (TDD, 9 tasks)
([c8ddfb2](c8ddfb2))
* **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as
diagnostic-only
([#174](#174))
([e8a3285](e8a3285))
* **eval:** spec for recall-quality optimization harness
([b1a1995](b1a1995))
* fix stale claims and document gated flags for 0.16.0
([b152d64](b152d64))
* note develop-branch contribution policy in README
([ccf02dd](ccf02dd))
* **positioning:** add scout reference
([#168](#168))
([922d23b](922d23b))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([3ff95bd](3ff95bd))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([#204](#204))
([89c30e0](89c30e0))
* refresh README and benchmark guidance
([#157](#157))
([bba31cc](bba31cc))
* **runtime:** align Docker viewer paths and setup guidance
([#155](#155))
([bbda79b](bbda79b))
* scripts catalog, recall quality lab guide, and 0.16.0 migration
runbook ([#199](#199))
([f190ae5](f190ae5))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants