docs(evals): mark Memora/FAMA/WRIT lifecycle diagnostics as diagnostic-only - #174
Merged
Conversation
…c-only Clarify across EVALS_CONTRACT, TESTING, and EXPERIMENT_LOG that Memora/FAMA and WRIT lifecycle diagnostics from automem-evals stay diagnostic signals until their scenario, harness, and artifacts are promoted into this repo's official benchmark policy. Keeps benchmark-claim ownership in automem. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Contributor
There was a problem hiding this comment.
Pull request overview
Documentation-only hygiene update clarifying that Memora/FAMA, WRIT, and similar lifecycle diagnostics run from automem-evals should be treated as diagnostic signals until their scenarios, harnesses, and artifacts are promoted into automem’s official benchmark policy/flow.
Changes:
- Clarifies in
docs/TESTING.mdthat Memora/FAMA and WRIT runs are diagnostic unless promoted into the official benchmark flow and recorded in the experiment log. - Adds an explicit contract rule in
docs/EVALS_CONTRACT.mdto treat these lifecycle diagnostics as diagnostic signals until promoted. - Notes in
benchmarks/EXPERIMENT_LOG.mdthat Memora/FAMA and WRIT lifecycle diagnostics remain non-canonical unless promoted through official benchmark policy.
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| docs/TESTING.md | Adds benchmark-ownership clarification for Memora/FAMA and WRIT runs. |
| docs/EVALS_CONTRACT.md | Adds a contract bullet classifying lifecycle diagnostics as diagnostic signals until promoted. |
| benchmarks/EXPERIMENT_LOG.md | Adds an experiment-log note to keep lifecycle diagnostics categorized as non-canonical unless promoted. |
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
jack-arturo
added a commit
that referenced
this pull request
Jun 26, 2026
🤖 I have created a release *beep* *boop* --- ## [0.16.0](v0.15.2...v0.16.0) (2026-06-26) ### Features * **api:** add admin backup endpoint ([#162](#162)) ([8b1f264](8b1f264)) * **api:** support bulk memory associations ([1221e36](1221e36)) * **api:** support bulk memory associations ([#198](#198)) ([28eb916](28eb916)) * **benchmarks:** LongMemEval failure-mode diagnosis harness + judge quota preflight ([#183](#183)) ([f99bece](f99bece)) * **consolidation:** expose cluster threshold and min size as env vars ([#163](#163)) ([7e731f3](7e731f3)) * **enrichment:** expose classification fallback-rate metrics in /enrichment/status ([#188](#188)) ([0b522a9](0b522a9)) * **entity:** harden identity cleanup and repair tooling ([#176](#176)) ([827dfbc](827dfbc)) * **eval:** recall-quality optimization harness — lab foundation + design ([#197](#197)) ([431433e](431433e)) * **graph:** support unbounded visualizer snapshots ([#141](#141)) ([c730128](c730128)) * **lab:** add aged labelled distractor injection ([cc5d546](cc5d546)) * **lab:** add config_complexity simplicity metric ([dfb10d9](dfb10d9)) * **lab:** add distractor_rate_at_k precision guardrail metric ([872eab2](872eab2)) * **lab:** add lab_corpus with parameterized recall ([5e1e071](5e1e071)) * **lab:** add pick_winner scorecard decision rule ([3187eac](3187eac)) * **lab:** add real consolidation pass helper ([48a7d4a](48a7d4a)) * **lab:** isolate production clone restores ([#171](#171)) ([aef90c0](aef90c0)) * **lab:** wire scorecard, distractors, recall params, consolidation into runner ([589ec30](589ec30)) * **recall:** add metadata sidecar search ([#177](#177)) ([4e7956e](4e7956e)) * **recall:** add state_mode=current|history recall alias ([#173](#173)) ([b1df86c](b1df86c)) * **recall:** cap tag-score denominator to fix query-length bias ([#193](#193)) ([cefa516](cefa516)) * **recall:** date-aware ranking + latest-fact selection ([#158](#158), [#159](#159)) ([#187](#187)) ([a6ed945](a6ed945)) * **recall:** make recency decay window and curve configurable ([#182](#182)) ([dbb933f](dbb933f)) * **recall:** ranking release — recency config, tag-score cap, relevance gate, date-aware ranking ([#182](#182), [#193](#193), [#186](#186), [#187](#187), [#183](#183), [#184](#184), [#188](#188)) ([#194](#194)) ([337fe98](337fe98)) * **scripts:** safer reclassify_with_llm.py with provider flags + tighter prompt ([#164](#164)) ([a742602](a742602)) ### Bug Fixes * **api:** address copilot review on PR [#198](#198) ([0466a1e](0466a1e)) * **api:** handle grouped association write failures ([cd93df9](cd93df9)) * **backup:** make backup_automem.py runnable as `python scripts/backup_automem.py` ([#175](#175)) ([edd9742](edd9742)) * **benchmarks:** add publication verification bundle ([#166](#166)) ([420d721](420d721)) * **consolidation:** skip eager first tick at startup to avoid FalkorDB load race ([#165](#165)) ([1b812cf](1b812cf)) * **docs:** keep dispatch payload arrays stable ([df6e9e8](df6e9e8)) * **embedding:** fall back to per-item real embeddings before placeholders in batch path ([#189](#189)) ([6e9c62c](6e9c62c)) * **entity:** restore person-shape exemption on the slug validation path ([#179](#179)) ([5e29960](5e29960)) * **entity:** stop validator over-rejecting real people, code tools, and event categories ([#178](#178)) ([193b730](193b730)) * **lab:** address copilot review on PR [#197](#197) ([45f80d6](45f80d6)) * **lab:** align scorecard key contract (build_scorecard -> pick_winner) ([7d91530](7d91530)) * **mcp-sse:** decouple /health liveness from upstream readiness ([#151](#151)) ([5bcfb8b](5bcfb8b)) * **mcp:** cap association failure summary ([ea4e08f](ea4e08f)) * **mcp:** surface stored metadata and updated_at in detailed recall format ([#184](#184)) ([230416e](230416e)) * **recall:** address copilot review on PR [#194](#194) ([50b1647](50b1647)) * **recall:** canonicalize / and : separators in context_tag matching ([3afd9d3](3afd9d3)) * **recall:** canonicalize / and : separators in context_tag matching ([#203](#203)) ([ba5e9ff](ba5e9ff)) * **recall:** gate query-independent scoring on topical evidence within tag scope ([#130](#130)) ([#186](#186)) ([c11b594](c11b594)) * **recall:** hydrate semantic recall summaries ([#192](#192)) ([76e845d](76e845d)) * **recall:** normalize graph keyword scores into the 0-1 component range ([#191](#191)) ([3653ddf](3653ddf)) * **recall:** respect current memory state ([#170](#170)) ([ed36b98](ed36b98)), closes [#169](#169) [#158](#158) [#159](#159) * **scripts:** add sys.path guard to reembed_embeddings.py ([d333cf0](d333cf0)) ### Documentation * add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations ([f20c664](f20c664)) * **bench:** log full judged 500q LongMemEval ship-config run with churn attribution ([41bf8d0](41bf8d0)) * **eval:** Plan A — lab metric foundation (TDD, 9 tasks) ([0087dda](0087dda)) * **eval:** Plan B — parallel matrix harness (TDD, 9 tasks) ([c8ddfb2](c8ddfb2)) * **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as diagnostic-only ([#174](#174)) ([e8a3285](e8a3285)) * **eval:** spec for recall-quality optimization harness ([b1a1995](b1a1995)) * fix stale claims and document gated flags for 0.16.0 ([b152d64](b152d64)) * note develop-branch contribution policy in README ([ccf02dd](ccf02dd)) * **positioning:** add scout reference ([#168](#168)) ([922d23b](922d23b)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([3ff95bd](3ff95bd)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([#204](#204)) ([89c30e0](89c30e0)) * refresh README and benchmark guidance ([#157](#157)) ([bba31cc](bba31cc)) * **runtime:** align Docker viewer paths and setup guidance ([#155](#155)) ([bbda79b](bbda79b)) * scripts catalog, recall quality lab guide, and 0.16.0 migration runbook ([#199](#199)) ([f190ae5](f190ae5)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Documentation-only hygiene. Clarifies across
EVALS_CONTRACT.md,TESTING.md, andEXPERIMENT_LOG.mdthat Memora/FAMA and WRIT lifecycle diagnostics fromautomem-evalsremain diagnostic signals until their scenario, harness, and result artifacts are promoted into this repo's official benchmark policy.Keeps benchmark-claim ownership in
automem, consistent with the existing Benchmark Ownership policy.No code changes.
🤖 Generated with Claude Code