feat(recall): make recency decay window and curve configurable - #182
Conversation
There was a problem hiding this comment.
Pull request overview
This PR makes the recall recency scoring configurable via environment variables, replacing the previously hardcoded 180-day linear decay. This supports tuning and benchmarking recency behavior (including an exponential “half-life” mode) while keeping defaults identical to legacy behavior.
Changes:
- Add
SEARCH_RECENCY_WINDOW_DAYSandSEARCH_RECENCY_CURVE(linear/exp) to config and recency scoring. - Update diagnostics tooling and lab harness to respect the new
SEARCH_RECENCY_*configuration. - Document the new environment variables and update supporting docs/examples.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 2 comments.
Show a summary per file
| File | Description |
|---|---|
automem/utils/scoring.py |
Use configured recency window/curve in _compute_recency_score(). |
automem/config.py |
Add validated config for SEARCH_RECENCY_WINDOW_DAYS and SEARCH_RECENCY_CURVE. |
tests/test_api_endpoints.py |
Add unit tests for recency scoring behavior and config guard. |
scripts/browse_memories.py |
Make diagnose recency simulation configurable via SEARCH_RECENCY_*. |
scripts/lab/run_recall_test.py |
Restart API container for any SEARCH_* changes during sweeps. |
docs/ENVIRONMENT_VARIABLES.md |
Document SEARCH_RECENCY_WINDOW_DAYS / SEARCH_RECENCY_CURVE and clarify weights text. |
docs/COMPARISON.md |
Update recency description to reflect configurability. |
CLAUDE.md |
Update env var example text to reference SEARCH_RECENCY_*. |
| recency_window = float(os.getenv("SEARCH_RECENCY_WINDOW_DAYS", "180")) | ||
| if recency_window <= 0: | ||
| recency_window = 180.0 |
There was a problem hiding this comment.
Fixed by validating math.isfinite(recency_window) before using the diagnose recency window — scripts/browse_memories.py:764. GitHub marked the original line outdated after the push.
| def _timestamp_days_ago(days: float) -> str: | ||
| return (datetime.now(timezone.utc) - timedelta(days=days)).isoformat() | ||
|
|
There was a problem hiding this comment.
Pinned recency-score tests to the legacy 180-day linear config via legacy_recency_config, so local .env overrides no longer affect them — tests/test_api_endpoints.py:433. GitHub marked the original line outdated after the push.
…nge (#191) Fixes #190. ## Problem `_graph_keyword_search` returned the **raw Cypher additive score** — +2 per keyword contained in content, +1 per keyword in any tag, summed over all extracted keywords, plus a +2/+1 whole-phrase bonus — so a K-keyword query can score up to **3K+3** while every other channel (vector cosine, metadata, trending importance) lives in 0–1. Observed during the 2026-06-11 production forensics: a tag-scoped exact-content match returned `keyword=11.0`, `final_score=4.03`. Consequences: - `SEARCH_WEIGHT_KEYWORD (0.35) × 11 = 3.85` — a keyword hit trumps any vector/metadata/importance combination, scaling with *query length* rather than match quality. - Defeats `RECALL_RELEVANCE_GATE` semantics (PR #186): `evidence = max(vector, keyword, metadata, exact)` assumes 0–1 components; `evidence = 11` sails past any gate. ## Fix 1. **Producer**: normalize the raw score by its per-query maximum (`3·len(keywords) + 3` when a phrase is present; `3` in the phrase-only branch) before it leaves `_graph_keyword_search`. This is a monotone per-query transform — within-channel ordering (and the Cypher `ORDER BY`) is unchanged; only cross-channel blending changes, which is the point. 2. **Consumer**: defensively clamp the keyword component to `min(1.0, …)` in `_compute_metadata_score`, so no future producer can break the 0–1 contract or the gate again. ## Verification - 4 new tests in `tests/test_keyword_score_normalization.py` (TDD'd against the bug, including the literal `keyword=11.0` repro). Full suite: **503 passed**. - **Production-corpus lab A/B** (10,142-memory snapshot, 200 queries, vs the pooled 3-run parity baseline from the 2026-06-11 release sweep): - Recall@5 −0.2pp, Recall@10 −0.7pp, MRR −0.008, NDCG@10 −0.007 (all within baseline run-to-run variance; paired t-test p=0.32) - Per-query: **196/197 unchanged**, 0 improved, 1 regressed - The single flip is the intended behavior change made visible: the expected memory held rank 1 *only* via the inflated keyword score (ranks 2–5 identical before/after). It's in the fallback-typed `Memory` cohort (MRR 0.15 baseline — the known data-quality cohort from #188's classification incident). ## Notes - This lands on main independently of the #182–#187 chain; #186's gate evidence check is the main beneficiary once the chain rebases over it. - Trending (`importance`), metadata (capped), and vector (cosine) channels were verified already bounded; the graph keyword channel was the only unbounded producer. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
5388e65 to
02f9bb9
Compare
Add SEARCH_RECENCY_WINDOW_DAYS (default 180) and SEARCH_RECENCY_CURVE (linear|exp, default linear) so the recency score's decay shape can be tuned via environment instead of the hardcoded 180-day linear decay. Defaults reproduce the previous behavior exactly. Also widen the Recall Quality Lab container-restart trigger from SEARCH_WEIGHT_ to SEARCH_ so sweeps of the new SEARCH_RECENCY_* vars actually restart the API container instead of silently testing the baseline. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Code-review follow-ups for the tunable recency change: - Guard SEARCH_RECENCY_WINDOW_DAYS in config.py: non-positive values now fall back to 180 via a small _positive_or_default helper (a window of 0 caused a request-time ZeroDivisionError on every recall; negative values produced unbounded scores). Unparseable values still raise, matching the neighboring float() parses. - Make scripts/browse_memories.py diagnose read SEARCH_RECENCY_WINDOW_DAYS and SEARCH_RECENCY_CURVE from env with the same validation semantics instead of hardcoding the old 180-day linear curve, and reflect the configured window/curve in the explanatory message. - De-flake test_compute_recency_score_defaults_match_legacy_behavior by monkeypatching window=180/curve=linear instead of asserting raw config values (config.py runs load_dotenv() at import, so tuned .env files leaked into the assertion). Add coverage for the non-positive window guard. - Sync stale docs: CLAUDE.md and docs/COMPARISON.md no longer claim a fixed 180-day linear decay; docs/ENVIRONMENT_VARIABLES.md weight paragraph now scopes itself to SEARCH_WEIGHT_* since the table gained non-weight rows. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s apply docker compose --env-file only affects compose-file interpolation, never container environment. Since flask-api's environment block lists no SEARCH_*/RECALL_*/CONSOLIDATION_* keys, every lab harness config override was silently dropped: all historical lab A/B runs (incl. the 2026-02-17 SEARCH_WEIGHT_RELEVANCE sweep) produced bit-identical metrics (MRR=0.7995787545787546 across baseline, candidate, and all sweep values). Wire .env.bench in via env_file with required:false so the file is optional outside lab runs. environment: still wins for keys it defines, which is fine: lab configs only use SEARCH_/RECALL_/CONSOLIDATION_ prefixes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
02f9bb9 to
8c7d737
Compare
Continuation of #185, which GitHub auto-closed (and refused to reopen) when its stacked base branch `feat/tunable-recency` was deleted on #182's merge. Identical branch and content, now rebased onto `develop` with #182 included. See #185 for the full description, validation evidence (production-corpus lab A/B that flipped the default to `SEARCH_TAG_SCORE_TOKEN_CAP=0`), and review history. Part of the develop-branch integration series for the ranking release. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…nce gate, date-aware ranking (#182, #193, #186, #187, #183, #184, #188) (#194) ## Release: ranking & recall series (develop → main)⚠️ **Merge with a MERGE COMMIT — do not squash.** release-please needs the individual conventional commits below to compute the version and changelog for PR #154. ### What's in this release | PR | Change | Default behavior | |---|---|---| | #182 | `feat(recall)`: configurable recency decay window/curve | unchanged (env-gated) | | #193 (replaces #185) | `feat(recall)`: tag-score denominator cap fixes query-length bias | unchanged (`SEARCH_TAG_SCORE_TOKEN_CAP=0`) | | #186 | `fix(recall)`: relevance gate — query-independent scoring gated on topical evidence (#130) | unchanged (gate off) | | #187 | `feat(recall)`: date-aware ranking, `recency_bias=off\|on\|auto`, latest-fact selection (#158, #159) | `RECALL_RECENCY_BIAS=off`; adds deterministic timestamp tiebreak for near-ties | | #183 | `feat(benchmarks)`: failure-mode diagnosis harness + judge quota preflight | tooling only | | #184 | `fix(mcp)`: surface stored metadata + `updated_at` in detailed recall format (#111) | additive | | #188 | `feat(enrichment)`: classification fallback-rate metrics in `/enrichment/status` | additive | Plus: CI now runs on `develop` pushes/PRs; benchmark experiment log + README contribution-policy note. ### Verification evidence - **Unit/lint/npm**: 625 pytest + 16 mcp-sse-server tests green on develop head; CI green. - **Default-preserve**: recall-lab baseline on the 10k-memory production snapshot — develop defaults vs main pooled baseline identical aggregates (R@5 0.655 / R@10 0.710 / MRR 0.434 / NDCG@10 0.501). Two-stack probe run (main vs develop, defaults): 11/12 preserve-exact, remaining diffs are near-tie reorders (top-1 score deltas ≤ 5.4e-5, the #187 timestamp tiebreak). - **Full judged 500q LongMemEval** (ship config: `RECALL_RECENCY_BIAS=auto` + `temporal-answer` harness): recall@5 96.6% (483/500), accuracy 86.0% (430/500), `judge_errors=0`, `memory_ingest_failures=0`. - **Churn attribution** (targeted re-runs of all 17 churned questions on current-main-at-defaults and develop-at-defaults): 15/17 moved with #191 (already on main) — the April canonical 97.2% floor is stale; current main measures ~97.0%. Develop-at-defaults differs from current main by **1 question in 500** (a near-tie rank-5/6 flip from #187's deterministic tiebreak). Accuracy is within answerer replicate noise (identical-config reference runs flip 28/500 answers). - Full detail: `benchmarks/EXPERIMENT_LOG.md` (2026-06-11 entry) and `benchmarks/results/lme_churn17_*` + `analyze_churn17.py`. ### Opt-in features shipped OFF `RECALL_RELEVANCE_GATE` (validated at 0.40 on lab corpus; improves negative-probe precision) and `RECALL_RECENCY_BIAS=auto` (current-state query re-ranking). Neither affects default behavior; see `docs/ENVIRONMENT_VARIABLES.md`. ### After merging release-please will update PR #154 (v0.16.0); merging *that* cuts the tag and publishes the `:stable` image — the actual user-facing deploy event for Railway template users. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
🤖 I have created a release *beep* *boop* --- ## [0.16.0](v0.15.2...v0.16.0) (2026-06-26) ### Features * **api:** add admin backup endpoint ([#162](#162)) ([8b1f264](8b1f264)) * **api:** support bulk memory associations ([1221e36](1221e36)) * **api:** support bulk memory associations ([#198](#198)) ([28eb916](28eb916)) * **benchmarks:** LongMemEval failure-mode diagnosis harness + judge quota preflight ([#183](#183)) ([f99bece](f99bece)) * **consolidation:** expose cluster threshold and min size as env vars ([#163](#163)) ([7e731f3](7e731f3)) * **enrichment:** expose classification fallback-rate metrics in /enrichment/status ([#188](#188)) ([0b522a9](0b522a9)) * **entity:** harden identity cleanup and repair tooling ([#176](#176)) ([827dfbc](827dfbc)) * **eval:** recall-quality optimization harness — lab foundation + design ([#197](#197)) ([431433e](431433e)) * **graph:** support unbounded visualizer snapshots ([#141](#141)) ([c730128](c730128)) * **lab:** add aged labelled distractor injection ([cc5d546](cc5d546)) * **lab:** add config_complexity simplicity metric ([dfb10d9](dfb10d9)) * **lab:** add distractor_rate_at_k precision guardrail metric ([872eab2](872eab2)) * **lab:** add lab_corpus with parameterized recall ([5e1e071](5e1e071)) * **lab:** add pick_winner scorecard decision rule ([3187eac](3187eac)) * **lab:** add real consolidation pass helper ([48a7d4a](48a7d4a)) * **lab:** isolate production clone restores ([#171](#171)) ([aef90c0](aef90c0)) * **lab:** wire scorecard, distractors, recall params, consolidation into runner ([589ec30](589ec30)) * **recall:** add metadata sidecar search ([#177](#177)) ([4e7956e](4e7956e)) * **recall:** add state_mode=current|history recall alias ([#173](#173)) ([b1df86c](b1df86c)) * **recall:** cap tag-score denominator to fix query-length bias ([#193](#193)) ([cefa516](cefa516)) * **recall:** date-aware ranking + latest-fact selection ([#158](#158), [#159](#159)) ([#187](#187)) ([a6ed945](a6ed945)) * **recall:** make recency decay window and curve configurable ([#182](#182)) ([dbb933f](dbb933f)) * **recall:** ranking release — recency config, tag-score cap, relevance gate, date-aware ranking ([#182](#182), [#193](#193), [#186](#186), [#187](#187), [#183](#183), [#184](#184), [#188](#188)) ([#194](#194)) ([337fe98](337fe98)) * **scripts:** safer reclassify_with_llm.py with provider flags + tighter prompt ([#164](#164)) ([a742602](a742602)) ### Bug Fixes * **api:** address copilot review on PR [#198](#198) ([0466a1e](0466a1e)) * **api:** handle grouped association write failures ([cd93df9](cd93df9)) * **backup:** make backup_automem.py runnable as `python scripts/backup_automem.py` ([#175](#175)) ([edd9742](edd9742)) * **benchmarks:** add publication verification bundle ([#166](#166)) ([420d721](420d721)) * **consolidation:** skip eager first tick at startup to avoid FalkorDB load race ([#165](#165)) ([1b812cf](1b812cf)) * **docs:** keep dispatch payload arrays stable ([df6e9e8](df6e9e8)) * **embedding:** fall back to per-item real embeddings before placeholders in batch path ([#189](#189)) ([6e9c62c](6e9c62c)) * **entity:** restore person-shape exemption on the slug validation path ([#179](#179)) ([5e29960](5e29960)) * **entity:** stop validator over-rejecting real people, code tools, and event categories ([#178](#178)) ([193b730](193b730)) * **lab:** address copilot review on PR [#197](#197) ([45f80d6](45f80d6)) * **lab:** align scorecard key contract (build_scorecard -> pick_winner) ([7d91530](7d91530)) * **mcp-sse:** decouple /health liveness from upstream readiness ([#151](#151)) ([5bcfb8b](5bcfb8b)) * **mcp:** cap association failure summary ([ea4e08f](ea4e08f)) * **mcp:** surface stored metadata and updated_at in detailed recall format ([#184](#184)) ([230416e](230416e)) * **recall:** address copilot review on PR [#194](#194) ([50b1647](50b1647)) * **recall:** canonicalize / and : separators in context_tag matching ([3afd9d3](3afd9d3)) * **recall:** canonicalize / and : separators in context_tag matching ([#203](#203)) ([ba5e9ff](ba5e9ff)) * **recall:** gate query-independent scoring on topical evidence within tag scope ([#130](#130)) ([#186](#186)) ([c11b594](c11b594)) * **recall:** hydrate semantic recall summaries ([#192](#192)) ([76e845d](76e845d)) * **recall:** normalize graph keyword scores into the 0-1 component range ([#191](#191)) ([3653ddf](3653ddf)) * **recall:** respect current memory state ([#170](#170)) ([ed36b98](ed36b98)), closes [#169](#169) [#158](#158) [#159](#159) * **scripts:** add sys.path guard to reembed_embeddings.py ([d333cf0](d333cf0)) ### Documentation * add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations ([f20c664](f20c664)) * **bench:** log full judged 500q LongMemEval ship-config run with churn attribution ([41bf8d0](41bf8d0)) * **eval:** Plan A — lab metric foundation (TDD, 9 tasks) ([0087dda](0087dda)) * **eval:** Plan B — parallel matrix harness (TDD, 9 tasks) ([c8ddfb2](c8ddfb2)) * **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as diagnostic-only ([#174](#174)) ([e8a3285](e8a3285)) * **eval:** spec for recall-quality optimization harness ([b1a1995](b1a1995)) * fix stale claims and document gated flags for 0.16.0 ([b152d64](b152d64)) * note develop-branch contribution policy in README ([ccf02dd](ccf02dd)) * **positioning:** add scout reference ([#168](#168)) ([922d23b](922d23b)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([3ff95bd](3ff95bd)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([#204](#204)) ([89c30e0](89c30e0)) * refresh README and benchmark guidance ([#157](#157)) ([bba31cc](bba31cc)) * **runtime:** align Docker viewer paths and setup guidance ([#155](#155)) ([bbda79b](bbda79b)) * scripts catalog, recall quality lab guide, and 0.16.0 migration runbook ([#199](#199)) ([f190ae5](f190ae5)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
Summary
SEARCH_RECENCY_WINDOW_DAYS(default 180) andSEARCH_RECENCY_CURVE(linear|exp, defaultlinear) replace the hardcoded 180-day linear decay in_compute_recency_score(). Defaults are bit-for-bit identical to current behavior;exptreats the window as a half-life.0would otherwise 500 every recall withZeroDivisionError; negatives produce unbounded scores).scripts/browse_memories.py diagnosenow simulates recency with the same env config instead of its own hardcoded copy.run_recall_test.pyrestarted the API container only forSEARCH_WEIGHT_*keys, so sweeps of any otherSEARCH_*var silently tested the baseline — prefix widened toSEARCH_.Why
Recency was dead for any memory older than 180 days (and for entire benchmark corpora, which are dated 2023) — zero temporal discrimination among conflicting facts. This is the substrate for the date-aware ranking work (#158/#159) and makes the window/curve sweepable via
make lab-sweep.Testing
🤖 Generated with Claude Code