Skip to content

feat(recall): make recency decay window and curve configurable - #182

Merged
jack-arturo merged 4 commits into
developfrom
feat/tunable-recency
Jun 11, 2026
Merged

jack-arturo merged 4 commits into
developfrom
feat/tunable-recency

Conversation

@jack-arturo

Copy link
Copy Markdown
Member

Summary

  • SEARCH_RECENCY_WINDOW_DAYS (default 180) and SEARCH_RECENCY_CURVE (linear|exp, default linear) replace the hardcoded 180-day linear decay in _compute_recency_score(). Defaults are bit-for-bit identical to current behavior; exp treats the window as a half-life.
  • Window values ≤ 0 fall back to 180 (a 0 would otherwise 500 every recall with ZeroDivisionError; negatives produce unbounded scores).
  • scripts/browse_memories.py diagnose now simulates recency with the same env config instead of its own hardcoded copy.
  • Lab harness fix: run_recall_test.py restarted the API container only for SEARCH_WEIGHT_* keys, so sweeps of any other SEARCH_* var silently tested the baseline — prefix widened to SEARCH_.

Why

Recency was dead for any memory older than 180 days (and for entire benchmark corpora, which are dated 2023) — zero temporal discrimination among conflicting facts. This is the substrate for the date-aware ranking work (#158/#159) and makes the window/curve sweepable via make lab-sweep.

Testing

  • 10 new unit tests (defaults, linear/exp curves, guards, monkeypatched windows, config domain validation)
  • Full suite: 497 passed, 12 skipped; black + flake8 clean

🤖 Generated with Claude Code

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR makes the recall recency scoring configurable via environment variables, replacing the previously hardcoded 180-day linear decay. This supports tuning and benchmarking recency behavior (including an exponential “half-life” mode) while keeping defaults identical to legacy behavior.

Changes:

  • Add SEARCH_RECENCY_WINDOW_DAYS and SEARCH_RECENCY_CURVE (linear / exp) to config and recency scoring.
  • Update diagnostics tooling and lab harness to respect the new SEARCH_RECENCY_* configuration.
  • Document the new environment variables and update supporting docs/examples.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
automem/utils/scoring.py Use configured recency window/curve in _compute_recency_score().
automem/config.py Add validated config for SEARCH_RECENCY_WINDOW_DAYS and SEARCH_RECENCY_CURVE.
tests/test_api_endpoints.py Add unit tests for recency scoring behavior and config guard.
scripts/browse_memories.py Make diagnose recency simulation configurable via SEARCH_RECENCY_*.
scripts/lab/run_recall_test.py Restart API container for any SEARCH_* changes during sweeps.
docs/ENVIRONMENT_VARIABLES.md Document SEARCH_RECENCY_WINDOW_DAYS / SEARCH_RECENCY_CURVE and clarify weights text.
docs/COMPARISON.md Update recency description to reflect configurability.
CLAUDE.md Update env var example text to reference SEARCH_RECENCY_*.

Comment on lines +763 to +765
recency_window = float(os.getenv("SEARCH_RECENCY_WINDOW_DAYS", "180"))
if recency_window <= 0:
recency_window = 180.0

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed by validating math.isfinite(recency_window) before using the diagnose recency window — scripts/browse_memories.py:764. GitHub marked the original line outdated after the push.

Comment on lines +429 to +431
def _timestamp_days_ago(days: float) -> str:
return (datetime.now(timezone.utc) - timedelta(days=days)).isoformat()

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pinned recency-score tests to the legacy 180-day linear config via legacy_recency_config, so local .env overrides no longer affect them — tests/test_api_endpoints.py:433. GitHub marked the original line outdated after the push.

jack-arturo added a commit that referenced this pull request Jun 11, 2026
…nge (#191)

Fixes #190.

## Problem

`_graph_keyword_search` returned the **raw Cypher additive score** — +2
per keyword contained in content, +1 per keyword in any tag, summed over
all extracted keywords, plus a +2/+1 whole-phrase bonus — so a K-keyword
query can score up to **3K+3** while every other channel (vector cosine,
metadata, trending importance) lives in 0–1.

Observed during the 2026-06-11 production forensics: a tag-scoped
exact-content match returned `keyword=11.0`, `final_score=4.03`.
Consequences:

- `SEARCH_WEIGHT_KEYWORD (0.35) × 11 = 3.85` — a keyword hit trumps any
vector/metadata/importance combination, scaling with *query length*
rather than match quality.
- Defeats `RECALL_RELEVANCE_GATE` semantics (PR #186): `evidence =
max(vector, keyword, metadata, exact)` assumes 0–1 components; `evidence
= 11` sails past any gate.

## Fix

1. **Producer**: normalize the raw score by its per-query maximum
(`3·len(keywords) + 3` when a phrase is present; `3` in the phrase-only
branch) before it leaves `_graph_keyword_search`. This is a monotone
per-query transform — within-channel ordering (and the Cypher `ORDER
BY`) is unchanged; only cross-channel blending changes, which is the
point.
2. **Consumer**: defensively clamp the keyword component to `min(1.0,
…)` in `_compute_metadata_score`, so no future producer can break the
0–1 contract or the gate again.

## Verification

- 4 new tests in `tests/test_keyword_score_normalization.py` (TDD'd
against the bug, including the literal `keyword=11.0` repro). Full
suite: **503 passed**.
- **Production-corpus lab A/B** (10,142-memory snapshot, 200 queries, vs
the pooled 3-run parity baseline from the 2026-06-11 release sweep):
- Recall@5 −0.2pp, Recall@10 −0.7pp, MRR −0.008, NDCG@10 −0.007 (all
within baseline run-to-run variance; paired t-test p=0.32)
  - Per-query: **196/197 unchanged**, 0 improved, 1 regressed
- The single flip is the intended behavior change made visible: the
expected memory held rank 1 *only* via the inflated keyword score (ranks
2–5 identical before/after). It's in the fallback-typed `Memory` cohort
(MRR 0.15 baseline — the known data-quality cohort from #188's
classification incident).

## Notes

- This lands on main independently of the #182–#187 chain; #186's gate
evidence check is the main beneficiary once the chain rebases over it.
- Trending (`importance`), metadata (capped), and vector (cosine)
channels were verified already bounded; the graph keyword channel was
the only unbounded producer.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@jack-arturo
jack-arturo force-pushed the feat/tunable-recency branch from 5388e65 to 02f9bb9 Compare June 11, 2026 18:46
@jack-arturo
jack-arturo changed the base branch from main to develop June 11, 2026 18:46
jack-arturo and others added 4 commits June 11, 2026 20:47
Add SEARCH_RECENCY_WINDOW_DAYS (default 180) and SEARCH_RECENCY_CURVE
(linear|exp, default linear) so the recency score's decay shape can be
tuned via environment instead of the hardcoded 180-day linear decay.
Defaults reproduce the previous behavior exactly.

Also widen the Recall Quality Lab container-restart trigger from
SEARCH_WEIGHT_ to SEARCH_ so sweeps of the new SEARCH_RECENCY_* vars
actually restart the API container instead of silently testing the
baseline.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Code-review follow-ups for the tunable recency change:

- Guard SEARCH_RECENCY_WINDOW_DAYS in config.py: non-positive values now
  fall back to 180 via a small _positive_or_default helper (a window of 0
  caused a request-time ZeroDivisionError on every recall; negative values
  produced unbounded scores). Unparseable values still raise, matching the
  neighboring float() parses.
- Make scripts/browse_memories.py diagnose read SEARCH_RECENCY_WINDOW_DAYS
  and SEARCH_RECENCY_CURVE from env with the same validation semantics
  instead of hardcoding the old 180-day linear curve, and reflect the
  configured window/curve in the explanatory message.
- De-flake test_compute_recency_score_defaults_match_legacy_behavior by
  monkeypatching window=180/curve=linear instead of asserting raw config
  values (config.py runs load_dotenv() at import, so tuned .env files leaked
  into the assertion). Add coverage for the non-positive window guard.
- Sync stale docs: CLAUDE.md and docs/COMPARISON.md no longer claim a fixed
  180-day linear decay; docs/ENVIRONMENT_VARIABLES.md weight paragraph now
  scopes itself to SEARCH_WEIGHT_* since the table gained non-weight rows.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…s apply

docker compose --env-file only affects compose-file interpolation, never
container environment. Since flask-api's environment block lists no
SEARCH_*/RECALL_*/CONSOLIDATION_* keys, every lab harness config override
was silently dropped: all historical lab A/B runs (incl. the 2026-02-17
SEARCH_WEIGHT_RELEVANCE sweep) produced bit-identical metrics
(MRR=0.7995787545787546 across baseline, candidate, and all sweep values).

Wire .env.bench in via env_file with required:false so the file is
optional outside lab runs. environment: still wins for keys it defines,
which is fine: lab configs only use SEARCH_/RECALL_/CONSOLIDATION_ prefixes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@jack-arturo
jack-arturo force-pushed the feat/tunable-recency branch from 02f9bb9 to 8c7d737 Compare June 11, 2026 18:48
@jack-arturo
jack-arturo merged commit dbb933f into develop Jun 11, 2026
7 checks passed
@jack-arturo
jack-arturo deleted the feat/tunable-recency branch June 11, 2026 18:53
jack-arturo added a commit that referenced this pull request Jun 11, 2026
Continuation of #185, which GitHub auto-closed (and refused to reopen)
when its stacked base branch `feat/tunable-recency` was deleted on
#182's merge. Identical branch and content, now rebased onto `develop`
with #182 included.

See #185 for the full description, validation evidence
(production-corpus lab A/B that flipped the default to
`SEARCH_TAG_SCORE_TOKEN_CAP=0`), and review history.

Part of the develop-branch integration series for the ranking release.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
jack-arturo added a commit that referenced this pull request Jun 12, 2026
…nce gate, date-aware ranking (#182, #193, #186, #187, #183, #184, #188) (#194)

## Release: ranking & recall series (develop → main)

⚠️ **Merge with a MERGE COMMIT — do not squash.** release-please needs
the individual conventional commits below to compute the version and
changelog for PR #154.

### What's in this release

| PR | Change | Default behavior |
|---|---|---|
| #182 | `feat(recall)`: configurable recency decay window/curve |
unchanged (env-gated) |
| #193 (replaces #185) | `feat(recall)`: tag-score denominator cap fixes
query-length bias | unchanged (`SEARCH_TAG_SCORE_TOKEN_CAP=0`) |
| #186 | `fix(recall)`: relevance gate — query-independent scoring gated
on topical evidence (#130) | unchanged (gate off) |
| #187 | `feat(recall)`: date-aware ranking,
`recency_bias=off\|on\|auto`, latest-fact selection (#158, #159) |
`RECALL_RECENCY_BIAS=off`; adds deterministic timestamp tiebreak for
near-ties |
| #183 | `feat(benchmarks)`: failure-mode diagnosis harness + judge
quota preflight | tooling only |
| #184 | `fix(mcp)`: surface stored metadata + `updated_at` in detailed
recall format (#111) | additive |
| #188 | `feat(enrichment)`: classification fallback-rate metrics in
`/enrichment/status` | additive |

Plus: CI now runs on `develop` pushes/PRs; benchmark experiment log +
README contribution-policy note.

### Verification evidence

- **Unit/lint/npm**: 625 pytest + 16 mcp-sse-server tests green on
develop head; CI green.
- **Default-preserve**: recall-lab baseline on the 10k-memory production
snapshot — develop defaults vs main pooled baseline identical aggregates
(R@5 0.655 / R@10 0.710 / MRR 0.434 / NDCG@10 0.501). Two-stack probe
run (main vs develop, defaults): 11/12 preserve-exact, remaining diffs
are near-tie reorders (top-1 score deltas ≤ 5.4e-5, the #187 timestamp
tiebreak).
- **Full judged 500q LongMemEval** (ship config:
`RECALL_RECENCY_BIAS=auto` + `temporal-answer` harness): recall@5 96.6%
(483/500), accuracy 86.0% (430/500), `judge_errors=0`,
`memory_ingest_failures=0`.
- **Churn attribution** (targeted re-runs of all 17 churned questions on
current-main-at-defaults and develop-at-defaults): 15/17 moved with #191
(already on main) — the April canonical 97.2% floor is stale; current
main measures ~97.0%. Develop-at-defaults differs from current main by
**1 question in 500** (a near-tie rank-5/6 flip from #187's
deterministic tiebreak). Accuracy is within answerer replicate noise
(identical-config reference runs flip 28/500 answers).
- Full detail: `benchmarks/EXPERIMENT_LOG.md` (2026-06-11 entry) and
`benchmarks/results/lme_churn17_*` + `analyze_churn17.py`.

### Opt-in features shipped OFF

`RECALL_RELEVANCE_GATE` (validated at 0.40 on lab corpus; improves
negative-probe precision) and `RECALL_RECENCY_BIAS=auto` (current-state
query re-ranking). Neither affects default behavior; see
`docs/ENVIRONMENT_VARIABLES.md`.

### After merging

release-please will update PR #154 (v0.16.0); merging *that* cuts the
tag and publishes the `:stable` image — the actual user-facing deploy
event for Railway template users.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
jack-arturo added a commit that referenced this pull request Jun 26, 2026
🤖 I have created a release *beep* *boop*
---


##
[0.16.0](v0.15.2...v0.16.0)
(2026-06-26)


### Features

* **api:** add admin backup endpoint
([#162](#162))
([8b1f264](8b1f264))
* **api:** support bulk memory associations
([1221e36](1221e36))
* **api:** support bulk memory associations
([#198](#198))
([28eb916](28eb916))
* **benchmarks:** LongMemEval failure-mode diagnosis harness + judge
quota preflight
([#183](#183))
([f99bece](f99bece))
* **consolidation:** expose cluster threshold and min size as env vars
([#163](#163))
([7e731f3](7e731f3))
* **enrichment:** expose classification fallback-rate metrics in
/enrichment/status
([#188](#188))
([0b522a9](0b522a9))
* **entity:** harden identity cleanup and repair tooling
([#176](#176))
([827dfbc](827dfbc))
* **eval:** recall-quality optimization harness — lab foundation +
design ([#197](#197))
([431433e](431433e))
* **graph:** support unbounded visualizer snapshots
([#141](#141))
([c730128](c730128))
* **lab:** add aged labelled distractor injection
([cc5d546](cc5d546))
* **lab:** add config_complexity simplicity metric
([dfb10d9](dfb10d9))
* **lab:** add distractor_rate_at_k precision guardrail metric
([872eab2](872eab2))
* **lab:** add lab_corpus with parameterized recall
([5e1e071](5e1e071))
* **lab:** add pick_winner scorecard decision rule
([3187eac](3187eac))
* **lab:** add real consolidation pass helper
([48a7d4a](48a7d4a))
* **lab:** isolate production clone restores
([#171](#171))
([aef90c0](aef90c0))
* **lab:** wire scorecard, distractors, recall params, consolidation
into runner
([589ec30](589ec30))
* **recall:** add metadata sidecar search
([#177](#177))
([4e7956e](4e7956e))
* **recall:** add state_mode=current|history recall alias
([#173](#173))
([b1df86c](b1df86c))
* **recall:** cap tag-score denominator to fix query-length bias
([#193](#193))
([cefa516](cefa516))
* **recall:** date-aware ranking + latest-fact selection
([#158](#158),
[#159](#159))
([#187](#187))
([a6ed945](a6ed945))
* **recall:** make recency decay window and curve configurable
([#182](#182))
([dbb933f](dbb933f))
* **recall:** ranking release — recency config, tag-score cap, relevance
gate, date-aware ranking
([#182](#182),
[#193](#193),
[#186](#186),
[#187](#187),
[#183](#183),
[#184](#184),
[#188](#188))
([#194](#194))
([337fe98](337fe98))
* **scripts:** safer reclassify_with_llm.py with provider flags +
tighter prompt
([#164](#164))
([a742602](a742602))


### Bug Fixes

* **api:** address copilot review on PR
[#198](#198)
([0466a1e](0466a1e))
* **api:** handle grouped association write failures
([cd93df9](cd93df9))
* **backup:** make backup_automem.py runnable as `python
scripts/backup_automem.py`
([#175](#175))
([edd9742](edd9742))
* **benchmarks:** add publication verification bundle
([#166](#166))
([420d721](420d721))
* **consolidation:** skip eager first tick at startup to avoid FalkorDB
load race
([#165](#165))
([1b812cf](1b812cf))
* **docs:** keep dispatch payload arrays stable
([df6e9e8](df6e9e8))
* **embedding:** fall back to per-item real embeddings before
placeholders in batch path
([#189](#189))
([6e9c62c](6e9c62c))
* **entity:** restore person-shape exemption on the slug validation path
([#179](#179))
([5e29960](5e29960))
* **entity:** stop validator over-rejecting real people, code tools, and
event categories
([#178](#178))
([193b730](193b730))
* **lab:** address copilot review on PR
[#197](#197)
([45f80d6](45f80d6))
* **lab:** align scorecard key contract (build_scorecard -&gt;
pick_winner)
([7d91530](7d91530))
* **mcp-sse:** decouple /health liveness from upstream readiness
([#151](#151))
([5bcfb8b](5bcfb8b))
* **mcp:** cap association failure summary
([ea4e08f](ea4e08f))
* **mcp:** surface stored metadata and updated_at in detailed recall
format ([#184](#184))
([230416e](230416e))
* **recall:** address copilot review on PR
[#194](#194)
([50b1647](50b1647))
* **recall:** canonicalize / and : separators in context_tag matching
([3afd9d3](3afd9d3))
* **recall:** canonicalize / and : separators in context_tag matching
([#203](#203))
([ba5e9ff](ba5e9ff))
* **recall:** gate query-independent scoring on topical evidence within
tag scope
([#130](#130))
([#186](#186))
([c11b594](c11b594))
* **recall:** hydrate semantic recall summaries
([#192](#192))
([76e845d](76e845d))
* **recall:** normalize graph keyword scores into the 0-1 component
range ([#191](#191))
([3653ddf](3653ddf))
* **recall:** respect current memory state
([#170](#170))
([ed36b98](ed36b98)),
closes [#169](#169)
[#158](#158)
[#159](#159)
* **scripts:** add sys.path guard to reembed_embeddings.py
([d333cf0](d333cf0))


### Documentation

* add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations
([f20c664](f20c664))
* **bench:** log full judged 500q LongMemEval ship-config run with churn
attribution
([41bf8d0](41bf8d0))
* **eval:** Plan A — lab metric foundation (TDD, 9 tasks)
([0087dda](0087dda))
* **eval:** Plan B — parallel matrix harness (TDD, 9 tasks)
([c8ddfb2](c8ddfb2))
* **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as
diagnostic-only
([#174](#174))
([e8a3285](e8a3285))
* **eval:** spec for recall-quality optimization harness
([b1a1995](b1a1995))
* fix stale claims and document gated flags for 0.16.0
([b152d64](b152d64))
* note develop-branch contribution policy in README
([ccf02dd](ccf02dd))
* **positioning:** add scout reference
([#168](#168))
([922d23b](922d23b))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([3ff95bd](3ff95bd))
* refresh benchmark currency for the neutral AMB run and prune stale
archive docs
([#204](#204))
([89c30e0](89c30e0))
* refresh README and benchmark guidance
([#157](#157))
([bba31cc](bba31cc))
* **runtime:** align Docker viewer paths and setup guidance
([#155](#155))
([bbda79b](bbda79b))
* scripts catalog, recall quality lab guide, and 0.16.0 migration
runbook ([#199](#199))
([f190ae5](f190ae5))

---
This PR was generated with [Release
Please](https://github.com/googleapis/release-please). See
[documentation](https://github.com/googleapis/release-please#release-please).
@jack-arturo jack-arturo mentioned this pull request Aug 28, 2026
1 task
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants