feat(consolidation): expose cluster threshold and min size as env vars - #163
Merged
jack-arturo merged 5 commits intoJun 6, 2026
Conversation
Adds two new env-var-backed knobs and tunes one default: - CONSOLIDATION_CLUSTER_SIMILARITY_THRESHOLD (default 0.65) - CONSOLIDATION_MIN_CLUSTER_SIZE (default 2) - CONSOLIDATION_CLUSTER_INTERVAL_SECONDS default: 2592000 (30d) -> 604800 (7d) Why --- similarity_threshold and min_cluster_size were hardcoded in MemoryConsolidator.__init__ at 0.75 and 3 respectively, with no configuration path. Embedding geometry is deployment-dependent (Voyage-4 1024d, OpenAI text-embedding-3-small Matryoshka-1024d, Ollama nomic 768d, FastEmbed bge 768d) and a single hardcoded threshold cannot suit all of them. The pattern of exposing similarity thresholds via env vars is already established for enrichment (ENRICHMENT_SIMILARITY_THRESHOLD). Default 0.65 (down from 0.75): at 1024d, 0.75 sits in near-duplicate territory; on a 62.7k-memory corpus with the 0.75 threshold the cluster tick produced zero clusters. 0.65 puts the threshold in the "topically related, not literally the same thing" band that clustering is meant to capture. Operators on tighter cosine distributions (e.g. text-embedding-3-large at 3072d) can raise it. Default min_cluster_size 2 (down from 3): with min=3, a strongly- related pair of memories never clusters. Connected-components on sparse semantic graphs typically produces pairs first; min=2 keeps those in scope. Operators on very large corpora can raise it. Default interval 7d (down from 30d): monthly clustering on a live, growing corpus means the graph is always up to a month stale. 7d matches the existing creative-consolidation cadence. The breaking default change is the one behavior shift in this PR; deployments wanting the old behavior set CONSOLIDATION_CLUSTER_INTERVAL_SECONDS=2592000. Wiring ------ - automem/config.py: add the two new env vars; flip the interval default. - automem/consolidation/runtime_helpers.py: build_consolidator_from_config accepts the two new params and sets them as attributes post-construction (the MemoryConsolidator constructor does not accept them; this avoids changing the public class signature). - automem/consolidation/runtime_bindings.py: forward through create_consolidation_runtime. - app.py: import and pass the new config values. All other behavior is purely additive — no env override required to preserve prior behavior except the interval. Tests ----- - tests/test_consolidation_engine.py: 7/7 pass. - Full unit suite: clean except two pre-existing tests/test_content_size.py failures (auth setup, unrelated).
4 tasks
Member
|
Scope update pushed in a551221: this PR now preserves the existing defaults and only exposes the config surface. Defaults remain similarity_threshold=0.75, min_cluster_size=3, and CONSOLIDATION_CLUSTER_INTERVAL_SECONDS=2592000. I updated the PR body with the measurement note: local diagnostics make lower thresholds / faster cadence plausible for some deployments, especially high-ingest graphs, but default retuning should be a separate measured follow-up. |
jack-arturo
pushed a commit
that referenced
this pull request
May 1, 2026
…B load race (#165) ## Why When `init_consolidation_scheduler()` runs a tick **immediately** after spawning the worker thread, FalkorDB can still be loading its RDB snapshot from disk. Every Redis command during that window returns: > `LOADING Redis is loading the dataset in memory` The eager tick catches the error, logs it, and bumps `last_run` timestamps — silently skipping the day's decay / creative / cluster work until tomorrow. The bigger the corpus, the longer the RDB load, the more reliably this fires. On any restart-on-deploy host (Railway, Docker, systemd) with a few thousand memories, it hits every deploy. ## What changes One line in `automem/consolidation/runtime_scheduler.py:100` — drop the eager `run_consolidation_tick_fn()` call after starting the worker thread, and add a comment explaining why. ```diff state.consolidation_thread.start() - run_consolidation_tick_fn() + # Skip eager first tick: FalkorDB may still be loading its RDB snapshot at + # startup and the "Redis is loading the dataset in memory" error poisons + # the day's decay/creative run. The worker loop will fire its first tick + # after consolidation_tick_seconds, which is plenty of warm-up time. logger.info("Consolidation scheduler initialized") ``` ## Why this is safe - The worker loop still fires within `CONSOLIDATION_TICK_SECONDS` (default 3600s = 1h). For decay/creative/cluster intervals measured in days, a one-tick startup delay is invisible. - The scheduler is timestamp-driven (`last_run` per task), not edge-triggered. Missed intervals get picked up by the next loop iteration — nothing is "lost" by deferring. - Failure mode flips from "silent broken run" to "no run yet, will run shortly" — strictly better. ## Out of scope - A more involved fix would actively probe FalkorDB readiness with retries before the first tick. That's a bigger change and arguably belongs at the FalkorDB-client layer, not here. This PR is the minimal, low-risk fix. - The `discover_creative_associations` / clustering improvements live in #163 and #164. ## Test plan - [ ] Service starts cleanly with no eager tick log entry - [ ] Worker loop fires its first tick after `CONSOLIDATION_TICK_SECONDS` - [ ] Forcing a tick via `POST /consolidate` still works immediately - [ ] On a restart with a large RDB, no `LOADING Redis is loading the dataset in memory` errors appear in consolidation logs Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
jack-arturo
added a commit
that referenced
this pull request
Jun 26, 2026
🤖 I have created a release *beep* *boop* --- ## [0.16.0](v0.15.2...v0.16.0) (2026-06-26) ### Features * **api:** add admin backup endpoint ([#162](#162)) ([8b1f264](8b1f264)) * **api:** support bulk memory associations ([1221e36](1221e36)) * **api:** support bulk memory associations ([#198](#198)) ([28eb916](28eb916)) * **benchmarks:** LongMemEval failure-mode diagnosis harness + judge quota preflight ([#183](#183)) ([f99bece](f99bece)) * **consolidation:** expose cluster threshold and min size as env vars ([#163](#163)) ([7e731f3](7e731f3)) * **enrichment:** expose classification fallback-rate metrics in /enrichment/status ([#188](#188)) ([0b522a9](0b522a9)) * **entity:** harden identity cleanup and repair tooling ([#176](#176)) ([827dfbc](827dfbc)) * **eval:** recall-quality optimization harness — lab foundation + design ([#197](#197)) ([431433e](431433e)) * **graph:** support unbounded visualizer snapshots ([#141](#141)) ([c730128](c730128)) * **lab:** add aged labelled distractor injection ([cc5d546](cc5d546)) * **lab:** add config_complexity simplicity metric ([dfb10d9](dfb10d9)) * **lab:** add distractor_rate_at_k precision guardrail metric ([872eab2](872eab2)) * **lab:** add lab_corpus with parameterized recall ([5e1e071](5e1e071)) * **lab:** add pick_winner scorecard decision rule ([3187eac](3187eac)) * **lab:** add real consolidation pass helper ([48a7d4a](48a7d4a)) * **lab:** isolate production clone restores ([#171](#171)) ([aef90c0](aef90c0)) * **lab:** wire scorecard, distractors, recall params, consolidation into runner ([589ec30](589ec30)) * **recall:** add metadata sidecar search ([#177](#177)) ([4e7956e](4e7956e)) * **recall:** add state_mode=current|history recall alias ([#173](#173)) ([b1df86c](b1df86c)) * **recall:** cap tag-score denominator to fix query-length bias ([#193](#193)) ([cefa516](cefa516)) * **recall:** date-aware ranking + latest-fact selection ([#158](#158), [#159](#159)) ([#187](#187)) ([a6ed945](a6ed945)) * **recall:** make recency decay window and curve configurable ([#182](#182)) ([dbb933f](dbb933f)) * **recall:** ranking release — recency config, tag-score cap, relevance gate, date-aware ranking ([#182](#182), [#193](#193), [#186](#186), [#187](#187), [#183](#183), [#184](#184), [#188](#188)) ([#194](#194)) ([337fe98](337fe98)) * **scripts:** safer reclassify_with_llm.py with provider flags + tighter prompt ([#164](#164)) ([a742602](a742602)) ### Bug Fixes * **api:** address copilot review on PR [#198](#198) ([0466a1e](0466a1e)) * **api:** handle grouped association write failures ([cd93df9](cd93df9)) * **backup:** make backup_automem.py runnable as `python scripts/backup_automem.py` ([#175](#175)) ([edd9742](edd9742)) * **benchmarks:** add publication verification bundle ([#166](#166)) ([420d721](420d721)) * **consolidation:** skip eager first tick at startup to avoid FalkorDB load race ([#165](#165)) ([1b812cf](1b812cf)) * **docs:** keep dispatch payload arrays stable ([df6e9e8](df6e9e8)) * **embedding:** fall back to per-item real embeddings before placeholders in batch path ([#189](#189)) ([6e9c62c](6e9c62c)) * **entity:** restore person-shape exemption on the slug validation path ([#179](#179)) ([5e29960](5e29960)) * **entity:** stop validator over-rejecting real people, code tools, and event categories ([#178](#178)) ([193b730](193b730)) * **lab:** address copilot review on PR [#197](#197) ([45f80d6](45f80d6)) * **lab:** align scorecard key contract (build_scorecard -> pick_winner) ([7d91530](7d91530)) * **mcp-sse:** decouple /health liveness from upstream readiness ([#151](#151)) ([5bcfb8b](5bcfb8b)) * **mcp:** cap association failure summary ([ea4e08f](ea4e08f)) * **mcp:** surface stored metadata and updated_at in detailed recall format ([#184](#184)) ([230416e](230416e)) * **recall:** address copilot review on PR [#194](#194) ([50b1647](50b1647)) * **recall:** canonicalize / and : separators in context_tag matching ([3afd9d3](3afd9d3)) * **recall:** canonicalize / and : separators in context_tag matching ([#203](#203)) ([ba5e9ff](ba5e9ff)) * **recall:** gate query-independent scoring on topical evidence within tag scope ([#130](#130)) ([#186](#186)) ([c11b594](c11b594)) * **recall:** hydrate semantic recall summaries ([#192](#192)) ([76e845d](76e845d)) * **recall:** normalize graph keyword scores into the 0-1 component range ([#191](#191)) ([3653ddf](3653ddf)) * **recall:** respect current memory state ([#170](#170)) ([ed36b98](ed36b98)), closes [#169](#169) [#158](#158) [#159](#159) * **scripts:** add sys.path guard to reembed_embeddings.py ([d333cf0](d333cf0)) ### Documentation * add scripts catalog, recall-quality-lab guide, and 0.16.0 migrations ([f20c664](f20c664)) * **bench:** log full judged 500q LongMemEval ship-config run with churn attribution ([41bf8d0](41bf8d0)) * **eval:** Plan A — lab metric foundation (TDD, 9 tasks) ([0087dda](0087dda)) * **eval:** Plan B — parallel matrix harness (TDD, 9 tasks) ([c8ddfb2](c8ddfb2)) * **evals:** mark Memora/FAMA/WRIT lifecycle diagnostics as diagnostic-only ([#174](#174)) ([e8a3285](e8a3285)) * **eval:** spec for recall-quality optimization harness ([b1a1995](b1a1995)) * fix stale claims and document gated flags for 0.16.0 ([b152d64](b152d64)) * note develop-branch contribution policy in README ([ccf02dd](ccf02dd)) * **positioning:** add scout reference ([#168](#168)) ([922d23b](922d23b)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([3ff95bd](3ff95bd)) * refresh benchmark currency for the neutral AMB run and prune stale archive docs ([#204](#204)) ([89c30e0](89c30e0)) * refresh README and benchmark guidance ([#157](#157)) ([bba31cc](bba31cc)) * **runtime:** align Docker viewer paths and setup guidance ([#155](#155)) ([bbda79b](bbda79b)) * scripts catalog, recall quality lab guide, and 0.16.0 migration runbook ([#199](#199)) ([f190ae5](f190ae5)) --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds configuration surface for cluster consolidation tuning while preserving current runtime defaults.
CONSOLIDATION_CLUSTER_SIMILARITY_THRESHOLDis now configurable; default remains0.75.CONSOLIDATION_MIN_CLUSTER_SIZEis now configurable; default remains3.CONSOLIDATION_CLUSTER_INTERVAL_SECONDSremains the existing env var with its existing2592000/ 30-day default.Scope update
This PR is intentionally scoped down to config exposure only. It no longer changes the default threshold, minimum cluster size, or cluster cadence.
The motivation for exposing these knobs still holds: embedding geometry and corpus shape vary by deployment, so hardcoded clustering parameters are hard to operate. But retuning the defaults should happen after measurement, not as part of this mechanical config-surface change.
Local graph diagnostics on May 1, 2026 suggest that lowering the threshold may be useful for some corpora, and Flint's 50k+ memory growth rate makes the 30-day cadence feel stale operationally. Those are real signals, but they are not yet enough to bless new global defaults because exact
cluster_similar_memories()is expensive at graph scale and nearest-neighbor samples need quality review before changing merge behavior for everyone.What changed
automem/config.pyadds the two cluster tuning env vars with behavior-preserving defaults.automem/consolidation/runtime_helpers.pyapplies optional overrides toMemoryConsolidatorafter construction.automem/consolidation/runtime_bindings.pyforwards the optional values through the runtime factory.app.pyimports and passes the new config values.tests/test_consolidation_engine.pycovers both default preservation and explicit override behavior.Out of scope
Follow-up PRs should handle runtime behavior separately:
PARALLEL_CONTEXTnormalization;Test plan
.venv/bin/python -m pytest tests/test_consolidation_engine.py— 9 passed.venv/bin/python -m compileall automem tests/test_consolidation_engine.py