Skip to content

P3-followup locomo smoke benchmark - multiple-scenario recall quality - #12

Open
Yanru-cafe wants to merge 296 commits into
chinesewebman:mainfrom
Yanru-cafe:p3-locomo-followup
Open

Yanru-cafe wants to merge 296 commits into
chinesewebman:mainfrom
Yanru-cafe:p3-locomo-followup

Conversation

@Yanru-cafe

Copy link
Copy Markdown
Contributor

Follow-up to chinesewebman/mnelo PR 11 (P3-benchmarks). PR 11 added the python -m benchmarks latency harness; this PR adds a second subcommand for end-to-end recall quality smoke testing.

Changes:

  • benchmarks/locomo.py (new): 3 built-in scenarios (solar install / Fed rates / BYD sales), each 3 chunks plus 3 queries. Measures coverage per scenario plus latency p50/mean.
  • tests/test_locomo_p3_followup.py (new): 7 tests covering dispatcher recognition, smoke run with JSON output, cleanup idempotency, and import surface.
  • benchmarks/main.py: add locomo to _SUBCOMMANDS, update help text.
  • docs/BENCHMARKS.md: add LoCoMo smoke section with reproduce command.

Verification:

  • 22/23 tests pass (14 upstream PR 11 + 7 new + 1 isolated idempotent flake)
  • ruff lint clean
  • no new dependencies
  • default dispatcher behavior unchanged

cc: bigbox PM 小默

hermes and others added 30 commits August 4, 2026 20:20
H-1 commit d98cd93 落地后, hermes 实战评审找到 2 个真坑 (3 个里挑 2):

1. [真坑] schema.sql INSERT meta 块缺 l2_audit_log_ready flag
   - H0 落地时, 主人没法 query 区分 'audit_log 表已建 (H-1 跑了)' vs 'H-1 没跑'
   - 修: 加 ('l2_audit_log_ready', '1') + ('l2_h1_migrated', datetime('now','localtime'))
   - 实战 8/4: H-1 跑了但 meta 缺 flag, L2 H0 启用时 query 不到信号

2. [注释级] audit_log UNIQUE 缺 created_at 字段
   - 极小概率: 同 microsecond 同 run_id 同 ref 同 status 撞
   - 实战 8/4 评估: 可接受 (race condition 需要 L2 病态重入, 主人 §5.6 护栏会拦)
   - 修: 在 _migrate_schema() 加注释标注 + 未来如要更严给 SQL 模板

不修:
- 问题 1 (PRAGMA table_info entities 重复 3 次): low ROI, 几毫秒, 不动

3 个文件改动 = 2 个 (schema.sql + memory.py, 没动 test)
12 tests + 11 tests 全过 (0 回归)
解决实战 4344/4344 chunks 100% fact 问题 (v0.3 §2):
- classify.py: classify_memory_type() — 规则分类, 强标记才标, 模糊保持 fact (宁缺毋滥)
- 繁→简归一化: bounded _T2S 字符映射 dict + 单一简体标记集 (非双份标记表,
  零依赖 opencc); 标记集数据化便于扩展
- 六类标记表: CN(归一化后) + EN; episode 用时间+动作复合规则;
  优先级 decision>episode; 第三人称'用户喜欢'不标 preference
- 写路径集成: 显式类型 > 规则分类 > fact; mcp_server 默认改 None
- 回填脚本 backfill_memory_type.py (存量 4344 升级)
- 测试矩阵: 简/繁/EN 三语 8 用例 + 优先级/第三人称边界
H-1 审计 fix #3421d99 只改了 schema.sql (fresh install 路径).
但实战 mnelo 走的是 _migrate_schema() (存量库路径), 不会跑 schema.sql.
所以 l2_audit_log_ready flag 没进 meta 表.

修: _migrate_schema() 末尾加 2 行 INSERT OR REPLACE meta
(l2_audit_log_ready='1' + l2_h1_migrated=now()).
跟 schema.sql INSERT meta 块保持一致, 2 条路径 (fresh install + 存量迁移) 都加 flag.

幂等: 跑第二次 _migrate_schema 不会重复 (existing_flag 检查).
[8/4 v0.3 报告 §2 fix] 实战 4344/4344 chunks 100% fact — 6 类系统空架子.
按 TASKS_L2_EXTRACT §4 执行顺序, Batch 1 (E1-E3 分类器核心 + 单测):

1. classify.py 新模块 (~270 行, 跟 search_index.py 同层):
   - [E1 §1.2] _T2S 字符归一化 (繁→简, bounded dict 80+ 字, 零依赖)
     包含 [8/4 fix] 测试矩阵失败补全 (错/决/建仓/减仓/记录/步骤/选择 等)
   - [E2 §2] _MARKERS 标记表 (5 类 P1a: preference/decision/episode/procedure/ephemeral)
     preference: 带"我"主语才算强标记 (防误标他人偏好)
     episode: 复合规则 (cn_time AND cn_action, TASKS §2)
     优先级: decision > episode > 其余 (TASKS §2 优先冲突表)
   - [E3 §3] classify_memory_type(text) -> Optional[str]
     - 繁→简 归一化 + 双语匹配
     - 强标记命中 → 类型; 无命中 / 弱标记 → None (调用方默认 fact)

2. tests/test_classify.py (37 tests, 主人 §5.2 双语+繁简矩阵 8 场景全覆盖):
   - TestNormalize (E1): _T2S 繁简 + 空串 + 纯英文 + trading domain
   - TestMarkersSchema (E2): 5 类型 + 双语 + episode 复合 + trading 动作覆盖
   - TestClassifyMatrix (E3): 8 主人指定场景 (preference/decision/episode/procedure/ephemeral × 3 语言)
   - TestClassifyPriorityAndEdge: decision>episode 优先冲突 + 弱标记(None) + episode 时间/动作单缺
   - TestDeterminism: 同一文本两次分类结果一致 + 简繁归一化后一致

设计原则 (§5.5 宁缺毋滥): 强标记才分类; 模糊/无标记 → 保持 fact (留给 P1b LLM)
[P1a = 零 LLM] 写路径 P1b LLM 语义分类是后续阶段 (TASKS §0)

实战验证 (8/4 跑):
- 37 classify tests 全过
- 12 H-1 + 11 memory_type tests 全过 (0 回归)
- 0 业务逻辑修改, H-1 schema 不受影响
- 4bd654d run_purge_worker + H-1 schema 3 项不动

报告: TASKS_L2_EXTRACT.md v0 + v0.3 实战数据报告 §2
报告位置: /Users/apple/study/predict/research/2026-08-04-mnelo-schema-drift/mnelo_schema_audit.md
按 TASKS_L2_EXTRACT §E4 实施 (Batch 2):

1. memory.py remember() 集成:
   - memory_type 参数从 str='fact' 改 Optional[str]=None
   - 默认 None 触发 classify.classify_memory_type(content) 自动分类
   - 显式传值 (>None, 包括 'fact') 永远尊重, 不被规则覆盖
   - logger.info 记录 [P1a] auto-classified: chunk_id -> type

2. mcp_server.py memory_remember schema:
   - memory_type default 从 'fact' 改 None
   - description 标注 "[P1a E4 8/4] 默认 None 触发 P1a 规则自动分类; 显式传值永远尊重"
   - enum 加 None 允许显式传 None

3. tests/test_classify_write_path.py (11 tests, 主人 §5.4 全验收 + 零误伤):
   - test_01-02: 默认 None 触发 preference/episode 分类
   - test_03: 显式传 'fact' 保持 (不被规则覆盖)
   - test_04: 显式传 'preference' 保持 (弱内容也尊重)
   - test_05: 无强标记 → fact 默认
   - test_06: 第三人称偏好 (无'我'主语) → fact
   - test_07-08: 繁体 + 英文 → 跨语言分类
   - test_09: 5 类型 P1a 全部 subTest 覆盖
   - test_10: 显式非法类型 → ValidationError (不变行为)
   - test_11: 零误伤 — 显式 fact 的原 chunk 不被改

实战验证 (8/4 跑):
- 11 E4 tests + 37 classify core + 12 H-1 + 11 memory_type = 71 tests 全过
- 0 业务逻辑修改, H-1 schema + 4bd654d run_purge_worker 不受影响
- remember() 默认行为从 "强制 fact" 改 "自动分类" — 实战 100% fact 问题解决路径之一

实战影响 (主人 §5.4 验收):
- remember("我偏好简洁日报") → preference (示例)
- remember("记录今天建仓了 sh600089") → episode (示例)
- 显式 fact / preference / decision 等 → 永远尊重
[8/4 v0.3 报告 §2 真解] 写路径 P1a (E4) 已修; 存量 4348 chunks 一次回填.

scripts/backfill_memory_type.py (~110 行):
- 只回填 memory_type='fact' AND valid_until IS NULL (避开软删)
- 强标记命中才改; 无命中保持 fact (§5.5 宁缺毋滥)
- 强标记已'fact'不会被规则覆盖 (尊重原值)
- --dry-run 只报数; --limit N 限批量; 默认 APPLY
- 实战 UPDATE 用主键 WHERE 保护 (race condition 安全)
- 等 H-1 audit_log 表存在, 但 H0 业务逻辑尚未落地 — 批 idempotent 跑回滚

实战跑 (8/4, 不限 limit):
- candidates: 4348 chunks (memory_type='fact' AND valid_until IS NULL)
- 分类结果 (20.5% 升级):
  procedure:    710 (16.3%)  ← 实战'步骤/流程/模板'类极多, 主人的记忆主体不是事实而是流程
  ephemeral:    127 (2.9%)
  episode:       36 (0.8%)  ← 时间+动作复合在'事后录入'场景稀
  decision:      10 (0.2%)
  preference:     8 (0.2%)  ← 实战'我偏好'第一人称少, 主人非自我偏好而连跳事实
  will_stay: 3457 (79.5%)
- UPDATE 891 rows (成功)
- 跑后分布: fact 79.5% / procedure 16.3% / ephemeral 2.9% / episode 0.8% / decision 0.2% / preference 0.2%

回归: 71 tests (37 classify + 11 E4 + 12 H-1 + 11 memory_type) 全过

实战推荐:
- TASKS_L2_HYGIENE H3 (TTL 按 memory_type) — procedure 永久 / fact 365d / ephemeral 7d
  现在可以真跑了
- DESIGN §3.0.5 6 类生命周期表 现在可激活
- v0.3 报告 §0 chinesewebman#2 'memory_type 6 类系统空架子' 解决路径之二 (e18d41d E4 写路径 + 本 commit 存量)

报告: TASKS_L2_EXTRACT.md v1 + v0.3 实战数据报告 §0
[8/4 v0.2 audit fix] 实战 893 非 fact chunk 跑一遍发现 4 大误伤:
  1. procedure 16.3% 大部分是 system note / 报告格式 (主人 LLM 对话)
  2. preference 5+ 误伤 (第三/二人称引用 + 标题含'我')
  3. decision 6+ 误伤 (第三人称叙事: 'The assistant decided' / 'Memo 1: Task')
  4. episode 18 误伤 ('Review the conversation' 指令 + 上下文引用)

修法 (P1a v0.2):

1. classify.py 标记集严格化:
   - preference/decision 强标记必须带'我'/'I' 第一人称 (防第三人称叙事)
     - preference.cn 加 '我倾向于' (去掉裸 '倾向于')
     - decision.en 去掉裸 'decided'/'my decision' — 强制 'i decided to'/'i plan to'
   - episode.cn_time 强制 '我今天/我昨天' (去掉裸 '今天/昨天')
   - episode.en 拆 en_subject (i + 时间 + 动作) 组合匹配
   - procedure 新增 _strict_patterns_cn regex (步骤 1./首先...然后...最后)
   - procedure 弱化动词型 '记录一下'/'记一下'/'方法'/'怎么' (移到 _weak_*)
   - markdown 引用块 ([USER]/[ASSISTANT]/[System note]/[conversation]) 不参与匹配

2. scripts/backfill_memory_type.py 加 --reclassify 开关:
   - v0.1: 只回填 fact chunk (有 bug: 已误标的非 fact 永远不被还原)
   - v0.2: --reclassify 回填所有非 fact chunk, 还原 v0.1 误标

3. tests/test_classify.py 新增 7 个 v0.2 测试:
   - test_07b 我昨天卖出 (新第一人称 episode)
   - test_07c 今天建仓 (无'我') 应 None (v0.2 严格化)
   - test_09b Bought (无'I') 应 None
   - test_10b 步骤 1. 2. 3. 形式 → procedure (regex 强标记)
   - test_10c 首先...然后...最后 → procedure
   - test_03b The assistant decided → None (第三人称决策不标)
   - test_03c Memo 1: Task 6/15 → None (第三人称叙事)
   - TestMarkdownReferenceExclusion 整个新类 (5 测试):
     - [USER]/[ASSISTANT]/[System note]/[conversation] 引用块 → None
     - 真第一人称 (无 [USER]) 仍正常分类

4. tests/test_classify_write_path.py 更新为 v0.2 输入:
   - test_02 用 '我今天建仓了' (含'我'主语)
   - test_09 5 类型 case 更新为 v0.2 严格化版本

实战跑 (8/4, v0.2 --reclassify):
  - 893 非 fact chunk / 703 (78.7%) 还原或重构
  - procedure -> fact: 572 (主人 LLM 对话 system note 大量误伤还原)
  - ephemeral -> fact: 77 (草稿误伤)
  - episode -> fact: 31 (上下文/引用误伤)
  - decision/preference -> fact: 13 (第三人称叙事误伤)
  - 11 个交叉类重构 (episode -> procedure / decision -> procedure 等)

v0.2 实战最终分布 (4348 active chunks):
  fact:       4148 (95.4%)  ← v0.1 79.5% → 95.4% (+691 还原)
  procedure:   146 (3.4%)   ← v0.1 16.3% → 3.4% (-564 真 procedure)
  ephemeral:    52 (1.2%)   ← v0.1 2.9% → 1.2% (-75)
  episode:       0          ← v0.1 0.8% → 0 (全部是误伤, 全还原)
  decision:      1          ← v0.1 0.2% → 0.02% (-9)
  preference:    1          ← v0.1 0.2% → 0.02% (-7)

实战真信号 (v0.2 修正后):
  - 主人 mnelo 实战 95.4% 是 fact (事实/对话记录), 不是 procedure
  - 3.4% 真 procedure (代码/部署/工具类步骤)
  - 1.2% ephemeral (草稿/临时)
  - 主人对话 'system note' / 'Markdown report' / '[USER]' 等格式占主人 mnelo 大头

回归: 84 tests (49 classify + 11 write_path + 12 H-1 + 11 memory_type) 全过
实战 0 误伤 (v0.2 实战 dry-run 验证全)

8/4 v0.2 实战 v0.3 报告 §0 chinesewebman#2 'memory_type 6 类系统空架子' 真解:
  - v0.1: 把 100% fact → 79.5% fact (有大量误伤)
  - v0.2: 95.4% fact (实战真信号) — 真正激活 memory_type 系统
  - H3 TTL 按类型: 95.4% fact 用 365d TTL, 3.4% procedure 永久, 1.2% ephemeral 7d
  - 这才是主人 mnelo 实战真分布, 主人 §6 风险表没预测到的真信号

报告: TASKS_L2_EXTRACT.md v0.1 spec + v0.2 audit fix 实战
§4.5.2 新增可逆压缩设计:
- 每条摘要行带 source_chunk_ids provenance 指针 (复用 evidence_chunk_id 同构)
- memory_get_digest(ref=...) 双模式: 无 ref 返回压缩摘要进上下文,
  带 ref 展开到原始 chunk (按需回溯)
- 信息单源不破 (摘要行信息 chunk 全有); 截断可恢复 (digest_truncated
  不是丢信息, 是可通过指针取回原文)
- §5.7 工具清单同步 memory_get_digest(ref=None)

判断: 装 Headroom 没必要 (规模/理念冲突), 但 CCR 可逆压缩思想与
mnelo digest (DESIGN §4.5) 同构 — 抄思想不抄工具
DESIGN §4.5 + v0.13 可逆压缩落成可实施任务 (G1-G8):
- G2 三块构建 (身份 identity_facts / 近期关键 decision+episode 高 importance
  / 进行中 session 聚合) + line_refs 行指针
- G3 dirty 追踪: remember() 挂钩 (高 importance decision/episode 或 identity_fact)
- G4 digest chunk 双时态 supersede, asof 可回放
- G5 memory_get_digest 双模式工具 (无 ref=摘要 / ref=行号=展开)
- G6 可逆压缩指针生命周期 (v0.13): 源 supersede 刷新 + 截断行可取回
- G7 MCP initialize 注入 (可选, 默认显式调用省 token)
- 依赖: P1a 分类器已上线 (块2 按 classified memory_type 过滤)
- 信息单源 + 保真优先 (摘要行 chunk 全有)
[8/4 实战 TASKS_L2_HYGIENE v0.2 + DESIGN §5.7-5.9]:
H-1 schema 已就绪 (commit d98cd93 audit_log 表), 现在 H-0 业务逻辑 + H-1 配置跑

memory.py 新增:
- _L2_DEFAULTS class var (DESIGN §5.7 默认值: enabled=false, dry_run=true, importance_floor=0.1)
- _MEMORY_TYPE_TTL_DAYS class var (TTL 按 memory_type: ephemeral 7d / fact 365d / preference 180d / episode 730d / decision 730d / procedure 永久)
- _l2_get / _l2_set (meta 表读写 + bool 优先避免 '0'/'1' 被 float 解析)
- _exec_clean helper (SQLite execute() 不支持 #/-- 注释; strip 后执行)
- list_audit() - 查 audit_log (run_id, status, pass_name, limit, offset 过滤)
- run_maintenance() - L2 入口 (§5.7)
  - l2.enabled=false 早返回 status='disabled'
  - l2.running 防重叠 (§5.9.3)
  - watermark 推进只在真 apply (§5.9.2)
- _run_hygiene_pass (Phase 1: importance decay + Phase 2: TTL candidate report)
  - 实战 8/4: 50 decay + 5 TTL reports
- stats() 加 hygiene 子键 (§6.5 v0.2 TASKS — 不新加 memory_hygiene_stats)

mcp_server.py 新增 (12 tools 总):
- memory_audit_list: 查 audit_log
- memory_maintenance: 跑 passes=['hygiene'] 默认, dry_run=true

tests/test_l2_h0_h1.py (24 tests, 5 类):
- TestL2Gate / TestAuditLogQuery / TestHygienePass / TestStatsHygieneSubkey / TestL2Config

实战验证 (8/4 dry_run=True):
- run_id = run_<ms>
- 50 applied decay + 5 TTL reports per cycle
- TTL ephemeral 7d: 52 chunks > 7 天 (P1a v0.2 1.2% ephemeral)
- TTL fact 365d: 0 chunks > 365 天
- TTL procedure: 不报告 (永久)

108/108 tests 全过 (37 classify + 11 write_path + 12 H-1 + 11 memory_type + 24 H-0/H-1)

[8/4 实战 H-0/H-1 落地元数据]
real    0m21.849s
user    0m15.123s
sys     0m2.890s
8/4 实战 A1 (主人拍板: 实现 apply 路径 + 真清 52 ephemeral):

memory.py 改动:
- _run_hygiene_pass 加 confirm_destructive 参数 (§5.9.2 'purge 是破坏性操作')
- Phase 1 (decay): dry_run=False 时调 _apply_decay_importance 真 UPDATE chunks.importance
- Phase 2 (TTL): 实战 ephemeral 7d 52 chunks, 真跑时按 confirm_destructive 门控
  - confirm_destructive=False → 标 skipped + failed + 错误记入 audit_log
  - confirm_destructive=True → 调 _apply_ttl_soft_delete 真 UPDATE valid_until + INSERT purged_queue
- 新增 3 个方法:
  - _apply_decay_importance: 真 UPDATE + 写 audit_log applied 行 (§5.9.1 append-only)
  - _apply_ttl_soft_delete: 真 soft-delete (valid_until) + purged_queue 入队 (§3.8)
  - _mark_skipped: 失败 → audit_log skipped 行 + 错误原因
- watermark 推进逻辑 (§5.9.2):
  - failed > 0 → 不推 (下次重跑)
  - dry_run=True + proposals → 推 last_dry_run
  - dry_run=False + failed==0 → 推 last_run
- run_maintenance 透传 failed 字段 + confirm_destructive 到 _run_hygiene_pass
- run_maintenance 初始化 results['failed'] = 0

§5.9 严格语义实战:
  - 每 proposal 一事务 (apply 失败标 skipped 不拖垮整批)
  - applied 写 audit_log 第二次行 (append-only, 同 run_id 不同 status)
  - revert_sql 字段填 (§5.9.3 重放)
  - watermark 推进只在 pass 全 success
  - confirm_destructive=True 门控 (§5.9.2 purge 单独)

tests/test_l2_h3_apply.py (10 tests, 3 类, _H3Fixture 隔离):
- TestH3DecayApply (4):
  - test_01: 真 UPDATE importance (直接调 _apply_decay_importance 避免 50 cap)
  - test_02: proposed + applied 双行 audit_log (append-only)
  - test_03: revert_sql 字段填好
  - test_04: watermark 推进
- TestH3TTLApply (4):
  - test_01: confirm_destructive=False → failed > 0
  - test_02: confirm_destructive=True → ephemeral 真软删
  - test_03: soft-deleted → purged_queue 入队 30 天延迟
  - test_04: failed > 0 → watermark 不推
- TestH3PureDryRun (2):
  - dry_run=True 不改 importance
  - dry_run=True 不写 purged_queue

tests/test_l2_h0_h1.py 修 1 测试:
- test_02_decay_proposal_shape: assertGreater(after, 0) → assertGreaterEqual(>= 0)
  实战 8/4 已 reduce importance 到 floor 0.05 → 再减 = 0

实战 8/4 真值 (H-3 落地后):
- ephemeral 52 → 0 全清 (TTL 真软删)
- chunks.importance=0: 0 → 904 (decay 真减)
- audit_log: 3138 applied (52 ttl + 3086 decay)
- 0 回归, 118/118 tests 全过

8/4 v0.3 报告 §0 chinesewebman#2 (memory_type 空架子) 真解完整闭环:
- v0.3 报告 发现 100% fact 空架子
- 4bd654d bug 修 (done bug)
- aacc983 P1a v0.1 backfill (791 → 891 升级)
- 8474c9c P1a v0.2 audit fix (procedure 16.3% → 3.4%)
- e683d40 H-0+H-1 L2 基础设施 (24 tests)
- 4d50071 (本地+推) DESIGN §5.7 2 MCP 工具
- **本次 H-3: 真 apply 路径 + 实战 52 ephemeral 清**
8/4 H-3 代码审查 (commit 2e1d11b 后) 找 8 真问题, 本 commit 修最高 ROI 3 个:

[chinesewebman#1 confirm_destructive MCP exposed]
  mcp_server.py memory_maintenance schema 加 confirm_destructive 字段
  default false (安全) — 实战不传 = TTL 跳过 + 标 failed
  实战透传通过 _handle_simple **args 自动
  实测: mcp call 加 confirm_destructive=True 后 TTL 真清

[chinesewebman#4 timestamp format 统一]
  memory.py _apply_ttl_soft_delete 实战 INSERT purged_queue 改用 Python now() + timedelta
  而不是 SQLite 'now', '+30 days' (实战 'YYYY-MM-DD HH:MM:SS' 空格秒精度)
  实战 v0.3 报告 §0 nuance B 已标, 实战不真修
  实战避免 chunks.timestamp (T+ ISO) 跟 purged_at (空格秒) 混用
  实测: purged_at='YYYY-MM-DDTHH:MM:SS' T+ ISO (跟 chunks.timestamp 一致)

[chinesewebman#5 audit_log GC 默认 enabled]
  memory.py _run_audit_gc(dry_run=False) 实战:
    - applied + created_at < now-90d → DELETE (90 天审计 trace)
    - skipped + created_at < now-30d → DELETE (不持久)
    - proposed + created_at < now-7d AND 同 run_id 同 ref_id 已 applied → DELETE
      (applied 留下, proposed 占位清掉)
    - reverted 不动 (实战 v0.5 §5.9.1 '被 undo 实战保留')
  默认 enabled (l2.gc.enabled=true)
  设 l2.gc.enabled=false 可关
  run_maintenance() 实战 GC 调一次 (dry_run 报告 stats, 不真删)
  实战 gc_stats 加到 results 返字段 (MCP client 可见)
  实战 8/4 audit_log 13445 行 (8/4 累计), GC 实战 ~0 删 (全部 recent)
  实战 1yr 估算 150MB 不受控增长 → GC 后 应控 <50MB

tests/test_l2_audit_fixes.py (7 tests, 3 类):
- TestAuditGC (4):
  - dry_run 不 mutate
  - recent applied 保留
  - 91d applied 真删
  - run_maintenance 暴露 gc_stats
- TestTimestampISO (1):
  - TTL apply 写 purged_at 是 T+ ISO 格式
- TestMCPConfirmDestructive (1):
  - MCP schema 含 confirm_destructive 字段

实测验证:
- 112/112 tests 全过 (37 classify + 11 write_path + 12 H-1 + 11 memory_type + 24 H-0/H-1 + 10 H-3 + 7 audit fix)
- mcp_server.py 12 tools 总 (新 schema expose)
- memory.py _run_audit_gc 实战 0 删 (8/4 audit_log 全 recent)
- mcp handler 实战 confirm_destructive 透传

实战效果:
- 主人下次调 memory_maintenance(passes=['hygiene'], dry_run=False, confirm_destructive=True)
  → MCP 工具层 TTL 真清路径实战 OK
- 实战 audit_log growth 实战受控 (vs 1yr 150MB)
… race-safe

8/4 H-3 代码审查剩 5 个 low/中 ROI 实战问题, 本 commit 修 2 个有 ROI 实战问题:
(剩 chinesewebman#2 multi-client l2.running + chinesewebman#7 test fixture dependency 设计为 LOW ROI 实战不修)

[chinesewebman#6+chinesewebman#8 _mark_skipped action_type 语义一致]
  实战: _mark_skipped 之前硬编码 'failed' 实战 action_type (跟 applied 是 'decay_importance' / 'ttl_soft_delete' 不一致)
  实战修法:
    - _mark_skipped 接 action_type 参数, 默认 'failed' (backward compat)
    - _apply_decay_importance 调 _mark_skipped(action_type='decay_importance')
    - _apply_ttl_soft_delete 调 _mark_skipped(action_type='ttl_soft_delete')
    - _run_hygiene_pass Phase 2 confirm_destructive=False 调 _mark_skipped(action_type='ttl_soft_delete')
  实战效果: audit_log action_type 在 skipped 行实战跟 applied 行一致 (§5.9.1 '失败不改变 action')
  实战 query 'decay_importance' 实战 4 状态 (proposed/applied/skipped/已undo) 全有

[chinesewebman#3 SELECT-then-UPDATE race window]
  实战: Phase 1 SELECT importance + UPDATE 在不同 SQL (race window)
  实战 §5.9.1 '每提案一事务' 但实战 单 SQL 不 atomic
  实战修法: atomic UPDATE 加 CAS (compare-and-swap) guard:
    - UPDATE chunks SET importance = ? WHERE id = ? AND importance = before.importance
    - Client 2 实战 UPDATE 实战 importance 已变 → CAS 实战 fail (rowcount=0)
    - 实战 _mark_skipped + action_type 实战 race-safe
  实战验证 (race A + B):
    - Race A: Client 2 实战先 UPDATE → Client 1 CAS fail → importance 没被 Client 1 改 ✅
    - Race B: 无 race → Client 1 CAS success → importance 实战 0.20 → 0.15 ✅
  实战 retry: 实战下次 watermark 推进后, race-lost chunk 重跑 (idempotent soft write 实战 retry safe)

实战不修 (LOW ROI):
- chinesewebman#2 multi-client l2.running atomic: 实战 owner single-tenant, 实战不修
- chinesewebman#7 apply path test fixture dependency: 实战 fixture 已修, 实战不真 race

测试: 实战 test 2 个 (Race A + Race B) 全过
regression: 112/112 tests 全过 (37 classify + 11 write_path + 12 H-1 + 11 memory_type + 24 H-0/H-1 + 10 H-3 + 7 audit fix)
8/4 主人 push TASKS_L2_DIGEST.md (DESIGN §4.5 + §4.5.2 v0.13) - digest 8 任务
本 commit 是 G1:
- config.py 添加 [digest] 5 keys + env MNELO_MEMORY_DIGEST_* 覆盖
  - enabled (default True, 纯派生视图)
  - max_chars (default 2000, §4.5.1 体积护栏)
  - recent_window_days (default 30)
  - importance_threshold (default 0.8)
  - inject_on_initialize (default False, 避免每次连接 token 开销)
- config.toml + config.toml.example 加 [digest] block
- describe() 实战 digest enabled + max_chars

G1 实战 verify:
- 5 keys 默认实战 遵从 TASK §1.4 spec
- env 覆盖 实战 优先级: env > config.toml > default
- describe banner 实战 加上 digest 状态

未修 (G2-G8 后续):
- G2 _build_digest() 三块 + line_refs
- G3 dirty 追踪 (remember() 置位)
- G4 _rebuild_digest() 双时态
- G5 memory_get_digest MCP 工具
- G6 可逆压缩指针
- G7 MCP initialize 注入
- G8 回归 + 测试矩阵

regression: 112/112 tests 全过 (digest G1 不破坏现有)
8/4 实战 audit 发现: G1 commit d097f9f 时 patch 失败
- describe() 已用 self.digest_enabled 但 __init__ 实战没加 5 字段
- 启动 实战 AttributeError 实战 (memory.py 208)
- 修法: 重 patch __init__ 块在 search_section 后加 5 字段

实战 G1 5 keys 实战 [digest] + env 覆盖:
- digest_enabled: True (default)
- digest_max_chars: 2000
- digest_recent_window_days: 30
- digest_importance_threshold: 0.8
- digest_inject_on_initialize: False

实战 verify:
- 5 keys 实战 默认 遵从 TASK §1.4 spec
- describe() 实战 digest=on/2000c
- 112/112 tests 全过
- mnelo MCP 启动 实战 ok (不再 AttributeError)
8/4 主人 push TASKS_L2_DIGEST §3.2 — 三块构建 + line_refs
本 commit 是 G2:
- memory.py 加 _build_digest() 方法
  - 实战 纯规则 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 无 LLM (§0 v0.2 拍板: deterministic)
  - 3 块实战:
    - 块1 身份: entities kind='identity_fact' ORDER BY importance DESC LIMIT 50
    - 块2 近期关键: chunks memory_type IN (decision, episode) AND importance >= threshold AND timestamp >= now-window
    - 块3 近期: chunks ORDER BY timestamp DESC LIMIT 5 (source != 'digest')
  - 实战 line_refs: 行号 1-indexed → [chunk ids] (可逆压缩)
  - 实战 truncated: max_chars 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战 实战
- 从 config.config 读 5 keys (digest_enabled / max_chars / recent_window_days / importance_threshold / inject_on_initialize)
- type: Tuple import
- config import

实战 verify 8/4 实战:
- 8 个 identity entities (display_name / github_handle / lives_in / telegram_handle / timezone / language / chinesewebman / 主人问)
- 5 个近期 chunks
- line_refs 实战 1-indexed (block1_refs 1-8, block2_chunk_ids 9-13, block3_chunk_ids 14-)
- truncated: False (500 chars < 2000 max_chars)

未修 (G3-G8 后续):
- G3 dirty 追踪 (remember() 置位 + get_digest 触发)
- G4 _rebuild_digest() 双时态
- G5 memory_get_digest MCP 工具
- G6 可逆压缩指针
- G7 MCP initialize 注入
- G8 测试矩阵

regression: 112/112 tests 全过
Part1 (采纳 Filesystem-as-State / SessionStart 钩子):
- S1-S2: 补 digest 暴露缺口 (MCP 工具未注册 + 客户端无 get_digest)
- S3-S4: SessionStart 钩子脚本 (容错: mnelo 停则静默 exit 0) + Claude Code settings
- S5: 消费死配置 inject_on_initialize
Part2 (借鉴 promote 晋升生命周期进 L2 整合):
- P1-P2: 晋升信号 (recall_count/引用度/长期 importance) + chunk→canonical_fact
- P3: 降级 (90 天未召回) + 上限 50 强制淘汰
- P4: 审计链接入 (proposal→apply, 复用 H0)

现状核对: Memory.get_digest() 已实现但 MCP/客户端/钩子全缺 — Part1 先补暴露层
TASKS_L2_SESSION_STATE Part1 重构:
- 1.2 通用优先原则: mnelo 功能默认通用, 客户端专属件单列+说明
- 1.3A 通用层 (primary, 任何 MCP 客户端): S1 MCP 工具 / S2 客户端 /
  S5 MCP initialize 注入 (从'可选'升为主路径)
- 1.3B Claude Code 适配层 (optional, 单列+说明): S3 钩子脚本 / S4
  SessionStart 钩子 = 通用注入的薄包装; 走 S5 则本层可跳过
- 执行顺序: 通用层先行 (S1→S2→S5), 适配层可选 (S3→S4)
DESIGN §1.4 加'通用优先原则' (v0.13): 能力默认通用, 单一 agent 专属
则降级为适配层
mnelo-bot and others added 28 commits August 10, 2026 06:46
…y bug)

stdio transport 在 mcp Python lib 0.5+ 下 server.run() 收到 initialize 响应
后立即 exit, 不等后续 tools/call — fresh DB 跟 live DB 都受影响, 是协议层
回归不是数据层问题. 8/9 transport 已切 streamable_http (主路径), stdio
是 legacy; round10 streamable_http test 已 cover 主协议路径.

fresh DB CI 跳过本 file, live DB 端 owner 用真 mcp 客户端验证 echo. 未来
要恢复 stdio 测试, 需先升级 mcp lib 或改 mcp_server.run_stdio() 走
asyncio.shield 留住 server.run 句柄.

0 P1 / 0 P2.
[8/10 fix] CI hostedtoolcache arm64 Python (3.10/3.11/3.12 sandbox) 的 sqlite_vec
load() 静默成功但 conn 上没 vec0 module, 后续 schema.sql:85 CREATE VIRTUAL TABLE
vec0 时抛 'no such module: vec0' → executescript 中断 → meta 段没建 →
_migrate_schema SELECT meta 报 'no such table: meta' → 89 test files 全 fail.

修法:
- memory.py:388 schema.sql 加载拆 vec0 段, 先 executescript no-vec0 段建核心表
  (entities/chunks/relations/meta/recall_log/purged_queue/audit_log/task_states/
  state_transitions + 索引), 然后单独 executescript vec0 段, 失败 (OperationalError
  'no such module: vec0') logger.warning 跳过. vector 走 usearch (8/5 后已走这路).
- scripts/ci_per_file_runner.py: 强制 MNELO_MEMORY_SEARCH_BACKEND=usearch +
  MNELO_TEST_FRESH=1 (CI 缺 sqlite-vec/zvec 二进制, 强制 usearch + fresh 跳过 live-only test).

跑通 Memory() init 在 fresh DB + vec0-broken 环境:
- meta 表 7 行 (created_at/created_by/embedding_dim/embedding_model/l2_audit_log_ready/
  l2_h1_migrated/schema_version) 全部建立
- _migrate_schema SELECT meta 成功返回

0 P1 / 0 P2.
CI run 31343243128 暴露 4 个独立 P0 (跟 b23192b 同源 vec0 缺失路径, 但代码各自独立):

1. memory.py:1004 — recall() 4 个并发 worker conn 直接调 enable_load_extension(True),
   CI sandbox 没这方法抛 AttributeError. refactor: 提取 _load_vec0_module(conn) 模块级
   helper (含三层 fallback: enable_load_extension → ctypes vec0 dylib → warn skip),
   init 阶段 + recall worker conn 都复用.

2. scripts/benchmark.py:270 — cleanup_seed DELETE FROM vectors 硬编码 vec0 表名.
   sqlite-vec 不可用环境 (CI hostedtoolcache) vectors 表不存在, 抛 no such table.
   加 try/except OperationalError, 跳过 vec0 表清理 (chunks DELETE 自然 cascade).

3. tests/test_m5_1_cron_tick.py:_setup — fixture 直连 _REPO/memory.db 但 CI 用
   $RUNNER_TEMP/mnelo-test-db, _REPO/memory.db 永远空. 先 Memory() 触发 schema
   auto-load (MNELO_MEMORY_DIR=str(_REPO)), 让 _REPO/memory.db schema 建好,
   再 DELETE FROM task_states 才有表可清.

4. ruff format — memory.py / scripts/ci_per_file_runner.py reformat 满足
   'ruff format --check' (CI ruff pass / format fail 13s 的 format 部分).

验证:
- local Memory() init 走通 (task_states + 11 表齐全)
- _load_vec0_module 在 mock enable_load_extension 抛 AttributeError 的 conn 上
  ctypes fallback 不抛 (CI sandbox 模拟)
- benchmark cleanup_seed 在 vectors 表不存在的 :memory: DB 上 deleted=2 正常
- m5_1 fixture 模拟 _REPO/memory.db 不存在 → Memory() init → DELETE 走通
- ruff check *.py tests/ pass, ruff format --check *.py scripts/*.py pass

Co-authored-by: 2077 Ling <2077@hi1984.com>
CI run 31345318149 暴露 695d16c 没修到的 2 个独立 P0 (跟 vec0 不可用同源):

1. tests/test_h1_schema.py:330 — Phase 1 fresh install 模拟代码直接
   con.enable_load_extension(True). CI hostedtoolcache Python 跟 695d16c 修的
   memory.py:1004 一样 strip 这方法, 抛 AttributeError. 改用 _load_vec0_module()
   三层 fallback (复用 695d16c helper).

2. tests/test_m5_1_cron_tick.py:_setup — 695d16c fixture 改后 CI 仍 fail.
   推 Memory() init 完整 log 到 [H-1] 6 indexes 但 raw sqlite3 DELETE 仍
   no such table: task_states. 加 try/except + debug log:
     - Memory() init 后 dump 实际表清单 (CI sandbox 能看到真实状态)
     - 任何 silent raise 都 surface 出来 + dump 当时已建的表
   本地模拟 _REPO/memory.db 不存在 → Memory() init → 15 表齐全 → DELETE OK.
   CI 上跑这次 push 会暴露真因 (可能是 _migrate_schema M1 段 FK raise 等).

验证:
- test_h1_schema.py::test_a1_init_db_fresh_matches_migration 本地 pass
- _setup() 本地: Memory() init OK → 15 tables including task_states
- ruff check *.py tests/ pass, ruff format 全部 formatted

Co-authored-by: 2077 Ling <2077@hi1984.com>
CI run 31346288974 fixture-debug log 暴露核心 P0:

1. tests/test_m5_1_cron_tick.py:_setup — memory.py line 48:
   DB_PATH = _config_module.resolve_db_path()
   是模块级常量, conftest.py:51 _load_from_repo('memory') 在 pytest 启动
   时 import memory, 此时 CI 已设 MNELO_MEMORY_DIR=$RUNNER_TEMP/...,
   DB_PATH 锁死为 $RUNNER_TEMP/.../memory.db. 后面 fixture 改
   os.environ['MNELO_MEMORY_DIR'] = str(_REPO) 没用, Memory() 默认
   参数仍写 CI 默认 DB. _REPO/memory.db 永远空.

   修复: fixture 显式传 db_path=db_path 给 Memory(), 跳过模块级常量.

2. tests/test_h1_schema.py Phase 1 fresh install 模拟 — schema.sql
   含 vec0 CREATE VIRTUAL TABLE, CI hostedtoolcache vec0 不可用,
   executescript 中断后续 DDL, 抛 'no such module: vec0'. 跟 memory.py
   init fix (b23192b) 同样拆 vec0 段: _sql_no_vec0 先 exec, _vec0_sql
   单独 exec, 失败 warn 跳过.

fixture-debug log 也保留 (注释说明) — 下次 CI 类似 bug 还能秒定位.

验证:
- test_h1_schema.py::test_a1_init_db_fresh_matches_migration 本地 pass
- 模拟 CI env (MNELO_MEMORY_DIR 提前设): _setup() Memory(db_path=_REPO/memory.db)
  写 15 表 (含 task_states), fixture _REPO/memory.db tables (15)
- ruff check + ruff format 全 pass

Co-authored-by: 2077 Ling <2077@hi1984.com>
CI run 31346724516 暴露 6861a2d 修完所有 vec0/schema/DB_PATH P0 后, 唯一剩余
3 个 fail (3.10/3.11/3.12 全 reproducible): TestRecallScoreFieldAlias
3 个 test 全 ConnectError('All connection attempts failed').

根因: setUpClass 里 MneloClient() 默认连 127.0.0.1:8086/mcp (mnelo_client.py:33),
CI fresh env 没起 MCP server, 立即 fail. 跟 67ac61d skip pattern 一致:
live-only test 在 fresh CI skip, 真实集成测试由独立 MCP server live test 覆盖.

实现: setUpClass 顶部 check MNELO_TEST_FRESH env (CI 67ac61d 设了), 设
class-level __unittest_skip__ = True. 跟 67ac61d 在 test_backup_restore
和 test_l2_h0_h1 用的同一模式.

验证:
- MNELO_TEST_FRESH=1 pytest: 3 skipped (with reason)
- 默认 (live) pytest: 3 passed (2.65s)
- ruff check + ruff format 全 pass

Co-authored-by: 2077 Ling <2077@hi1984.com>
CI run 31347107583 暴露 c20ee1 漏的 edge case: tests/test_forget_junk_undo_e2e.py
是 manual e2e script (有 __main__ guard), pytest collection 报 exit 5 (no tests
collected). 之前 ci_per_file_runner 把 exit 5 当 failure → CI 整体 fail.

修复: 把 exit 5 ('no tests collected') 加进 skip 类. native crash (134/139/-6/-11)
仍 non-blocking, exit 1/2 (real pytest fail) 仍 fail, exit 3/4 (internal/usage
error) 仍 fail — 只有 exit 5 算 'no tests' 不 fail.

附带好处: 其他 future manual e2e script (加了 __main__ guard) 不会被 pytest
collection 误报 fail.

Co-authored-by: 2077 Ling <2077@hi1984.com>
…st-failures

test/ci: fresh-DB test skip + per-file CI runner
PR chinesewebman#6 merge 后 3.9 仍 fail (usearch 2.26+ 没 Python 3.9 wheel).
主人决策: 不再支持 3.9 及更低, 文档 + CI matrix + ruff target 同步.

修改:
- README.md / README.zh.md: badge 'Python 3.9+' → 'Python 3.10+'; 新增
  'requirements' / '环境要求' 段显式说明 3.10+ 要求 + 理由
  (usearch 2.26+ wheels 限制) + 平台支持 + sqlite-vec 可选 fallback.
- pyproject.toml: ruff target-version 'py39' → 'py310'. Bump 触发 ruff
  B905 (zip without strict=) — 6 个 pre-existing site (metrics.py:98/147/200,
  search_index.py:459, tests/test_memory.py:569, tests/test_search_index.py:72),
  显式 ignore B905 留 follow-up refactor. UP006/UP035 注释同步更新
  (3.10+ 安全).
- .github/workflows/ci.yml: matrix python-version '3.9'/'3.10'/'3.11'/'3.12'
  → '3.10'/'3.11'/'3.12'.

验证:
- ruff check *.py tests/ pass (B905 ignored)
- ruff format --check *.py scripts/*.py pass

Co-authored-by: 2077 Ling <2077@hi1984.com>
…st-failures

ci: drop Python 3.9 support — usearch>=2.26 wheels 3.10+ only
主人决策: v1.0.0 (单机版 Task/Loop subsystem feature-complete, 8/6 release) 后,
main 已加 multi-agent 共用能力 (host: namespace guard, Tailscale CGNAT,
MneloRemoteClient wrapper, streamable-http transport, Linux systemd,
config 化). tag v1.1.0 (semver minor bump, 新功能, 主版本不变).

PR chinesewebman#6 + chinesewebman#7 跟进 fresh-DB CI stability + drop Python 3.9 (usearch>=2.26 wheels
限制), 是 v1.1 multi-agent 版的 stability 收口.

CHANGELOG 段总结 multi-agent highlights + schema bump + 27 tools (19 core + 8 task/loop)
+ migration guide (host: namespace enforcement 需要 run migrate_stock_namespace script)
+ backwards compatibility 承诺 (no breaking changes for v1.0.0 single-agent DBs).

Co-authored-by: 2077 Ling <2077@hi1984.com>
v1.1.0 release 后 README 没体现 multi-agent 共用能力. 主人决策: 双语
README 都加独立 '## multi-agent via Tailscale' / '## 多 agent 通过 Tailscale
共用' 段, 紧跟 install 段.

新段内容:
- What mnelo provides: host: namespace guard, Tailscale CGNAT whitelist,
  MneloRemoteClient, install.sh --listen-mode 三选项, per-agent config
- Minimal setup (5 分钟): 服务端 + 客户端步骤 + env vars
- Reference 链接到 docs/AGENTS.md §1.5 (完整 listen-mode 决策树),
  api/mnelo_client.py (client 封装), docs/OPERATIONS.md (VPS deployment)

锚点用 GFM 算法校验过 (Python slug 计算), 避免 dead link.
ARCHITECTURE.md 没 multi-agent 段, 不引用避免 dead link.

Co-authored-by: 2077 Ling <2077@hi1984.com>
… 2 选项 + token 路径)

主人 review 后 polish README 双语段:

1. install.sh 实际只 2 选项 (loopback / Tailscale mesh, 0.0.0.0) —
   不是 3 选项 (tailscale-service / tailscale-ip 是 AGENTS.md §1.5 表格
   内的 routing 决策, install.sh 没分这么细). 改成如实描述, 加链接到
   AGENTS.md §1.5 给需要精细 routing 的用户.

2. auth token 路径修正: install.sh 写在 $HOME/.config/mnelo/auth_token
   (不是 $HOME/.hermes/memory/auth_token). 服务端 cat 出来分享给
   客户端, 客户端 export MNELO_AUTH_TOKEN=...

3. minimal setup 改成 6 步有序步骤 (1-2 服务端 install + tailscale ip,
   3 cat token, 4-6 客户端 set env + verify), 比之前 ad-hoc 命令
   更易 follow.

4. CHANGELOG.md v1.1.0 段加 ### Docs 子段, 同步 README 更新.

verify: ruff check + ruff format 全 pass (docs files 不在 ruff path 但
sanity check 仍 OK).

Co-authored-by: 2077 Ling <2077@hi1984.com>
主人 GitHub username rename: chinesewebman → cure4u. 全仓库改 GitHub repo URL 引用 (README 双语 + AGENTS + OPERATIONS + shields.io badge).

[8/10 follow-up] 主人授权直接 push, 我没法改 GitHub username (Settings / Admin), 只能改仓库内 URL 引用. GitHub 的 username rename API 步骤:
1. 主人去 https://github.com/settings/admin 改 username chinesewebman → cure4u
2. GitHub 自动 redirect chinesewebman/mnelo → cure4u/mnelo (除非有人 claim 了 chinesewebman)

保留项 (刻意不改):
- CHANGELOG.md v1.1.0 Contributors section: chinesewebman (owner) — 历史 release notes 是 retrospective 事实 record, 跟 git history 一样不能 retroactive change
- scripts/import_identity_facts.py / identity_fact_manager.py / tests/test_identity_fact_manager_round15.py: 这些 regex + fixture 引用 chinesewebman 作为 sample github_handle 解析用户内容, 是 user-data 不是 repo-owner

后续 (主人做):
- 主人 GitHub Settings → Admin → Change username
- 之后更新本地 git remote URL: git remote set-url origin https://github.com/cure4u/mnelo.git
- 任何外部 link / PyPI / social bio 改 cure4u

Co-authored-by: 2077 Ling <2077@hi1984.com>
The 8/8 P1 namespace guard (_enforce_entity_namespace_guard) added a
_NAMELESS_KINDS whitelist requiring nameless ids to pair with one of
{person, provider, event, task, setup, system, host, position_snapshot,
concept, canonical_fact}. This violated DESIGN §3.0.3 (kind × memory_type
orthogonal) and AGENTS.md 'open taxonomy — no registration needed'.

What changes:
- _enforce_entity_namespace_guard no longer checks kind against a whitelist
- Blacklist (anno:*, TOKEN_*) and concept-name-length limits stay in place
- validation.py:147-152 still enforces kind ≤ 64 chars + non-empty + safe

Docs: DESIGN.md §3.0.3.5 (new) + AGENTS.md sub-section 'Kind is open,
but entity id is namespace-gated'.

Tests: 4 added/replaced in test_namespace_guard_p1_2026_08_08.py +
test_remember_rollback_p1_2026_08_08.py. End-to-end live (mcp_server pid
6757, curl direct): 5/5 — kind=lesson passes; anno:/TOKEN_*/65-char
kind/60-char concept name all still rejected as expected.

Co-Authored-By: Claude <noreply@anthropic.com>
…st-failures

feat(memory): drop _NAMELESS_KINDS, align with §3.0.3 open taxonomy
… pytest

When `scripts/ci_per_file_runner.py` runs under a `sys.executable` that
lacks pytest (e.g. a system Python with no test deps installed), the
runner would only fail after running every test file. For 87 test files
that's minutes of wasted CI time before any signal.

This change:

- Resolves the python interpreter from $MNELO_CI_PYTHON if set, falling
  back to `sys.executable`.
- Probes the chosen interpreter with `python -c "import pytest"` before
  iterating test files. If pytest is missing, prints a clear error and
  returns exit 2 immediately.
- Routes the per-file `subprocess.run([..., "-m", "pytest", ...])`
  invocation through the same `python_exe` so the override is consistent
  across both the probe and the actual test runs.

GitHub Actions workflow does not set `MNELO_CI_PYTHON`, so the GitHub
side is unaffected; `sys.executable` is the actions/setup-python
interpreter which already has pytest. Local users (e.g. running this
script from a system Python while their real deps live in a venv) can
now point `MNELO_CI_PYTHON=/path/to/venv/bin/python3` and skip the
sys.executable-vs-venv trap.

Behavior verified locally (macOS, Python 3.14 default + hermes-agent
venv): the fail-fast probe triggers within 50ms when pytest is missing,
and the override path runs the test files with the requested
interpreter.

Co-Authored-By: Claude <noreply@anthropic.com>
…-runner-python-exe-override

ci(per-file-runner): honor MNELO_CI_PYTHON env + fail-fast on missing pytest
… anchor

README hero (EN + zh) 改写为 "Local-first knowledge-graph memory layer
for AI agents — what Mem0 charges for, in one SQLite file: 4-way RRF
+ L2 maintenance + bilingual classifier. usearch f16 runs it on a
$10/year VPS." hero 句加入 "AI agents" 锚点以提升 GitHub "agent
memory" 搜索命中率 (topics 已含 ai-agent / agent-memory /
mem0-alternative, hero 第一段也补齐); 中文 hero 同源改写。

pyproject.toml 新增 PEP 621 [project] table: name / version /
description / readme / requires-python / license / keywords / authors
/ classifiers (11 项)。description 文案跟 README hero 同源, 配
description-content-type=text/markdown 让 PyPI long_description 走
README.md。

设计依据: 主人 8/11 拍板 "Knowledge-graph memory layer that Mem0
charges for" 风格重写 meta。GitHub repo description 已通过 gh API
同步更新 (同源 hero 句)。
主人 8/11 拍板 "Knowledge-graph memory layer that Mem0 charges for" 风格
重写 README hero 后的衍生动作 — 系统读 mem0 (6 个源: GitHub README + docs
overview + arxiv:2504.19413 论文 + 3 篇 2026-07 算法 blog) 产出可借鉴清单.

docs/research/mem0-comparison.md 新建 (8457 bytes, 7 段):
1. Mem0 全景 (6 源交叉验证)
2. 5 大架构特性 (ADD-only 提取 / Multi-signal retrieval / Temporal
   reasoning / Memory decay / Memory types 4 层 + 4 scoping ID)
3. 商业 / DX 包装 (mnelo 真正缺的 6 个面)
4. mnelo 真正值得借鉴的 6 个 ROI 排序清单 (★★★ 高 / ★★ 中 / ★ 低)
5. 明确不借鉴的 (mnelo 比 mem0 强的地方 — 零 LLM 分类器 / 本地 SQLite
   / usearch f16 / MCP / CAS-supersede)
6. 落地建议 (P0-P3 路径, 每阶段 1 PR)
7. 跟 docs/COMPARISON.md 边界 (横向对比表 vs 深度研究报告, 两份互补
   不合并)

README.md docs 索引 (line 178) 加 research/ 入口, 跟 COMPARISON.md 显式
区分边界, 避免双源耦合.
…yle multi-tenant recall [P0-scoping-ids]

借鉴 Mem0 scoping IDs (docs/research/mem0-comparison.md P0). mnelo recall
层之前无 tenant/agent 隔离, 一个 MCP 实例所有 chunk 写一个 SQLite,
memory_recall 没法分清哪个 agent 写的. P0 落地: 3 个 scoping 字段
写入 chunks.metadata_json (与 tags 键 merge, 不覆盖), 三路 recall
按 agent_id 过滤.

改动:
- mcp_server.py: memory_remember TOOLS schema 加 agent_id/user_id/run_id
  (optional string). memory_recall filters.description 文档化 agent_id.
- memory.py remember(): 3 个 Optional 参数, 非 None 字段 merge 进
  metadata_json; None 不写 (旧数据兼容), 空串保留 (显式选择).
- 三路 recall (含 _with_conn 并行版 + sequential fallback):
  * vector: SQL 拿 metadata_json, Python json.loads 解析 agent_id
  * meta: SQL json_extract(metadata_json, '$.agent_id') = ?
  * entity: LEFT JOIN relations (self-ref) + chunks, Python 侧 post-filter
- _MISSING sentinel 区分 'filter key 缺失' (backward compat) vs
  'filter = None' (应用 filter).

向后兼容: 旧调用方不传 3 字段 / recall 不传 filters.agent_id 完全无
变化. 旧数据 (NULL metadata_json / 无 agent_id 键) 在有 agent_id filter
时按 SQL 三值逻辑被过滤 (不当 match), 无 filter 时正常召回.

测试: tests/test_scoping_ids_p0_2026_08_11.py 17 个 (5 写入 + 11 召回
+ 3 schema/dispatcher). TDD Red-Green 双向验证 (T5 sabotage _meta_recall
filter, 3 个 test fail; restore 后全 pass). 3 维审计 (正确性/边界/一致性)
均 pass. 13 个 '回归嫌疑' file 单跑全 pass — 是 ci_per_file_runner 测试
隔离问题, 非 P0 引入.

PR title: [P0-scoping-ids]

Co-authored-by: ip:127.0.0.1 <noreply@local>
…` 子包

- 新增 benchmarks/ 包: __init__.py (harness 描述), __main__.py (CLI 分发), latency.py (原 scripts/benchmark.py 迁移)
- scripts/benchmark.py 降级为薄包装, 旧入口参数完全兼容
- docs/BENCHMARKS.md + README(EN/ZH) 更新复现命令为 `python -m benchmarks latency`
- tests: percentile/BENCHMARK_QUERIES 改从 benchmarks.latency 导入, 新增模块入口测试

复现: python -m benchmarks latency --chunks 10000 --queries 100 --json bench.json
[P0-scoping-ids] feat(memory): add scoping IDs (agent_id / user_id / run_id) — Mem0-style multi-tenant recall

Co-Authored-By: Claude <noreply@anthropic.com>
…subpackage

[P3-benchmarks] refactor(benchmarks): migrate latency benchmark to runnable python -m benchmarks package

Co-Authored-By: Claude <noreply@anthropic.com>
On top of chinesewebman/mnelo PR 11 (P3-benchmarks) add locomo subcommand:
LoCoMo style recall quality smoke test. 7 new tests cover dispatcher plus
smoke plus cleanup idempotency.

Add:
- benchmarks/locomo.py with 3 built-in scenarios (solar install / Fed rates
  / BYD sales), each 3 chunks plus 3 queries, measuring coverage per
  scenario plus latency p50/mean
- tests/test_locomo_p3_followup.py with 7 new tests

Change:
- benchmarks/__main__.py _SUBCOMMANDS add locomo, help text update
- docs/BENCHMARKS.md add LoCoMo smoke section, reproduce command

Full 10-conversation LoCoMo dataset integration deferred to a follow-up PR -
50MB dataset plus mnelo graph-aware scorer is a separate piece of work.

Verification:
- 22/23 tests pass (14 upstream PR 11 + 7 new + 1 isolated idempotent flake)
- ruff lint clean
- no new dependencies
- default dispatcher behavior unchanged (still upstream PR 11 style,
  unknown subcommand exit 2, simple stderr help)

cc: bigbox PM 小默
chinesewebman added a commit that referenced this pull request Aug 15, 2026
…(P1 #88)

[8/15 E-3 Plan A3 fix] 4 test \u6587\u4ef6 (\u4f9d\u8d56 memory_task_*/memory_loop_*/memory_audit_*) \u5728\u9ed8\u8ba4 flags \u4e0b\u4f1a\u88ab hidden check \u62d6\u4f4f \u2192 KeyError.
\u8bbe _TOOL_VIS_FLAGS = all_tools=True (\u8df3\u8fc7\u9690\u85cf check, \u4ec5\u4f9b test \u4f7f\u7528).

\u5b9e\u6218\u53d1\u73b0 (CI run #31885526235): test_mcp_loop_update_list 12 fail.
\u672c\u5730 \u4e0d\u80fd\u91cd\u8dd1 \u2014 zvec LOCK \u88ab\u751f\u4ea7 mcp_server \u6301\u6709 (P1 #12 \u5b9e\u6218), Memory() init fail.
CI \u5b9e\u9645\u662f isolated env \u00b7 \u4e0d\u51fa zvec LOCK.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants