Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -146,7 +146,7 @@ Two layers, one seamless experience.
Curated facts injected into every system prompt. Active projects, current deadlines, operational lessons. Tagged with dates, automatically evicted when stale.

**L2 — Deep Memory (memU)**
Semantic search over everything — conversations, facts, preferences, events. SQLite-persisted with `text-embedding-3-small` embeddings.
Semantic search over everything — conversations, facts, preferences, events. SQLite-persisted. Uses vector embeddings when an OpenAI key is configured, or LLM-based ranking with Anthropic models only.

- Four memory types: `profile`, `event`, `knowledge`, `behavior`
- Automatic conversation indexing on session close
Expand Down Expand Up @@ -254,7 +254,7 @@ See [docs/config.md](docs/config.md) for all options.
- [Node.js](https://nodejs.org/) 18+ (for web UI build)
- [Claude Code CLI](https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/overview) (bundled with `claude-agent-sdk`)
- Anthropic API key **or** Claude subscription via [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) proxy
- Optional: OpenAI API key (for memU embeddings), Telegram bot token, [gog](https://github.com/googleworkspace/cli) CLI, [gh](https://cli.github.com/) CLI
- Optional: OpenAI API key (for vector-based memory search — without it, LLM-based recall is used), Telegram bot token, [gog](https://github.com/googleworkspace/cli) CLI, [gh](https://cli.github.com/) CLI

## Documentation

Expand Down
2 changes: 1 addition & 1 deletion config.example.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -46,7 +46,7 @@ sync:
# Memory (memU)
memory:
chat_model: claude-sonnet-4-6
embed_model: text-embedding-3-small
# embed_model: text-embedding-3-small # Only needed with openai_api_key

# Cron
cron:
Expand Down
4 changes: 2 additions & 2 deletions docs/config.md
Original file line number Diff line number Diff line change
Expand Up @@ -99,7 +99,7 @@ Sources pull data from external services on a schedule. See [sources.md](sources
| `memory.recall_model` | string | `claude-sonnet-4-6` | Model for recall routing |
| `memory.memorize_model` | string | `claude-sonnet-4-6` | Model for extraction & preprocessing |
| `memory.fast_model` | string | `claude-haiku-4-5-20251001` | Model for categorization, date resolution, knowledge filtering |
| `memory.embed_model` | string | `text-embedding-3-small` | Embedding model |
| `memory.embed_model` | string | *(empty)* | Embedding model (only used when `openai_api_key` is set, e.g. `text-embedding-3-small`) |
| `memory.semantic_dedup_threshold` | float | `0.85` | Cosine similarity threshold for semantic deduplication (0 to disable) |
| `memory.knowledge_filter` | bool | `false` | Post-extraction LLM filter that deletes generic knowledge items (extra Haiku API call per memorize) |
| `memory.categories` | list | `[]` | Seed categories — each entry has `name` and `description` fields. Used for semantic routing when memorizing and recalling facts. `nerve init` populates mode-appropriate defaults (personal: relationships, finances, health, etc.; worker: patterns, procedures, approvals, etc.). |
Expand Down Expand Up @@ -174,7 +174,7 @@ Nerve automatically discovers MCP servers from Claude Code's enabled plugins. An
| Key | Type | Description |
|-----|------|-------------|
| `anthropic_api_key` | string | Anthropic API key (agent + memU chat). Not required when proxy is enabled. |
| `openai_api_key` | string | OpenAI API key (memU embeddings only) |
| `openai_api_key` | string | OpenAI API key (optional — enables vector-based memory search via embeddings; without it, LLM-based recall is used) |
| `brave_search_api_key` | string | Brave Search API key (optional) |

## Proxy (CLIProxyAPI)
Expand Down
15 changes: 9 additions & 6 deletions docs/memory.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,17 +78,19 @@ Reinforced items rank higher in search results via salience-aware ranking: `simi

### Configuration

memU uses three LLM profiles:
memU uses two or three LLM profiles depending on configuration:
- **Chat** — Anthropic API for recall routing (claude-sonnet-4-6)
- **Fast** — Anthropic API for fact extraction and categorization (claude-haiku-4-5)
- **Embedding** — OpenAI text-embedding-3-small for vector search
- **Embedding** *(optional)* — OpenAI text-embedding-3-small for vector search. Only active when `openai_api_key` is set.

When no OpenAI key is configured, memU uses **LLM-based recall** instead of vector search — the Chat profile ranks memories directly, requiring no embeddings. This uses more Anthropic API tokens per recall but removes the OpenAI dependency entirely.

Config in `config.yaml`:
```yaml
memory:
chat_model: claude-sonnet-4-6 # recall routing
fast_model: claude-haiku-4-5-20251001 # extraction & categorization
embed_model: text-embedding-3-small
# embed_model: text-embedding-3-small # only needed with openai_api_key
categories: [...] # see Categories section above
```

Expand All @@ -102,10 +104,11 @@ This uses the `anthropic` Python SDK directly (not the OpenAI-compatible endpoin

### Performance Optimizations

- **Vector cache** — All item and category embeddings are preloaded into memory at startup, eliminating repeated SQLite JSON parsing (~2s per search saved)
- **Vector cache** *(with OpenAI key)* — All item and category embeddings are preloaded into memory at startup, eliminating repeated SQLite JSON parsing (~2s per search saved)
- **Fast model** — Extraction and category summary updates use Haiku instead of Sonnet
- **Disabled pipeline steps** — Route intention, sufficiency checks, and resource retrieval are disabled in the retrieve pipeline (saves 3+ LLM calls per recall)
- **Category embedding reuse** — Category ranking uses stored embeddings instead of re-embedding summaries on every recall
- **Category embedding reuse** *(with OpenAI key)* — Category ranking uses stored embeddings instead of re-embedding summaries on every recall
- **LLM-based fallback** — When no embedding provider is configured, retrieval and memorization work without embeddings; semantic deduplication falls back to content-hash only
- **Client warmup** — Anthropic LLM clients are pinged during startup to force HTTP/2 connection establishment (avoids a cold-start hang on the first memorize call)
- **Memorize timeout** — Each memorize call is capped at 300s; if it hangs, it is cancelled, LLM clients are evicted (cache cleared + HTTP transport closed), a fresh client is created, and one retry is attempted
- **Per-call LLM timeout** — Base LLM client `.chat()` methods are wrapped with a 120s `asyncio.wait_for()` at init time (instance attribute shadowing). A single dead HTTP/2 connection fails fast instead of consuming the entire 300s pipeline budget. Hung calls log `memU LLM HUNG [profile]: no response after 120s (prompt=N chars)`
Expand Down Expand Up @@ -149,7 +152,7 @@ Delete a memory item. Use for wrong, duplicate, or stale memories.
```
category_update(category_id="abc123", summary="Updated summary", description="New description")
```
Update a category's summary and/or description. Re-embeds the category after update to keep vector search in sync.
Update a category's summary and/or description. Re-embeds the category after update to keep vector search in sync (when an embedding provider is configured).

All recall and history results include memory IDs (`id:abc123...`), enabling the agent to target specific items for update or deletion.

Expand Down
2 changes: 1 addition & 1 deletion docs/setup.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,7 +131,7 @@ The wizard handles all of this automatically, but you can also configure manuall
# Create secrets file (gitignored)
cat > config.local.yaml << 'EOF'
anthropic_api_key: sk-ant-...
openai_api_key: sk-... # For memU embeddings (optional)
openai_api_key: sk-... # Optional — enables vector-based memory search

telegram:
bot_token: "123456:ABC..."
Expand Down
7 changes: 4 additions & 3 deletions nerve/bootstrap.py
Original file line number Diff line number Diff line change
Expand Up @@ -673,9 +673,10 @@ def _prompt_openai_key(self) -> None:
"""Prompt for optional OpenAI API key (used by both auth paths)."""
click.echo()
click.secho(
"Optionally, an OpenAI key enables better memory search\n"
"(text-embedding-3-small for vector embeddings). Nerve works\n"
"without it but recall quality improves significantly.",
"Optionally, an OpenAI key enables vector-based memory search\n"
"(text-embedding-3-small for semantic embeddings). Nerve works\n"
"without it using LLM-based recall, which uses more API tokens\n"
"per query but requires no additional API key.",
dim=True,
)
click.echo()
Expand Down
4 changes: 2 additions & 2 deletions nerve/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -603,9 +603,9 @@ def doctor_report(config) -> str:
errors.append("[ERR] Anthropic API key not set and proxy not enabled (config.local.yaml)")

if config.openai_api_key:
lines.append(f"[OK] OpenAI API key: ...{config.openai_api_key[-4:]}")
lines.append(f"[OK] OpenAI API key: ...{config.openai_api_key[-4:]} (vector embeddings enabled)")
else:
warnings.append("[WARN] OpenAI API key not set (needed for memU embeddings)")
lines.append("[--] OpenAI API key not set (using LLM-based memory recall)")

# Check Telegram
if config.telegram.enabled:
Expand Down
4 changes: 2 additions & 2 deletions nerve/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -251,7 +251,7 @@ class MemoryConfig:
recall_model: str = "claude-sonnet-4-6" # Recall routing
memorize_model: str = "claude-sonnet-4-6" # Extraction & preprocessing
fast_model: str = "claude-haiku-4-5-20251001" # Category summaries, date resolution
embed_model: str = "text-embedding-3-small"
embed_model: str = ""
sqlite_dsn: str = ""
semantic_dedup_threshold: float = 0.85 # Cosine similarity threshold for semantic dedup
knowledge_filter: bool = False # Post-extraction LLM filter for generic knowledge (extra API call)
Expand All @@ -266,7 +266,7 @@ def from_dict(cls, d: dict) -> MemoryConfig:
recall_model=d.get("recall_model", "claude-sonnet-4-6"),
memorize_model=d.get("memorize_model", "claude-sonnet-4-6"),
fast_model=d.get("fast_model", "claude-haiku-4-5-20251001"),
embed_model=d.get("embed_model", "text-embedding-3-small"),
embed_model=d.get("embed_model", ""),
sqlite_dsn=d.get("sqlite_dsn", default_dsn),
semantic_dedup_threshold=float(d.get("semantic_dedup_threshold", 0.85)),
knowledge_filter=bool(d.get("knowledge_filter", False)),
Expand Down
147 changes: 136 additions & 11 deletions nerve/memory/memu_bridge.py
Original file line number Diff line number Diff line change
Expand Up @@ -829,7 +829,13 @@ async def initialize(self) -> bool:
# conversations as instructions, generating long off-topic
# responses that exceed the per-call timeout.
"preprocess_llm_profile": fast_profile,
"memory_extract_llm_profile": memorize_profile,
# When no embedding provider is configured, all memory
# work goes through Anthropic — use Haiku for extraction
# too to avoid saturating the rate-limit budget.
"memory_extract_llm_profile": (
fast_profile if not self.config.openai_api_key
else memorize_profile
),
"category_update_llm_profile": fast_profile,
# Pass Nerve's configured categories to memU so the LLM
# prompt only shows categories that actually exist in the
Expand All @@ -855,9 +861,14 @@ async def initialize(self) -> bool:
},
},
retrieve_config={
"method": "llm" if not self.config.openai_api_key else "rag",
"route_intention": False,
"sufficiency_check": False,
"resource": {"enabled": False},
# Use Haiku for LLM-based ranking — cheaper and avoids
# sharing Sonnet's rate-limit budget with the main agent.
**({"llm_ranking_llm_profile": fast_profile}
if not self.config.openai_api_key else {}),
},
)
self._available = True
Expand Down Expand Up @@ -888,6 +899,104 @@ async def initialize(self) -> bool:
len(ctx.category_name_to_id),
)

# When no embedding provider is configured, replace the
# memorize pipeline's "categorize_items" step with one that
# stores items and resources with embedding=None. This
# avoids KeyError on the missing "embedding" LLM profile.
if not self.config.openai_api_key:
from memu.workflow.step import WorkflowStep as _WfStep

_svc = self._service

async def _categorize_no_embed(
state: dict, step_context: Any,
) -> dict:
svc_ctx = state["ctx"]
store = state["store"]
modality = state["modality"]
local_path = state["local_path"]
resources: list = []
items: list = []
relations: list = []
category_updates: dict[str, list[tuple[str, str]]] = {}
user_scope = state.get("user", {})

for plan in state.get("resource_plans", []):
caption_text = (plan.get("caption") or "").strip() or None
res = store.resource_repo.create_resource(
url=plan["resource_url"],
modality=modality,
local_path=local_path,
caption=caption_text,
embedding=None,
user_data=dict(user_scope or {}),
)
resources.append(res)

entries = plan.get("entries") or []
if not entries:
continue

reinforce = _svc.memorize_config.enable_item_reinforcement
for memory_type, summary_text, cat_names in entries:
item = store.memory_item_repo.create_item(
resource_id=res.id,
memory_type=memory_type,
summary=summary_text,
embedding=None,
user_data=dict(user_scope or {}),
reinforce=reinforce,
)
items.append(item)
if reinforce and item.extra.get(
"reinforcement_count", 1,
) > 1:
continue
mapped = _svc._map_category_names_to_ids(
cat_names, svc_ctx,
)
for cid in mapped:
relations.append(
store.category_item_repo.link_item_category(
item.id, cid,
user_data=dict(user_scope or {}),
)
)
category_updates.setdefault(cid, []).append(
(item.id, summary_text),
)

state.update({
"resources": resources,
"items": items,
"relations": relations,
"category_updates": category_updates,
})
return state

self._service.replace_step(
target_step_id="categorize_items",
new_step=_WfStep(
step_id="categorize_items",
role="categorize",
handler=_categorize_no_embed,
requires={
"resource_plans", "ctx", "store",
"local_path", "modality", "user",
},
produces={
"resources", "items",
"relations", "category_updates",
},
capabilities={"db"},
),
pipeline="memorize",
)
logger.info(
"No embedding provider — replaced memorize categorize_items "
"step (embeddings disabled, using LLM-based recall)"
)

# Warm up the Anthropic LLM clients. The first HTTP request on
# a fresh AsyncOpenAI→httpx connection to Anthropic's Cloudflare
# endpoint often hangs indefinitely (HTTP/2 negotiation issue).
Expand Down Expand Up @@ -1859,10 +1968,15 @@ async def create_category(self, name: str, description: str, source: str = "brid
if not self._available or not self._service:
return False
try:
# Generate embedding for the category
embed_text = f"{name}: {description}" if description else name
vecs = await self._service._get_llm_client("embedding").embed([embed_text])
embedding = vecs[0]
# Generate embedding for the category (requires OpenAI key)
embedding = None
if self._has_embeddings:
try:
embed_text = f"{name}: {description}" if description else name
vecs = await self._service._get_llm_client("embedding").embed([embed_text])
embedding = vecs[0]
except Exception as e:
logger.warning("Could not embed category %s: %s", name, e)

# Create in the DB repo
cat = self._service.database.memory_category_repo.get_or_create_category(
Expand Down Expand Up @@ -1912,17 +2026,23 @@ async def update_category(
new_desc = description if description is not None else cat.description
new_summary = summary if summary is not None else cat.summary

# Re-embed from the updated text
embed_text = f"{new_name}: {new_desc}"
if new_summary:
embed_text += f"\n{new_summary}"
vecs = await self._service._get_llm_client("embedding").embed([embed_text])
# Re-embed from the updated text (requires OpenAI key)
embedding = None
if self._has_embeddings:
try:
embed_text = f"{new_name}: {new_desc}"
if new_summary:
embed_text += f"\n{new_summary}"
vecs = await self._service._get_llm_client("embedding").embed([embed_text])
embedding = vecs[0]
except Exception as e:
logger.warning("Could not re-embed category %s: %s", category_id, e)

repo.update_category(
category_id=category_id,
description=new_desc if description is not None else None,
summary=new_summary if summary is not None else None,
embedding=vecs[0],
embedding=embedding,
)
logger.info("Updated category: %s", category_id)
await self._audit("category_updated", "category", category_id, source, {
Expand Down Expand Up @@ -2097,3 +2217,8 @@ def metrics(self) -> MemUMetrics:
@property
def available(self) -> bool:
return self._available

@property
def _has_embeddings(self) -> bool:
"""Whether an embedding provider (e.g. OpenAI) is configured."""
return bool(self.config.openai_api_key)
2 changes: 1 addition & 1 deletion tests/test_memu_bridge.py
Original file line number Diff line number Diff line change
Expand Up @@ -311,7 +311,7 @@ def test_defaults(self):
assert config.recall_model == "claude-sonnet-4-6"
assert config.memorize_model == "claude-sonnet-4-6"
assert config.fast_model == "claude-haiku-4-5-20251001"
assert config.embed_model == "text-embedding-3-small"
assert config.embed_model == ""

def test_from_dict(self):
config = MemoryConfig.from_dict({
Expand Down