diff --git a/README.md b/README.md index b78002de9..cfd86aa5a 100644 --- a/README.md +++ b/README.md @@ -146,7 +146,7 @@ Two layers, one seamless experience. Curated facts injected into every system prompt. Active projects, current deadlines, operational lessons. Tagged with dates, automatically evicted when stale. **L2 — Deep Memory (memU)** -Semantic search over everything — conversations, facts, preferences, events. SQLite-persisted with `text-embedding-3-small` embeddings. +Semantic search over everything — conversations, facts, preferences, events. SQLite-persisted. Uses vector embeddings when an OpenAI key is configured, or LLM-based ranking with Anthropic models only. - Four memory types: `profile`, `event`, `knowledge`, `behavior` - Automatic conversation indexing on session close @@ -254,7 +254,7 @@ See [docs/config.md](docs/config.md) for all options. - [Node.js](https://nodejs.org/) 18+ (for web UI build) - [Claude Code CLI](https://docs.anthropic.com/en/docs/agents-and-tools/claude-code/overview) (bundled with `claude-agent-sdk`) - Anthropic API key **or** Claude subscription via [CLIProxyAPI](https://github.com/router-for-me/CLIProxyAPI) proxy -- Optional: OpenAI API key (for memU embeddings), Telegram bot token, [gog](https://github.com/googleworkspace/cli) CLI, [gh](https://cli.github.com/) CLI +- Optional: OpenAI API key (for vector-based memory search — without it, LLM-based recall is used), Telegram bot token, [gog](https://github.com/googleworkspace/cli) CLI, [gh](https://cli.github.com/) CLI ## Documentation diff --git a/config.example.yaml b/config.example.yaml index e06974b01..bd9c98c34 100644 --- a/config.example.yaml +++ b/config.example.yaml @@ -46,7 +46,7 @@ sync: # Memory (memU) memory: chat_model: claude-sonnet-4-6 - embed_model: text-embedding-3-small + # embed_model: text-embedding-3-small # Only needed with openai_api_key # Cron cron: diff --git a/docs/config.md b/docs/config.md index 1f6166e97..d8e8df642 100644 --- a/docs/config.md +++ b/docs/config.md @@ -99,7 +99,7 @@ Sources pull data from external services on a schedule. See [sources.md](sources | `memory.recall_model` | string | `claude-sonnet-4-6` | Model for recall routing | | `memory.memorize_model` | string | `claude-sonnet-4-6` | Model for extraction & preprocessing | | `memory.fast_model` | string | `claude-haiku-4-5-20251001` | Model for categorization, date resolution, knowledge filtering | -| `memory.embed_model` | string | `text-embedding-3-small` | Embedding model | +| `memory.embed_model` | string | *(empty)* | Embedding model (only used when `openai_api_key` is set, e.g. `text-embedding-3-small`) | | `memory.semantic_dedup_threshold` | float | `0.85` | Cosine similarity threshold for semantic deduplication (0 to disable) | | `memory.knowledge_filter` | bool | `false` | Post-extraction LLM filter that deletes generic knowledge items (extra Haiku API call per memorize) | | `memory.categories` | list | `[]` | Seed categories — each entry has `name` and `description` fields. Used for semantic routing when memorizing and recalling facts. `nerve init` populates mode-appropriate defaults (personal: relationships, finances, health, etc.; worker: patterns, procedures, approvals, etc.). | @@ -174,7 +174,7 @@ Nerve automatically discovers MCP servers from Claude Code's enabled plugins. An | Key | Type | Description | |-----|------|-------------| | `anthropic_api_key` | string | Anthropic API key (agent + memU chat). Not required when proxy is enabled. | -| `openai_api_key` | string | OpenAI API key (memU embeddings only) | +| `openai_api_key` | string | OpenAI API key (optional — enables vector-based memory search via embeddings; without it, LLM-based recall is used) | | `brave_search_api_key` | string | Brave Search API key (optional) | ## Proxy (CLIProxyAPI) diff --git a/docs/memory.md b/docs/memory.md index 36e2053e9..f85ac2dea 100644 --- a/docs/memory.md +++ b/docs/memory.md @@ -78,17 +78,19 @@ Reinforced items rank higher in search results via salience-aware ranking: `simi ### Configuration -memU uses three LLM profiles: +memU uses two or three LLM profiles depending on configuration: - **Chat** — Anthropic API for recall routing (claude-sonnet-4-6) - **Fast** — Anthropic API for fact extraction and categorization (claude-haiku-4-5) -- **Embedding** — OpenAI text-embedding-3-small for vector search +- **Embedding** *(optional)* — OpenAI text-embedding-3-small for vector search. Only active when `openai_api_key` is set. + +When no OpenAI key is configured, memU uses **LLM-based recall** instead of vector search — the Chat profile ranks memories directly, requiring no embeddings. This uses more Anthropic API tokens per recall but removes the OpenAI dependency entirely. Config in `config.yaml`: ```yaml memory: chat_model: claude-sonnet-4-6 # recall routing fast_model: claude-haiku-4-5-20251001 # extraction & categorization - embed_model: text-embedding-3-small + # embed_model: text-embedding-3-small # only needed with openai_api_key categories: [...] # see Categories section above ``` @@ -102,10 +104,11 @@ This uses the `anthropic` Python SDK directly (not the OpenAI-compatible endpoin ### Performance Optimizations -- **Vector cache** — All item and category embeddings are preloaded into memory at startup, eliminating repeated SQLite JSON parsing (~2s per search saved) +- **Vector cache** *(with OpenAI key)* — All item and category embeddings are preloaded into memory at startup, eliminating repeated SQLite JSON parsing (~2s per search saved) - **Fast model** — Extraction and category summary updates use Haiku instead of Sonnet - **Disabled pipeline steps** — Route intention, sufficiency checks, and resource retrieval are disabled in the retrieve pipeline (saves 3+ LLM calls per recall) -- **Category embedding reuse** — Category ranking uses stored embeddings instead of re-embedding summaries on every recall +- **Category embedding reuse** *(with OpenAI key)* — Category ranking uses stored embeddings instead of re-embedding summaries on every recall +- **LLM-based fallback** — When no embedding provider is configured, retrieval and memorization work without embeddings; semantic deduplication falls back to content-hash only - **Client warmup** — Anthropic LLM clients are pinged during startup to force HTTP/2 connection establishment (avoids a cold-start hang on the first memorize call) - **Memorize timeout** — Each memorize call is capped at 300s; if it hangs, it is cancelled, LLM clients are evicted (cache cleared + HTTP transport closed), a fresh client is created, and one retry is attempted - **Per-call LLM timeout** — Base LLM client `.chat()` methods are wrapped with a 120s `asyncio.wait_for()` at init time (instance attribute shadowing). A single dead HTTP/2 connection fails fast instead of consuming the entire 300s pipeline budget. Hung calls log `memU LLM HUNG [profile]: no response after 120s (prompt=N chars)` @@ -149,7 +152,7 @@ Delete a memory item. Use for wrong, duplicate, or stale memories. ``` category_update(category_id="abc123", summary="Updated summary", description="New description") ``` -Update a category's summary and/or description. Re-embeds the category after update to keep vector search in sync. +Update a category's summary and/or description. Re-embeds the category after update to keep vector search in sync (when an embedding provider is configured). All recall and history results include memory IDs (`id:abc123...`), enabling the agent to target specific items for update or deletion. diff --git a/docs/setup.md b/docs/setup.md index 5f23119f7..4c4e93553 100644 --- a/docs/setup.md +++ b/docs/setup.md @@ -131,7 +131,7 @@ The wizard handles all of this automatically, but you can also configure manuall # Create secrets file (gitignored) cat > config.local.yaml << 'EOF' anthropic_api_key: sk-ant-... -openai_api_key: sk-... # For memU embeddings (optional) +openai_api_key: sk-... # Optional — enables vector-based memory search telegram: bot_token: "123456:ABC..." diff --git a/nerve/bootstrap.py b/nerve/bootstrap.py index 9de169442..cfb0cd945 100644 --- a/nerve/bootstrap.py +++ b/nerve/bootstrap.py @@ -673,9 +673,10 @@ def _prompt_openai_key(self) -> None: """Prompt for optional OpenAI API key (used by both auth paths).""" click.echo() click.secho( - "Optionally, an OpenAI key enables better memory search\n" - "(text-embedding-3-small for vector embeddings). Nerve works\n" - "without it but recall quality improves significantly.", + "Optionally, an OpenAI key enables vector-based memory search\n" + "(text-embedding-3-small for semantic embeddings). Nerve works\n" + "without it using LLM-based recall, which uses more API tokens\n" + "per query but requires no additional API key.", dim=True, ) click.echo() diff --git a/nerve/cli.py b/nerve/cli.py index 84251ffc4..9cc96cab3 100644 --- a/nerve/cli.py +++ b/nerve/cli.py @@ -603,9 +603,9 @@ def doctor_report(config) -> str: errors.append("[ERR] Anthropic API key not set and proxy not enabled (config.local.yaml)") if config.openai_api_key: - lines.append(f"[OK] OpenAI API key: ...{config.openai_api_key[-4:]}") + lines.append(f"[OK] OpenAI API key: ...{config.openai_api_key[-4:]} (vector embeddings enabled)") else: - warnings.append("[WARN] OpenAI API key not set (needed for memU embeddings)") + lines.append("[--] OpenAI API key not set (using LLM-based memory recall)") # Check Telegram if config.telegram.enabled: diff --git a/nerve/config.py b/nerve/config.py index 27efd6ff5..d4fdb961b 100644 --- a/nerve/config.py +++ b/nerve/config.py @@ -251,7 +251,7 @@ class MemoryConfig: recall_model: str = "claude-sonnet-4-6" # Recall routing memorize_model: str = "claude-sonnet-4-6" # Extraction & preprocessing fast_model: str = "claude-haiku-4-5-20251001" # Category summaries, date resolution - embed_model: str = "text-embedding-3-small" + embed_model: str = "" sqlite_dsn: str = "" semantic_dedup_threshold: float = 0.85 # Cosine similarity threshold for semantic dedup knowledge_filter: bool = False # Post-extraction LLM filter for generic knowledge (extra API call) @@ -266,7 +266,7 @@ def from_dict(cls, d: dict) -> MemoryConfig: recall_model=d.get("recall_model", "claude-sonnet-4-6"), memorize_model=d.get("memorize_model", "claude-sonnet-4-6"), fast_model=d.get("fast_model", "claude-haiku-4-5-20251001"), - embed_model=d.get("embed_model", "text-embedding-3-small"), + embed_model=d.get("embed_model", ""), sqlite_dsn=d.get("sqlite_dsn", default_dsn), semantic_dedup_threshold=float(d.get("semantic_dedup_threshold", 0.85)), knowledge_filter=bool(d.get("knowledge_filter", False)), diff --git a/nerve/memory/memu_bridge.py b/nerve/memory/memu_bridge.py index b3170a0f8..4e2330fdb 100644 --- a/nerve/memory/memu_bridge.py +++ b/nerve/memory/memu_bridge.py @@ -829,7 +829,13 @@ async def initialize(self) -> bool: # conversations as instructions, generating long off-topic # responses that exceed the per-call timeout. "preprocess_llm_profile": fast_profile, - "memory_extract_llm_profile": memorize_profile, + # When no embedding provider is configured, all memory + # work goes through Anthropic — use Haiku for extraction + # too to avoid saturating the rate-limit budget. + "memory_extract_llm_profile": ( + fast_profile if not self.config.openai_api_key + else memorize_profile + ), "category_update_llm_profile": fast_profile, # Pass Nerve's configured categories to memU so the LLM # prompt only shows categories that actually exist in the @@ -855,9 +861,14 @@ async def initialize(self) -> bool: }, }, retrieve_config={ + "method": "llm" if not self.config.openai_api_key else "rag", "route_intention": False, "sufficiency_check": False, "resource": {"enabled": False}, + # Use Haiku for LLM-based ranking — cheaper and avoids + # sharing Sonnet's rate-limit budget with the main agent. + **({"llm_ranking_llm_profile": fast_profile} + if not self.config.openai_api_key else {}), }, ) self._available = True @@ -888,6 +899,104 @@ async def initialize(self) -> bool: len(ctx.category_name_to_id), ) + # When no embedding provider is configured, replace the + # memorize pipeline's "categorize_items" step with one that + # stores items and resources with embedding=None. This + # avoids KeyError on the missing "embedding" LLM profile. + if not self.config.openai_api_key: + from memu.workflow.step import WorkflowStep as _WfStep + + _svc = self._service + + async def _categorize_no_embed( + state: dict, step_context: Any, + ) -> dict: + svc_ctx = state["ctx"] + store = state["store"] + modality = state["modality"] + local_path = state["local_path"] + resources: list = [] + items: list = [] + relations: list = [] + category_updates: dict[str, list[tuple[str, str]]] = {} + user_scope = state.get("user", {}) + + for plan in state.get("resource_plans", []): + caption_text = (plan.get("caption") or "").strip() or None + res = store.resource_repo.create_resource( + url=plan["resource_url"], + modality=modality, + local_path=local_path, + caption=caption_text, + embedding=None, + user_data=dict(user_scope or {}), + ) + resources.append(res) + + entries = plan.get("entries") or [] + if not entries: + continue + + reinforce = _svc.memorize_config.enable_item_reinforcement + for memory_type, summary_text, cat_names in entries: + item = store.memory_item_repo.create_item( + resource_id=res.id, + memory_type=memory_type, + summary=summary_text, + embedding=None, + user_data=dict(user_scope or {}), + reinforce=reinforce, + ) + items.append(item) + if reinforce and item.extra.get( + "reinforcement_count", 1, + ) > 1: + continue + mapped = _svc._map_category_names_to_ids( + cat_names, svc_ctx, + ) + for cid in mapped: + relations.append( + store.category_item_repo.link_item_category( + item.id, cid, + user_data=dict(user_scope or {}), + ) + ) + category_updates.setdefault(cid, []).append( + (item.id, summary_text), + ) + + state.update({ + "resources": resources, + "items": items, + "relations": relations, + "category_updates": category_updates, + }) + return state + + self._service.replace_step( + target_step_id="categorize_items", + new_step=_WfStep( + step_id="categorize_items", + role="categorize", + handler=_categorize_no_embed, + requires={ + "resource_plans", "ctx", "store", + "local_path", "modality", "user", + }, + produces={ + "resources", "items", + "relations", "category_updates", + }, + capabilities={"db"}, + ), + pipeline="memorize", + ) + logger.info( + "No embedding provider — replaced memorize categorize_items " + "step (embeddings disabled, using LLM-based recall)" + ) + # Warm up the Anthropic LLM clients. The first HTTP request on # a fresh AsyncOpenAI→httpx connection to Anthropic's Cloudflare # endpoint often hangs indefinitely (HTTP/2 negotiation issue). @@ -1859,10 +1968,15 @@ async def create_category(self, name: str, description: str, source: str = "brid if not self._available or not self._service: return False try: - # Generate embedding for the category - embed_text = f"{name}: {description}" if description else name - vecs = await self._service._get_llm_client("embedding").embed([embed_text]) - embedding = vecs[0] + # Generate embedding for the category (requires OpenAI key) + embedding = None + if self._has_embeddings: + try: + embed_text = f"{name}: {description}" if description else name + vecs = await self._service._get_llm_client("embedding").embed([embed_text]) + embedding = vecs[0] + except Exception as e: + logger.warning("Could not embed category %s: %s", name, e) # Create in the DB repo cat = self._service.database.memory_category_repo.get_or_create_category( @@ -1912,17 +2026,23 @@ async def update_category( new_desc = description if description is not None else cat.description new_summary = summary if summary is not None else cat.summary - # Re-embed from the updated text - embed_text = f"{new_name}: {new_desc}" - if new_summary: - embed_text += f"\n{new_summary}" - vecs = await self._service._get_llm_client("embedding").embed([embed_text]) + # Re-embed from the updated text (requires OpenAI key) + embedding = None + if self._has_embeddings: + try: + embed_text = f"{new_name}: {new_desc}" + if new_summary: + embed_text += f"\n{new_summary}" + vecs = await self._service._get_llm_client("embedding").embed([embed_text]) + embedding = vecs[0] + except Exception as e: + logger.warning("Could not re-embed category %s: %s", category_id, e) repo.update_category( category_id=category_id, description=new_desc if description is not None else None, summary=new_summary if summary is not None else None, - embedding=vecs[0], + embedding=embedding, ) logger.info("Updated category: %s", category_id) await self._audit("category_updated", "category", category_id, source, { @@ -2097,3 +2217,8 @@ def metrics(self) -> MemUMetrics: @property def available(self) -> bool: return self._available + + @property + def _has_embeddings(self) -> bool: + """Whether an embedding provider (e.g. OpenAI) is configured.""" + return bool(self.config.openai_api_key) diff --git a/tests/test_memu_bridge.py b/tests/test_memu_bridge.py index 5426b13fc..f848ca5b6 100644 --- a/tests/test_memu_bridge.py +++ b/tests/test_memu_bridge.py @@ -311,7 +311,7 @@ def test_defaults(self): assert config.recall_model == "claude-sonnet-4-6" assert config.memorize_model == "claude-sonnet-4-6" assert config.fast_model == "claude-haiku-4-5-20251001" - assert config.embed_model == "text-embedding-3-small" + assert config.embed_model == "" def test_from_dict(self): config = MemoryConfig.from_dict({