Repository navigation
fix(smart-extractor): only learn noise from a genuine empty extraction; recover JSON from the reasoning field - #914
Merged
rwmjhb merged 4 commits intoJul 10, 2026
Conversation
extractCandidates() returned an empty array for three different reasons (LLM/gateway failure, malformed response shape, or a genuinely empty memories list), and the caller treated all three as the same "strongest noise signal" and fed the conversation text into the noise-prototype bank. A failing gateway or a malformed response would silently train the bank on real conversational content, degrading future extraction/recall. extractCandidates() now returns a discriminated result so the caller can gate noise-bank learning on status === "ok" (a genuine parse, whether or not any candidates survived validation), never on llm_failure or malformed.
Some gateways (e.g. vLLM serving Qwen3 behind a custom endpoint) don't fully honor chat_template_kwargs.enable_thinking:false, so the model's answer lands only in a reasoning field while the regular content comes back empty. completeJson() treated that identically to a hard failure and returned null, which (before the previous commit) also mistrained the noise bank. Before giving up, the api-key client now checks for reasoning text (message.reasoning_content / message.reasoning, covering both OpenAI/DeepSeek and vLLM naming conventions) and tries to parse JSON out of it. If that succeeds, it's used and logged distinctly; otherwise it falls through to the original null return unchanged.
test/smart-extractor-noise-gating.test.mjs existed but wasn't run by npm test or included in the CI group manifest, so the noise-bank gating regression it covers wasn't actually enforced anywhere.
rwmjhb
approved these changes
Jul 9, 2026
rwmjhb
left a comment
Collaborator
There was a problem hiding this comment.
Approved on head 8ca3844. Orchestrator verdict: approve. Note: adversarial R4 returned invalid JSON twice, so orchestrator confidence was reduced; I completed independent verification on the same head before approving.
Independent verification run:
- npm ci --silent
- npm run build --if-present
- git diff --exit-code -- dist src index.ts package.json package-lock.json openclaw.plugin.json README.md scripts test
- node --test test/smart-extractor-noise-gating.test.mjs
- node --test test/llm-api-key-client.test.mjs
- node test/plugin-manifest-regression.mjs
- npm test
- git diff --exit-code -- dist src index.ts package.json package-lock.json openclaw.plugin.json README.md scripts test
All passed; build left committed dist/source artifacts clean.
Non-blocking follow-up: consider distinguishing a truly empty raw memories: [] response from a non-empty raw memories array where every candidate is dropped by local validation. The latter still trains the noise bank today, which may be worth revisiting if malformed per-candidate output can represent real memory content rather than true noise.
This was referenced Jul 10, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
extractCandidatesreturns an empty list for three different reasons: the LLM/gateway call failed (empty or blank response), the response had an unexpected shape, or the model genuinely returned an emptymemorieslist. The caller treats all three the same and feeds the conversation to the noise-prototype bank as a "nothing worth remembering" signal. So a transient model or gateway failure teaches the bank that real conversation content is noise, which then suppresses legitimate future extraction and recall. On a run where every call fails identically (see #913, Qwen3 served via a gateway that returns emptycontent), the bank fills with false prototypes and accuracy degrades.Changes
extractCandidatesnow returns a discriminated result (ok/llm_failure/malformed), andlearnAsNoiseis called only on a genuine parsed empty list. Failures and malformed responses no longer poison the noise bank.content, recover the JSON answer from the reasoning field before giving up. Checks bothreasoning_content(OpenAI and DeepSeek convention) andreasoning(vLLM's Qwen3 parser) and runs it through the existingextractJsonFromResponse. This salvages the common case where a reasoning model in thinking mode routes its answer into the reasoning channel and leavescontentempty.Tests
smart-extractor-noise-gatingsuite: null and malformed responses do not learn noise; a genuine empty list and validation-dropped candidates do.Refs #913.