Skip to content

fix(smart-extractor): only learn noise from a genuine empty extraction; recover JSON from the reasoning field - #914

Merged
rwmjhb merged 4 commits into
CortexReach:masterfrom
gorkem2020:fix/noise-bank-only-learn-on-genuine-empty-list
Jul 10, 2026
Merged

rwmjhb merged 4 commits into
CortexReach:masterfrom
gorkem2020:fix/noise-bank-only-learn-on-genuine-empty-list

Conversation

@gorkem2020

Copy link
Copy Markdown
Contributor

Problem

extractCandidates returns an empty list for three different reasons: the LLM/gateway call failed (empty or blank response), the response had an unexpected shape, or the model genuinely returned an empty memories list. The caller treats all three the same and feeds the conversation to the noise-prototype bank as a "nothing worth remembering" signal. So a transient model or gateway failure teaches the bank that real conversation content is noise, which then suppresses legitimate future extraction and recall. On a run where every call fails identically (see #913, Qwen3 served via a gateway that returns empty content), the bank fills with false prototypes and accuracy degrades.

Changes

  1. extractCandidates now returns a discriminated result (ok / llm_failure / malformed), and learnAsNoise is called only on a genuine parsed empty list. Failures and malformed responses no longer poison the noise bank.
  2. When the model returns empty content, recover the JSON answer from the reasoning field before giving up. Checks both reasoning_content (OpenAI and DeepSeek convention) and reasoning (vLLM's Qwen3 parser) and runs it through the existing extractJsonFromResponse. This salvages the common case where a reasoning model in thinking mode routes its answer into the reasoning channel and leaves content empty.

Tests

  • New smart-extractor-noise-gating suite: null and malformed responses do not learn noise; a genuine empty list and validation-dropped candidates do.
  • Extended the api-key client tests to cover reasoning recovery from both field names.
  • Wired the new suite into the CI test manifest.

Refs #913.

extractCandidates() returned an empty array for three different reasons
(LLM/gateway failure, malformed response shape, or a genuinely empty
memories list), and the caller treated all three as the same "strongest
noise signal" and fed the conversation text into the noise-prototype
bank. A failing gateway or a malformed response would silently train the
bank on real conversational content, degrading future extraction/recall.

extractCandidates() now returns a discriminated result so the caller can
gate noise-bank learning on status === "ok" (a genuine parse, whether or
not any candidates survived validation), never on llm_failure or
malformed.
Some gateways (e.g. vLLM serving Qwen3 behind a custom endpoint) don't
fully honor chat_template_kwargs.enable_thinking:false, so the model's
answer lands only in a reasoning field while the regular content comes
back empty. completeJson() treated that identically to a hard failure
and returned null, which (before the previous commit) also mistrained
the noise bank.

Before giving up, the api-key client now checks for reasoning text
(message.reasoning_content / message.reasoning, covering both
OpenAI/DeepSeek and vLLM naming conventions) and tries to parse JSON
out of it. If that succeeds, it's used and logged distinctly;
otherwise it falls through to the original null return unchanged.
test/smart-extractor-noise-gating.test.mjs existed but wasn't run by
npm test or included in the CI group manifest, so the noise-bank
gating regression it covers wasn't actually enforced anywhere.

@rwmjhb rwmjhb left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approved on head 8ca3844. Orchestrator verdict: approve. Note: adversarial R4 returned invalid JSON twice, so orchestrator confidence was reduced; I completed independent verification on the same head before approving.

Independent verification run:

  • npm ci --silent
  • npm run build --if-present
  • git diff --exit-code -- dist src index.ts package.json package-lock.json openclaw.plugin.json README.md scripts test
  • node --test test/smart-extractor-noise-gating.test.mjs
  • node --test test/llm-api-key-client.test.mjs
  • node test/plugin-manifest-regression.mjs
  • npm test
  • git diff --exit-code -- dist src index.ts package.json package-lock.json openclaw.plugin.json README.md scripts test

All passed; build left committed dist/source artifacts clean.

Non-blocking follow-up: consider distinguishing a truly empty raw memories: [] response from a non-empty raw memories array where every candidate is dropped by local validation. The latter still trains the noise bank today, which may be worth revisiting if malformed per-candidate output can represent real memory content rather than true noise.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants