Skip to content

feat(engine): render backtest score and comparison reports as Markdown - #8116

Merged
loopover-orb[bot] merged 1 commit into
JSONbored:mainfrom
RealDiligent:fix/critical-issue-backtest-report-8088-v2
Jul 22, 2026
Merged

feat(engine): render backtest score and comparison reports as Markdown#8116
loopover-orb[bot] merged 1 commit into
JSONbored:mainfrom
RealDiligent:fix/critical-issue-backtest-report-8088-v2

Conversation

@RealDiligent

Copy link
Copy Markdown
Contributor

Summary

  • Closes calibration: render a backtest score/comparison report as Markdown #8088
  • BacktestScoreReport (calibration: pure confusion-matrix scorer for a candidate rule classifier against the backtest corpus #8085) and BacktestComparison (calibration: Pareto-floor comparator between two BacktestScoreReports #8086) are plain data with no human-readable rendering. Adds packages/loopover-engine/src/calibration/backtest-report.ts: renderBacktestScoreReport (a Markdown table of the rule ID, case count, all four confusion-matrix counts, and precision/recall) and renderBacktestComparison (regressed axes under a Regressed heading, improved axes under a visually separate Improved heading — an empty section is omitted entirely, so nothing ever reads as regressed when it isn't — plus a closing verdict line).
  • Null precision/recall render as the literal N/A — never 0, the word null, or an empty cell — mirroring the reports' own null-is-not-zero discipline. The regressed closing line is exactly Verdict: REGRESSED — do not merge. (containing the literal case-sensitive word REGRESSED and a do-not-merge statement), so the follow-up CI wiring can detect the regressed case by string match without re-implementing the comparison logic. Both functions are pure (string in, string out; no IO, no wall-clock reads) and byte-identical for byte-identical input.
  • Barrel export added on its own line immediately after the backtest-compare.js line in packages/loopover-engine/src/index.ts, exactly where the issue requires.

Scope

  • The PR title follows type(scope): short summary Conventional Commit format, for example fix(api): restore profile access checks.
  • This PR is focused and does not mix unrelated backend, UI, MCP, docs, dependency, and deploy changes.
  • This follows CONTRIBUTING.md and does not reintroduce GitHub Pages, VitePress, site/, or CNAME.
  • I linked a currently open issue this PR resolves (e.g. Closes #123) — a linked open issue is required for every contributor PR.

Validation

  • git diff --check
  • npm run typecheck
  • npm run actionlint
  • npm run test:coverage locally; codecov/patch requires ≥99% coverage of the lines AND branches you changed (aim for 100% on your diff so CI variance does not fail near the threshold). Global coverage is a non-blocking trend with a loose 90% backstop, not the gate.
  • npm run test:workers
  • npm run build:mcp
  • npm run test:mcp-pack
  • npm run ui:openapi:check
  • npm run ui:lint
  • npm run ui:typecheck
  • npm run ui:build
  • npm audit --audit-level=moderate
  • New or changed behavior has unit/integration tests for new branches, fallback paths, and sanitizer boundaries

If any required check was skipped, explain why:

  • Ran everything this change can affect: the engine deliverable suite via npm run build && npm run test --workspace @loopover/engine (all green, including the new backtest-report.test.ts), the root test/unit/backtest-report.test.ts (8 tests green), and the full root npm run typecheck. Verified per-diff-line coverage via lcov: every changed line and branch in backtest-report.ts is covered by the root suite — the snapshot-exact table render, N/A for both null axes, the literal REGRESSED + do-not-merge line with section ordering asserted, improved-only with no regressed claim, unchanged with neither section, and byte-identical determinism for both renderers — and simulated the scoped-CI shard condition with the exact CI invocation (--changed=origin/main --coverage.all=false): the lcov is non-empty and contains the changed instrumented file. actionlint/workers/mcp/ui checks are untouched surfaces; CI runs them all.

Safety

  • No secrets, wallet details, hotkeys, coldkeys, user PATs, private keys, raw trust scores, private rankings, or private maintainer evidence are exposed.
  • Public GitHub text stays sanitized, low-noise, and does not imply compensation guarantees or optimization tactics.
  • Auth, cookie, CORS, GitHub App, Cloudflare, or session changes include negative-path tests.
  • API/OpenAPI/MCP behavior is updated and tested where needed.
  • UI changes use live API data or real empty/error/loading states, not production mock/demo fallbacks.
  • Visible UI changes include a UI Evidence section below with JPG/JPEG or PNG screenshots arranged as organized, captioned, clickable thumbnails. SVG screenshots are not used as review evidence. Review-only screenshots or recordings are not committed to the repository.
  • Public docs/changelogs are updated where needed; changelogs are only edited for release-prep PRs.

UI Evidence

Not applicable — engine-only pure-function addition (no UI, docs, or extension surface touched).

Notes

  • Tests intentionally land in BOTH suites: packages/loopover-engine/test/backtest-report.test.ts is the issue's deliverable (node:test against dist/), and test/unit/backtest-report.test.ts imports the source directly so the repo's Codecov patch gate measures every changed line and branch — the engine workspace's own node --test run does not feed Codecov.

JSONbored#8088)

BacktestScoreReport (JSONbored#8085) and BacktestComparison (JSONbored#8086) are plain data with
no human-readable rendering. Add calibration/backtest-report.ts:
renderBacktestScoreReport (a Markdown table of the rule ID, case count, all
four confusion-matrix counts, and precision/recall) and
renderBacktestComparison (regressed axes under a Regressed heading, improved
axes under a visually separate Improved heading -- an empty section is omitted
entirely so nothing ever reads as regressed when it isn't -- plus a closing
verdict line). Null precision/recall render as the literal N/A, never 0 or the
word null, mirroring the reports' own null-is-not-zero discipline. The
regressed closing line is exactly "Verdict: REGRESSED — do not merge." so the
follow-up CI wiring can detect it by string match without re-implementing the
comparison. Both functions are pure and byte-identical for identical input.
Barrel export added directly after the backtest-compare line, per the issue's
placement requirement.

Tests in both suites (engine node:test deliverable + root vitest for the
coverage gate): snapshot-exact table render, N/A for both null axes, the
literal REGRESSED + do-not-merge line with section ordering asserted,
improved-only with no regressed claim, unchanged with neither section, and
byte-identical determinism for both renderers.
@codecov

codecov Bot commented Jul 22, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 92.01%. Comparing base (b51103f) to head (2d7d3d2).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main    #8116      +/-   ##
==========================================
- Coverage   92.02%   92.01%   -0.01%     
==========================================
  Files         759      760       +1     
  Lines       77328    77348      +20     
  Branches    23376    23382       +6     
==========================================
+ Hits        71159    71175      +16     
  Misses       5061     5061              
- Partials     1108     1112       +4     
Flag Coverage Δ
shard-1 57.40% <0.00%> (+3.42%) ⬆️
shard-2 51.09% <0.00%> (-3.72%) ⬇️
shard-3 54.46% <100.00%> (+0.59%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
...loopover-engine/src/calibration/backtest-report.ts 100.00% <100.00%> (ø)

... and 1 file with indirect coverage changes

@loopover-orb loopover-orb Bot added the gittensor:feature Gittensor-scored feature linked to a feature issue — scores a 0.25x multiplier. label Jul 22, 2026
@loopover-orb

loopover-orb Bot commented Jul 22, 2026

Copy link
Copy Markdown
Contributor

Tip

✅ LoopOver review result - approve/merge recommended

Review updated: 2026-07-22 23:32:57 UTC

4 files · 1 AI reviewer · no blockers · readiness 100/100 · CI green · clean

✅ Suggested Action - Approve/Merge

  • safe to merge

Review summary
This PR adds two pure Markdown-rendering functions for BacktestScoreReport and BacktestComparison, plus a barrel export, matching the description closely. The implementation is straightforward string-building with correct null-handling (N/A), correct section omission logic, and the exact verdict wording claimed. Tests are duplicated across two harnesses (vitest under test/unit and node:test under packages/loopover-engine/test using dist/index.js), which is consistent with this repo's dual-test convention seen elsewhere in the codebase, and both suites exercise the real exported functions rather than fabricated payloads.

Nits — 5 non-blocking
  • The external brief flags a 'magic number' at backtest-report.ts:13, but that's just an issue number in a comment (calibration: pure confusion-matrix scorer for a candidate rule classifier against the backtest corpus #8085), not a literal used in logic — not an actual concern.
  • packages/loopover-engine/src/calibration/backtest-report.ts: renderBacktestComparison's axis lines index into comparison.baseline[axis]/candidate[axis] without any guard if an axis key ever mismatches the report's shape, though this is only reachable if BacktestComparison's producer supplies an invalid axis name.
  • The node:test file imports from '../dist/index.js' rather than the source, so it depends on a fresh build being present before running — verify the test pipeline always rebuilds first.
  • Consider consolidating the near-identical vitest and node:test suites into shared fixture builders if this duplication pattern isn't already established elsewhere in the repo.
  • packages/loopover-engine/src/calibration/backtest-report.ts:44-58 could extract the repeated axis-line-rendering loop (used for both Regressed and Improved) into a small local helper to avoid the duplicated `for` block, though this is optional given the small size.

Decision drivers

  • ✅ Code review — No blockers (1 reviewer)
  • ✅ Gate result — Passing (No configured blocker found.)
Context & advisory signals — never blocks the verdict
Signal Result Evidence
Linked issue ✅ Linked #8088
Related work ✅ No active overlap found No same-issue or scoped active PR overlap found.
Change scope ✅ 20/20 Low review scope from cached public metadata (1 linked issue).
Validation posture ✅ 25/25 PR body includes validation/test evidence.
Contributor workload ✅ 10/10 Author activity: 361 registered-repo PR(s), 144 merged, 38 issue(s).
Contributor context ✅ Confirmed Gittensor contributor RealDiligent; Gittensor profile; 361 PR(s), 38 issue(s).
Improvement ✅ Minor risk: clean · value: minor · LLM: moderate
Linked issue satisfaction

Addressed
The diff adds backtest-report.ts with both required functions, correctly renders null precision/recall as literal N/A, produces the exact 'Verdict: REGRESSED — do not merge.' wording, cleanly separates Regressed/Improved sections (omitting empty ones), and adds the barrel export on its own line immediately after the backtest-compare.js export as required. Tests cover all specified cases including

Review context
  • Author: RealDiligent
  • Role context: outside_contributor
  • Public audience mode: oss maintainer
  • Lane context: Repository is configured for direct PR review.
  • Public profile languages: Python, Ruby, JavaScript, Svelte, TypeScript, Markdown, MDX, Rust
  • Official Gittensor activity: 361 PR(s), 38 issue(s).
  • PR-specific overlap: none found.
Contributor next steps
  • Keep the PR focused and include validation evidence before maintainer review.
Signal definitions
  • Related work = same linked issue, overlapping active PRs, or title/path similarity.
  • Change scope = cached public metadata such as size labels, draft state, and review-burden hints.
  • Validation posture = whether the PR provides enough public validation/test evidence for maintainer review.
  • Contributor workload = public contributor activity and cleanup pressure, not a repo-wide quality failure.
  • Contributor context = public GitHub/Gittensor identity context; non-Gittensor status is not a blocker.
🧪 Chat with LoopOver

Ask LoopOver a question about this PR directly in a comment — grounded only in the same cached, public-safe facts shown above, never a new claim.

  • @loopover ask &lt;question&gt; answers contribution-quality Q&A with source citations and freshness.
  • @loopover chat &lt;question&gt; answers in natural prose from cached decision-pack facts via local inference (maintainer/collaborator; read-only).
  • A plain-language @loopover mention with a real question is routed to the closest matching read-only command automatically — no exact syntax required.

Full command reference: https://loopover.ai/docs/loopover-commands

🧪 Experimental — new and may change.

🟩 Safe / merged · 🟦 Advisory · 🟨 Held for review · 🟥 Blocked / closed


💰 Earn for open-source contributions like this. Gittensor lets GitHub contributors earn for the work they already do — register to start earning →.

Checked by LoopOver, a quiet PR intelligence layer for OSS maintainers.

  • Re-run LoopOver review

@loopover-orb loopover-orb Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LoopOver approves — the gate is satisfied and CI is green.

@loopover-orb
loopover-orb Bot merged commit 5553138 into JSONbored:main Jul 22, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

gittensor:feature Gittensor-scored feature linked to a feature issue — scores a 0.25x multiplier.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

calibration: render a backtest score/comparison report as Markdown

1 participant