Skip to content

R5 Batch 1: engineering refactor — server/engine/frontend/cli modularization - #30

Merged
LeoLin990405 merged 25 commits into
mainfrom
refactor/r5-modularization
Jul 29, 2026
Merged

LeoLin990405 merged 25 commits into
mainfrom
refactor/r5-modularization

Conversation

@LeoLin990405

Copy link
Copy Markdown
Owner

Stacked on #29. Pure engineering refactor — no new features, no behavior change. Built in parallel by four domestic-model agents in isolated worktrees (path-disjoint ownership), verified on the home cluster.

What changed

server — every route module now uses the createXxxRouter({rootDir}) factory shape that tournaments.mjs already had, so tests inject a temp state dir instead of depending on the real ~/.civagent. GET /api/regimes no longer re-reads all 57 regimes' IDENTITY.md/SOUL.md/skills on every request (server/services/regimes.mjs, mtime-invalidated cache) and gains an additive ?summary=1 mode. Repeated try/catch + error-JSON collapsed into server/http.mjs. New route tests for matches/history/regimes.

engine — tournament.mjs 710 → 462 lines, split into judge-rubric.mjs (rubric prompt, both score parsers, pass aggregation) and deterministic-grading.mjs (T3 spec loader, grader runner, score mixer). All 17 previously-exported symbols still resolve from tournament.mjs.

frontend — deleted four components that were bundled but imported by nothing (CodexBrowser, JudgeLeaderboard, TerminalPanel, old RegimeBrowser). HistoryExplorer had real backing endpoints so it is revived behind a new MatchArchive component and sidebar tab. React-Compiler-readiness warnings 15 → 6 (0 errors).

cli — civagent list spawned four python3 processes per regime to read one metadata.json (228 for 57 regimes). Now one read per regime; python call sites 13 → 3, civagent list 4.1s → 1.2s. Shared helpers extracted to bin/lib/common.sh.

Verification (beyond CI)

Refactors claim "no behavior change", so that claim was tested directly rather than trusted:

  • HTTP contract: baseline and refactored servers booted side by side; all 15 endpoints byte-identical, including the 400/404 path-traversal rejections.
  • Engine split is a pure move: the six relocated function bodies are byte-identical to their originals after whitespace/comment normalization.
  • CLI output: list/modes/info china/tang/info global/athens/info china/qin/skills byte-identical; regime-id whitelisting still rejects injection and traversal.
  • No dangling imports after the component deletions; MatchArchive hits endpoints that return 200 with real local data.

Test plan

  • npm run ci — 268 tests (was 252), 57 regimes validated
  • tsc -b, frontend eslint 0 errors, vite build
  • Home-cluster CI on three workers, all bound to commit e5c945e: leo-03 test+cli-smoke, leo-01 lint+syntax+regimes, leo-02 frontend types+lint+build

…ding modules

judge-rubric.mjs takes the rubric prompt, both score parsers and the pass
aggregator; deterministic-grading.mjs takes the T3 spec loader, grader runner
and score mixer. tournament.mjs keeps orchestration + judge() + the CLI and
re-exports every symbol it exported before, so existing imports are unchanged.

Pure move: the six relocated function bodies are byte-identical to their
originals; all 17 previously-exported symbols still resolve.
…or helper

All route modules follow the createXxxRouter({rootDir}) factory shape that
tournaments.mjs already used, so tests can inject a temp state dir instead of
depending on the real ~/.civagent. GET /api/regimes no longer re-reads every
IDENTITY.md/SOUL.md/skills file on each request (server/services/regimes.mjs,
mtime-invalidated cache) and gains an additive ?summary=1 mode; without the
param the response is byte-identical to before. Repeated try/catch + error
JSON collapsed into a helper (server/http.mjs).

Adds route tests for matches/history/regimes (252 -> 268).
Verified: all 15 endpoints byte-identical vs baseline, incl. traversal rejects.
… a tab

CodexBrowser/JudgeLeaderboard/TerminalPanel/RegimeBrowser were built and
bundled but imported by nothing; deleted. HistoryExplorer had real backing
endpoints (GET /api/matches, /api/matches/:id) so it is revived behind a new
MatchArchive component and a Match Archive sidebar tab.

React-Compiler-readiness warnings 15 -> 6 (0 errors) by moving derived state
out of effects and replacing in-place mutation.
…calls

civagent list spawned four python3 processes per regime to read one
metadata.json (228 for 57 regimes). Metadata is now read once per regime into
tab-separated fields; python3 call sites drop 13 -> 3 and list runs 4.1s ->
1.2s. Colors, die/usage and validate_regime move into bin/lib/common.sh.

CLI surface unchanged: list/modes/info/skills output byte-identical to before;
regime-id whitelisting still rejects injection and traversal attempts.
The four refactor branches were told not to touch package.json so they could
not collide; this adds the syntax gate for what they introduced —
judge-rubric.mjs, deterministic-grading.mjs, server/http.mjs,
server/services/regimes.mjs and bin/lib/common.sh.
…root mtime

Codex review P1 on #30: the cache was keyed by the regimes/ directory mtime,
which does not change when a nested file is edited in place — so an edited
metadata.json/IDENTITY.md/SOUL.md, or an added/removed skill, was served stale
for the life of the process. The key is now a sha256 over every served file's
path + mtimeMs + size (one stat per file, no reads).

Also closes the related P2: ?summary=1 was response-light but not I/O-light,
because the cold path still read every markdown body and mapped it away. The
summary view now reads only metadata.json; the full bodies are filled lazily,
on first request for the full catalog.

test/routes-regimes.test.mjs had asserted the stale read as intended behavior,
locking the bug in; it now asserts the edit is visible. Adds
test/services-regimes-cache.test.mjs covering in-place edits, skill add/remove,
regime add, cache reuse, and that summary never touches the bodies (268 -> 281).

Verified the full GET /api/regimes payload is still byte-identical to baseline.
Also lands tasks/REVIEW-r5-codex.md, the review this batch of fixes answers.
Backs the Skill Library UI. The aggregate fields (total, uniqueTopics,
duplicateGroups, stats) come straight from engine/v5/skill-quality.mjs
analyzeSkillsDir() rather than a second implementation of the dedup logic; the
route adds per-file detail (frontmatter name/description, provenance
content_hash/audited_by, size, mtime).

A regime with no skills/ dir returns 200 with total:0, not 404. analyzeSkillsDir
omits the date fields entirely in that case, so the service normalizes stats to
always carry all three keys (null when unknown) — the empty case is the common
one and the response contract promises the keys.
Adds 30 scenarios: 10 shaped by recurring problems of Chinese imperial
statecraft (trade-route severance, examination reform, grain-transport
collapse, regional militarization, court-faction interference, river breach
relief, frontier-market crisis, state-monopoly corruption, military-vs-civil
power struggle, succession ritual law) and 20 general (war financing, colonial
independence, religious reform, abolition transition, federal dissolution,
currency devaluation, inheritance law, border-city autonomy, naval blockade,
industrial displacement, export ban, refugee influx, spy defection, education
reform, water rights, post-coup legitimacy, debt default, assimilation
resistance, urban poverty, disaster blame).

Every new prompt is 60-150 words, carries a valid category, and names no
dynasty, country, place, person or year, so any of the 57 regimes can answer in
character. The original 10 entries are untouched.
Launcher: pick up to 6 civs from GET /api/regimes?summary=1, type a task or
draw one at random from GET /api/scenarios, choose a backend, optionally enable
multi-judge and blind civs, then POST /api/tournaments and follow the returned
id into Live Court. Server-side validation errors are surfaced verbatim in the
form instead of being swallowed.

Skill Library: per-regime view of GET /api/skills/:region/:id/stats with
duplicate-group badges and a total/topics/duplicates summary bar.

Contract fixes found while cross-checking the two sides against each other:
stats.firstSedimented/lastSedimented are ISO date strings (or null), not
numbers, and a skill entry name/description/contentHash/auditedBy are nullable.
The components already rendered those defensively; only the types were wrong.
Both PRs came back APPROVE with the four P1s confirmed closed; these are the
remaining P2s, and the first is a user-visible defect my Batch 2 acceptance
missed — I verified the POST body shape and that the endpoints answered 200,
but never checked what the launcher actually renders.

TournamentLauncher typed GET /api/regimes?summary=1 rows as RegimeMetadata[]
and read r.name, but the endpoint returns { id, metadata } — so every civ in
the picker showed its id twice instead of 'Tang Dynasty' over 'china/tang'.
SkillLibrary already consumed the shape correctly. Adds a RegimeSummary type
so the envelope is named once and both components share it.

The write API matched BACKEND_COMMANDS keys exactly while the engine's
resolveBackend() also tries the lowercased id, so 'NATIVE' was rejected for a
run the engine would have executed. Validation now mirrors the resolver.

ci: 288 tests green, frontend tsc/lint/build clean.
…etriever)

The RAG layer that records tournament outcomes and feeds historical context
back into later matches had zero tests, despite running inside the match tail
where a throw would destroy the whole run.

Adds 16 tests over the real SQLite (no fake), with HOME pointed at a temp dir
before the dynamic import so ~/.civagent is never touched: field mapping,
retrieval hits/misses/empty-db/regime filtering, parameterized-binding checks,
and the degradation path (an unopenable DB must not throw).

Pins the idempotency behaviour found while testing: re-recording the same
tournamentId leaves tournaments at one row (INSERT OR IGNORE) but stacks
duplicate match_results and episodic_memory rows. Left as-is and asserted, so
any future idempotency change trips the test deliberately.
…nents

The frontend had 15 components and no test framework at all, including the
launcher whose validation logic had just shipped. Adds vitest +
@testing-library/react + jsdom and 14 tests over TournamentLauncher (civ cap,
task validation, POST body shape, server error surfacing, navigation on 202,
random scenario), SkillLibrary (loading/error/empty, duplicate badges, stats
bar), MatchArchive (list, detail, failure) and RankingsPanel (empty dataset).

fetch is mocked; no test touches the network.
… plan R6

CHANGELOG v6.1.0 gains the Batch 1 refactor, the Batch 2 skills endpoint /
scenario expansion / launcher + library tabs, and the Codex-review regression
tests. ITERATION_PLAN records R5 against reality (PR #29/#30 green, review
closed) and decomposes R6 — regime-editing API and the historical-accuracy
review loop — into per-party tasks with explicit boundaries and acceptance
commands. Adds tasks/TASK-r6-*.md skeletons; marks the R5 task files done.
…d a fiction

The launcher test mocked GET /api/regimes?summary=1 as a flat RegimeMetadata[]
— a shape the endpoint has never returned. That mock is how the component
shipped reading r.name: the test was green against a fiction, so it could not
have caught the display bug the review found. The mock now mirrors the real
{ id, metadata } envelope.

Verified it actually pins the contract: reverting the mock to the flat shape
turns 5 of the 6 launcher cases red; restoring it returns the suite to green.

Wires the suite in as a gate — 'npm run test:frontend' joins 'npm run ci', and
the GitHub frontend job now runs vitest alongside build and lint. Without this
the 14 new tests would never have run in CI.

ci: 304 backend tests + 14 frontend tests green.
@LeoLin990405
LeoLin990405 changed the base branch from refactor/backend-arg-contract to main July 29, 2026 17:03
@LeoLin990405
LeoLin990405 merged commit bdcca25 into main Jul 29, 2026
2 checks passed
LeoLin990405 pushed a commit that referenced this pull request Aug 6, 2026
LeoLin990405 pushed a commit that referenced this pull request Aug 6, 2026
基于 PR #9 (@782042369) 的思路,在 PR #30 的 Docker/root 框架上重写:

新增:
- detect_os() 系统检测(debian/redhat/alpine/macos)
- pkg_update() / pkg_install() 封装函数,替代 eval + 死代码变量
- install_nodejs() 封装,智能检测版本 >= 22 跳过
- macOS: Homebrew 安装,.zshrc 配置,跳过防火墙/Swap/systemd
- CentOS/RHEL: dnf/yum 双兜底,EPEL 源,chromium-headless
- Alpine: edge 仓库获取较新 Node.js
- pnpm 优先安装 OpenClaw

修复 PR #9 中的问题:
- PKG_INSTALL 变量不再是死代码(改为函数)
- RedHat 命令加了 -y -q flag
- Alpine Node.js 版本问题(用 edge 仓库)
- 保留 $SUDO/$IN_DOCKER 框架(Docker 兼容)
- Puppeteer 路径按系统分别配置

Co-authored-by: waibuzheng <782042369@qq.com>
Closes #9
LeoLin990405 added a commit that referenced this pull request Aug 6, 2026
…root mtime

Codex review P1 on #30: the cache was keyed by the regimes/ directory mtime,
which does not change when a nested file is edited in place — so an edited
metadata.json/IDENTITY.md/SOUL.md, or an added/removed skill, was served stale
for the life of the process. The key is now a sha256 over every served file's
path + mtimeMs + size (one stat per file, no reads).

Also closes the related P2: ?summary=1 was response-light but not I/O-light,
because the cold path still read every markdown body and mapped it away. The
summary view now reads only metadata.json; the full bodies are filled lazily,
on first request for the full catalog.

test/routes-regimes.test.mjs had asserted the stale read as intended behavior,
locking the bug in; it now asserts the edit is visible. Adds
test/services-regimes-cache.test.mjs covering in-place edits, skill add/remove,
regime add, cache reuse, and that summary never touches the bodies (268 -> 281).

Verified the full GET /api/regimes payload is still byte-identical to baseline.
Also lands tasks/REVIEW-r5-codex.md, the review this batch of fixes answers.
LeoLin990405 added a commit that referenced this pull request Aug 6, 2026
… plan R6

CHANGELOG v6.1.0 gains the Batch 1 refactor, the Batch 2 skills endpoint /
scenario expansion / launcher + library tabs, and the Codex-review regression
tests. ITERATION_PLAN records R5 against reality (PR #29/#30 green, review
closed) and decomposes R6 — regime-editing API and the historical-accuracy
review loop — into per-party tasks with explicit boundaries and acceptance
commands. Adds tasks/TASK-r6-*.md skeletons; marks the R5 task files done.
LeoLin990405 added a commit that referenced this pull request Aug 6, 2026
R5 Batch 1: engineering refactor — server/engine/frontend/cli modularization
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant