R5 Batch 1: engineering refactor — server/engine/frontend/cli modularization - #30
Merged
Merged
Conversation
…ding modules judge-rubric.mjs takes the rubric prompt, both score parsers and the pass aggregator; deterministic-grading.mjs takes the T3 spec loader, grader runner and score mixer. tournament.mjs keeps orchestration + judge() + the CLI and re-exports every symbol it exported before, so existing imports are unchanged. Pure move: the six relocated function bodies are byte-identical to their originals; all 17 previously-exported symbols still resolve.
…or helper
All route modules follow the createXxxRouter({rootDir}) factory shape that
tournaments.mjs already used, so tests can inject a temp state dir instead of
depending on the real ~/.civagent. GET /api/regimes no longer re-reads every
IDENTITY.md/SOUL.md/skills file on each request (server/services/regimes.mjs,
mtime-invalidated cache) and gains an additive ?summary=1 mode; without the
param the response is byte-identical to before. Repeated try/catch + error
JSON collapsed into a helper (server/http.mjs).
Adds route tests for matches/history/regimes (252 -> 268).
Verified: all 15 endpoints byte-identical vs baseline, incl. traversal rejects.
… a tab CodexBrowser/JudgeLeaderboard/TerminalPanel/RegimeBrowser were built and bundled but imported by nothing; deleted. HistoryExplorer had real backing endpoints (GET /api/matches, /api/matches/:id) so it is revived behind a new MatchArchive component and a Match Archive sidebar tab. React-Compiler-readiness warnings 15 -> 6 (0 errors) by moving derived state out of effects and replacing in-place mutation.
…calls civagent list spawned four python3 processes per regime to read one metadata.json (228 for 57 regimes). Metadata is now read once per regime into tab-separated fields; python3 call sites drop 13 -> 3 and list runs 4.1s -> 1.2s. Colors, die/usage and validate_regime move into bin/lib/common.sh. CLI surface unchanged: list/modes/info/skills output byte-identical to before; regime-id whitelisting still rejects injection and traversal attempts.
The four refactor branches were told not to touch package.json so they could not collide; this adds the syntax gate for what they introduced — judge-rubric.mjs, deterministic-grading.mjs, server/http.mjs, server/services/regimes.mjs and bin/lib/common.sh.
…root mtime Codex review P1 on #30: the cache was keyed by the regimes/ directory mtime, which does not change when a nested file is edited in place — so an edited metadata.json/IDENTITY.md/SOUL.md, or an added/removed skill, was served stale for the life of the process. The key is now a sha256 over every served file's path + mtimeMs + size (one stat per file, no reads). Also closes the related P2: ?summary=1 was response-light but not I/O-light, because the cold path still read every markdown body and mapped it away. The summary view now reads only metadata.json; the full bodies are filled lazily, on first request for the full catalog. test/routes-regimes.test.mjs had asserted the stale read as intended behavior, locking the bug in; it now asserts the edit is visible. Adds test/services-regimes-cache.test.mjs covering in-place edits, skill add/remove, regime add, cache reuse, and that summary never touches the bodies (268 -> 281). Verified the full GET /api/regimes payload is still byte-identical to baseline. Also lands tasks/REVIEW-r5-codex.md, the review this batch of fixes answers.
Backs the Skill Library UI. The aggregate fields (total, uniqueTopics, duplicateGroups, stats) come straight from engine/v5/skill-quality.mjs analyzeSkillsDir() rather than a second implementation of the dedup logic; the route adds per-file detail (frontmatter name/description, provenance content_hash/audited_by, size, mtime). A regime with no skills/ dir returns 200 with total:0, not 404. analyzeSkillsDir omits the date fields entirely in that case, so the service normalizes stats to always carry all three keys (null when unknown) — the empty case is the common one and the response contract promises the keys.
Adds 30 scenarios: 10 shaped by recurring problems of Chinese imperial statecraft (trade-route severance, examination reform, grain-transport collapse, regional militarization, court-faction interference, river breach relief, frontier-market crisis, state-monopoly corruption, military-vs-civil power struggle, succession ritual law) and 20 general (war financing, colonial independence, religious reform, abolition transition, federal dissolution, currency devaluation, inheritance law, border-city autonomy, naval blockade, industrial displacement, export ban, refugee influx, spy defection, education reform, water rights, post-coup legitimacy, debt default, assimilation resistance, urban poverty, disaster blame). Every new prompt is 60-150 words, carries a valid category, and names no dynasty, country, place, person or year, so any of the 57 regimes can answer in character. The original 10 entries are untouched.
Launcher: pick up to 6 civs from GET /api/regimes?summary=1, type a task or draw one at random from GET /api/scenarios, choose a backend, optionally enable multi-judge and blind civs, then POST /api/tournaments and follow the returned id into Live Court. Server-side validation errors are surfaced verbatim in the form instead of being swallowed. Skill Library: per-regime view of GET /api/skills/:region/:id/stats with duplicate-group badges and a total/topics/duplicates summary bar. Contract fixes found while cross-checking the two sides against each other: stats.firstSedimented/lastSedimented are ISO date strings (or null), not numbers, and a skill entry name/description/contentHash/auditedBy are nullable. The components already rendered those defensively; only the types were wrong.
Both PRs came back APPROVE with the four P1s confirmed closed; these are the
remaining P2s, and the first is a user-visible defect my Batch 2 acceptance
missed — I verified the POST body shape and that the endpoints answered 200,
but never checked what the launcher actually renders.
TournamentLauncher typed GET /api/regimes?summary=1 rows as RegimeMetadata[]
and read r.name, but the endpoint returns { id, metadata } — so every civ in
the picker showed its id twice instead of 'Tang Dynasty' over 'china/tang'.
SkillLibrary already consumed the shape correctly. Adds a RegimeSummary type
so the envelope is named once and both components share it.
The write API matched BACKEND_COMMANDS keys exactly while the engine's
resolveBackend() also tries the lowercased id, so 'NATIVE' was rejected for a
run the engine would have executed. Validation now mirrors the resolver.
ci: 288 tests green, frontend tsc/lint/build clean.
…etriever) The RAG layer that records tournament outcomes and feeds historical context back into later matches had zero tests, despite running inside the match tail where a throw would destroy the whole run. Adds 16 tests over the real SQLite (no fake), with HOME pointed at a temp dir before the dynamic import so ~/.civagent is never touched: field mapping, retrieval hits/misses/empty-db/regime filtering, parameterized-binding checks, and the degradation path (an unopenable DB must not throw). Pins the idempotency behaviour found while testing: re-recording the same tournamentId leaves tournaments at one row (INSERT OR IGNORE) but stacks duplicate match_results and episodic_memory rows. Left as-is and asserted, so any future idempotency change trips the test deliberately.
…nents The frontend had 15 components and no test framework at all, including the launcher whose validation logic had just shipped. Adds vitest + @testing-library/react + jsdom and 14 tests over TournamentLauncher (civ cap, task validation, POST body shape, server error surfacing, navigation on 202, random scenario), SkillLibrary (loading/error/empty, duplicate badges, stats bar), MatchArchive (list, detail, failure) and RankingsPanel (empty dataset). fetch is mocked; no test touches the network.
… plan R6 CHANGELOG v6.1.0 gains the Batch 1 refactor, the Batch 2 skills endpoint / scenario expansion / launcher + library tabs, and the Codex-review regression tests. ITERATION_PLAN records R5 against reality (PR #29/#30 green, review closed) and decomposes R6 — regime-editing API and the historical-accuracy review loop — into per-party tasks with explicit boundaries and acceptance commands. Adds tasks/TASK-r6-*.md skeletons; marks the R5 task files done.
…d a fiction
The launcher test mocked GET /api/regimes?summary=1 as a flat RegimeMetadata[]
— a shape the endpoint has never returned. That mock is how the component
shipped reading r.name: the test was green against a fiction, so it could not
have caught the display bug the review found. The mock now mirrors the real
{ id, metadata } envelope.
Verified it actually pins the contract: reverting the mock to the flat shape
turns 5 of the 6 launcher cases red; restoring it returns the suite to green.
Wires the suite in as a gate — 'npm run test:frontend' joins 'npm run ci', and
the GitHub frontend job now runs vitest alongside build and lint. Without this
the 14 new tests would never have run in CI.
ci: 304 backend tests + 14 frontend tests green.
LeoLin990405
pushed a commit
that referenced
this pull request
Aug 6, 2026
基于 PR #9 (@782042369) 的思路,在 PR #30 的 Docker/root 框架上重写: 新增: - detect_os() 系统检测(debian/redhat/alpine/macos) - pkg_update() / pkg_install() 封装函数,替代 eval + 死代码变量 - install_nodejs() 封装,智能检测版本 >= 22 跳过 - macOS: Homebrew 安装,.zshrc 配置,跳过防火墙/Swap/systemd - CentOS/RHEL: dnf/yum 双兜底,EPEL 源,chromium-headless - Alpine: edge 仓库获取较新 Node.js - pnpm 优先安装 OpenClaw 修复 PR #9 中的问题: - PKG_INSTALL 变量不再是死代码(改为函数) - RedHat 命令加了 -y -q flag - Alpine Node.js 版本问题(用 edge 仓库) - 保留 $SUDO/$IN_DOCKER 框架(Docker 兼容) - Puppeteer 路径按系统分别配置 Co-authored-by: waibuzheng <782042369@qq.com> Closes #9
LeoLin990405
added a commit
that referenced
this pull request
Aug 6, 2026
…root mtime Codex review P1 on #30: the cache was keyed by the regimes/ directory mtime, which does not change when a nested file is edited in place — so an edited metadata.json/IDENTITY.md/SOUL.md, or an added/removed skill, was served stale for the life of the process. The key is now a sha256 over every served file's path + mtimeMs + size (one stat per file, no reads). Also closes the related P2: ?summary=1 was response-light but not I/O-light, because the cold path still read every markdown body and mapped it away. The summary view now reads only metadata.json; the full bodies are filled lazily, on first request for the full catalog. test/routes-regimes.test.mjs had asserted the stale read as intended behavior, locking the bug in; it now asserts the edit is visible. Adds test/services-regimes-cache.test.mjs covering in-place edits, skill add/remove, regime add, cache reuse, and that summary never touches the bodies (268 -> 281). Verified the full GET /api/regimes payload is still byte-identical to baseline. Also lands tasks/REVIEW-r5-codex.md, the review this batch of fixes answers.
LeoLin990405
added a commit
that referenced
this pull request
Aug 6, 2026
… plan R6 CHANGELOG v6.1.0 gains the Batch 1 refactor, the Batch 2 skills endpoint / scenario expansion / launcher + library tabs, and the Codex-review regression tests. ITERATION_PLAN records R5 against reality (PR #29/#30 green, review closed) and decomposes R6 — regime-editing API and the historical-accuracy review loop — into per-party tasks with explicit boundaries and acceptance commands. Adds tasks/TASK-r6-*.md skeletons; marks the R5 task files done.
LeoLin990405
added a commit
that referenced
this pull request
Aug 6, 2026
R5 Batch 1: engineering refactor — server/engine/frontend/cli modularization
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #29. Pure engineering refactor — no new features, no behavior change. Built in parallel by four domestic-model agents in isolated worktrees (path-disjoint ownership), verified on the home cluster.
What changed
server — every route module now uses the
createXxxRouter({rootDir})factory shape thattournaments.mjsalready had, so tests inject a temp state dir instead of depending on the real~/.civagent.GET /api/regimesno longer re-reads all 57 regimes'IDENTITY.md/SOUL.md/skills on every request (server/services/regimes.mjs, mtime-invalidated cache) and gains an additive?summary=1mode. Repeated try/catch + error-JSON collapsed intoserver/http.mjs. New route tests for matches/history/regimes.engine —
tournament.mjs710 → 462 lines, split intojudge-rubric.mjs(rubric prompt, both score parsers, pass aggregation) anddeterministic-grading.mjs(T3 spec loader, grader runner, score mixer). All 17 previously-exported symbols still resolve fromtournament.mjs.frontend — deleted four components that were bundled but imported by nothing (CodexBrowser, JudgeLeaderboard, TerminalPanel, old RegimeBrowser). HistoryExplorer had real backing endpoints so it is revived behind a new
MatchArchivecomponent and sidebar tab. React-Compiler-readiness warnings 15 → 6 (0 errors).cli —
civagent listspawned fourpython3processes per regime to read onemetadata.json(228 for 57 regimes). Now one read per regime; python call sites 13 → 3,civagent list4.1s → 1.2s. Shared helpers extracted tobin/lib/common.sh.Verification (beyond CI)
Refactors claim "no behavior change", so that claim was tested directly rather than trusted:
list/modes/info china/tang/info global/athens/info china/qin/skillsbyte-identical; regime-id whitelisting still rejects injection and traversal.MatchArchivehits endpoints that return 200 with real local data.Test plan
npm run ci— 268 tests (was 252), 57 regimes validatedtsc -b, frontend eslint 0 errors,vite builde5c945e: leo-03 test+cli-smoke, leo-01 lint+syntax+regimes, leo-02 frontend types+lint+build