P3: skill 学习闭环供应链加固 - #21
Merged
Merged
Conversation
Static scan + staging gate + provenance pinning for extracted skills
(cf. Snyk ToxicSkills, SkillSieve two-layer design):
- engine/v5/skill-scan.mjs (new, zero-dep, deterministic):
scanSkillText() → {verdict: pass|flag|reject, findings[]}.
reject: pipe-to-shell (curl|wget … |sh), base64-decode+exec, runtime
remote instruction fetch (eval/source of curl/wget), credential path
access (~/.ssh, .env, .aws/credentials, netrc, npmrc, kubeconfig).
flag: external URLs, install/network/destructive verbs, expanded
prompt-injection shapes (superset of the legacy INJECTION_PATTERNS,
plus role-override and chat-delimiter patterns).
pinSkillFrontmatter(): inserts content_hash (sha256[:16] of the
extracted body) + schema_version into the skill frontmatter,
idempotently.
- engine/v5/skill-sediment.mjs: gated pipeline (default on;
CIVAGENT_SKILL_GATE=off restores legacy behavior exactly, including
the old hard injection reject). reject → discarded with findings;
flag → written to skills/staging/ (audit skipped — a human decides);
pass → independent audit → promoted into skills/. All writes pinned.
- run-v5.mjs buildSkillEvent: new 'staged' status + contentHash field;
schema status enum extended ('staged'); frontend type + TerminalPanel
rendering updated.
- CLI: 'civagent skills pending [regime]' lists staged skills;
'civagent skills approve <regime> <file>' promotes staging → active
(basename-validated, no traversal). CIVAGENT_REGIMES_DIR env override
for sandboxed runs.
- Note on ACE-style incrementality: sediment already appends one new
file per match (learned-<date>-<topic>-<matchId>.md) and never
rewrites existing skills, so no change was needed there.
- test/skill-scan.test.mjs (14 cases): scanner reject/flag/pass +
precedence, pin idempotence, gated staging/promotion with fake
extractor+auditor, gate-off legacy parity, CLI pending/approve/traversal.
…ak=flag Strong markers (ignore/disregard previous instructions, <system> and [INST] delimiters, jailbreak/DAN, role-persona override via system:) are unambiguous attacks → reject, never staged. Weak shapes (you are now / from now on / new instructions:) also appear in benign governance prose → flag for human review instead of silently discarding. Tests split accordingly.
LeoLin990405
pushed a commit
that referenced
this pull request
Aug 6, 2026
LeoLin990405
added a commit
that referenced
this pull request
Aug 6, 2026
* feat(skills): supply-chain hardening for the learning loop (P3)
Static scan + staging gate + provenance pinning for extracted skills
(cf. Snyk ToxicSkills, SkillSieve two-layer design):
- engine/v5/skill-scan.mjs (new, zero-dep, deterministic):
scanSkillText() → {verdict: pass|flag|reject, findings[]}.
reject: pipe-to-shell (curl|wget … |sh), base64-decode+exec, runtime
remote instruction fetch (eval/source of curl/wget), credential path
access (~/.ssh, .env, .aws/credentials, netrc, npmrc, kubeconfig).
flag: external URLs, install/network/destructive verbs, expanded
prompt-injection shapes (superset of the legacy INJECTION_PATTERNS,
plus role-override and chat-delimiter patterns).
pinSkillFrontmatter(): inserts content_hash (sha256[:16] of the
extracted body) + schema_version into the skill frontmatter,
idempotently.
- engine/v5/skill-sediment.mjs: gated pipeline (default on;
CIVAGENT_SKILL_GATE=off restores legacy behavior exactly, including
the old hard injection reject). reject → discarded with findings;
flag → written to skills/staging/ (audit skipped — a human decides);
pass → independent audit → promoted into skills/. All writes pinned.
- run-v5.mjs buildSkillEvent: new 'staged' status + contentHash field;
schema status enum extended ('staged'); frontend type + TerminalPanel
rendering updated.
- CLI: 'civagent skills pending [regime]' lists staged skills;
'civagent skills approve <regime> <file>' promotes staging → active
(basename-validated, no traversal). CIVAGENT_REGIMES_DIR env override
for sandboxed runs.
- Note on ACE-style incrementality: sediment already appends one new
file per match (learned-<date>-<topic>-<matchId>.md) and never
rewrites existing skills, so no change was needed there.
- test/skill-scan.test.mjs (14 cases): scanner reject/flag/pass +
precedence, pin idempotence, gated staging/promotion with fake
extractor+auditor, gate-off legacy parity, CLI pending/approve/traversal.
* refactor(skill-scan): two-tier injection severity — strong=reject, weak=flag
Strong markers (ignore/disregard previous instructions, <system> and
[INST] delimiters, jailbreak/DAN, role-persona override via system:)
are unambiguous attacks → reject, never staged. Weak shapes (you are
now / from now on / new instructions:) also appear in benign governance
prose → flag for human review instead of silently discarding. Tests
split accordingly.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
P3: skill 学习闭环供应链加固
外部依据:Snyk ToxicSkills(公开 skill 13.4% 含严重问题,91% 恶意样本组合注入+代码)、SkillSieve(静态层+LLM jury 双层)、ACE(增量追加防 context collapse——本仓库现状已是每场一文件,天然增量,无需改)。
改动
engine/v5/skill-scan.mjs(新增,零依赖):确定性扫描 →{verdict: pass|flag|reject, findings[]}skills/staging/;reject → 丢弃+事件;pass → 异源审计链(不变)→ 正式写入。CIVAGENT_SKILL_GATE=off恢复旧行为civagent skills pending [regime]、civagent skills approve <regime> <file>(basename 防穿越)content_hash(sha256[:16],幂等)+schema_version;skill_commit 事件带contentHash,新增staged状态(schema/前端类型同步,TerminalPanel 金色渲染 "Staged for human review")测试
待决策
--log