Skip to content

P3: skill 学习闭环供应链加固 - #21

Merged
LeoLin990405 merged 3 commits into
mainfrom
feature/p3-skill-supply-chain
Jul 22, 2026
Merged

LeoLin990405 merged 3 commits into
mainfrom
feature/p3-skill-supply-chain

Conversation

@LeoLin990405

Copy link
Copy Markdown
Owner

P3: skill 学习闭环供应链加固

外部依据:Snyk ToxicSkills(公开 skill 13.4% 含严重问题,91% 恶意样本组合注入+代码)、SkillSieve(静态层+LLM jury 双层)、ACE(增量追加防 context collapse——本仓库现状已是每场一文件,天然增量,无需改)。

改动

  • engine/v5/skill-scan.mjs(新增,零依赖):确定性扫描 → {verdict: pass|flag|reject, findings[]}
    • reject:pipe-to-shell(curl|bash)、base64-exec、remote-instruction-fetch、credential-path(.ssh/.aws/.env 等)
    • flag(留 staging 待人工):external-url、fs-network-verbs(rm -rf/install/ssh 等)、prompt-injection(扩充正则;注意从旧硬 reject 降级为 flag)
  • staging + 人工 gate:flag → skills/staging/;reject → 丢弃+事件;pass → 异源审计链(不变)→ 正式写入。CIVAGENT_SKILL_GATE=off 恢复旧行为
  • CLI:civagent skills pending [regime]、civagent skills approve <regime> <file>(basename 防穿越)
  • 溯源:skill frontmatter 自动 pin content_hash(sha256[:16],幂等)+ schema_version;skill_commit 事件带 contentHash,新增 staged 状态(schema/前端类型同步,TerminalPanel 金色渲染 "Staged for human review")

测试

  • 93 pass / 0 fail(基线 79 + 新增 14);lint 通过;前端 build 通过
  • 覆盖:reject/flag/pass 分级、pin 幂等、全流程集成、gate-off 等价、CLI 路径穿越拒绝

待决策

  1. 注入模式从硬 reject 降级为 flag(staging 等人工)——是否接受
  2. external-url 一刀切 flag 会误伤史料引用链接——是否加白名单
  3. credential-path 不区分"指令"与"禁令"语境(宁严勿宽)
  4. approve(staging→正式)暂无审计事件,可加 --log
  5. skill schema_version 独立 "1.0"

Static scan + staging gate + provenance pinning for extracted skills
(cf. Snyk ToxicSkills, SkillSieve two-layer design):

- engine/v5/skill-scan.mjs (new, zero-dep, deterministic):
  scanSkillText() → {verdict: pass|flag|reject, findings[]}.
  reject: pipe-to-shell (curl|wget … |sh), base64-decode+exec, runtime
  remote instruction fetch (eval/source of curl/wget), credential path
  access (~/.ssh, .env, .aws/credentials, netrc, npmrc, kubeconfig).
  flag: external URLs, install/network/destructive verbs, expanded
  prompt-injection shapes (superset of the legacy INJECTION_PATTERNS,
  plus role-override and chat-delimiter patterns).
  pinSkillFrontmatter(): inserts content_hash (sha256[:16] of the
  extracted body) + schema_version into the skill frontmatter,
  idempotently.
- engine/v5/skill-sediment.mjs: gated pipeline (default on;
  CIVAGENT_SKILL_GATE=off restores legacy behavior exactly, including
  the old hard injection reject). reject → discarded with findings;
  flag → written to skills/staging/ (audit skipped — a human decides);
  pass → independent audit → promoted into skills/. All writes pinned.
- run-v5.mjs buildSkillEvent: new 'staged' status + contentHash field;
  schema status enum extended ('staged'); frontend type + TerminalPanel
  rendering updated.
- CLI: 'civagent skills pending [regime]' lists staged skills;
  'civagent skills approve <regime> <file>' promotes staging → active
  (basename-validated, no traversal). CIVAGENT_REGIMES_DIR env override
  for sandboxed runs.
- Note on ACE-style incrementality: sediment already appends one new
  file per match (learned-<date>-<topic>-<matchId>.md) and never
  rewrites existing skills, so no change was needed there.
- test/skill-scan.test.mjs (14 cases): scanner reject/flag/pass +
  precedence, pin idempotence, gated staging/promotion with fake
  extractor+auditor, gate-off legacy parity, CLI pending/approve/traversal.
…ak=flag

Strong markers (ignore/disregard previous instructions, <system> and
[INST] delimiters, jailbreak/DAN, role-persona override via system:)
are unambiguous attacks → reject, never staged. Weak shapes (you are
now / from now on / new instructions:) also appear in benign governance
prose → flag for human review instead of silently discarding. Tests
split accordingly.
@LeoLin990405
LeoLin990405 merged commit cb6d622 into main Jul 22, 2026
1 check passed
@LeoLin990405
LeoLin990405 deleted the feature/p3-skill-supply-chain branch July 22, 2026 11:00
LeoLin990405 pushed a commit that referenced this pull request Aug 6, 2026
- install.sh: 自动检测 root/Docker 环境,跳过 sudo
- install.sh: Docker 内跳过防火墙和 Swap 配置
- doctor.sh: 新增 Docker/sandbox 权限检查(第5步)
- doctor.sh: 非 root 用户提示 usermod -aG docker 修复方案

Fixes #13, Fixes #21
LeoLin990405 pushed a commit that referenced this pull request Aug 6, 2026
LeoLin990405 added a commit that referenced this pull request Aug 6, 2026
* feat(skills): supply-chain hardening for the learning loop (P3)

Static scan + staging gate + provenance pinning for extracted skills
(cf. Snyk ToxicSkills, SkillSieve two-layer design):

- engine/v5/skill-scan.mjs (new, zero-dep, deterministic):
  scanSkillText() → {verdict: pass|flag|reject, findings[]}.
  reject: pipe-to-shell (curl|wget … |sh), base64-decode+exec, runtime
  remote instruction fetch (eval/source of curl/wget), credential path
  access (~/.ssh, .env, .aws/credentials, netrc, npmrc, kubeconfig).
  flag: external URLs, install/network/destructive verbs, expanded
  prompt-injection shapes (superset of the legacy INJECTION_PATTERNS,
  plus role-override and chat-delimiter patterns).
  pinSkillFrontmatter(): inserts content_hash (sha256[:16] of the
  extracted body) + schema_version into the skill frontmatter,
  idempotently.
- engine/v5/skill-sediment.mjs: gated pipeline (default on;
  CIVAGENT_SKILL_GATE=off restores legacy behavior exactly, including
  the old hard injection reject). reject → discarded with findings;
  flag → written to skills/staging/ (audit skipped — a human decides);
  pass → independent audit → promoted into skills/. All writes pinned.
- run-v5.mjs buildSkillEvent: new 'staged' status + contentHash field;
  schema status enum extended ('staged'); frontend type + TerminalPanel
  rendering updated.
- CLI: 'civagent skills pending [regime]' lists staged skills;
  'civagent skills approve <regime> <file>' promotes staging → active
  (basename-validated, no traversal). CIVAGENT_REGIMES_DIR env override
  for sandboxed runs.
- Note on ACE-style incrementality: sediment already appends one new
  file per match (learned-<date>-<topic>-<matchId>.md) and never
  rewrites existing skills, so no change was needed there.
- test/skill-scan.test.mjs (14 cases): scanner reject/flag/pass +
  precedence, pin idempotence, gated staging/promotion with fake
  extractor+auditor, gate-off legacy parity, CLI pending/approve/traversal.

* refactor(skill-scan): two-tier injection severity — strong=reject, weak=flag

Strong markers (ignore/disregard previous instructions, <system> and
[INST] delimiters, jailbreak/DAN, role-persona override via system:)
are unambiguous attacks → reject, never staged. Weak shapes (you are
now / from now on / new instructions:) also appear in benign governance
prose → flag for human review instead of silently discarding. Tests
split accordingly.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant