Skip to content

Add training-data-evaluator skill - #147

Open
huangneng wants to merge 1 commit into
AgenticAIPlan:devfrom
huangneng:feat/training-data-evaluator
Open

huangneng wants to merge 1 commit into
AgenticAIPlan:devfrom
huangneng:feat/training-data-evaluator

Conversation

@huangneng

@huangneng huangneng commented Apr 30, 2026 •

Copy link
Copy Markdown
  • PR 类型: 新 Skill 提交到 dev
  • 目标分支: dev
  • 源分支: feat/training-data-evaluator
  • Skill 名称: training-data-evaluator
  • Skill 路径: skills/training-data-evaluator/
  • 业务场景: 大模型训练数据质量评估工具,用于全自动化评估合作伙伴提供的样例数据价值,帮助数据引入负责人快速判断数据采购决策。
  • 分支名: feat/training-data-evaluator
  • 本次是否由 Agent 辅助提交: 是

功能说明

支持的数据类型:

  • 预训练数据:技术文档、教程、Q&A、博客、Demo示例
  • 中训练数据:SWE工程数据、Agent轨迹数据
  • 后训练数据:RL数据
  • 基础数据:源代码、推理数据、GUI交互数据

核心能力:

  1. 数据类型自动分类(预训练/中训练/后训练)
  2. 模型检测(Claude、GPT-4、Llama 3、DeepSeek等)
  3. 智能体检测(cline、Terminal、Thunderbird等)
  4. 统一维度展示(Tokens/Size/Samples/Language/Models/Agents)
  5. Vim配色风格的HTML评估报告
  6. 累加模式(自动检测新数据集)

触发场景:

  • 数据评估、数据质量检查、样例评估
  • 需要检测模型/智能体信息
  • 生成数据评估报告

输出形式:Vim配色风格的HTML评估报告

…sment

This skill provides automated evaluation of training data samples from partners,
supporting multiple data types including pre-training, mid-training, post-training,
code data, agent trajectories, thinking data, and GUI interaction data.

Features:
- Automatic data type classification (pre-train/mid-train/post-train)
- Model detection (Claude, GPT-4, Llama 3, DeepSeek, etc.)
- Agent detection (cline, Terminal, Thunderbird, LibreOffice, etc.)
- Unified dimension display (Tokens/Size/Samples/Language/Models/Agents)
- Vim color scheme style HTML report
- Incremental mode for cumulative reporting
@nlp-zn
nlp-zn self-requested a review May 7, 2026 04:01

@nlp-zn nlp-zn left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

按 skill-creator 原则 review,这个 PR 目前还不能合。

Blockers:

  1. scripts/evaluate_data.py 把 WORKSPACE 硬编码成作者本机路径 /Users/huangneng/...,OUTPUT_DIR 和 HISTORY_FILE 也随之写到该路径。skill 应该可被任何调用者复用,不能依赖作者机器。请改为 CLI 参数,例如 python3 scripts/evaluate_data.py --workspace <path> --output <path>,并让 SKILL.md 的使用方法与脚本参数一致。

  2. SKILL.md 声明支持 zip, rar, txt, json, jsonl,但 main() 只扫描 *.zip, *.rar, *.txt。这会让用户按描述传 JSON/JSONL 时静默漏评。请补齐扫描和解析路径,或收窄 frontmatter/使用说明中的支持格式。

  3. 解压逻辑多处用 os.system(f'ditto ... "{file_path}" ...') / unrar 拼 shell 命令。训练数据文件名来自用户输入,应该避免 shell 拼接,改用 zipfile/tarfile 或 subprocess.run([...], check=True);RAR 不可用时也要在报告中明确提示,而不是静默失败。

  4. 缺少 skill-creator 要求的 evals/test prompts。建议补 evals/evals.json,覆盖至少:JSONL agent 轨迹、zip 数据集、普通 txt 文档、无可解析文件时的失败提示,并为“报告是否包含数据集数量/类型/来源/建议”等写可检查的 assertions。

整体方向有价值,但先把可移植性、格式承诺和 eval 覆盖补上,不然这个 skill 在作者机器之外大概率跑不起来或输出误导结果。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants