Conversation
…sment This skill provides automated evaluation of training data samples from partners, supporting multiple data types including pre-training, mid-training, post-training, code data, agent trajectories, thinking data, and GUI interaction data. Features: - Automatic data type classification (pre-train/mid-train/post-train) - Model detection (Claude, GPT-4, Llama 3, DeepSeek, etc.) - Agent detection (cline, Terminal, Thunderbird, LibreOffice, etc.) - Unified dimension display (Tokens/Size/Samples/Language/Models/Agents) - Vim color scheme style HTML report - Incremental mode for cumulative reporting
nlp-zn
left a comment
There was a problem hiding this comment.
按 skill-creator 原则 review,这个 PR 目前还不能合。
Blockers:
-
scripts/evaluate_data.py把WORKSPACE硬编码成作者本机路径/Users/huangneng/...,OUTPUT_DIR和HISTORY_FILE也随之写到该路径。skill 应该可被任何调用者复用,不能依赖作者机器。请改为 CLI 参数,例如python3 scripts/evaluate_data.py --workspace <path> --output <path>,并让SKILL.md的使用方法与脚本参数一致。 -
SKILL.md声明支持zip, rar, txt, json, jsonl,但main()只扫描*.zip,*.rar,*.txt。这会让用户按描述传 JSON/JSONL 时静默漏评。请补齐扫描和解析路径,或收窄 frontmatter/使用说明中的支持格式。 -
解压逻辑多处用
os.system(f'ditto ... "{file_path}" ...')/unrar拼 shell 命令。训练数据文件名来自用户输入,应该避免 shell 拼接,改用zipfile/tarfile或subprocess.run([...], check=True);RAR 不可用时也要在报告中明确提示,而不是静默失败。 -
缺少 skill-creator 要求的 evals/test prompts。建议补
evals/evals.json,覆盖至少:JSONL agent 轨迹、zip 数据集、普通 txt 文档、无可解析文件时的失败提示,并为“报告是否包含数据集数量/类型/来源/建议”等写可检查的 assertions。
整体方向有价值,但先把可移植性、格式承诺和 eval 覆盖补上,不然这个 skill 在作者机器之外大概率跑不起来或输出误导结果。
功能说明
支持的数据类型:
核心能力:
触发场景:
输出形式:Vim配色风格的HTML评估报告