Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .github/aw/engines.json
Original file line number Diff line number Diff line change
Expand Up @@ -27,6 +27,10 @@
{
"id": "opencode",
"import": "github/gh-aw/.github/workflows/shared/opencode.md"
},
{
"id": "pydantic-ai",
"import": "github/gh-aw/.github/workflows/shared/pydantic.md"
}
]
}
7 changes: 4 additions & 3 deletions .github/workflows/agentic_commands.yml

Large diffs are not rendered by default.

140 changes: 140 additions & 0 deletions .github/workflows/shared/pydantic.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,140 @@
---
runtimes:
uv:
version: latest
pre-agent-steps:
- name: Predownload Pydantic AI CLI
run: uv run pai --version
engine:
id: pydantic-ai
version: "0.1.0"
display-name: Pydantic AI
description: Pydantic AI headless coding agent CLI with MCP support
experimental: true
mcp: true
provider:
name: github
behaviors:
secret-strategy: universal-llm-consumer
manifest:
files:
- AGENTS.md
path-prefixes:
- .pydantic-ai/
network:
defaults:
- host.docker.internal

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@copilot none of these should be needed since pydantic is already pre-installed. keep network as "defaults"

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated in d36df30. I removed the Pydantic install pre-step and switched the shared engine networking back to defaults (dropping the extra PyPI domains).

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[/codebase-design] Network defaults includes pypi.org and files.pythonhosted.org — these are only needed during pre-agent-steps installation, not agent runtime.

💡 Suggestion

PyPI hosts are consumed by pip install in pre-agent-steps, which runs before the agent sandbox. If the runtime network policy cannot be scoped per-phase, document why these hosts must remain open during agent execution. Also, the PR comment from @pelikhan suggests pydantic-ai-harness may already be pre-installed, which would make the pip install step and these entries unnecessary entirely.

@copilot please address this.

- github.com
- raw.githubusercontent.com
- api.github.com
- objects.githubusercontent.com
provider-domains:
copilot: api.githubcopilot.com
anthropic: api.anthropic.com
openai: api.openai.com
execution:
command-name: uv
args:
- run
- pai
- run
step-name: Execute Pydantic AI CLI
model-env-var: PAI_MODEL
mcp-config-env-var: GH_AW_MCP_CONFIG
write-timestamp: true
provider-env-mode: universal-llm-consumer
log-parser: |
function parseLog(logContent) {
const lines = logContent.split("\n");
const logEntries = [];
const mcpFailures = [];
let maxTurnsHit = false;
const AWF_INFRA_RE = /^\[(INFO|WARN|SUCCESS|ERROR|entrypoint|health-check)\]|^ (?:Container|Network|Volume) |^Process exiting with code:/;
let inputTokens = 0;
let outputTokens = 0;
let toolCallIndex = 0;
let turnCount = 0;
let pendingText = [];

function flushText() {
if (pendingText.length === 0) return;
const text = pendingText.join("\n").trim();
if (text) {
logEntries.push({ type: "assistant", message: { content: [{ type: "text", text }] } });
turnCount++;
}
pendingText = [];
}

logEntries.push({ type: "system", subtype: "init", model: null, session_id: null });

for (const line of lines) {
if (!line.trim()) continue;
if (AWF_INFRA_RE.test(line)) continue;
if (/max.?turns|maximum.*turns.*reached|turn limit/i.test(line)) maxTurnsHit = true;
if (/MCP server .* failed|MCP.*connection.*error|Failed to connect to MCP/i.test(line)) {
const serverMatch = line.match(/MCP server ['"]?([^\s'"]+)['"]?/i);
mcpFailures.push(serverMatch ? serverMatch[1] : line.trim());
}

let parsed = null;
try {
if (line.trim().startsWith("{")) parsed = JSON.parse(line.trim());
} catch (e) { /* not JSON */ }

if (parsed) {
if (parsed.input_tokens) inputTokens += parsed.input_tokens;
if (parsed.output_tokens) outputTokens += parsed.output_tokens;
const entryType = parsed.type != null ? String(parsed.type) : "log";
const msg = parsed.msg || parsed.message || parsed.content || "";

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[/tdd] The log-parser only accumulates input_tokens / output_tokens from JSON lines that have those exact keys — if Pydantic AI emits them under a different key (e.g. usage.input_tokens) the totals silently remain zero.

💡 Suggestion

Add a test fixture (a small JSONL sample) and a unit test that asserts:

  • known token-bearing lines are counted correctly
  • lines with a nested usage object are also handled
  • the final result entry reflects the correct totals

Without this, a Pydantic AI output format change will produce invisible zero-token audit entries.

@copilot please address this.


if (/tool[._]call|tool[._]use/i.test(entryType)) {
flushText();
const toolId = `pai_tool_${toolCallIndex++}`;
const toolName = parsed.tool || parsed.name || entryType;
logEntries.push({ type: "assistant", message: { content: [{ type: "tool_use", id: toolId, name: toolName, input: {} }] } });
logEntries.push({ type: "user", message: { content: [{ type: "tool_result", tool_use_id: toolId, content: msg }] } });
} else if (msg) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[/codebase-design] Tool-result entries always set content: msg where msg may be an empty string — the harness may reject a tool_result with an empty content array.

💡 Suggestion

Use a fallback: content: msg || '(no output)' to guarantee a non-empty string. Downstream log consumers that validate against the Anthropic message schema will reject empty content.

@copilot please address this.

pendingText.push(msg);
}
} else {
pendingText.push(line.trim());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

MCP failure lines match the mcpFailures push above but are not skipped afterward — they fall through to pendingText.push(line.trim()) and are emitted as synthetic assistant text in the log transcript. Every MCP error message ends up double-counted: once in mcpFailures and once as an assistant turn.

Fix by adding continue after the mcpFailures.push(...) call:

if (/MCP server .* failed|MCP.*connection.*error|Failed to connect to MCP/i.test(line)) {
  const serverMatch = line.match(/MCP server ['"x]?([^\s'"]+)['"x]?/i);
  mcpFailures.push(serverMatch ? serverMatch[1] : line.trim());
  continue; // prevent double-emit as transcript text
}

@copilot please address this.

}
}
flushText();

const usage = {};
if (inputTokens) usage.input_tokens = inputTokens;
if (outputTokens) usage.output_tokens = outputTokens;
logEntries.push({ type: "result", num_turns: turnCount, usage });
const parts = [`**Turns:** ${turnCount}`, `**Tool calls:** ${toolCallIndex}`];
if (inputTokens || outputTokens) parts.push(`**Tokens:** ${((inputTokens ?? 0) + (outputTokens ?? 0)).toLocaleString()}`);
if (mcpFailures.length) parts.push(`**MCP failures:** ${mcpFailures.length}`);
if (maxTurnsHit) parts.push("**Max turns reached**");
return { markdown: parts.join(" · "), logEntries, mcpFailures, maxTurnsHit };
}
---

<!--
# Pydantic AI

Shared engine definition for [Pydantic AI](https://ai.pydantic.dev), the
headless AI coding agent. Import this file and set `engine: id: pydantic-ai`
to use it:

```yaml
engine:
id: pydantic-ai
model: copilot/claude-sonnet-4-5
imports:
- shared/pydantic.md
```

`model` must use `provider/model` format. Supported providers are `copilot`,
`anthropic`, and `openai`. Requests are routed through the AWF proxy.

The engine reads MCP server configuration from `GH_AW_MCP_CONFIG`, so
safe outputs flow through the standard `safeoutputs` server automatically.

Pydantic AI is preinstalled in the runtime image.
-->
Loading
Loading