Skip to content

feat(agents): add experimental agents/models/pi-ai - #2446

Merged
aron-cf merged 5 commits into
models-corefrom
models-pi-ai
Oct 2, 2026
Merged

aron-cf merged 5 commits into
models-corefrom
models-pi-ai

Conversation

@aron-cf

@aron-cf aron-cf commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Add createAI() for pi-ai over the model core. ai("@cf/...") runs Workers AI on env.AI.run through the core's chat-completions compatibility layer while ai("<provider>/<id>") resolves an id from pi-ai's generated Cloudflare AI Gateway catalog.

Vendor models go through AI Gateway's universal request with pi-ai's own converters building and parsing the vendor's body.

ai.provider is a pi-ai Provider (id "cloudflare") for a Models registry for use with harnesses.

The gateway catalog in pi-ai 0.99.2 uses the same spelling as AI Gateway (e.g. anthropic/claude-opus-4.8, openai/gpt-5.4) so the tests and docs use those ids.

The pi harness example takes its models from createAI instead of its own Workers AI provider, with provider id "cloudflare".

@earendil-works/pi-ai ^0.99.2 is an optional peer dependency.


Devin Review

@aron-cf
aron-cf added this pull request to stack #2447 October 2, 2026 08:53
@changeset-bot

changeset-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: f8d985a

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 2 packages
Name Type
agents Patch
@cloudflare/agent-think Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 5 potential issues.

Devin Review

Comment on lines +78 to +81
if (isGatewayShapedUrl(model.baseUrl)) {
const endpoint = GATEWAY_ENDPOINTS[model.api];
if (endpoint === undefined) throw noEndpoint(model.api);
return { endpoint, provider };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Groq catalog requests hit wrong endpoint

For a Groq catalog model, routeForModel sends v1/chat/completions instead of Groq's chat/completions. The gateway routes the request to the wrong path, so generation fails.

Learn more

Gateway catalog entries carry a gateway-shaped base URL, so this branch uses the same endpoint for all openai-completions providers. The Groq provider mapping strips openai/v1/ from Groq's URL, producing chat/completions, while this branch sends v1/chat/completions. The existing vendor routing test only exercises a Groq model object with a vendor URL, not the catalog-string form.

Example: ai("groq/llama-3.1-8b-instant") uses a gateway-shaped base URL. Its universal request targets Groq with v1/chat/completions; the vendor-URL form targets chat/completions.

Recommended fix: Resolve the gateway-shaped provider slug to its provider-specific endpoint, rather than using a common completions path. Add a catalog-string Groq request test alongside the vendor-URL test.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +405 to +412
if (!sawFinishReason) {
if (sawDone || this.blocks.length > 0) {
this.output.stopReason = hasToolCalls ? "toolUse" : "stop";
} else {
this.output.stopReason = "error";
this.output.errorMessage =
"The stream ended before the model produced any output.";
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Truncated completions become successful answers

When a chat stream closes after a text delta without a finish marker, finish assigns stop because content exists. Callers receive a partial answer as successful, so fallback and retry cannot recover it.

Learn more

A chat-completions response normally signals completion with a finish reason or [DONE]. The stream reader passes both indicators to finish. When a connection ends after delivering content but before either indicator, this branch uses blocks.length > 0 as proof of success and the dispatcher treats the resulting done event as final.

Example: A stream sends data: {"choices":[{"delta":{"content":"Hel"}}]} then closes. The result has stopReason: "stop" and text Hel, rather than an error for the incomplete response.

Recommended fix: Require a terminal finish reason or [DONE] before reporting success on SSE streams. Distinguish non-streaming JSON bodies, which are complete without SSE sentinels, and add an interrupted-stream test.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +110 to +121
if (model.reasoning === true) {
const effort = request.reasoningEffort;
if (typeof effort === "string") {
body.reasoning = {
effort: mappedEffort(model, effort),
summary: "auto"
};
// Encrypted reasoning items keep `store: false` multi-turn replay
// stateless, mirroring pi-ai's own Responses request shape.
body.include = ["reasoning.encrypted_content"];
} else if (model.thinkingLevelMap?.off !== null) {
body.reasoning = { effort: model.thinkingLevelMap?.off ?? "none" };

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Default reasoning loses stateless replay

When a Responses model reasons by default, include omits encrypted reasoning unless reasoningEffort is set. With store: false, subsequent turns cannot replay that reasoning item.

Learn more

OpenAI Responses reasoning items need reasoning.encrypted_content when store is false so a later turn can send them back. This request always sets store to false. A reasoning model such as gpt-5-mini declares thinkingLevelMap.off: null and can reason without an explicit level, but this branch requests encrypted items only when the effort is a string. The message converter replays the reasoning item on the next turn, which needs that encrypted content.

Example: Run ai.stream(ai("openai/gpt-5-mini"), context) without reasoning options. Its default reasoning item lacks encrypted content; the next turn cannot replay it with store: false.

Recommended fix: Request reasoning.encrypted_content for any reasoning model that can produce reasoning while store is false, including the no-explicit-effort case; test a two-turn default-reasoning exchange.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +92 to +103
if (!committed) {
if (event.type === "error" && isAbandonable(event) && !isLast) {
attempts.push({
errorMessage: event.error.errorMessage,
model: leg.model.id
});
lastError = event.error;
break;
}
committed = true;
if (pendingStart !== undefined) outer.push(pendingStart);
pendingStart = undefined;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Empty block prevents model fallback

If a leg emits text_start then errors before its first delta, streamWithFallback commits that leg. The next model never runs, although the failing leg produced no output.

Learn more

A fallback leg can announce a text, thinking, or tool-call block before delivering the first delta. The completion assembler emits text_start before text_delta, while a later read can fail. The dispatcher commits on the start event and refuses to try another leg, even though no content was emitted.

Example: The primary emits start, text_start, then error when its connection drops. The outer stream reports the error without attempting the configured fallback model.

Recommended fix: Buffer block-start events with start until an actual text, thinking, or tool-call delta arrives. If an error arrives first, discard the buffered events and try the next leg.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +210 to +221
const reasoningEffort =
resolved.reasoningEffort !== undefined
? resolved.reasoningEffort
: isThinkingLevel(explicitEffort)
? effortForLevel(model, explicitEffort)
: explicitEffort === "off"
? undefined
: typeof explicitEffort === "string"
? explicitEffort
: simple
? effortForLevel(model, options.reasoning)
: undefined;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Workers-only effort overrides vendor reasoning

For a Responses model with reasoningEffort, buildRequest prefers that setting over the call's reasoning. The Responses wire sends it to the vendor, changing the requested reasoning level.

Learn more

The model option reasoningEffort is designated for Workers AI, while pi-ai's call-level reasoning controls vendor thinking. resolveOptions includes the model's reasoningEffort, and this branch selects it before the call's reasoning level. The Responses request then puts it into body.reasoning.effort; the same option is only warned about and ignored by the Anthropic wire.

Example: ai.streamSimple(ai("openai/gpt-5.4", { reasoningEffort: "low" }), context, { reasoning: "high" }) sends reasoning.effort: "low" instead of "high".

Recommended fix: Ignore resolved.reasoningEffort for non-Workers models unless an explicit vendor call-level effort is supported. Preserve precedence for the pi reasoning call option and warn consistently when a Workers-only option is discarded.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@agent-think

agent-think Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

🟢 agents import sizes: 1 entry point changed, no growth

Entry point Exports Largest gzip change Size now
🆕 agents/models/pi-ai 9 new — 154.6 KiB
Changed exports (9)
Import Gzip change Size now
🆕 agents/models/pi-ai#createAI — 154.6 KiB
🆕 agents/models/pi-ai#CloudflareAIError — 1.2 KiB
🆕 agents/models/pi-ai#GATEWAY_PROVIDER_NAMES — 1.1 KiB
🆕 agents/models/pi-ai#FALLBACK_DIAGNOSTIC — 1 KiB
🆕 agents/models/pi-ai#CLOUDFLARE_AI_IMAGES_API — 1 KiB
🆕 agents/models/pi-ai#CLOUDFLARE_AI_API — 1 KiB
🆕 agents/models/pi-ai#CLOUDFLARE_ERROR_DIAGNOSTIC — 1 KiB
🆕 agents/models/pi-ai#CLOUDFLARE_PROVIDER_ID — 1 KiB
🆕 agents/models/pi-ai#CLOUDFLARE_DIAGNOSTIC — 1 KiB
How this works

Each runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement.

Compared f057bc7c → d09e2f09 · workflow run · reported by agent-think[bot]

@pkg-pr-new

pkg-pr-new Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

agents

npm i https://pkg.pr.new/agents@2446

@cloudflare/ai-chat

npm i https://pkg.pr.new/@cloudflare/ai-chat@2446

@cloudflare/codemode

npm i https://pkg.pr.new/@cloudflare/codemode@2446

hono-agents

npm i https://pkg.pr.new/hono-agents@2446

@cloudflare/shell

npm i https://pkg.pr.new/@cloudflare/shell@2446

@cloudflare/think

npm i https://pkg.pr.new/@cloudflare/think@2446

@cloudflare/voice

npm i https://pkg.pr.new/@cloudflare/voice@2446

@cloudflare/worker-bundler

npm i https://pkg.pr.new/@cloudflare/worker-bundler@2446

commit: f8d985a

aron-cf and others added 5 commits October 2, 2026 09:50
Add createAI for pi-ai over the model core. ai("@cf/...") runs Workers AI
on env.AI.run through the core's chat-completions compatibility layer;
ai("<provider>/<id>") resolves an id from pi-ai's generated Cloudflare
AI Gateway catalog; ai(model) takes any pi-ai model whose api is
anthropic-messages, openai-responses or openai-completions. Vendor
models go through AI Gateway's universal request with pi-ai's own
converters building and parsing the vendor's body, so the Worker holds
no vendor credential. ai.provider is a pi-ai Provider (id "cloudflare")
for a Models registry, which is what pi-durable's Harness takes.

Ported to pi-ai 0.99.2. Provider code now receives a TranscriptContext,
so the wires read tools from the transcript with getCurrentTools and
ai.stream normalizes a raw Context first. pi-ai's Anthropic
implementation calls client.beta.messages.create and chooses its own
beta features, which the wire now sends as anthropic-beta instead of
recomputing them. Diagnostic details are JSON, the strict OpenAI compat
profile gains 0.99's new fields, "max" maps to Workers AI's "high", and
ImagesModel is ImageModel. The gateway catalog in 0.99.2 spells ids as
AI Gateway does (anthropic/claude-opus-4.8, openai/gpt-5.4), and the
tests and docs use those ids.

The pi harness example takes its models from createAI instead of its own
Workers AI provider, with provider id "cloudflare".

Client-side fallback always ends its stream: a leg that throws, a stream
that rejects, a leg that ends without a done or error event, and a last
leg that produces nothing all end it with an error event. A gateway id
that differs from a catalog id only in separators, such as
anthropic/claude-opus-4-8 for anthropic/claude-opus-4.8, gets the catalog
id suggested in its error.

@earendil-works/pi-ai ^0.99.2 is an optional peer dependency.

Co-authored-by: Matt Carey <mcarey@cloudflare.com>
The chat-completions assembler treated any output as proof of a
complete stream, so an SSE connection that closed after some text but
before a finish reason or [DONE] was reported as a successful stop. A
complete stream ends with one of the two; without either, the stream
now ends with an error that keeps the partial content. A non-streaming
JSON body is complete as before.
…easons

The Responses wire sends store: false, so a reasoning item can only be
replayed on the next turn if it carries its encrypted content. The wire
asked for reasoning.encrypted_content only when a reasoning effort was
set, so a model that cannot turn reasoning off, such as gpt-5-mini,
reasoned without it by default. The include is now sent whenever the
model reasons.
streamWithFallback committed to a leg on its first event after start, so
a leg that emitted text_start, thinking_start or toolcall_start and then
failed before any content was never retried on the next model. Those
block starts carry no content, so they are now held back with start and
replayed in order once the leg produces output or a terminal event.
…odels

reasoningEffort in the pi-ai model options is documented as a Workers AI
setting, but buildRequest preferred it over the call's reasoning level
for every model, so the Responses and chat-completions wires sent it to
vendors: ai("openai/gpt-5.4", { reasoningEffort: "low" }) called with
reasoning: "high" asked for low. The setting now applies to Workers AI
models only. Vendor models follow the call's level and their own
metadata, and every wire records the dropped setting as a
cloudflare-compat diagnostic.
@aron-cf
aron-cf merged commit 5cf86f8 into main Oct 2, 2026
16 of 30 checks passed
@aron-cf
aron-cf deleted the models-pi-ai branch October 2, 2026 10:17
@github-actions github-actions Bot mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant