Repository navigation
feat(agents): add experimental agents/models/pi-ai - #2446
Conversation
🦋 Changeset detectedLatest commit: f8d985a The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
| if (isGatewayShapedUrl(model.baseUrl)) { | ||
| const endpoint = GATEWAY_ENDPOINTS[model.api]; | ||
| if (endpoint === undefined) throw noEndpoint(model.api); | ||
| return { endpoint, provider }; |
There was a problem hiding this comment.
🔴 Groq catalog requests hit wrong endpoint
For a Groq catalog model, routeForModel sends v1/chat/completions instead of Groq's chat/completions. The gateway routes the request to the wrong path, so generation fails.
Learn more
Gateway catalog entries carry a gateway-shaped base URL, so this branch uses the same endpoint for all openai-completions providers. The Groq provider mapping strips openai/v1/ from Groq's URL, producing chat/completions, while this branch sends v1/chat/completions. The existing vendor routing test only exercises a Groq model object with a vendor URL, not the catalog-string form.
Example: ai("groq/llama-3.1-8b-instant") uses a gateway-shaped base URL. Its universal request targets Groq with v1/chat/completions; the vendor-URL form targets chat/completions.
Recommended fix: Resolve the gateway-shaped provider slug to its provider-specific endpoint, rather than using a common completions path. Add a catalog-string Groq request test alongside the vendor-URL test.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (!sawFinishReason) { | ||
| if (sawDone || this.blocks.length > 0) { | ||
| this.output.stopReason = hasToolCalls ? "toolUse" : "stop"; | ||
| } else { | ||
| this.output.stopReason = "error"; | ||
| this.output.errorMessage = | ||
| "The stream ended before the model produced any output."; | ||
| } |
There was a problem hiding this comment.
🔴 Truncated completions become successful answers
When a chat stream closes after a text delta without a finish marker, finish assigns stop because content exists. Callers receive a partial answer as successful, so fallback and retry cannot recover it.
Learn more
A chat-completions response normally signals completion with a finish reason or [DONE]. The stream reader passes both indicators to finish. When a connection ends after delivering content but before either indicator, this branch uses blocks.length > 0 as proof of success and the dispatcher treats the resulting done event as final.
Example: A stream sends data: {"choices":[{"delta":{"content":"Hel"}}]} then closes. The result has stopReason: "stop" and text Hel, rather than an error for the incomplete response.
Recommended fix: Require a terminal finish reason or [DONE] before reporting success on SSE streams. Distinguish non-streaming JSON bodies, which are complete without SSE sentinels, and add an interrupted-stream test.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (model.reasoning === true) { | ||
| const effort = request.reasoningEffort; | ||
| if (typeof effort === "string") { | ||
| body.reasoning = { | ||
| effort: mappedEffort(model, effort), | ||
| summary: "auto" | ||
| }; | ||
| // Encrypted reasoning items keep `store: false` multi-turn replay | ||
| // stateless, mirroring pi-ai's own Responses request shape. | ||
| body.include = ["reasoning.encrypted_content"]; | ||
| } else if (model.thinkingLevelMap?.off !== null) { | ||
| body.reasoning = { effort: model.thinkingLevelMap?.off ?? "none" }; |
There was a problem hiding this comment.
🔴 Default reasoning loses stateless replay
When a Responses model reasons by default, include omits encrypted reasoning unless reasoningEffort is set. With store: false, subsequent turns cannot replay that reasoning item.
Learn more
OpenAI Responses reasoning items need reasoning.encrypted_content when store is false so a later turn can send them back. This request always sets store to false. A reasoning model such as gpt-5-mini declares thinkingLevelMap.off: null and can reason without an explicit level, but this branch requests encrypted items only when the effort is a string. The message converter replays the reasoning item on the next turn, which needs that encrypted content.
Example: Run ai.stream(ai("openai/gpt-5-mini"), context) without reasoning options. Its default reasoning item lacks encrypted content; the next turn cannot replay it with store: false.
Recommended fix: Request reasoning.encrypted_content for any reasoning model that can produce reasoning while store is false, including the no-explicit-effort case; test a two-turn default-reasoning exchange.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (!committed) { | ||
| if (event.type === "error" && isAbandonable(event) && !isLast) { | ||
| attempts.push({ | ||
| errorMessage: event.error.errorMessage, | ||
| model: leg.model.id | ||
| }); | ||
| lastError = event.error; | ||
| break; | ||
| } | ||
| committed = true; | ||
| if (pendingStart !== undefined) outer.push(pendingStart); | ||
| pendingStart = undefined; |
There was a problem hiding this comment.
🟡 Empty block prevents model fallback
If a leg emits text_start then errors before its first delta, streamWithFallback commits that leg. The next model never runs, although the failing leg produced no output.
Learn more
A fallback leg can announce a text, thinking, or tool-call block before delivering the first delta. The completion assembler emits text_start before text_delta, while a later read can fail. The dispatcher commits on the start event and refuses to try another leg, even though no content was emitted.
Example: The primary emits start, text_start, then error when its connection drops. The outer stream reports the error without attempting the configured fallback model.
Recommended fix: Buffer block-start events with start until an actual text, thinking, or tool-call delta arrives. If an error arrives first, discard the buffered events and try the next leg.
Was this helpful? React with 👍 or 👎 to provide feedback.
| const reasoningEffort = | ||
| resolved.reasoningEffort !== undefined | ||
| ? resolved.reasoningEffort | ||
| : isThinkingLevel(explicitEffort) | ||
| ? effortForLevel(model, explicitEffort) | ||
| : explicitEffort === "off" | ||
| ? undefined | ||
| : typeof explicitEffort === "string" | ||
| ? explicitEffort | ||
| : simple | ||
| ? effortForLevel(model, options.reasoning) | ||
| : undefined; |
There was a problem hiding this comment.
🟡 Workers-only effort overrides vendor reasoning
For a Responses model with reasoningEffort, buildRequest prefers that setting over the call's reasoning. The Responses wire sends it to the vendor, changing the requested reasoning level.
Learn more
The model option reasoningEffort is designated for Workers AI, while pi-ai's call-level reasoning controls vendor thinking. resolveOptions includes the model's reasoningEffort, and this branch selects it before the call's reasoning level. The Responses request then puts it into body.reasoning.effort; the same option is only warned about and ignored by the Anthropic wire.
Example: ai.streamSimple(ai("openai/gpt-5.4", { reasoningEffort: "low" }), context, { reasoning: "high" }) sends reasoning.effort: "low" instead of "high".
Recommended fix: Ignore resolved.reasoningEffort for non-Workers models unless an explicit vendor call-level effort is supported. Preserve precedence for the pi reasoning call option and warn consistently when a Workers-only option is discarded.
Was this helpful? React with 👍 or 👎 to provide feedback.
🟢 agents import sizes: 1 entry point changed, no growth
Changed exports (9)
How this worksEach runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement. Compared |
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
Add createAI for pi-ai over the model core. ai("@cf/...") runs Workers AI
on env.AI.run through the core's chat-completions compatibility layer;
ai("<provider>/<id>") resolves an id from pi-ai's generated Cloudflare
AI Gateway catalog; ai(model) takes any pi-ai model whose api is
anthropic-messages, openai-responses or openai-completions. Vendor
models go through AI Gateway's universal request with pi-ai's own
converters building and parsing the vendor's body, so the Worker holds
no vendor credential. ai.provider is a pi-ai Provider (id "cloudflare")
for a Models registry, which is what pi-durable's Harness takes.
Ported to pi-ai 0.99.2. Provider code now receives a TranscriptContext,
so the wires read tools from the transcript with getCurrentTools and
ai.stream normalizes a raw Context first. pi-ai's Anthropic
implementation calls client.beta.messages.create and chooses its own
beta features, which the wire now sends as anthropic-beta instead of
recomputing them. Diagnostic details are JSON, the strict OpenAI compat
profile gains 0.99's new fields, "max" maps to Workers AI's "high", and
ImagesModel is ImageModel. The gateway catalog in 0.99.2 spells ids as
AI Gateway does (anthropic/claude-opus-4.8, openai/gpt-5.4), and the
tests and docs use those ids.
The pi harness example takes its models from createAI instead of its own
Workers AI provider, with provider id "cloudflare".
Client-side fallback always ends its stream: a leg that throws, a stream
that rejects, a leg that ends without a done or error event, and a last
leg that produces nothing all end it with an error event. A gateway id
that differs from a catalog id only in separators, such as
anthropic/claude-opus-4-8 for anthropic/claude-opus-4.8, gets the catalog
id suggested in its error.
@earendil-works/pi-ai ^0.99.2 is an optional peer dependency.
Co-authored-by: Matt Carey <mcarey@cloudflare.com>
The chat-completions assembler treated any output as proof of a complete stream, so an SSE connection that closed after some text but before a finish reason or [DONE] was reported as a successful stop. A complete stream ends with one of the two; without either, the stream now ends with an error that keeps the partial content. A non-streaming JSON body is complete as before.
…easons The Responses wire sends store: false, so a reasoning item can only be replayed on the next turn if it carries its encrypted content. The wire asked for reasoning.encrypted_content only when a reasoning effort was set, so a model that cannot turn reasoning off, such as gpt-5-mini, reasoned without it by default. The include is now sent whenever the model reasons.
streamWithFallback committed to a leg on its first event after start, so a leg that emitted text_start, thinking_start or toolcall_start and then failed before any content was never retried on the next model. Those block starts carry no content, so they are now held back with start and replayed in order once the leg produces output or a terminal event.
…odels
reasoningEffort in the pi-ai model options is documented as a Workers AI
setting, but buildRequest preferred it over the call's reasoning level
for every model, so the Responses and chat-completions wires sent it to
vendors: ai("openai/gpt-5.4", { reasoningEffort: "low" }) called with
reasoning: "high" asked for low. The setting now applies to Workers AI
models only. Vendor models follow the call's level and their own
metadata, and every wire records the dropped setting as a
cloudflare-compat diagnostic.
Add
createAI()for pi-ai over the model core.ai("@cf/...")runs Workers AI on env.AI.run through the core's chat-completions compatibility layer whileai("<provider>/<id>")resolves an id from pi-ai's generated Cloudflare AI Gateway catalog.Vendor models go through AI Gateway's universal request with pi-ai's own converters building and parsing the vendor's body.
ai.provideris a pi-ai Provider (id "cloudflare") for a Models registry for use with harnesses.The gateway catalog in pi-ai 0.99.2 uses the same spelling as AI Gateway (e.g. anthropic/claude-opus-4.8, openai/gpt-5.4) so the tests and docs use those ids.
The pi harness example takes its models from createAI instead of its own Workers AI provider, with provider id "cloudflare".
@earendil-works/pi-ai ^0.99.2 is an optional peer dependency.