Repository navigation
feat(agents): add experimental agents/models/ai-sdk - #2448
Conversation
🦋 Changeset detectedLatest commit: a5cbdda The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
| providerMetadata: { | ||
| [PROVIDER_OPTIONS_KEY]: { | ||
| ...(first.providerMetadata[PROVIDER_OPTIONS_KEY] as JSONObject), | ||
| images: generated.map((image) => ({ | ||
| mediaType: image.mediaType | ||
| })) as JSONArray |
There was a problem hiding this comment.
🟡 Generated images lose gateway correlation
When generateImage merges image results, providerMetadata.cloudflare loses the gateway and log ID. Only images entries survive that merge, and those entries contain mediaType alone.
Learn more
The AI SDK combines image calls using the per-image entries in each provider metadata block. This result stores the gateway fields at the top level, but places only mediaType inside images. The merged result therefore retains the media type while dropping the model, gateway ID, log ID, and other correlation fields. The routed image model already accounts for this shape in metadataFor.
Example: generateImage({ model: ai.image('@cf/black-forest-labs/flux-1-schnell'), prompt: 'fox' }) receives a response with cf-aig-log-id: L1; its merged providerMetadata.cloudflare.images[0] contains only { mediaType: 'image/jpeg' }, not logId: 'L1'.
Recommended fix: Copy each answer's providerMetadata.cloudflare into the matching images element along with its media type. Preserve per-answer data for direct doGenerate({ n: 2 }) calls, including cases where different fallback legs answer the two images.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (UNSUPPORTED.has(this.modelId)) { | ||
| throw new CloudflareAIError({ | ||
| code: "bad-request", | ||
| isRetryable: false, | ||
| message: | ||
| `${this.modelId} takes the audio as an object body, which the JSON ` + | ||
| "run path cannot carry. Use @cf/openai/whisper-large-v3-turbo.", | ||
| model: this.modelId, | ||
| requestBodyValues: undefined, | ||
| url: this.endpoint | ||
| }); |
There was a problem hiding this comment.
🟡 Unsupported transcription blocks its fallback
When @cf/deepgram/nova-3 has a Whisper fallback, doGenerate rejects before send can try it. The configured fallback never transcribes the audio.
Learn more
Transcription models use send to try a configured fallback after a failed leg. The Deepgram guard runs before send, so its exception bypasses that chain entirely. The same guard needs to apply to each attempted model, because Deepgram can also be configured as a fallback leg.
Example: ai.transcription('@cf/deepgram/nova-3', { fallback: ['@cf/openai/whisper-large-v3-turbo'] }) rejects before making a request, rather than transcribing with Whisper.
Recommended fix: Reject unsupported model IDs inside the per-leg request builder passed to send, or otherwise move the guard into the fallback loop. Keep the no-request guarantee for unsupported Deepgram legs while allowing the next leg to run.
Was this helpful? React with 👍 or 👎 to provide feedback.
| const first = answers[0]; | ||
| const generated = answers.map(readImage); |
There was a problem hiding this comment.
🟡 Malformed image and speech answers skip fallback
When a 200 image answer lacks image, readImage fails after send completes its fallback loop. Speech answers without audio fail the same way, leaving configured fallback models unused.
Learn more
send considers a successful HTTP response a successful fallback leg after decoding its JSON or bytes. The image and speech subclasses validate required result fields only after send returns. If a model returns 200 with an incomplete JSON answer, the exception happens outside the fallback loop. The failure occurs before any image or audio is delivered, but a configured fallback cannot answer.
Example: A primary image model returns HTTP 200 with {}; readImage raises CloudflareAIError even if a configured secondary image model can generate the requested image.
Recommended fix: Perform modality-specific response validation inside each fallback attempt, for instance by letting send accept a per-leg answer parser. Apply the same pattern to speech so each invalid answer can advance to its fallback.
Was this helpful? React with 👍 or 👎 to provide feedback.
| function cloneWithFetch<T extends LanguageModelV4>( | ||
| model: T, | ||
| fetchImpl: typeof globalThis.fetch | ||
| ): T { | ||
| const config = requireConfig(model); | ||
| const clone = Object.create(Object.getPrototypeOf(model)) as T; | ||
| Object.defineProperties(clone, Object.getOwnPropertyDescriptors(model)); | ||
| Object.defineProperty(clone, "config", { | ||
| configurable: true, | ||
| enumerable: true, | ||
| value: { ...config, fetch: fetchImpl }, | ||
| writable: true | ||
| }); | ||
| return clone; |
There was a problem hiding this comment.
🔍 Vendor routing relies on an undocumented model property
The config.fetch swap depends on vendor model internals rather than the LanguageModelV4 contract. Provider updates or other vendor implementations need compatibility checks before the broad ai(model) guarantee holds.
Was this helpful? React with 👍 or 👎 to provide feedback.
| ## Why a model object and not an id | ||
|
|
||
| An id is only useful if someone keeps its catalog current. We keep exactly one: | ||
| Workers AI. Every other vendor already ships a maintained catalog inside its | ||
| own `@ai-sdk/*` package — ids, wire format, thinking levels, tool shapes, error | ||
| types — and that package is updated the day the vendor ships something. Copying | ||
| any of it here would mean a second, slower copy that is wrong every time a | ||
| vendor moves. | ||
|
|
||
| So `ai(model)` takes what the vendor built, clones it per call, swaps the | ||
| `fetch` its provider would have used, and sends the very same bytes to AI | ||
| Gateway. What comes back is the vendor's own response, parsed by the vendor's | ||
| own code. Your model object is never mutated — two concurrent calls never see | ||
| each other's transport. |
There was a problem hiding this comment.
| export default { | ||
| async fetch(request: Request, env: Env): Promise<Response> { | ||
| const url = new URL(request.url); | ||
|
|
||
| // The pi-ai twin of every route below lives under `/pi/*`. | ||
| if (url.pathname === "/pi" || url.pathname.startsWith("/pi/")) { | ||
| return handlePi(url, env); | ||
| } | ||
|
|
||
| // One provider over the whole catalog. Inside a Worker the binding is | ||
| // keyless; `gateway` here is a provider-wide default that per-model and | ||
| // per-call options override. | ||
| const ai = createAI({ | ||
| binding: env.AI, | ||
| gateway: { id: "default", metadata: { app: "next-models" } } | ||
| }); |
There was a problem hiding this comment.
| function modelOf(url: URL, fallback = DEFAULT_MODEL): WorkersAIModelId { | ||
| // Validated by `ai()`, which throws a TypeError naming the fix for anything | ||
| // that is not a `@cf/` id. | ||
| return (url.searchParams.get("model") ?? fallback) as WorkersAIModelId; | ||
| } | ||
|
|
||
| /** | ||
| * `?vendor=<slug>:<id>` builds a third-party model with its own provider. A | ||
| * vendor id is never a string this provider accepts: only its own package | ||
| * knows that catalog. | ||
| */ | ||
| function vendorOf(url: URL): LanguageModelV4 | undefined { | ||
| const spec = url.searchParams.get("vendor"); | ||
| if (spec === null) return undefined; | ||
| const separator = spec.indexOf(":"); | ||
| const slug = separator === -1 ? spec : spec.slice(0, separator); | ||
| const id = separator === -1 ? "" : spec.slice(separator + 1); | ||
| if (slug === "anthropic") return anthropic(id || "claude-opus-4-8"); | ||
| if (slug === "openai") return openai.responses(id || "gpt-5-mini"); |
There was a problem hiding this comment.
🟢 agents import sizes: 1 entry point changed, no growth
Changed exports (3)
How this worksEach runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement. Compared |
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
Add createAI for the AI SDK over the model core. ai("@cf/...") runs
Workers AI on env.AI.run with the core's chat-completions compatibility
layer; ai(model) takes any AI SDK v4 model the caller built, clones it
per call, and swaps its fetch for AI Gateway's universal request, so the
vendor's own provider builds and parses the request while the gateway
holds the credential. The provider is a full ProviderV4 with Workers AI
embedding, image, transcription, speech and reranking models, gateway
options at three layers, client-side fallback, and
providerMetadata.cloudflare on every result.
@ai-sdk/provider ^4.0.0 is an optional peer dependency.
@ai-sdk/anthropic and @ai-sdk/openai are dev dependencies for the tests,
and scripts/use-ai-sdk-major.mjs moves @ai-sdk/provider with the other
AI SDK packages.
Also add docs/agents/models.md, link it from the pi-ai doc, and add
examples/next/models, which exercises every call form of both providers.
The example takes pi-ai ^0.99.2 and uses AI Gateway's spelling for
catalog ids.
Co-authored-by: Matt Carey <mcarey@cloudflare.com>
generateImage merges the calls it fans an n out into by spreading each result's providerMetadata.cloudflare.images, and only those entries survive. The Workers AI image model put the gateway correlation at the top of the block and only mediaType inside images, so the merged result lost the model, gateway and log id. Each images entry now carries its own answer's metadata, which may come from a different fallback leg.
…odel The Deepgram recognition models take audio as an object body that the JSON run path cannot carry, and the transcription model refused them before its fallback loop ran, so a fallback such as Whisper never got the chance to answer. The check now runs per leg: an unsupported leg is skipped without a request and the next one runs.
The modality fallback loop treated any 2xx answer as success, and the image and speech models checked for their image or audio only after the loop had returned. A 200 answer without one failed the call even when a fallback model was configured. send now takes a per-leg check that runs inside the loop, and the image and speech models use it, so an incomplete answer moves on to the next leg.
## pnpm-workspace.yaml (default) ## Dependency Updates | Package | From | To | Type | | --- | --- | --- | --- | | `agents` | 0.24.0 | 0.25.0 | minor | ## Release Notes <details> <summary><b>agents</b> (0.24.0 → 0.25.0)</summary> ### Minor Changes - [#2005](cloudflare/agents#2005) [`c2f7672`](cloudflare/agents@c2f7672) Thanks [@<!---->cjol](https://github.com/cjol)! - Native RPC calls to async Agent and Think methods now start lifecycle initialization first; address Agents by name because raw IDs from `newUniqueId()` and `idFromString()` now fail their first async RPC. See [Lifecycle](https://github.com/cloudflare/agents/blob/main/docs/agents/lifecycle.md). ### Patch Changes - [#2390](cloudflare/agents#2390) [`c55ec80`](cloudflare/agents@c55ec80) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Report a failed agent-tool child as failed even when it was evicted before recording the failure. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2384](cloudflare/agents#2384) [`f904999`](cloudflare/agents@f904999) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Fix agent-tool chunks being duplicated or dropped on reconnect, child re-attach, and fiber recovery. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2364](cloudflare/agents#2364) [`5e0507e`](cloudflare/agents@5e0507e) Thanks [@<!---->threepointone](https://github.com/threepointone)! - Add `eventDelivery: "terminal"` to `runAgentTool` to forward only lifecycle, progress, and milestone events for a run. See [Agent tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md). - [#2448](cloudflare/agents#2448) [`54f9ca7`](cloudflare/agents@54f9ca7) Thanks [@<!---->aron-cf](https://github.com/aron-cf)! …[full notes](https://github.com/cloudflare/agents/releases/tag/agents%400.25.0) </details> --- *This PR was auto-generated by [catalog-update-action](https://github.com/brandhaug/catalog-update-action).* Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Add createAI for the AI SDK over the model core.
ai("@cf/...")runs Workers AI on env.AI.run with the core's chat-completions compatibility layer;ai(model)takes any AI SDK v4 model.The provider is a full ProviderV4 with Workers AI embedding, image, transcription, speech and reranking models.
@ai-sdk/provider ^4.0.0 is an optional peer dependency. @ai-sdk/anthropic and @ai-sdk/openai are dev dependencies for the tests.
Also add docs/agents/models.md, link it from the pi-ai doc, and add examples/next/models, which exercises every call form of both providers. The example takes pi-ai ^0.99.2 and uses AI Gateway's spelling for catalog ids.