Skip to content

feat(agents): add experimental agents/models/ai-sdk - #2448

Merged
aron-cf merged 4 commits into
models-pi-aifrom
models-ai-sdk
Oct 2, 2026
Merged

aron-cf merged 4 commits into
models-pi-aifrom
models-ai-sdk

Conversation

@aron-cf

@aron-cf aron-cf commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Add createAI for the AI SDK over the model core. ai("@cf/...") runs Workers AI on env.AI.run with the core's chat-completions compatibility layer; ai(model) takes any AI SDK v4 model.

The provider is a full ProviderV4 with Workers AI embedding, image, transcription, speech and reranking models.

@ai-sdk/provider ^4.0.0 is an optional peer dependency. @ai-sdk/anthropic and @ai-sdk/openai are dev dependencies for the tests.

Also add docs/agents/models.md, link it from the pi-ai doc, and add examples/next/models, which exercises every call form of both providers. The example takes pi-ai ^0.99.2 and uses AI Gateway's spelling for catalog ids.


Devin Review

@aron-cf
aron-cf added this pull request to stack #2447 October 2, 2026 08:55
@changeset-bot

changeset-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: a5cbdda

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 2 packages
Name Type
agents Patch
@cloudflare/agent-think Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 7 potential issues.

Devin Review

Comment on lines +250 to +255
providerMetadata: {
[PROVIDER_OPTIONS_KEY]: {
...(first.providerMetadata[PROVIDER_OPTIONS_KEY] as JSONObject),
images: generated.map((image) => ({
mediaType: image.mediaType
})) as JSONArray

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Generated images lose gateway correlation

When generateImage merges image results, providerMetadata.cloudflare loses the gateway and log ID. Only images entries survive that merge, and those entries contain mediaType alone.

Learn more

The AI SDK combines image calls using the per-image entries in each provider metadata block. This result stores the gateway fields at the top level, but places only mediaType inside images. The merged result therefore retains the media type while dropping the model, gateway ID, log ID, and other correlation fields. The routed image model already accounts for this shape in metadataFor.

Example: generateImage({ model: ai.image('@cf/black-forest-labs/flux-1-schnell'), prompt: 'fox' }) receives a response with cf-aig-log-id: L1; its merged providerMetadata.cloudflare.images[0] contains only { mediaType: 'image/jpeg' }, not logId: 'L1'.

Recommended fix: Copy each answer's providerMetadata.cloudflare into the matching images element along with its media type. Preserve per-answer data for direct doGenerate({ n: 2 }) calls, including cases where different fallback legs answer the two images.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +161 to +171
if (UNSUPPORTED.has(this.modelId)) {
throw new CloudflareAIError({
code: "bad-request",
isRetryable: false,
message:
`${this.modelId} takes the audio as an object body, which the JSON ` +
"run path cannot carry. Use @cf/openai/whisper-large-v3-turbo.",
model: this.modelId,
requestBodyValues: undefined,
url: this.endpoint
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Unsupported transcription blocks its fallback

When @cf/deepgram/nova-3 has a Whisper fallback, doGenerate rejects before send can try it. The configured fallback never transcribes the audio.

Learn more

Transcription models use send to try a configured fallback after a failed leg. The Deepgram guard runs before send, so its exception bypasses that chain entirely. The same guard needs to apply to each attempted model, because Deepgram can also be configured as a fallback leg.

Example: ai.transcription('@cf/deepgram/nova-3', { fallback: ['@cf/openai/whisper-large-v3-turbo'] }) rejects before making a request, rather than transcribing with Whisper.

Recommended fix: Reject unsupported model IDs inside the per-leg request builder passed to send, or otherwise move the guard into the fallback loop. Keep the no-request guarantee for unsupported Deepgram legs while allowing the next leg to run.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +239 to +240
const first = answers[0];
const generated = answers.map(readImage);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Malformed image and speech answers skip fallback

When a 200 image answer lacks image, readImage fails after send completes its fallback loop. Speech answers without audio fail the same way, leaving configured fallback models unused.

Learn more

send considers a successful HTTP response a successful fallback leg after decoding its JSON or bytes. The image and speech subclasses validate required result fields only after send returns. If a model returns 200 with an incomplete JSON answer, the exception happens outside the fallback loop. The failure occurs before any image or audio is delivered, but a configured fallback cannot answer.

Example: A primary image model returns HTTP 200 with {}; readImage raises CloudflareAIError even if a configured secondary image model can generate the requested image.

Recommended fix: Perform modality-specific response validation inside each fallback attempt, for instance by letting send accept a per-leg answer parser. Apply the same pattern to speech so each invalid answer can advance to its fallback.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +101 to +114
function cloneWithFetch<T extends LanguageModelV4>(
model: T,
fetchImpl: typeof globalThis.fetch
): T {
const config = requireConfig(model);
const clone = Object.create(Object.getPrototypeOf(model)) as T;
Object.defineProperties(clone, Object.getOwnPropertyDescriptors(model));
Object.defineProperty(clone, "config", {
configurable: true,
enumerable: true,
value: { ...config, fetch: fetchImpl },
writable: true
});
return clone;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Vendor routing relies on an undocumented model property

The config.fetch swap depends on vendor model internals rather than the LanguageModelV4 contract. Provider updates or other vendor implementations need compatibility checks before the broad ai(model) guarantee holds.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread docs/agents/models.md
Comment on lines +136 to +149
## Why a model object and not an id

An id is only useful if someone keeps its catalog current. We keep exactly one:
Workers AI. Every other vendor already ships a maintained catalog inside its
own `@ai-sdk/*` package — ids, wire format, thinking levels, tool shapes, error
types — and that package is updated the day the vendor ships something. Copying
any of it here would mean a second, slower copy that is wrong every time a
vendor moves.

So `ai(model)` takes what the vendor built, clones it per call, swaps the
`fetch` its provider would have used, and sends the very same bytes to AI
Gateway. What comes back is the vendor's own response, parsed by the vendor's
own code. Your model object is never mutated — two concurrent calls never see
each other's transport.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Separate design rationale from the API reference

The extended vendor-catalog rationale reads as design explanation. The docs convention places design decisions in /design and keeps new API pages focused on usage and behavior.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +114 to +129
export default {
async fetch(request: Request, env: Env): Promise<Response> {
const url = new URL(request.url);

// The pi-ai twin of every route below lives under `/pi/*`.
if (url.pathname === "/pi" || url.pathname.startsWith("/pi/")) {
return handlePi(url, env);
}

// One provider over the whole catalog. Inside a Worker the binding is
// keyless; `gateway` here is a provider-wide default that per-model and
// per-call options override.
const ai = createAI({
binding: env.AI,
gateway: { id: "default", metadata: { app: "next-models" } }
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟥 Public model routes allow unbounded billing

Without authentication, callers can invoke fetch routes that spend the Worker's AI account balance. Image, speech, vendor and text requests all reach billable model calls.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment on lines +63 to +81
function modelOf(url: URL, fallback = DEFAULT_MODEL): WorkersAIModelId {
// Validated by `ai()`, which throws a TypeError naming the fix for anything
// that is not a `@cf/` id.
return (url.searchParams.get("model") ?? fallback) as WorkersAIModelId;
}

/**
* `?vendor=<slug>:<id>` builds a third-party model with its own provider. A
* vendor id is never a string this provider accepts: only its own package
* knows that catalog.
*/
function vendorOf(url: URL): LanguageModelV4 | undefined {
const spec = url.searchParams.get("vendor");
if (spec === null) return undefined;
const separator = spec.indexOf(":");
const slug = separator === -1 ? spec : spec.slice(0, separator);
const id = separator === -1 ? "" : spec.slice(separator + 1);
if (slug === "anthropic") return anthropic(id || "claude-opus-4-8");
if (slug === "openai") return openai.responses(id || "gpt-5-mini");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟨 Unvalidated model requests exhaust inference budget

Callers control model, vendor, and prompt without size or catalog limits. Oversized prompts and costly model choices can spend the deployed Worker's AI budget.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@agent-think

agent-think Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

🟢 agents import sizes: 1 entry point changed, no growth

Entry point Exports Largest gzip change Size now
🆕 agents/models/ai-sdk 3 new — 19.3 KiB
Changed exports (3)
Import Gzip change Size now
🆕 agents/models/ai-sdk#createAI — 19.3 KiB
🆕 agents/models/ai-sdk#GatewayLanguageModel — 14.9 KiB
🆕 agents/models/ai-sdk#CloudflareAIError — 3 KiB
How this works

Each runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement.

Compared f8d985a5 → a5cbddaf · workflow run · reported by agent-think[bot]

@pkg-pr-new

pkg-pr-new Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Open in StackBlitz

agents

npm i https://pkg.pr.new/agents@2448

@cloudflare/ai-chat

npm i https://pkg.pr.new/@cloudflare/ai-chat@2448

@cloudflare/codemode

npm i https://pkg.pr.new/@cloudflare/codemode@2448

hono-agents

npm i https://pkg.pr.new/hono-agents@2448

@cloudflare/shell

npm i https://pkg.pr.new/@cloudflare/shell@2448

@cloudflare/think

npm i https://pkg.pr.new/@cloudflare/think@2448

@cloudflare/voice

npm i https://pkg.pr.new/@cloudflare/voice@2448

@cloudflare/worker-bundler

npm i https://pkg.pr.new/@cloudflare/worker-bundler@2448

commit: a5cbdda

aron-cf and others added 4 commits October 2, 2026 09:50
Add createAI for the AI SDK over the model core. ai("@cf/...") runs
Workers AI on env.AI.run with the core's chat-completions compatibility
layer; ai(model) takes any AI SDK v4 model the caller built, clones it
per call, and swaps its fetch for AI Gateway's universal request, so the
vendor's own provider builds and parses the request while the gateway
holds the credential. The provider is a full ProviderV4 with Workers AI
embedding, image, transcription, speech and reranking models, gateway
options at three layers, client-side fallback, and
providerMetadata.cloudflare on every result.

@ai-sdk/provider ^4.0.0 is an optional peer dependency.
@ai-sdk/anthropic and @ai-sdk/openai are dev dependencies for the tests,
and scripts/use-ai-sdk-major.mjs moves @ai-sdk/provider with the other
AI SDK packages.

Also add docs/agents/models.md, link it from the pi-ai doc, and add
examples/next/models, which exercises every call form of both providers.
The example takes pi-ai ^0.99.2 and uses AI Gateway's spelling for
catalog ids.

Co-authored-by: Matt Carey <mcarey@cloudflare.com>
generateImage merges the calls it fans an n out into by spreading each
result's providerMetadata.cloudflare.images, and only those entries
survive. The Workers AI image model put the gateway correlation at the
top of the block and only mediaType inside images, so the merged result
lost the model, gateway and log id. Each images entry now carries its
own answer's metadata, which may come from a different fallback leg.
…odel

The Deepgram recognition models take audio as an object body that the
JSON run path cannot carry, and the transcription model refused them
before its fallback loop ran, so a fallback such as Whisper never got
the chance to answer. The check now runs per leg: an unsupported leg is
skipped without a request and the next one runs.
The modality fallback loop treated any 2xx answer as success, and the
image and speech models checked for their image or audio only after the
loop had returned. A 200 answer without one failed the call even when a
fallback model was configured. send now takes a per-leg check that runs
inside the loop, and the image and speech models use it, so an
incomplete answer moves on to the next leg.
@aron-cf
aron-cf merged commit 54f9ca7 into main Oct 2, 2026
16 of 30 checks passed
@aron-cf
aron-cf deleted the models-ai-sdk branch October 2, 2026 10:17
@github-actions github-actions Bot mentioned this pull request Oct 2, 2026
brandhaug added a commit to brandhaug/b2b-saas-starter that referenced this pull request Oct 5, 2026
## pnpm-workspace.yaml (default)

## Dependency Updates

| Package | From | To | Type |
| --- | --- | --- | --- |
| `agents` | 0.24.0 | 0.25.0 | minor |

## Release Notes

<details>
<summary><b>agents</b> (0.24.0 → 0.25.0)</summary>

### Minor Changes

- [#2005](cloudflare/agents#2005)
[`c2f7672`](cloudflare/agents@c2f7672)
Thanks [@<!---->cjol](https://github.com/cjol)! - Native RPC calls to
async Agent and Think methods now start lifecycle initialization first;
address Agents by name because raw IDs from `newUniqueId()` and
`idFromString()` now fail their first async RPC. See
[Lifecycle](https://github.com/cloudflare/agents/blob/main/docs/agents/lifecycle.md).

### Patch Changes

- [#2390](cloudflare/agents#2390)
[`c55ec80`](cloudflare/agents@c55ec80)
Thanks [@<!---->threepointone](https://github.com/threepointone)! -
Report a failed agent-tool child as failed even when it was evicted
before recording the failure. See [Agent
tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md).

- [#2384](cloudflare/agents#2384)
[`f904999`](cloudflare/agents@f904999)
Thanks [@<!---->threepointone](https://github.com/threepointone)! - Fix
agent-tool chunks being duplicated or dropped on reconnect, child
re-attach, and fiber recovery. See [Agent
tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md).

- [#2364](cloudflare/agents#2364)
[`5e0507e`](cloudflare/agents@5e0507e)
Thanks [@<!---->threepointone](https://github.com/threepointone)! - Add
`eventDelivery: "terminal"` to `runAgentTool` to forward only lifecycle,
progress, and milestone events for a run. See [Agent
tools](https://github.com/cloudflare/agents/blob/main/docs/agents/agent-tools.md).

- [#2448](cloudflare/agents#2448)
[`54f9ca7`](cloudflare/agents@54f9ca7)
Thanks [@<!---->aron-cf](https://github.com/aron-cf)!

…[full
notes](https://github.com/cloudflare/agents/releases/tag/agents%400.25.0)

</details>

---
*This PR was auto-generated by
[catalog-update-action](https://github.com/brandhaug/catalog-update-action).*

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant