Add a model provider
Chat and the generator depend only on AiProvider (api/src/ai/provider.ts). A new model is one new adapter file plus one line in makeAiProvider. The turn loop, the tools, the events and the limits stay the same.
1. Implement the port
Create api/src/ai/<vendor>.ts:
api/src/ai/acme.ts
import type { AiProvider, JsonRequest, TurnEvent, TurnRequest } from './provider';
export function makeAcmeProvider(cfg: { apiKey?: string; model: string }): AiProvider {
return {
configured: Boolean(cfg.apiKey),
model: cfg.model,
generatorModel: cfg.model,
liveModel: 'none',
async *streamTurn(req: TurnRequest): AsyncIterable<TurnEvent> {
// 1. Translate req.system, req.messages and req.tools into the vendor's request.
// 2. Stream: yield { type: 'text', text } for each text chunk as it arrives.
// 3. Collect tool calls; yield { type: 'tool_calls', calls } once, after the stream.
// 4. Yield { type: 'done', usage, providerState } last.
},
async generateJson(req: JsonRequest): Promise<unknown> {
// One call in the vendor's JSON mode. Parse and return; throw with a clear message if it is not JSON.
},
};
}
What the loop relies on:
| Contract | Detail |
|---|---|
| Event order | Any number of text, at most one tool_calls, exactly one done, last. |
| Tool call ids | Return the vendor's call id in ToolCall.id so results can be matched; an empty string is allowed if the vendor has none. |
| Message translation | user text; assistant with optional text, tool calls and providerState; tool with results, where isError results should reach the model as errors. |
providerState | Anything the vendor needs replayed verbatim on the next round (Gemini's signed parts). The loop stores it and hands it back untouched. If the vendor needs nothing, omit it. |
| Usage | Input, output and cache-read tokens. Use 0 where the vendor does not report one. |
| Cancellation | Honour req.signal. |
thinkingLevel | Map low, medium, high to whatever the vendor offers, or ignore it. |
configured | false when there is no key: chat then answers with an error event instead of failing. |
Tool parameters arrive as JSON Schema objects. Most vendors accept that directly; translate if yours does not.
2. Select it
makeAiProvider in api/src/ai/index.ts is the one place an adapter is chosen. Extend AI_PROVIDER:
api/src/ai/index.ts
if (env.AI_PROVIDER === 'scripted') return makeScriptedProvider();
if (env.AI_PROVIDER === 'acme') return makeAcmeProvider({ apiKey: env.ACME_API_KEY, model: env.ACME_MODEL || 'acme-large' });
Add the new variables to Env in api/src/env.ts, to wrangler.jsonc (non-secret) or as secrets, and to the configuration reference.
3. Test it
- Unit-test the translation both ways with the vendor's SDK mocked: messages in, events out,
providerStatereplayed. - The rest of the pipeline is already covered against the scripted adapter, so there is no need to re-test the loop.
- Run a real conversation through the Prompt Lab's Try It with a few capabilities, and watch the trace for tool-calling quirks.
What does not come for free
- Voice. The relay in
api/src/voice/relay.tsspeaks the Gemini Live protocol directly. Another realtime model needs its own relay. - AI Gateway. If the vendor is supported by Cloudflare AI Gateway, add its provider path in
gateway.tsrather than building URLs in the adapter. - Prompt tuning. The playbooks were written and tuned against Gemini. Expect to revisit some, and use the feedback report to see where.