Skip to main content

AI providers

Chat, voice set-up and the claim generator depend on one interface, AiProvider in api/src/ai/provider.ts, never on a vendor SDK. The one place an adapter is chosen is makeAiProvider(env) in api/src/ai/index.ts.

The port​

api/src/ai/provider.ts (abridged)
export interface AiProvider {
readonly configured: boolean; // false without a key: chat answers with an error event
readonly model: string; // reported by /api/health and in the done event
readonly generatorModel: string;
readonly liveModel: string;
readonly thinkingLevel?: ThinkingLevel;

/** One model call with function calling, streamed. The caller runs the tools and calls again. */
streamTurn(req: TurnRequest): AsyncIterable<TurnEvent>;

/** One JSON-mode call, for the generator. */
generateJson(req: JsonRequest): Promise<unknown>;
}

export type TurnEvent =
| { type: 'text'; text: string }
| { type: 'tool_calls'; calls: ToolCall[] }
| { type: 'done'; usage: Usage; providerState?: unknown };

Messages are provider-neutral: user text, assistant turns (text, tool calls and an opaque providerState), and tool results. Tool definitions are plain JSON Schema. The adapter translates both ways.

The port does not run tools or loop. That is runChatTurn's job, so the loop, the limits and the events are the same whatever the model.

Gemini​

api/src/ai/gemini.ts uses the official @google/genai SDK, which is fetch-based and runs on Workers.

  • Streaming with function calling. generateContentStream with the system instruction, the tools as function declarations (parameters passed as JSON Schema; parameterless tools omit them), maxOutputTokens and the abort signal.
  • Thought parts are skipped. Text parts stream out as text events. Function calls are collected and emitted once as tool_calls.
  • Thought signatures are kept. Gemini 3 signs function-call parts with a thoughtSignature that must be echoed back when the turn is replayed. The adapter rebuilds the model's own parts as it streams and returns them as providerState. When an assistant turn has providerState, toContents replays those parts verbatim instead of rebuilding them from text and calls.
  • Usage maps promptTokenCount, candidatesTokenCount + thoughtsTokenCount and cachedContentTokenCount to input, output and cache-read tokens.
  • Thinking levels map low | medium | high to Gemini's levels. Unset leaves the model default.
  • JSON mode (generateJson) sets responseMimeType: 'application/json', adds responseJsonSchema only when a schema is passed, strips a stray code fence, and throws with the finish reason if the reply is not JSON.
note
Why the Claim schema is not sent as responseJsonSchema

Gemini's constrained decoding has a complexity cap that the Claim schema exceeds: it fails at roughly 20 of its 33 top-level sections. So the generator puts the full JSON Schema in the prompt, and zod validates the reply. Constrained decoding is for small shapes only. Keep it that way.

The scripted adapter​

api/src/ai/scripted.ts is selected with AI_PROVIDER=scripted. It plays the model's part deterministically so the API integration tests and the Playwright suite exercise the real Worker without Gemini:

  1. On a new question it calls get_claim_section for the section the question is about (parties → involvements, money words → financials, anything else → incident), plus highlight_panel for the matching panel when that tool is offered.
  2. Given the tool results back, it streams "Scripted answer about {claim id} — {title}. I looked up: {tools}." in three chunks.
  3. A question containing [scripted:fail] makes it throw.
  4. generateJson always throws: the scripted adapter does not generate claims.

It reads the claim's id and title from the system prompt's brief, so a test can tell the right claim reached the model. Never set it on a deployment. See Testing.

AI Gateway​

api/src/ai/gateway.ts builds every gateway URL. With both AI_GATEWAY_ACCOUNT_ID and AI_GATEWAY_ID set:

PathUsed forBase URL
google-ai-studioChat and generation, as the SDK's baseUrlhttps://gateway.ai.cloudflare.com/v1/<account>/<id>/google-ai-studio
googleThe voice relay's upstream sockethttps://gateway.ai.cloudflare.com/v1/<account>/<id>/google

gatewayHeaders() adds cf-aig-authorization: Bearer <CF_AIG_TOKEN> when a token is set. The request body, tools and streaming are untouched: the gateway is a base-URL swap. See Configuration for what to set where.

Thinking levels​

Three places can set reasoning depth, and the most specific wins:

  1. A capability's thinkingLevel (for example high for policy validation, medium for next actions).
  2. GEMINI_THINKING_LEVEL for the deployment.
  3. The model's own default, when neither is set.

Adding a provider​

See Add a model provider. The voice relay speaks the Gemini Live protocol directly and would need its own equivalent.