The chat pipeline
A chat turn is one POST /api/chat that streams Server-Sent Events back. The Worker runs the whole agentic loop server-side: it calls the model, runs any tools it asks for, calls it again with the results, and streams text and activity to the browser as they happen.
The code is api/src/chat/service.ts (runChatTurn, an async generator of ChatEvents) and api/src/chat/routes.ts (which writes each event as an SSE frame).
1. Validate and resolve
ChatRequest (zod, packages/shared/src/chat.ts) is parsed first. A malformed body is a 400 before any streaming starts. The request must carry either:
claim: the document itself. The demo and every embedding host send this, because the page already has the claim it is rendering.claimId: a reference the Worker looks up in its own store. The Prompt Lab's runner uses this.
From here on, everything sees a Claim and nothing knows which way it arrived.
These become error events rather than HTTP errors, because the stream has already started:
| Condition | Event message |
|---|---|
| No model key | GEMINI_API_KEY is not configured on the API. |
Unknown claimId | Claim <id> not found. |
Unknown capability | Unknown capability "<id>" |
| The model or a provider call throws | the error's message |
2. Assemble the system prompt
assemblePrompt() in api/src/chat/system.ts builds the system instruction in a fixed order. The order is part of the safety design: nothing read from the claim can come before the boundaries.
| # | Block | Source | When |
|---|---|---|---|
| 1 | Identity and working style | system.base | always |
| 2 | Pointing at the screen | system.screen_actions | only if screenActions |
| 3 | Safety and boundaries | system.safety | always |
| 4 | Spelling, dates and terminology | system.locale_uk or system.locale_us | by locale |
| 5 | The claim brief | buildClaimBrief(claim, today) | always |
| 6a | ## Task: <label> + the capability's playbook | capability.<id> | a quick action |
| 6b | Free-chat guidance + a list of the specialist tasks | capability.claim_qa + buildCapabilityGuide() | free chat |
| 7 | Spoken-style tail | system.voice | voice only |
Every prompt block goes through resolvePrompt(id, overrides): the request's Prompt Lab override if it is usable, otherwise the code default. Then applyBrand() fills in {{assistant}}, {{product}} and {{pronunciation}} for the requested brand. The claim brief is never branded or overridden.
The claim brief
buildClaimBrief (api/src/chat/context.ts) is a dozen lines: the claim and today's date, line of business, status, client and handler, policy period and state, claimant, incident, loss estimate, coverage and liability position, money totals, risk flags, and the size of the schedule and file. It ends with the list of sections get_claim_section can read.
It is small and stable on purpose. It only changes when the claim changes, so it caches well, and it is enough for the model to orient itself and answer simple questions without a tool call. Detail comes from tools.
3. Offer the tools
allTools({ screenActions }) returns every tool, shared readers first so the prompt-cache prefix stays stable. All tools are offered on every turn, whatever the capability. A capability's playbook tells the model which to call first.
highlight_panel is a screen action. It is offered only when the request sets screenActions: true. A caller that cannot highlight anything is never told it can, so the model cannot claim to have done it.
4. Run the loop
For up to 8 iterations:
- Call
ai.streamTurn({ system, messages, tools, thinkingLevel, maxOutputTokens: 8000, signal }). The thinking level is the capability's own if it sets one, otherwise the deployment default. - Forward each text chunk as a
text_deltaimmediately. - On
done, add the usage and keep the adapter'sproviderState. - Append the assistant turn to the in-request history, with its
providerState. - If there were no tool calls, stop.
- Otherwise, for each call:
- an unknown tool name gets an error result (
Unknown tool <name>) and no activity chip; - a known one emits
activitywith the tool's present-tense label, runs throughrunTool(which turns a throw into an error result rather than a crash), emitstool_resultwith the result truncated to 400 characters, and emits aclient_actionfor each screen action it returned.
- an unknown tool name gets an error result (
- Append the tool results to the history and go round again.
If the eighth round still ends in tool calls, the Worker appends "I’ve reached my limit of lookups for this question — ask me to continue if you need more." and finishes.
Why providerState is replayed
Gemini 3 signs its turns: function-call parts carry a thoughtSignature that must be echoed back verbatim when the turn is replayed, or the next call is rejected. The Gemini adapter rebuilds the model's own parts as it streams and returns them as opaque providerState. The turn loop stores it on the assistant message and hands it back unchanged. Never rebuild an assistant turn from its text and calls when replay parts exist.
History on the wire
Between turns the browser keeps only text: user messages and the assistant's final answers. Tool round-trips stay inside one request. toAiMessages drops any leading assistant messages so the model always starts with the user. Requests are capped at 60 messages of 20,000 characters each.
5. Finish
The last event is always done (unless an error ended the stream):
{
"type": "done",
"model": "gemini-3.1-flash-lite",
"usage": { "inputTokens": 5210, "outputTokens": 412, "cacheReadTokens": 3968 },
"turnId": "4f6b2c1e-8a9d-4c3b-9e2f-1d7a5b0c6e93",
"prompts": { "fingerprint": "default", "overridden": [] }
}
turnIdis a fresh UUID. Nothing is stored against it on the Worker. It comes back with a thumbs up or down, which is how a vote is tied to one answer (see Answer feedback).prompts.fingerprintisdefault, or the first 12 hex characters of a SHA-256 over the overrides that actually applied, in id order. The same Prompt Lab versions in any browser give the same fingerprint.
Cancellation
The route passes an AbortSignal into the loop and aborts it when the client disconnects. The browser's Stop button aborts its fetch, the stream closes, and the model call stops with it.
Reading the stream in the browser
api.chat() in frontend/src/lib/api.ts is an async generator over the response body. It splits frames on blank lines and parses each data: line as a ChatEvent, however the bytes are chunked. useChat then:
| Event | Effect |
|---|---|
text_delta | Appends to the streaming answer |
activity | Inserts a chip before the answer, so the answer reads as the conclusion |
tool_result | Adds to the answer's lookups trace |
client_action | highlight_panel → highlightPanel(panel); navigate → the host's callback, if any |
done | Records usage, the turn id, the model and the prompt fingerprint on the answer |
error | Adds an error line to the transcript |
The full event schema is in the chat protocol reference.