Skip to main content

The chat pipeline

A chat turn is one POST /api/chat that streams Server-Sent Events back. The Worker runs the whole agentic loop server-side: it calls the model, runs any tools it asks for, calls it again with the results, and streams text and activity to the browser as they happen.

The code is api/src/chat/service.ts (runChatTurn, an async generator of ChatEvents) and api/src/chat/routes.ts (which writes each event as an SSE frame).

1. Validate and resolve​

ChatRequest (zod, packages/shared/src/chat.ts) is parsed first. A malformed body is a 400 before any streaming starts. The request must carry either:

  • claim: the document itself. The demo and every embedding host send this, because the page already has the claim it is rendering.
  • claimId: a reference the Worker looks up in its own store. The Prompt Lab's runner uses this.

From here on, everything sees a Claim and nothing knows which way it arrived.

These become error events rather than HTTP errors, because the stream has already started:

ConditionEvent message
No model keyGEMINI_API_KEY is not configured on the API.
Unknown claimIdClaim <id> not found.
Unknown capabilityUnknown capability "<id>"
The model or a provider call throwsthe error's message

2. Assemble the system prompt​

assemblePrompt() in api/src/chat/system.ts builds the system instruction in a fixed order. The order is part of the safety design: nothing read from the claim can come before the boundaries.

#BlockSourceWhen
1Identity and working stylesystem.basealways
2Pointing at the screensystem.screen_actionsonly if screenActions
3Safety and boundariessystem.safetyalways
4Spelling, dates and terminologysystem.locale_uk or system.locale_usby locale
5The claim briefbuildClaimBrief(claim, today)always
6a## Task: <label> + the capability's playbookcapability.<id>a quick action
6bFree-chat guidance + a list of the specialist taskscapability.claim_qa + buildCapabilityGuide()free chat
7Spoken-style tailsystem.voicevoice only

Every prompt block goes through resolvePrompt(id, overrides): the request's Prompt Lab override if it is usable, otherwise the code default. Then applyBrand() fills in {{assistant}}, {{product}} and {{pronunciation}} for the requested brand. The claim brief is never branded or overridden.

The claim brief​

buildClaimBrief (api/src/chat/context.ts) is a dozen lines: the claim and today's date, line of business, status, client and handler, policy period and state, claimant, incident, loss estimate, coverage and liability position, money totals, risk flags, and the size of the schedule and file. It ends with the list of sections get_claim_section can read.

It is small and stable on purpose. It only changes when the claim changes, so it caches well, and it is enough for the model to orient itself and answer simple questions without a tool call. Detail comes from tools.

3. Offer the tools​

allTools({ screenActions }) returns every tool, shared readers first so the prompt-cache prefix stays stable. All tools are offered on every turn, whatever the capability. A capability's playbook tells the model which to call first.

highlight_panel is a screen action. It is offered only when the request sets screenActions: true. A caller that cannot highlight anything is never told it can, so the model cannot claim to have done it.

4. Run the loop​

For up to 8 iterations:

  1. Call ai.streamTurn({ system, messages, tools, thinkingLevel, maxOutputTokens: 8000, signal }). The thinking level is the capability's own if it sets one, otherwise the deployment default.
  2. Forward each text chunk as a text_delta immediately.
  3. On done, add the usage and keep the adapter's providerState.
  4. Append the assistant turn to the in-request history, with its providerState.
  5. If there were no tool calls, stop.
  6. Otherwise, for each call:
    • an unknown tool name gets an error result (Unknown tool <name>) and no activity chip;
    • a known one emits activity with the tool's present-tense label, runs through runTool (which turns a throw into an error result rather than a crash), emits tool_result with the result truncated to 400 characters, and emits a client_action for each screen action it returned.
  7. Append the tool results to the history and go round again.

If the eighth round still ends in tool calls, the Worker appends "I’ve reached my limit of lookups for this question — ask me to continue if you need more." and finishes.

Why providerState is replayed​

Gemini 3 signs its turns: function-call parts carry a thoughtSignature that must be echoed back verbatim when the turn is replayed, or the next call is rejected. The Gemini adapter rebuilds the model's own parts as it streams and returns them as opaque providerState. The turn loop stores it on the assistant message and hands it back unchanged. Never rebuild an assistant turn from its text and calls when replay parts exist.

History on the wire​

Between turns the browser keeps only text: user messages and the assistant's final answers. Tool round-trips stay inside one request. toAiMessages drops any leading assistant messages so the model always starts with the user. Requests are capped at 60 messages of 20,000 characters each.

5. Finish​

The last event is always done (unless an error ended the stream):

{
"type": "done",
"model": "gemini-3.1-flash-lite",
"usage": { "inputTokens": 5210, "outputTokens": 412, "cacheReadTokens": 3968 },
"turnId": "4f6b2c1e-8a9d-4c3b-9e2f-1d7a5b0c6e93",
"prompts": { "fingerprint": "default", "overridden": [] }
}
  • turnId is a fresh UUID. Nothing is stored against it on the Worker. It comes back with a thumbs up or down, which is how a vote is tied to one answer (see Answer feedback).
  • prompts.fingerprint is default, or the first 12 hex characters of a SHA-256 over the overrides that actually applied, in id order. The same Prompt Lab versions in any browser give the same fingerprint.

Cancellation​

The route passes an AbortSignal into the loop and aborts it when the client disconnects. The browser's Stop button aborts its fetch, the stream closes, and the model call stops with it.

Reading the stream in the browser​

api.chat() in frontend/src/lib/api.ts is an async generator over the response body. It splits frames on blank lines and parses each data: line as a ChatEvent, however the bytes are chunked. useChat then:

EventEffect
text_deltaAppends to the streaming answer
activityInserts a chip before the answer, so the answer reads as the conclusion
tool_resultAdds to the answer's lookups trace
client_actionhighlight_panel → highlightPanel(panel); navigate → the host's callback, if any
doneRecords usage, the turn id, the model and the prompt fingerprint on the answer
errorAdds an error line to the transcript

The full event schema is in the chat protocol reference.