Voice
Voice mode is a spoken conversation about the same claim, with the same prompts and the same tools as text chat. The browser talks to the Worker over a WebSocket; the Worker talks to the Gemini Live API over another, and relays between them. Transcripts of both sides stream into the same chat transcript.
The relay is api/src/voice/relay.ts; the browser side is frontend/src/chat/useVoice.ts.
Opening the session
The upgrade is returned before anything is awaited. The browser's socket only opens once the 101 arrives, so waiting for its first frame beforehand would deadlock. Everything after the upgrade runs in the background.
The browser's first frame is a clientConfig:
{ "clientConfig": { "promptOverrides": { "system.base": "…" }, "claim": { "id": "CP-AUT-260381", "…": "…" }, "screenActions": true } }
promptOverridesare the browser's active Prompt Lab versions, so text and voice use the same prompts.claimis the document on screen, the same way text chat sends it.screenActionsdecides whetherhighlight_panelis offered.brand(optional) sets the names the assistant goes by, as in text chat.
A claim that fails validation is dropped rather than throwing, so a malformed one falls back to any stored claim.
If no clientConfig arrives within 3 seconds the relay carries on with the code-default prompts. A claimId in the query string is the older path: the Worker looks it up before upgrading and returns 404 early if it does not exist.
The Worker then assembles the system prompt exactly as text chat does, plus the system.voice tail, and opens the upstream socket with our key (the key never reaches the browser). The setup frame carries the system prompt, the same tool declarations, input and output transcription, the configured voice (GEMINI_VOICE, default Charon) and the locale's language code.
The language code is sent, but native-audio Live models ignore it and choose their own language. The accent and spelling are held by the locale block and the voice tail in the prompt, which is where to change them.
Once Gemini replies setupComplete, the Worker sends the opening turn (system.voice_greeting) so the call does not start in silence, and flushes any frames the browser sent while the upstream was connecting.
During the session
| Direction | What | Handling |
|---|---|---|
| Browser → Gemini | Microphone audio, 16 kHz PCM16, only after setupComplete | Relayed verbatim |
| Gemini → browser | Audio, 24 kHz PCM16, and transcripts of both sides | Relayed verbatim |
| Gemini → Worker | toolCall frames | Intercepted. Each call runs in the Worker through runTool; the response goes back upstream |
| Worker → browser | { "chat": ChatEvent } frames | Activity chips and screen actions, the same events text chat uses |
| Worker → browser | { "error": { "message" } } | Then both sockets close |
Tools are server-side in voice exactly as in text. The browser never sees a tool call it could tamper with, and never supplies a system prompt: the relay injects its own setup.
In the browser, useVoice schedules returned audio through an analyser that drives the level meters and the assistant's glow. Barge-in flushes queued audio. If the microphone is silent for a few seconds the panel says so, because a muted microphone or the wrong input device is by far the most common fault.
Each voice answer gets a turn id minted in the browser, so voice answers can be rated like text ones.
In development and in production
The Vite proxy forwards the socket (ws: true in frontend/vite.config.ts). Without that, the browser sits on Connecting voice….
On a deployed Worker, voice depends on the Worker opening an outbound WebSocket to Gemini Live, through the AI Gateway's realtime provider when one is configured (with cf-aig-authorization sent as a header on the upgrade). It works, and it is the one thing that behaves differently from local development: spend thirty seconds checking voice after the first deploy of an environment.
Limits
- A voice session keeps its own memory in Gemini Live. Its transcript is shown in the chat but is not fed into the text-chat history, and the reverse.
- Voice is not stateless in the way text is: it holds two open connections for the length of a call, and the provider resets connections after about ten minutes.
- On speakers rather than headphones, echo cancellation can stop a user interrupting mid-sentence. That is the browser's audio stack.