Skip to main content

Voice

Voice mode is a spoken conversation about the same claim, with the same prompts and the same tools as text chat. The browser talks to the Worker over a WebSocket; the Worker talks to the Gemini Live API over another, and relays between them. Transcripts of both sides stream into the same chat transcript.

The relay is api/src/voice/relay.ts; the browser side is frontend/src/chat/useVoice.ts.

Opening the session​

The upgrade is returned before anything is awaited. The browser's socket only opens once the 101 arrives, so waiting for its first frame beforehand would deadlock. Everything after the upgrade runs in the background.

The browser's first frame is a clientConfig:

{ "clientConfig": { "promptOverrides": { "system.base": "…" }, "claim": { "id": "CP-AUT-260381", "…": "…" }, "screenActions": true } }
  • promptOverrides are the browser's active Prompt Lab versions, so text and voice use the same prompts.
  • claim is the document on screen, the same way text chat sends it.
  • screenActions decides whether highlight_panel is offered.
  • brand (optional) sets the names the assistant goes by, as in text chat.

A claim that fails validation is dropped rather than throwing, so a malformed one falls back to any stored claim.

If no clientConfig arrives within 3 seconds the relay carries on with the code-default prompts. A claimId in the query string is the older path: the Worker looks it up before upgrading and returns 404 early if it does not exist.

The Worker then assembles the system prompt exactly as text chat does, plus the system.voice tail, and opens the upstream socket with our key (the key never reaches the browser). The setup frame carries the system prompt, the same tool declarations, input and output transcription, the configured voice (GEMINI_VOICE, default Charon) and the locale's language code.

Accent comes from the prompt

The language code is sent, but native-audio Live models ignore it and choose their own language. The accent and spelling are held by the locale block and the voice tail in the prompt, which is where to change them.

Once Gemini replies setupComplete, the Worker sends the opening turn (system.voice_greeting) so the call does not start in silence, and flushes any frames the browser sent while the upstream was connecting.

During the session​

DirectionWhatHandling
Browser → GeminiMicrophone audio, 16 kHz PCM16, only after setupCompleteRelayed verbatim
Gemini → browserAudio, 24 kHz PCM16, and transcripts of both sidesRelayed verbatim
Gemini → WorkertoolCall framesIntercepted. Each call runs in the Worker through runTool; the response goes back upstream
Worker → browser{ "chat": ChatEvent } framesActivity chips and screen actions, the same events text chat uses
Worker → browser{ "error": { "message" } }Then both sockets close

Tools are server-side in voice exactly as in text. The browser never sees a tool call it could tamper with, and never supplies a system prompt: the relay injects its own setup.

In the browser, useVoice schedules returned audio through an analyser that drives the level meters and the assistant's glow. Barge-in flushes queued audio. If the microphone is silent for a few seconds the panel says so, because a muted microphone or the wrong input device is by far the most common fault.

Each voice answer gets a turn id minted in the browser, so voice answers can be rated like text ones.

In development and in production​

The Vite proxy forwards the socket (ws: true in frontend/vite.config.ts). Without that, the browser sits on Connecting voice….

On a deployed Worker, voice depends on the Worker opening an outbound WebSocket to Gemini Live, through the AI Gateway's realtime provider when one is configured (with cf-aig-authorization sent as a header on the upgrade). It works, and it is the one thing that behaves differently from local development: spend thirty seconds checking voice after the first deploy of an environment.

Limits​

  • A voice session keeps its own memory in Gemini Live. Its transcript is shown in the chat but is not fed into the text-chat history, and the reverse.
  • Voice is not stateless in the way text is: it holds two open connections for the length of a call, and the provider resets connections after about ten minutes.
  • On speakers rather than headphones, echo cancellation can stop a user interrupting mid-sentence. That is the browser's audio stack.