Skip to main content

Runbook

Each entry is a symptom, how to confirm the cause, and the fix. For problems on a host's page, see Troubleshooting an embed.

AI Gateway 2009 Unauthorized​

Symptom. Every chat answer is an error mentioning AiGatewayError 2009 Unauthorized. Voice never connects.

Cause. The gateway ids are set (so calls go through the gateway) but CF_AIG_TOKEN is not set for this environment. A missing token does not fall back to Google.

Fix.

  • Deployed: npx wrangler secret put CF_AIG_TOKEN (add --env neon for the Neon Worker). Secrets are per environment.
  • Locally: add CF_AIG_TOKEN to api/.dev.vars, or blank AI_GATEWAY_ACCOUNT_ID and AI_GATEWAY_ID there to go straight to Google. Restart wrangler dev: it reads .dev.vars only at start-up.

Confirm. /api/health shows the gateway id; the gateway's own log shows the rejected calls.

"GEMINI_API_KEY is not configured on the API."​

Cause. No key for this environment. /api/health shows configured: false.

Fix. wrangler secret put GEMINI_API_KEY, or add it to api/.dev.vars and restart.

"Claim … not found."​

Cause. The request sent a claimId the store does not have: a stale bookmark, a different store (D1 versus Neon), or DATABASE_PROVIDER=none.

Fix. Check /api/health → database. Load claims into that store, or open a claim from the list.

The claims list is empty​

Cause. Nothing loaded, the schema missing, or the wrong store.

Fix. npm run db:migrate (local) or db:migrate:remote, then generate or copy claims. Check /api/health. Locally, deleting api/.wrangler/state resets D1 to nothing.

Voice sits on "Connecting voice…"​

WhereCauseFix
Local, through ViteThe proxy is not forwarding WebSocketsws: true on the /api proxy in frontend/vite.config.ts (it is set; check it has not been removed)
AnywhereNo key, or the gateway is rejectingSee the two entries above
Deployed onlyThe outbound WebSocket to Gemini Live failedCheck the Worker log and the gateway's realtime log; confirm the Live model name in wrangler.jsonc
A host pageOrigin not allowedThe voice handshake is checked against ALLOWED_ORIGINS too

Voice connects but hears nothing​

Almost always a muted microphone or the wrong input device; the panel says so after a few seconds of silence. Voice also needs HTTPS or localhost.

The Worker will not start: "Capability catalogue mismatch"​

Cause. A capability id exists in packages/shared/src/capabilities.ts and not in the server registry, or the reverse.

Fix. Add the missing half. See Add a capability. The test suite catches this before deploy.

A 500 naming the database​

DATABASE_PROVIDER is "d1" but the DB binding is missing or … "neon" but DATABASE_URL is not set.

Fix. For D1, check d1_databases in wrangler.jsonc for this environment (named environments inherit nothing). For Neon, set the DATABASE_URL secret for this environment.

A host reports "CORS error"​

Cause. Their origin is not in ALLOWED_ORIGINS, or it is listed with a path or trailing slash.

Confirm.

curl -i -X OPTIONS https://<worker>/api/chat \
-H "Origin: https://<their-origin>" \
-H "Access-Control-Request-Method: POST"

A 403 body names the origin. Fix: add the exact origin (scheme, host, port) to ALLOWED_ORIGINS in wrangler.jsonc and deploy. Ask them to reload, because browsers cache preflight results.

A host reports 401​

The api-key attribute does not match EMBED_KEY for that deployment.

Generated claims fail or are thin​

  • "Model returned non-JSON output": the generator retries up to four times with the rejection fed back. Persistent failure usually means the output was cut off; try a model with a larger output limit via GENERATOR_MODEL.
  • Depth rejections: the minimums (8 diary entries, 6 documents, 5 tasks, 5 involvements) are enforced. Thin output is rejected and regenerated.
  • Weak records: read a sample per scenario; delete weak ones with SQL against the local store.

Playwright cannot start its server​

:8799 is taken, usually by a leftover test Worker. On Windows the wrangler child can outlive its parent.

Get-NetTCPConnection -LocalPort 8799 -State Listen | Select-Object OwningProcess

Stop that process (check it is the e2e Worker first: its /api/health reports scripted).