Observability
/api/health
The first thing to check, in every environment:
{
"ok": true,
"database": "d1",
"ai": {
"configured": true,
"model": "gemini-3.5-flash-lite",
"liveModel": "gemini-3.8-live",
"thinkingLevel": "default",
"gateway": "claimpilot-chat"
}
}
| Field | Tells us |
|---|---|
database | Which store the Worker is using: d1, neon or none |
ai.configured | Whether a model key is present. false means chat will answer with an error event |
ai.model, ai.liveModel | The models actually in use, after wrangler.jsonc and secrets |
ai.thinkingLevel | The deployment default, or default for the model's own |
ai.gateway | The AI Gateway id, or null when calls go straight to Google |
The Claims Data page (/data) and the unlisted /status page show the same, with what is loaded.
/api/health is not gated by EMBED_KEY, but a cross-origin caller not in ALLOWED_ORIGINS gets a 403.
AI Gateway
With the gateway on, every model call (chat, generation and voice) is logged in Cloudflare AI Gateway: the request, the response, tokens, cost, latency and errors. That is the place to answer "what did the model actually see" and "what did that cost".
Prompts carry the claim brief and tool results carry claim detail, so the gateway log is a second copy of claim content. Fine for synthetic claims; a data-protection question for real ones. See Security.
The Worker log
observability.enabled is on in wrangler.jsonc, so Workers Logs collects the Worker's console output. What it writes:
| When | What |
|---|---|
| A refused origin | Refused origin <origin> on <path> (warning) |
| An unhandled route error | The error, then a JSON 500 |
| A chat turn fails mid-stream | chat turn failed with the error; the browser gets an error event |
DATABASE_PROVIDER=none and a vote | One line: {"feedback": {…}}, or {"feedback": {"turnId", "rating": null}} for a withdrawal |
Locally, the same lines appear in the wrangler dev terminal.
Usage per turn
Every done event carries usage (input, output and cache-read tokens, summed over the turn's rounds) and the model. The chat footer shows the last turn's tokens; the Prompt Lab trace shows them per run. Prompt caching shows up as cacheReadTokens: the stable prefix (system prompts, the claim brief and the tool list, shared tools first) is what makes it work.
Answer quality
The feedback report in the Prompt Lab is the quality signal: the helpful rate by capability and by prompt set, and the reasons behind each thumbs down. GET /api/feedback returns the same as JSON.
Front-end versions
Each build writes version.json. Open tabs poll it every minute and offer a refresh when it changes, so a long-lived tab does not keep running an old build against a new API.