Skip to main content

Observability

/api/health​

The first thing to check, in every environment:

{
"ok": true,
"database": "d1",
"ai": {
"configured": true,
"model": "gemini-3.5-flash-lite",
"liveModel": "gemini-3.8-live",
"thinkingLevel": "default",
"gateway": "claimpilot-chat"
}
}
FieldTells us
databaseWhich store the Worker is using: d1, neon or none
ai.configuredWhether a model key is present. false means chat will answer with an error event
ai.model, ai.liveModelThe models actually in use, after wrangler.jsonc and secrets
ai.thinkingLevelThe deployment default, or default for the model's own
ai.gatewayThe AI Gateway id, or null when calls go straight to Google

The Claims Data page (/data) and the unlisted /status page show the same, with what is loaded.

/api/health is not gated by EMBED_KEY, but a cross-origin caller not in ALLOWED_ORIGINS gets a 403.

AI Gateway​

With the gateway on, every model call (chat, generation and voice) is logged in Cloudflare AI Gateway: the request, the response, tokens, cost, latency and errors. That is the place to answer "what did the model actually see" and "what did that cost".

The gateway log holds claim content

Prompts carry the claim brief and tool results carry claim detail, so the gateway log is a second copy of claim content. Fine for synthetic claims; a data-protection question for real ones. See Security.

The Worker log​

observability.enabled is on in wrangler.jsonc, so Workers Logs collects the Worker's console output. What it writes:

WhenWhat
A refused originRefused origin <origin> on <path> (warning)
An unhandled route errorThe error, then a JSON 500
A chat turn fails mid-streamchat turn failed with the error; the browser gets an error event
DATABASE_PROVIDER=none and a voteOne line: {"feedback": {…}}, or {"feedback": {"turnId", "rating": null}} for a withdrawal

Locally, the same lines appear in the wrangler dev terminal.

Usage per turn​

Every done event carries usage (input, output and cache-read tokens, summed over the turn's rounds) and the model. The chat footer shows the last turn's tokens; the Prompt Lab trace shows them per run. Prompt caching shows up as cacheReadTokens: the stable prefix (system prompts, the claim brief and the tool list, shared tools first) is what makes it work.

Answer quality​

The feedback report in the Prompt Lab is the quality signal: the helpful rate by capability and by prompt set, and the reasons behind each thumbs down. GET /api/feedback returns the same as JSON.

Front-end versions​

Each build writes version.json. Open tabs poll it every minute and offer a refresh when it changes, so a long-lived tab does not keep running an old build against a new API.