Design decisions
A short record of the decisions that are load-bearing: change one and much else changes with it. Each says what we decided, why, and what it costs.
1. Integrate at the claim, not at the tools
Decision. Hosts give us one thing, the claim on screen. Every tool runs on our side over that one document.
Why. All but one of the tools are calculations over one document, not calls to a system. Integrating at the tools would mean sixteen endpoints per host and every host reimplementing our checks slightly differently.
Cost. One mapping per host, and answers that track the quality of that mapping.
2. The claim travels with the request
Decision. The browser sends the claim it is displaying with every turn. claimId remains for the Prompt Lab and older callers.
Why. The assistant can then only ever see the claim on screen. Nothing is stored, no access decision is made twice, and the same Worker can serve hosts with no claims of ours at all.
Cost. About 10 KB per request. Nothing, at this scale.
3. Eight required fields, honest empty defaults
Decision. The Claim contract requires eight fields. Every other section is defaulted to an empty value by parsing.
Why. A host can map the minimum in twenty minutes and add sections as it goes. The parsed type stays complete, so no consumer had to change. Empty (not plausible) defaults mean absence reads as absence.
Cost. Nothing warns when a host's mapping is thin.
4. Read through tools, do arithmetic in code
Decision. The system prompt carries a compact brief. Detail comes through read-only tools; counting, dates and money are computed by deterministic tools.
Why. Small, cacheable context; traceable answers ("lookups"); no arithmetic mistakes from the model.
Cost. More round-trips per answer, bounded at eight.
5. Safety before data, structurally
Decision. system.safety is always assembled before the claim brief, in a fixed order the client cannot change. Tool results and record contents are data.
Why. Nothing read from a record can precede the rules about how to treat what is read from a record.
Cost. None we have found.
6. Prompt edits live in the browser
Decision. The Prompt Lab stores versions in localStorage and sends the active ones with each request. The server holds only code defaults.
Why. Without logins, a server-side save would be everyone's. This gives every person their own sandbox and lets text and voice share prompts.
Cost. Edits do not travel between browsers; making one the default is a code change. Hosts get no Prompt Lab.
7. The schema goes in the prompt, not in constrained decoding
Decision. The generator puts the Claim's JSON Schema in its prompt and validates the reply with zod, retrying with the rejection fed back.
Why. Gemini's constrained decoding fails on a schema this complex.
Cost. Occasional retries.
8. Replay the model's own parts
Decision. Assistant turns are replayed from the adapter's providerState when it exists, never rebuilt from text and calls.
Why. Gemini 3 rejects a replayed function call without its original thought signature.
Cost. An opaque field on assistant messages that only the adapter understands.
9. Ports for the model and the stores
Decision. AiProvider, ClaimRepository and FeedbackRepository are the only ways in. One factory per port chooses the adapter.
Why. D1 locally and in one deployment, Neon in another, "none" for embed-only; Gemini in production, a scripted model in tests; all without touching routes or the turn loop.
Cost. A little indirection.
10. URL-based AI Gateway, our own key
Decision. The gateway is a base-URL swap in the SDK, with our key sent on each call, rather than the env.AI binding or gateway-side keys.
Why. The generator runs in plain Node, where bindings do not exist; one integration style for one provider. Gateway-side key injection was unreliable with the Gemini 3.x previews.
Cost. Gateway logs hold claim content. See Security.
11. Generation is CLI-only
Decision. No in-app endpoints create claims.
Why. Generation is slow, costs money and writes to shared stores. It belongs behind a deliberate command, not a button.
Cost. Demo data refreshes need a developer.
12. A scripted model, not mocks, for tests
Decision. The test suites run the real Worker against a deterministic AiProvider selected by AI_PROVIDER=scripted.
Why. Everything except the language model runs for real (prompt assembly, tools, SSE, feedback) with no key, no network and the same answer every time.
Cost. The model itself is untested by the suites. Gemini and voice are checked by hand.
13. Feedback stays out of the conversation
Decision. Thumbs and reasons go to /api/feedback, never into the chat history.
Why. A complaint in the conversation would colour later answers and could be read as an instruction.
Cost. The model never learns from feedback within a conversation, which is the point.
14. A shadow root for the embed
Decision. The custom element renders into a shadow root, hoisting only Tailwind's @property registrations to the host document.
Why. Isolation both ways. Cascade layers alone let an unlayered host rule beat our utilities.
Cost. @property must be hoisted; fixed positioning depends on the tag's ancestors.