Skip to main content

Design decisions

A short record of the decisions that are load-bearing: change one and much else changes with it. Each says what we decided, why, and what it costs.

1. Integrate at the claim, not at the tools​

Decision. Hosts give us one thing, the claim on screen. Every tool runs on our side over that one document.

Why. All but one of the tools are calculations over one document, not calls to a system. Integrating at the tools would mean sixteen endpoints per host and every host reimplementing our checks slightly differently.

Cost. One mapping per host, and answers that track the quality of that mapping.

2. The claim travels with the request​

Decision. The browser sends the claim it is displaying with every turn. claimId remains for the Prompt Lab and older callers.

Why. The assistant can then only ever see the claim on screen. Nothing is stored, no access decision is made twice, and the same Worker can serve hosts with no claims of ours at all.

Cost. About 10 KB per request. Nothing, at this scale.

3. Eight required fields, honest empty defaults​

Decision. The Claim contract requires eight fields. Every other section is defaulted to an empty value by parsing.

Why. A host can map the minimum in twenty minutes and add sections as it goes. The parsed type stays complete, so no consumer had to change. Empty (not plausible) defaults mean absence reads as absence.

Cost. Nothing warns when a host's mapping is thin.

4. Read through tools, do arithmetic in code​

Decision. The system prompt carries a compact brief. Detail comes through read-only tools; counting, dates and money are computed by deterministic tools.

Why. Small, cacheable context; traceable answers ("lookups"); no arithmetic mistakes from the model.

Cost. More round-trips per answer, bounded at eight.

5. Safety before data, structurally​

Decision. system.safety is always assembled before the claim brief, in a fixed order the client cannot change. Tool results and record contents are data.

Why. Nothing read from a record can precede the rules about how to treat what is read from a record.

Cost. None we have found.

6. Prompt edits live in the browser​

Decision. The Prompt Lab stores versions in localStorage and sends the active ones with each request. The server holds only code defaults.

Why. Without logins, a server-side save would be everyone's. This gives every person their own sandbox and lets text and voice share prompts.

Cost. Edits do not travel between browsers; making one the default is a code change. Hosts get no Prompt Lab.

7. The schema goes in the prompt, not in constrained decoding​

Decision. The generator puts the Claim's JSON Schema in its prompt and validates the reply with zod, retrying with the rejection fed back.

Why. Gemini's constrained decoding fails on a schema this complex.

Cost. Occasional retries.

8. Replay the model's own parts​

Decision. Assistant turns are replayed from the adapter's providerState when it exists, never rebuilt from text and calls.

Why. Gemini 3 rejects a replayed function call without its original thought signature.

Cost. An opaque field on assistant messages that only the adapter understands.

9. Ports for the model and the stores​

Decision. AiProvider, ClaimRepository and FeedbackRepository are the only ways in. One factory per port chooses the adapter.

Why. D1 locally and in one deployment, Neon in another, "none" for embed-only; Gemini in production, a scripted model in tests; all without touching routes or the turn loop.

Cost. A little indirection.

10. URL-based AI Gateway, our own key​

Decision. The gateway is a base-URL swap in the SDK, with our key sent on each call, rather than the env.AI binding or gateway-side keys.

Why. The generator runs in plain Node, where bindings do not exist; one integration style for one provider. Gateway-side key injection was unreliable with the Gemini 3.x previews.

Cost. Gateway logs hold claim content. See Security.

11. Generation is CLI-only​

Decision. No in-app endpoints create claims.

Why. Generation is slow, costs money and writes to shared stores. It belongs behind a deliberate command, not a button.

Cost. Demo data refreshes need a developer.

12. A scripted model, not mocks, for tests​

Decision. The test suites run the real Worker against a deterministic AiProvider selected by AI_PROVIDER=scripted.

Why. Everything except the language model runs for real (prompt assembly, tools, SSE, feedback) with no key, no network and the same answer every time.

Cost. The model itself is untested by the suites. Gemini and voice are checked by hand.

13. Feedback stays out of the conversation​

Decision. Thumbs and reasons go to /api/feedback, never into the chat history.

Why. A complaint in the conversation would colour later answers and could be read as an instruction.

Cost. The model never learns from feedback within a conversation, which is the point.

14. A shadow root for the embed​

Decision. The custom element renders into a shadow root, hoisting only Tailwind's @property registrations to the host document.

Why. Isolation both ways. Cascade layers alone let an unlayered host rule beat our utilities.

Cost. @property must be hoisted; fixed positioning depends on the tag's ancestors.