Known limits and next steps
This is a pilot on synthetic data. These are the limits we know about, stated plainly.
Before real claims
| Limit | Why it matters | What would change |
|---|---|---|
| No authentication | There are no users. GET /api/feedback and GET /api/claims answer anyone who can reach the API. | Put the report and claims routes behind a login; record who voted. |
| Claim content in gateway logs | Prompts and replies, including claim detail, are logged in AI Gateway. | Agree retention and access; consider turning response logging off; keep caching off. |
| No audit trail | We keep no record of what the assistant said on a claim. | The host logs interactions against the claim in its own system. |
| No scored evaluation | The Prompt Lab runs prompts but does not grade them. Nobody measures whether a thin host mapping produces thin answers. | A small scored eval set per capability. |
EMBED_KEY is not access control | It is visible in the host page. | Real deployments rely on the host's own access decisions and on ALLOWED_ORIGINS. |
Product limits
- Voice and text keep separate memories. A voice session's transcript appears in the chat but is not fed into the text history, and the reverse.
- Documents are summaries. The assistant reads a document's written summary, not the file. Real document text needs a host endpoint and extraction (v0.2).
- One claim, no history. No cross-claim questions ("how many claims like this?"), by design; the safety prompt keeps the assistant on the claim on screen. Prior claims by the same claimant need a host interface (v0.2).
- Read-only. No diary notes, tasks, letters or payments. Writes would be host-provided tools, one per action (v0.3).
- The embed has no Prompt Lab. A host cannot tune tone; a change is a release on our side.
- Prompts assume our claim shape. The playbooks were written against generated data. Hosts with terser records may get quietly worse answers.
Data limits
- Generated-data quality. Flash Lite produces good records when held to the depth minimums; a richer
GENERATOR_MODELcosts more and produces richer records. Read a sample per scenario before a demo. - No delete endpoint. Remove weak claims with SQL against the local store, deliberately.
- Product palettes are approximations taken from screenshots. Swap the tokens in
styles.cssandregions.tswhen product brand guidance arrives.
Next steps
- Evals. A scored set per capability, run in CI against a recorded or live model.
- Documents (v0.2). One host endpoint for document text, with extraction and chunking.
- Cross-claim history (v0.2). A host interface returning the claimant's other claims, for the fraud check.
- Authentication for the feedback report and claims routes.
- Writes (v0.3). Host-provided action tools with the host enforcing authority limits and audit.