Skip to main content

Known limits and next steps

This is a pilot on synthetic data. These are the limits we know about, stated plainly.

Before real claims​

LimitWhy it mattersWhat would change
No authenticationThere are no users. GET /api/feedback and GET /api/claims answer anyone who can reach the API.Put the report and claims routes behind a login; record who voted.
Claim content in gateway logsPrompts and replies, including claim detail, are logged in AI Gateway.Agree retention and access; consider turning response logging off; keep caching off.
No audit trailWe keep no record of what the assistant said on a claim.The host logs interactions against the claim in its own system.
No scored evaluationThe Prompt Lab runs prompts but does not grade them. Nobody measures whether a thin host mapping produces thin answers.A small scored eval set per capability.
EMBED_KEY is not access controlIt is visible in the host page.Real deployments rely on the host's own access decisions and on ALLOWED_ORIGINS.

Product limits​

  • Voice and text keep separate memories. A voice session's transcript appears in the chat but is not fed into the text history, and the reverse.
  • Documents are summaries. The assistant reads a document's written summary, not the file. Real document text needs a host endpoint and extraction (v0.2).
  • One claim, no history. No cross-claim questions ("how many claims like this?"), by design; the safety prompt keeps the assistant on the claim on screen. Prior claims by the same claimant need a host interface (v0.2).
  • Read-only. No diary notes, tasks, letters or payments. Writes would be host-provided tools, one per action (v0.3).
  • The embed has no Prompt Lab. A host cannot tune tone; a change is a release on our side.
  • Prompts assume our claim shape. The playbooks were written against generated data. Hosts with terser records may get quietly worse answers.

Data limits​

  • Generated-data quality. Flash Lite produces good records when held to the depth minimums; a richer GENERATOR_MODEL costs more and produces richer records. Read a sample per scenario before a demo.
  • No delete endpoint. Remove weak claims with SQL against the local store, deliberately.
  • Product palettes are approximations taken from screenshots. Swap the tokens in styles.css and regions.ts when product brand guidance arrives.

Next steps​

  1. Evals. A scored set per capability, run in CI against a recorded or live model.
  2. Documents (v0.2). One host endpoint for document text, with extraction and chunking.
  3. Cross-claim history (v0.2). A host interface returning the claimant's other claims, for the fraud check.
  4. Authentication for the feedback report and claims routes.
  5. Writes (v0.3). Host-provided action tools with the host enforcing authority limits and audit.