Unit and integration tests
vitest.config.ts at the repository root defines one project per workspace. Tests sit beside the code they test as *.test.ts(x).
| Project | Files | Environment | Covers |
|---|---|---|---|
shared | packages/shared/src/**/*.test.ts | Node | The Claim schema's required core and defaults, the chat and feedback protocols, the capability catalogue, the fixtures |
api | api/src/**/*.test.ts, api/test/**/*.test.ts | Node | Every capability tool, the registry, prompt assembly order, the turn loop, prompts and fingerprints, both stores, provider and gateway selection, the whole Worker over HTTP |
frontend | frontend/src/**/*.test.{ts,tsx} | jsdom + Testing Library | The API client and SSE parser, deriveClaimView, region and brand, the dock and highlighting, the Prompt Lab store, the capability menu, the chat widget driven like a user, the embed element |
npx vitest run --project api
npx vitest run api/src/chat
npx vitest --project frontend # watch one project
API tests
Capability tools
api/src/capabilities/tools.test.ts and capabilities.test.ts run each tool against a claim built for the case and assert on its JSON: the policy period boundaries, deadline ordering, duplicate payments and bills, vulnerability drivers, complaint clocks, handling metrics, the timeline order. today is always passed explicitly.
Prompt assembly and the turn loop
chat/system.test.ts checks the order of the system prompt blocks, that safety precedes any claim data even when the base prompt is overridden, the screen-actions and locale blocks, task playbooks, voice, and that no placeholder survives branding.
chat/service.test.ts drives runChatTurn with an in-test scripted model (scriptedModel in api/test/builders.ts) that plays back one array of events per round and records every request. It covers the error events, claim resolution, the event sequence of a tool round, providerState replay, usage summing, turn ids, fingerprints, unknown tools, the eight-round limit and history trimming.
Real SQL without Cloudflare: api/test/d1.ts
createTestD1() returns a D1 binding backed by an in-memory SQLite database (Node's built-in node:sqlite) with every migration in db/migrations/d1 applied. D1 is SQLite, so the adapters' SQL (json_extract, json_each, on conflict … do update, datetime('now')) runs exactly as written.
As in D1, binding undefined throws. A missing ?? null in an adapter fails a test rather than passing silently. Each test gets a fresh database.
Node prints ExperimentalWarning: SQLite is an experimental feature once per run. That is expected.
The Neon adapters cannot run without Postgres, so their tests mock @neondatabase/serverless and check what matters: the SQL text and the $n placeholder numbering against the parameter list.
The whole Worker in process: api/test/app.test.ts
app.request(path, init, env) drives the real Hono app (middleware, routing, CORS, EMBED_KEY, error handling, the SPA fallback) with env set to the test D1, AI_PROVIDER: 'scripted' and a stub ASSETS. Chat responses are read as a whole SSE body and parsed with parseSse. It covers health, every claims route, prompts, chat streaming and its error paths, feedback record, replace, withdraw, filter and validation, the key, CORS, and the 404 and SPA fallbacks.
Helpers
| Helper | In | Does |
|---|---|---|
claimWith(overrides) | api/test/builders.ts | A deep copy of the auto fixture with overrides, re-parsed so defaults fill in |
minimalClaim() | builders.ts | The eight-field core and nothing else |
toolContext(claim, today) | builders.ts | A tool context with a pinned date |
runJson(tool, ctx, input) | builders.ts | Run a tool and parse its JSON |
scriptedModel(rounds) | builders.ts | An in-test AiProvider that records requests |
collect(events) · parseSse(body) | builders.ts | Drain an event stream · parse an SSE body |
payment, bill, diary, task, document | api/test/records.ts | One record with sensible defaults, so a test states only what it is about |
createTestD1() | api/test/d1.ts | The in-memory D1 binding |
Front-end tests
frontend/test/setup.ts loads the jest-dom matchers, stubs the scrolling jsdom lacks, and resets localStorage, the region attribute, timers and mocks after each test.
chat/ChatWidget.test.tsxrenders the widget inside its real providers with the network replaced by a scripted event stream (vi.mock('../lib/api')), then types, sends, and clicks like a user: the streamed answer, activity chips and lookups, the request contents, multi-turn history, error lines, thumbs up and down, the reason card and its validation, skip, switching, embed votes without content, a failed send, and that feedback never enters the history the model sees.embed/element.test.tsxcovers the custom element: a claim assigned before definition survives, the shadow root, the API base and key from attributes, following a new claim, and keeping its dock state apart from the demo's.claimpilot/claimView.test.tssets fake timers before importing the module, becauseTODAYis read at import.
Coverage
npm run test:coverage
Writes a text summary, coverage/index.html and coverage/lcov.info, and enforces the per-layer gates.