Skip to main content

Unit and integration tests

vitest.config.ts at the repository root defines one project per workspace. Tests sit beside the code they test as *.test.ts(x).

ProjectFilesEnvironmentCovers
sharedpackages/shared/src/**/*.test.tsNodeThe Claim schema's required core and defaults, the chat and feedback protocols, the capability catalogue, the fixtures
apiapi/src/**/*.test.ts, api/test/**/*.test.tsNodeEvery capability tool, the registry, prompt assembly order, the turn loop, prompts and fingerprints, both stores, provider and gateway selection, the whole Worker over HTTP
frontendfrontend/src/**/*.test.{ts,tsx}jsdom + Testing LibraryThe API client and SSE parser, deriveClaimView, region and brand, the dock and highlighting, the Prompt Lab store, the capability menu, the chat widget driven like a user, the embed element
npx vitest run --project api
npx vitest run api/src/chat
npx vitest --project frontend # watch one project

API tests​

Capability tools​

api/src/capabilities/tools.test.ts and capabilities.test.ts run each tool against a claim built for the case and assert on its JSON: the policy period boundaries, deadline ordering, duplicate payments and bills, vulnerability drivers, complaint clocks, handling metrics, the timeline order. today is always passed explicitly.

Prompt assembly and the turn loop​

chat/system.test.ts checks the order of the system prompt blocks, that safety precedes any claim data even when the base prompt is overridden, the screen-actions and locale blocks, task playbooks, voice, and that no placeholder survives branding.

chat/service.test.ts drives runChatTurn with an in-test scripted model (scriptedModel in api/test/builders.ts) that plays back one array of events per round and records every request. It covers the error events, claim resolution, the event sequence of a tool round, providerState replay, usage summing, turn ids, fingerprints, unknown tools, the eight-round limit and history trimming.

Real SQL without Cloudflare: api/test/d1.ts​

createTestD1() returns a D1 binding backed by an in-memory SQLite database (Node's built-in node:sqlite) with every migration in db/migrations/d1 applied. D1 is SQLite, so the adapters' SQL (json_extract, json_each, on conflict … do update, datetime('now')) runs exactly as written.

As in D1, binding undefined throws. A missing ?? null in an adapter fails a test rather than passing silently. Each test gets a fresh database.

Node prints ExperimentalWarning: SQLite is an experimental feature once per run. That is expected.

The Neon adapters cannot run without Postgres, so their tests mock @neondatabase/serverless and check what matters: the SQL text and the $n placeholder numbering against the parameter list.

The whole Worker in process: api/test/app.test.ts​

app.request(path, init, env) drives the real Hono app (middleware, routing, CORS, EMBED_KEY, error handling, the SPA fallback) with env set to the test D1, AI_PROVIDER: 'scripted' and a stub ASSETS. Chat responses are read as a whole SSE body and parsed with parseSse. It covers health, every claims route, prompts, chat streaming and its error paths, feedback record, replace, withdraw, filter and validation, the key, CORS, and the 404 and SPA fallbacks.

Helpers​

HelperInDoes
claimWith(overrides)api/test/builders.tsA deep copy of the auto fixture with overrides, re-parsed so defaults fill in
minimalClaim()builders.tsThe eight-field core and nothing else
toolContext(claim, today)builders.tsA tool context with a pinned date
runJson(tool, ctx, input)builders.tsRun a tool and parse its JSON
scriptedModel(rounds)builders.tsAn in-test AiProvider that records requests
collect(events) · parseSse(body)builders.tsDrain an event stream · parse an SSE body
payment, bill, diary, task, documentapi/test/records.tsOne record with sensible defaults, so a test states only what it is about
createTestD1()api/test/d1.tsThe in-memory D1 binding

Front-end tests​

frontend/test/setup.ts loads the jest-dom matchers, stubs the scrolling jsdom lacks, and resets localStorage, the region attribute, timers and mocks after each test.

  • chat/ChatWidget.test.tsx renders the widget inside its real providers with the network replaced by a scripted event stream (vi.mock('../lib/api')), then types, sends, and clicks like a user: the streamed answer, activity chips and lookups, the request contents, multi-turn history, error lines, thumbs up and down, the reason card and its validation, skip, switching, embed votes without content, a failed send, and that feedback never enters the history the model sees.
  • embed/element.test.tsx covers the custom element: a claim assigned before definition survives, the shadow root, the API base and key from attributes, following a new claim, and keeping its dock state apart from the demo's.
  • claimpilot/claimView.test.ts sets fake timers before importing the module, because TODAY is read at import.

Coverage​

npm run test:coverage

Writes a text summary, coverage/index.html and coverage/lcov.info, and enforces the per-layer gates.