End-to-end tests
Playwright drives Chromium against the real Worker under wrangler dev, serving the built SPA, with the scripted model and a throwaway local D1 seeded with the fixture claims.
npm run test:e2e # build, then every spec
npx playwright test e2e/feedback.spec.ts --headed
npx playwright test --repeat-each=4 # hunt for flakiness
npx playwright show-report # traces, screenshots and video of any failure
The test Worker: e2e/server.ts
playwright.config.ts starts e2e/server.ts as its web server. The script:
- requires a built frontend (
npm run test:e2ebuilds first); - deletes and recreates
e2e/.state/(gitignored), so every run starts from the same data; - applies the D1 migrations to a local database persisted there, not to
api/.wrangler/state; - seeds the three fixture claims;
- starts
wrangler devon 127.0.0.1:8799 (inspector on 9799) with--env-file e2e/worker.env, plus--var AI_PROVIDER:scripted --var DATABASE_PROVIDER:d1for belt and braces.
Passing an env file stops wrangler reading api/.dev.vars, so a developer's Gemini key and Neon URL are never loaded. e2e/worker.env holds no secrets: it blanks the key, the database URL and the gateway ids, and allows one test origin.
e2e/api.spec.ts asserts up front that /api/health reports d1, scripted and no gateway. If that test ever fails, the suite is talking to something real. Stop and fix e2e/worker.env before anything else.
Locally, an already-running e2e Worker on :8799 is reused, which makes iterating on specs quick. Its database then accumulates votes between runs, so specs never assume an empty feedback table; they find their own vote by turn id. On CI (CI set) a fresh Worker is always started.
Configuration
| Setting | Value | Why |
|---|---|---|
workers | 1, not fully parallel | One shared Worker and database |
retries | 2 on CI, 0 locally | CI absorbs a rare network blip; locally a flake should be seen |
trace, screenshot, video | kept on failure | Debug from the report |
| Viewport | 1440 × 900 | The claim screens are dense |
reporter | github + HTML on CI, list + HTML locally |
Each test gets a fresh browser context, so its localStorage starts empty: UK region, chat closed, no Prompt Lab versions.
The specs
| Spec | Covers |
|---|---|
api.spec.ts | The offline guard; seeded data; CORS refusal and preflight; request validation; API 404 versus the SPA fallback |
claims.spec.ts | The random landing claim; My Claims filters and search; panel anchors; the UK/US switch (chrome, claims, dates, persistence); unknown routes |
chat.spec.ts | A full turn (activity chips, answer, lookups, panel highlight, request contents); multi-turn history and clear; a demo-tip quick action; the error line; docking across reloads; switching claim resets the chat; the US locale |
feedback.spec.ts | Thumbs up stored with content, then withdrawn; thumbs down with reason-card validation and storage; down then up drops reasons; feedback never enters the model history; the Prompt Lab panel and Replay in Try It |
promptlab.spec.ts | Saving a version, its use in chat from the claim screen, reset to default, history; Try It with unsaved edits and its trace |
embed.spec.ts | <claimpilot-assistant> on Claimpanion: the shadow root, the host's claim (never stored), the claimpilot-feedback event with content versus the stored vote without, style isolation |
Helpers: e2e/support.ts
| Helper | Does |
|---|---|
UK_AUTO, UK_PROPERTY, US_WC | The seeded fixture ids and titles |
scriptedAnswer(claim, tools?) | The exact text the scripted model produces |
chatPanel(page) | The chat panel, wherever it is, including inside the embed's shadow root |
openChat(page, assistant?) | Click the launcher and wait for the panel |
ask(page, question) | Type, send, and resolve with the chat request as the browser sent it |
voteSentBy(page, action) | Run an action and resolve with the vote the browser posted (asserting 204) |
storedVote(request, turnId) | Read one stored vote back from GET /api/feedback |
Playwright's locators pierce open shadow roots, so the same helpers work inside the embed.
Writing a spec
- Prefer role and label locators:
getByRole('button', { name: 'Helpful', exact: true }). Useexactwhere one name contains another (Helpful is inside Not helpful). - Use the
data-panelanchors for claim-screen panels. - Assert on what the browser sent (
ask,voteSentBy) and what the server stored (storedVote), not only on what is visible. - Never assume an empty feedback table; identify your own data by turn id or a unique question.
- Run new specs with
--repeat-each=4before committing.