Skip to main content

End-to-end tests

Playwright drives Chromium against the real Worker under wrangler dev, serving the built SPA, with the scripted model and a throwaway local D1 seeded with the fixture claims.

npm run test:e2e # build, then every spec
npx playwright test e2e/feedback.spec.ts --headed
npx playwright test --repeat-each=4 # hunt for flakiness
npx playwright show-report # traces, screenshots and video of any failure

The test Worker: e2e/server.ts​

playwright.config.ts starts e2e/server.ts as its web server. The script:

  1. requires a built frontend (npm run test:e2e builds first);
  2. deletes and recreates e2e/.state/ (gitignored), so every run starts from the same data;
  3. applies the D1 migrations to a local database persisted there, not to api/.wrangler/state;
  4. seeds the three fixture claims;
  5. starts wrangler dev on 127.0.0.1:8799 (inspector on 9799) with --env-file e2e/worker.env, plus --var AI_PROVIDER:scripted --var DATABASE_PROVIDER:d1 for belt and braces.

Passing an env file stops wrangler reading api/.dev.vars, so a developer's Gemini key and Neon URL are never loaded. e2e/worker.env holds no secrets: it blanks the key, the database URL and the gateway ids, and allows one test origin.

The offline guard

e2e/api.spec.ts asserts up front that /api/health reports d1, scripted and no gateway. If that test ever fails, the suite is talking to something real. Stop and fix e2e/worker.env before anything else.

Locally, an already-running e2e Worker on :8799 is reused, which makes iterating on specs quick. Its database then accumulates votes between runs, so specs never assume an empty feedback table; they find their own vote by turn id. On CI (CI set) a fresh Worker is always started.

Configuration​

SettingValueWhy
workers1, not fully parallelOne shared Worker and database
retries2 on CI, 0 locallyCI absorbs a rare network blip; locally a flake should be seen
trace, screenshot, videokept on failureDebug from the report
Viewport1440 × 900The claim screens are dense
reportergithub + HTML on CI, list + HTML locally

Each test gets a fresh browser context, so its localStorage starts empty: UK region, chat closed, no Prompt Lab versions.

The specs​

SpecCovers
api.spec.tsThe offline guard; seeded data; CORS refusal and preflight; request validation; API 404 versus the SPA fallback
claims.spec.tsThe random landing claim; My Claims filters and search; panel anchors; the UK/US switch (chrome, claims, dates, persistence); unknown routes
chat.spec.tsA full turn (activity chips, answer, lookups, panel highlight, request contents); multi-turn history and clear; a demo-tip quick action; the error line; docking across reloads; switching claim resets the chat; the US locale
feedback.spec.tsThumbs up stored with content, then withdrawn; thumbs down with reason-card validation and storage; down then up drops reasons; feedback never enters the model history; the Prompt Lab panel and Replay in Try It
promptlab.spec.tsSaving a version, its use in chat from the claim screen, reset to default, history; Try It with unsaved edits and its trace
embed.spec.ts<claimpilot-assistant> on Claimpanion: the shadow root, the host's claim (never stored), the claimpilot-feedback event with content versus the stored vote without, style isolation

Helpers: e2e/support.ts​

HelperDoes
UK_AUTO, UK_PROPERTY, US_WCThe seeded fixture ids and titles
scriptedAnswer(claim, tools?)The exact text the scripted model produces
chatPanel(page)The chat panel, wherever it is, including inside the embed's shadow root
openChat(page, assistant?)Click the launcher and wait for the panel
ask(page, question)Type, send, and resolve with the chat request as the browser sent it
voteSentBy(page, action)Run an action and resolve with the vote the browser posted (asserting 204)
storedVote(request, turnId)Read one stored vote back from GET /api/feedback

Playwright's locators pierce open shadow roots, so the same helpers work inside the embed.

Writing a spec​

  • Prefer role and label locators: getByRole('button', { name: 'Helpful', exact: true }). Use exact where one name contains another (Helpful is inside Not helpful).
  • Use the data-panel anchors for claim-screen panels.
  • Assert on what the browser sent (ask, voteSentBy) and what the server stored (storedVote), not only on what is visible.
  • Never assume an empty feedback table; identify your own data by turn id or a unique question.
  • Run new specs with --repeat-each=4 before committing.