Capabilities and tools
A capability is something the assistant can do for a claim: summarise it, validate the policy, run a fraud check. A tool is a function the model can call to read or compute something from the claim. Capabilities own playbooks; tools do the reading and the arithmetic.
A capability has two halves
The public half (CapabilityMeta) is what the UI needs: the label and one-line description, the group, whether it is a pinned chip, and the message a quick action sends so it reads naturally in the transcript. The assistant reads the same rows when asked what it can do.
The server half (Capability in api/src/capabilities/types.ts):
export interface Capability {
id: string; // must match the shared catalogue
promptId: string; // the playbook in prompts/catalogue.ts
tools?: ChatTool[]; // tools this capability brings
thinkingLevel?: ThinkingLevel; // 'low' | 'medium' | 'high' for this task
}
Server capabilities are declared one object per capability in five group files, understand.ts, decide.ts, protect.ts, produce.ts and help.ts, and gathered in registry.ts:
export const decideCapabilities: Capability[] = [
{ id: 'policy_validation', promptId: 'capability.policy_validation', tools: [checkPolicyInForce], thinkingLevel: 'high' },
{ id: 'next_actions', promptId: 'capability.next_actions', tools: [listDeadlines], thinkingLevel: 'medium' },
// …
];
assertCatalogueMatches() runs at Worker start-up and throws if an id exists on one side and not the other. A unit test checks the same thing, and that every promptId exists in the prompt catalogue.
Quick actions and free chat
A quick action (a chip, the More menu, or a demo hint) sends the capability's quickActionMessage with capability: '<id>'. The system prompt then carries ## Task: <label> and that capability's playbook, and the capability's thinking level applies.
Free chat sends no capability. The prompt carries the claim_qa playbook plus a short guide listing what each specialist task does, so when a free question amounts to one of them the model can do it to the same standard.
Tools
export interface ChatTool {
definition: ToolDefinition; // name, description, JSON-schema parameters (provider-neutral)
screenAction?: boolean; // acts on the page, not the claim
activity: (input) => string; // the chip: "Reading the policy…"
run: (input, { claim, today }) => ToolOutcome | Promise<ToolOutcome>;
}
export interface ToolOutcome {
result: string; // back to the model; keep it compact
clientActions?: ClientAction[]; // for the browser to perform
}
Rules every tool follows:
- Read-only. A tool reads the one claim it is given. It cannot write, fetch or reach another claim.
- Deterministic. Same claim and same
today, same result.todayis passed in rather than read from the clock, so tests can pin it. - Compact JSON out. The result lives in the model's context for the rest of the turn. Return what the model needs to reason with, not the whole section again.
- Honest about absence. Say "No diagnosis recorded on this claim." rather than returning an empty object the model might misread.
- Errors are results.
runToolcatches a throw and returnsTool failed: <message>withisError, which the model can read and work around.
Shared tools and capability tools
Five shared tools are the general readers, always first in the list: get_claim_section, search_claim, get_document, financial_summary, and the screen action highlight_panel.
Eleven capability tools do the work a model should not do in its head: check_policy_in_force, list_deadlines, settlement_position, run_risk_cross_checks, scan_vulnerability_signals, scan_expressions_of_dissatisfaction, run_handling_checks, build_timeline, claim_metrics, get_party and list_capabilities.
Every tool is offered on every turn, deduplicated by name. A capability "owns" a tool in the sense that its playbook tells the model to call it first; the model can still use any tool in any task. The tools reference lists every one with its parameters.
Screen actions
highlight_panel returns a ClientAction instead of reading anything. The Worker forwards it to the browser as a client_action event and the browser performs it. It is flagged screenAction: true, so allTools() leaves it out unless the request says the caller can act on the screen. The panel names are the shared MATTER_PANELS list, so the tool's enum, the prompt and the data-panel anchors in the DOM cannot drift apart.
Dates
api/src/capabilities/dates.ts holds the date arithmetic the tools share. daysBetween and addDays work on YYYY-MM-DD strings, which Date.parse reads as UTC midnight, so daylight saving cannot shift a deadline by a day. ageInYears is the one helper that reads the clock, because an age is "as of now" by definition.