← Insights

A data readiness checklist for claims AI

Before you fund FNOL triage, claims document intelligence or a fraud model, answer these questions about your claims data. Your own team can check most of them in a few days.

Emerald Isle Consulting · October 2026 · 6 min read

Claims is where most carriers start with AI, for good reason: high volumes, expensive decisions and a lot of manual work. It's also where pilots most often stall, because claims data is created under time pressure at first notice and corrected for months afterward.

This checklist is organized by the systems a claims use case reads from. For each check there's a quick way to test it and what a failing answer looks like. A failing answer doesn't mean you can't build. It means you know what to fix first.

Start with the decision, not the data

Write down the use case leadership has funded and the exact moment the model decides. For triage it's first notice of loss. For a document model it might be when a file is assigned. For fraud it's often the first payment. Every check below is measured against that moment, because a model can only use what exists then.

If you can't name the decision point, you can't test whether the data is ready for it.

FNOL intake

CheckA failing answer looks likeHow to test it
Is the loss captured in structured fields?Most of the detail is in a free-text descriptionPull 100 recent first notices; count how many have cause, location and damage type as fields
Is every channel captured the same way?Each channel has its own form and its own codesCompare fields from phone, web, app and agent notices
Does anyone know what each intake field means?Different answers, and no data dictionary to settle itAsk two teams to define five intake fields

Claims system

CheckA failing answer looks likeHow to test it
Are key terms defined once?Three definitions, so training labels disagreeAsk claims, finance and actuarial what an "open claim" is
Is history kept?Changes overwrite the old valuesTry to show a claim's status and reserve as of a past date
Where do the outcome labels come from?Set by hand, late, or only on some claimsTrace how severity, fraud or litigation flags are set

Policy and coverage

CheckA failing answer looks likeHow to test it
Can you see coverage as of the date of loss?Only today's policy, or endorsements overwrittenPick 20 claims; retrieve coverage as it stood that day
How fresh is coverage data for the claims team?A weekly batch, when triage needs it at first noticeCheck how policy data reaches claims
Is each insured one record?The same insured under several IDsProfile duplicate policyholders

Documents and notes

CheckA failing answer looks likeHow to test it
How much of the claim lives only in documents?Estimates, statements and repair details exist only as imagesSample claim files; list facts found only in PDFs, photos or adjuster notes
Can documents be read by a machine today?Scanned images with no extracted textCheck whether text is extracted and stored with the claim

Across systems

CheckA failing answer looks likeHow to test it
Is there one claim key across channels and systems?Matches by name and date, or not at allJoin intake, claims and payments for 100 claims
Can a reported figure be traced to its source?The trail ends at the warehouse loadTrace paid loss on one report back to the claims system
How long does an AI team wait for access?Months, or one administrator approves everythingTime the last approved data request
Are personal fields classified?Nobody's sure which fields hold PIICheck whether claimant PII is tagged and masked where it isn't needed

What to do with the answers

Sort every failing answer into three groups: it blocks the use case, it degrades the model's accuracy, or it creates audit risk. Then fix only the blockers on the use case's path before you build.

  1. 01

    Fix the blockers first, even small ones. A missing claim key or unstructured loss capture at intake will stop a triage model regardless of everything else.

  2. 02

    Retest on first-notice data. Rebuild the test set from data as it existed at the decision point, and agree a written accuracy target on that basis.

  3. 03

    Schedule the rest. Degraders and audit risks go on a roadmap, sequenced by the next use case that needs them.

F

In our sample read-out for a fictional carrier, the FNOL triage pilot fails on two checks from this list: unstructured loss capture and an undocumented intake system. See the report card →

If you'd rather get a first read in five minutes, the free scorecard asks ten questions across the same five dimensions an assessment grades.