A data readiness checklist for claims AI
Before you fund FNOL triage, claims document intelligence or a fraud model, answer these questions about your claims data. Your own team can check most of them in a few days.
Emerald Isle Consulting · October 2026 · 6 min read
Claims is where most carriers start with AI, for good reason: high volumes, expensive decisions and a lot of manual work. It's also where pilots most often stall, because claims data is created under time pressure at first notice and corrected for months afterward.
This checklist is organized by the systems a claims use case reads from. For each check there's a quick way to test it and what a failing answer looks like. A failing answer doesn't mean you can't build. It means you know what to fix first.
Start with the decision, not the data
Write down the use case leadership has funded and the exact moment the model decides. For triage it's first notice of loss. For a document model it might be when a file is assigned. For fraud it's often the first payment. Every check below is measured against that moment, because a model can only use what exists then.
If you can't name the decision point, you can't test whether the data is ready for it.
FNOL intake
| Check | A failing answer looks like | How to test it |
|---|---|---|
| Is the loss captured in structured fields? | Most of the detail is in a free-text description | Pull 100 recent first notices; count how many have cause, location and damage type as fields |
| Is every channel captured the same way? | Each channel has its own form and its own codes | Compare fields from phone, web, app and agent notices |
| Does anyone know what each intake field means? | Different answers, and no data dictionary to settle it | Ask two teams to define five intake fields |
Claims system
| Check | A failing answer looks like | How to test it |
|---|---|---|
| Are key terms defined once? | Three definitions, so training labels disagree | Ask claims, finance and actuarial what an "open claim" is |
| Is history kept? | Changes overwrite the old values | Try to show a claim's status and reserve as of a past date |
| Where do the outcome labels come from? | Set by hand, late, or only on some claims | Trace how severity, fraud or litigation flags are set |
Policy and coverage
| Check | A failing answer looks like | How to test it |
|---|---|---|
| Can you see coverage as of the date of loss? | Only today's policy, or endorsements overwritten | Pick 20 claims; retrieve coverage as it stood that day |
| How fresh is coverage data for the claims team? | A weekly batch, when triage needs it at first notice | Check how policy data reaches claims |
| Is each insured one record? | The same insured under several IDs | Profile duplicate policyholders |
Documents and notes
| Check | A failing answer looks like | How to test it |
|---|---|---|
| How much of the claim lives only in documents? | Estimates, statements and repair details exist only as images | Sample claim files; list facts found only in PDFs, photos or adjuster notes |
| Can documents be read by a machine today? | Scanned images with no extracted text | Check whether text is extracted and stored with the claim |
Across systems
| Check | A failing answer looks like | How to test it |
|---|---|---|
| Is there one claim key across channels and systems? | Matches by name and date, or not at all | Join intake, claims and payments for 100 claims |
| Can a reported figure be traced to its source? | The trail ends at the warehouse load | Trace paid loss on one report back to the claims system |
| How long does an AI team wait for access? | Months, or one administrator approves everything | Time the last approved data request |
| Are personal fields classified? | Nobody's sure which fields hold PII | Check whether claimant PII is tagged and masked where it isn't needed |
What to do with the answers
Sort every failing answer into three groups: it blocks the use case, it degrades the model's accuracy, or it creates audit risk. Then fix only the blockers on the use case's path before you build.
- 01
Fix the blockers first, even small ones. A missing claim key or unstructured loss capture at intake will stop a triage model regardless of everything else.
- 02
Retest on first-notice data. Rebuild the test set from data as it existed at the decision point, and agree a written accuracy target on that basis.
- 03
Schedule the rest. Degraders and audit risks go on a roadmap, sequenced by the next use case that needs them.
In our sample read-out for a fictional carrier, the FNOL triage pilot fails on two checks from this list: unstructured loss capture and an undocumented intake system. See the report card →
If you'd rather get a first read in five minutes, the free scorecard asks ten questions across the same five dimensions an assessment grades.

