Insights

Why FNOL triage pilots die

The model is rarely the problem. The data at first notice of loss usually is, and you can see it coming before you fund the pilot.

Emerald Isle Consulting · October 2026 · 4 min read

The pattern is familiar to anyone who has sat on a claims AI steering committee. A vendor or internal team builds a model to triage first notices of loss: flag the complex claims, fast-track the simple ones, route the likely litigation to the right adjuster. The demo is excellent. Funding follows. Six months after go-live, adjusters have stopped looking at the scores, and the project is quietly shelved.

It is a common story. Capgemini's 2026 research on P&C insurers found that 60% are still stuck in AI pilot or proof-of-concept mode, and only 12% report high maturity in AI data readiness. Those two numbers are related. Most of the pilots that stall don't fail on the model. They fail on the data underneath it.

What a triage model actually needs

Triage happens at first notice, which is the whole point: the earlier a complex claim reaches the right handler, the more it saves. That means the model has to work with what exists at the moment the claim is reported, not what exists six weeks later.

Here is the problem in one table.

What the model usesAt first noticeIn the historical data
Loss type and causeFree text from the callerCoded by the adjuster
Loss locationOften missing or free textVerified and geocoded
Injuries and partiesPartial, unconfirmedConfirmed, structured
Coverage in forceWeekly batch, may be staleVerified by the adjuster
Vehicle or property detailsScanned PDFs, photosEstimate entered

Why the demo lies

Demos are trained and tested on historical claims. Historical claims are clean in exactly the way first notices are not: by the time a claim is closed, an adjuster has coded the cause of loss, confirmed the parties, verified coverage, and entered the estimate. The model learns from all of it.

So the demo is quietly answering an easier question than the one production asks. It isn't predicting how a claim will develop from a phone call. It is predicting how a claim developed, using details that were written down weeks after the call. Accuracy in the 90s is real, and also irrelevant.

A triage model trained on closed claims is being tested on information it will never see at first notice.

What production looks like

On go-live, the model meets the real intake stream. Call-center notes typed under time pressure. Web forms where the loss description is a paragraph of free text. Agent submissions as scanned PDFs. Call-center and web notices landing in separate tables with no shared claim key. Coverage pulled from a weekly batch, so recently changed policies look wrong.

None of that is unusual. It is just what first-notice data looks like at most carriers. But a model trained on structured, adjuster-coded fields has almost nothing it recognizes, and accuracy slides accordingly.

The grades that predict it

When we assess the systems behind a use case, we grade each one on five dimensions: data quality, documentation, lineage, access and security, and AI consumability. For FNOL triage, two of those decide the outcome before a line of model code is written.

Documentation. If nobody can say what each intake field means, nobody can map it to what the model expects. An F here usually means the field meanings live with the original vendor team.

AI consumability. Whether a model can read the data at the moment it needs to. Free-text loss descriptions and scanned PDFs at first notice are an F, however good the historical data looks.

F

In our sample report card for a fictional carrier, FNOL Intake scores F on both. That's where the triage pilot dies. See the report card →

What to do instead

  1. 01

    Grade the systems before you fund the model. Score every system the use case depends on. It takes weeks, costs a fraction of a stalled pilot, and tells you whether the pilot can work at all.

  2. 02

    Test only on what exists at first notice. Rebuild the training and test sets from data as it looked at the moment of reporting. If accuracy holds, you have a real model. If it collapses, you've learned that cheaply.

  3. 03

    Fix the narrow path, not the whole estate. Structured loss capture at intake, extraction from documents, a shared claim key across channels, and current coverage lookup. That's the path triage needs; the rest of the roadmap can wait.

  4. 04

    Write the accuracy target down before you build. Agree what "good enough" means, measured on first-notice data, before work starts. It turns a demo into a commitment.

None of this is exotic. It is simply doing the data work before the model work, in the order the use case needs it. The carriers that fix the data first will ship AI first.