← Insights

Why fraud models can't see provider networks

Fraud, waste and abuse detection rarely fails on the algorithm. It fails on provider data that splits one clinician into many, and you can grade that risk before you fund the model.

Emerald Isle Consulting · October 2026 · 5 min read

Payment integrity teams know the arc. A model is trained to spot fraud, waste and abuse: billing patterns that don't fit a provider's peers, services that cluster in suspicious ways, networks of providers that refer to each other more than they should. Trained on closed investigations, it scores beautifully in the steering committee demo. In production, investigators drown in false positives, the real schemes slip through, and within a few quarters the scores are ignored.

The usual post-mortem blames the model. Usually the model was fine. It was trained on data that had been cleaned by investigators, and deployed on data that hadn't.

What a fraud model actually needs

FWA detection is a comparison problem. To call a provider unusual, the model has to see everything that provider billed, compare it with true peers, and judge it against what was covered and authorized at the time. Every one of those depends on the data agreeing about who the provider is and what really happened.

Here is the gap in one table.

What the model usesIn live claimsIn the training cases
Provider identitySeveral NPIs, TINs and addressesStitched together by investigators
Specialty and peer groupTaxonomy codes mapped three waysConfirmed for the case
What was billedAdjustments overwrite the original lineReconstructed in the case file
Member eligibilityMonthly snapshotVerified for the date of service
Prior authorizationMatched to claims by hand, if at allPulled into the case file

Why the demo lies

Closed investigations are the obvious training set: they come with labels. But by the time a case is closed, the special investigations unit has done the hardest data work by hand. They have worked out that three NPIs, two tax IDs and four addresses belong to one clinician. They have rebuilt the claim lines as originally billed. They know which peer group the provider really belongs to.

So the model learns to recognize schemes in data where one provider looks like one provider. Live claims don't look like that, and the demo never tested whether the model could cope.

A fraud model trained on investigated cases is being tested on a provider network it has never seen assembled.

What production looks like

On go-live, the same clinician bills under an individual NPI, a group NPI and more than one tax ID. Directory addresses don't verify. Specialty codes arrive mapped differently from each source feed, so peer comparisons put providers in the wrong groups. Adjustments overwrite the original claim line, so the model can't see what was first billed.

The result cuts both ways. Billing that is unremarkable gets compared with the wrong peers and flagged. Billing that is genuinely suspicious is split across records that never meet, so no single record looks unusual. Investigators chase the first and miss the second.

The grades that predict it

We grade every system behind a use case on five dimensions: data quality, documentation, lineage, access and security, and AI consumability. For FWA detection, provider data decides the outcome before a line of model code is written.

AI consumability. Whether a model can read a provider as one entity at the moment it scores a claim. If identity is only resolved inside case files, it's an F, however good the investigated data looks.

Documentation and lineage. If taxonomy codes mean different things in different feeds, and credentialing updates arrive by spreadsheet with no trace to their source, nobody can build or defend a peer group. That is usually a D.

F

In our sample report card for a fictional health plan, Provider Data scores F on AI consumability and D on three other dimensions. That's where the fraud model dies. See the report card →

What to do instead

  1. 01

    Grade the systems before you fund the model. Score provider data, claims, eligibility and authorizations against what the use case needs. It takes weeks, costs a fraction of a stalled pilot, and tells you whether the model can work at all.

  2. 02

    Resolve provider identity first. One record per provider across NPI, TIN and address, with specialty mapped one way. It is the narrow path FWA detection depends on, and it pays off far beyond fraud.

  3. 03

    Keep claims as billed, and eligibility as of the day. Preserve original claim lines through adjustments, and reconstruct eligibility for the date of service. The model has to see what the plan saw when it paid.

  4. 04

    Write the precision target down, sized to your investigators. Agree how many alerts the team can work, and how many must be right, before work starts. It turns a demo into a commitment investigators can trust.

None of this is exotic. It is the data work investigators already do by hand, done once and upstream, before the model is built. The plans that fix provider data first will find the fraud their models were bought to find.