Why your AI demo is more accurate than your AI will ever be
Demos are scored on data that was cleaned after the fact. Production runs on data as it arrives. Here's how to tell, before you fund the build, which number you're looking at.
Emerald Isle Consulting · October 2026 · 5 min read
Every AI steering committee has seen the slide: a model scoring in the low 90s on historical data, a confident recommendation to fund the build. Months later, in production, the same model is somewhere in the 50s or 60s, and nobody can quite explain what changed.
Usually nothing changed about the model. What changed is the data. The demo was scored on one kind of data, and production runs on another.
The two datasets behind every AI decision
Insurance claims and healthcare claims share a pattern: information keeps arriving after the moment a decision has to be made. By the time a record is closed, someone has corrected it, coded it, verified it and filled in its gaps.
Models are almost always built from closed records, because closed records come with known outcomes to learn from. But they're deployed at the start of the process, where records are open, partial and messy.
| What the model uses | At decision time | In the closed record |
|---|---|---|
| Free-text descriptions | Typed under time pressure | Coded by an expert |
| Key identifiers | Missing, duplicated or split | Reconciled and linked |
| Coverage or eligibility | A stale batch extract | Verified for the date in question |
| Documents | Scanned PDFs and photos | Abstracted into fields |
| The outcome | Unknown, which is the point | Known, and leaking into the inputs |
Why the demo number is real, and irrelevant
The demo isn't faked. It is genuinely accurate at the question it was asked. The trouble is that the question was easier than the one production asks. It isn't predicting how a case will develop from what's known on day one. It's recognizing how a case developed, using details written down weeks later.
Data scientists call one form of this leakage: information from after the decision point creeping into the training data. A field that only gets filled in once a claim turns complex will look like a brilliant predictor of complexity. In production, it's empty at exactly the moment the model needs it.
A model tested on closed records is being graded on information it will never see when it matters.
The same pattern, two industries
In P&C insurance, a first-notice triage model is trained on closed claims, where an adjuster has coded the cause of loss, confirmed the parties and entered the estimate. At first notice, the model gets a call-center note and a scanned PDF.
In healthcare, a fraud model is trained on closed investigations, where investigators have stitched one clinician's several provider IDs into one identity. In live claims, that clinician looks like several unrelated providers, and the pattern the model learned is split across them.
Different industries, same mistake: the training data had already been fixed by hand, and production data hadn't.
How to get a number you can trust
- 01
Rebuild the test set as of the decision point. For every record in the test set, use only the data that existed at the moment the decision would have been made. This single change is usually where the demo number falls.
- 02
Audit every input for when it gets filled in. Any field that's populated after the decision point is a leakage risk. Drop it, or replace it with what was actually known then.
- 03
Grade the systems that feed the model. Point-in-time testing needs history the systems may not keep. If prior states are overwritten, that's a lineage gap, and it has to be fixed before you can even measure the model honestly.
- 04
Write the target down before you build. Agree in writing what accuracy or precision is good enough, measured on decision-time data. It turns a demo into a commitment.
If the honest number is still good, you have a real model and a stronger case for funding it. If it isn't, you've learned it in weeks rather than after go-live, and you know exactly which data to fix first.

