The five grades that predict whether an AI pilot ships
AI pilots rarely fail on the model. They fail on one of five things about the data underneath it, and each one can be graded before you fund the build.
Emerald Isle Consulting · October 2026 · 6 min read
Ask a steering committee why an AI pilot stalled and you'll hear about the model: it wasn't accurate enough, it drifted, the users didn't trust it. Look underneath and the cause is usually the data. Not "bad data" in general, but one specific weakness that was there before the first line of model code was written.
When we assess the systems behind an AI use case, we grade each one on five dimensions. Each is a different way a pilot fails in production. Together they answer the only question that matters before funding: can this data support this use case?
1. Data quality
The question: are the values right, and does anyone measure whether they are? Models learn whatever errors the data contains, at scale. A field that's wrong in 6% of records isn't a rounding error to a model; it's a pattern it will learn.
An F here isn't just messy data. It's messy data nobody measures, so nobody knows how messy it is. An A means the critical fields have automated checks, and someone is alerted when they fail.
2. Documentation
The question: is what each field means written down? If no one can say what a field means, no one can map it to what a model expects. It sounds basic, and it's one of the most common failing grades we see.
The usual symptom: three teams with three definitions of the same thing. In our sample read-out for a fictional carrier, claims, finance and actuarial each define an "open claim" differently, so a model's training labels disagree with each other before training starts.
3. Lineage
The question: can each number be traced to its source, and each change traced to what it affects? Regulators and auditors ask where every model input came from. In insurance and healthcare, being able to answer decides whether a model can be used at all, not just how well it performs.
Lineage failures are quiet. Endorsement history gets overwritten nightly. Retroactive eligibility changes replace the past. Nothing looks broken, until someone asks what the data looked like on the day a decision was made, and nobody can say.
4. Access and security
The question: can the right people reach the data, quickly and safely? AI teams that wait months for data never get past the pilot. Teams that get it fast, but without controls on personal and medical fields, create a different problem.
An F looks like access granted request by request through one administrator, with sensitive fields unclassified. An A looks like a standard, fast path to approved access, with sensitive fields classified and masked where they aren't needed.
5. AI consumability
The question: can a model read the data at the moment it needs to decide? This is the dimension most often missing from data assessments, and the one that most often kills a specific use case.
Data can be accurate, documented and traceable and still be unusable: locked in a weekly batch when the model scores in real time, or sitting in free-text notes and scanned PDFs when the model needs structured fields.
Good data that a model can't read at decision time is, for that model, no data at all.
What an F looks like, side by side
| Dimension | What an F looks like | What an A looks like |
|---|---|---|
| Data quality | Errors nobody measures | Automated checks on critical fields |
| Documentation | Meanings live in people's heads | A glossary people actually use |
| Lineage | History overwritten, sources untraceable | Every reported figure traceable |
| Access & security | Months to access; sensitive fields unclassified | Days to access; sensitive fields controlled |
| AI consumability | Batch extracts, free text, scanned PDFs | Governed feeds a model can read live |
Which grades decide the pilot
Not all five matter equally for every use case. The grade that decides a pilot is the one on the path that use case depends on. A first-notice-of-loss triage model lives or dies on the intake system's documentation and AI consumability. A fraud model lives or dies on whether provider data can be read as one identity per provider.
That's why we grade every system on all five, then weight the findings toward the use case leadership funded. A D on a system the use case never touches can wait. An F on the one it depends on can't.
In both of our sample read-outs, the pilot fails on AI consumability: FNOL Intake for the carrier, Provider Data for the health plan. See the report card →
How to use the five grades
- 01
Name the use case first. Grades only mean something against a decision. Start with the AI use case leadership has funded, and list every system it depends on.
- 02
Grade every one of those systems on all five. Use evidence, not opinion: interviews with the people who run each system, plus profiling on real extracts.
- 03
Find the grade that blocks. Look for the D or F on the path the use case depends on. That's the gap that decides whether the pilot ships.
- 04
Fix that path before you build. Fund the narrow fixes the use case needs, and let the rest of the estate wait its turn.
You can get a first read on your own data in five minutes: the free scorecard asks ten questions across the same five dimensions. It isn't an audit, but it shows you where to look first.

