Glossary
Plain-language definitions of the terms we use in assessments, read-outs and articles, each with why it matters when the data has to feed an AI model.
Whether the values in a system are correct, complete and consistent, and whether anyone measures that on a schedule.
WhyA model learns whatever errors the data holds, at scale. Unmeasured quality means nobody knows how wrong the model's inputs are.
Whether the meaning of each field is written down, kept current and agreed across the teams that use it.
WhyIf nobody can say what a field means, nobody can map it to what a model expects, and training labels end up disagreeing with each other.
The record of where each value came from, what transformed it on the way, and what depends on it downstream.
WhyRegulators and auditors ask where every model input came from. Without lineage, you can't answer, and you can't rebuild what the data looked like on a past date.
Whether the right people can reach the data through a defined, reasonably fast path, with sensitive fields classified and protected.
WhyAI teams that wait months for data never leave the pilot. Teams that get it without controls on personal or medical fields create a different problem.
Whether a model can read the data it needs, in a structured form, at the moment it has to make a decision.
WhyAccurate, documented data is still useless to a model if it sits in free text, scanned PDFs or a weekly batch when the model scores in real time.
Measuring real data, field by field: how complete it is, which values are invalid, where duplicates are, and how records link across systems.
WhyIt turns opinions from interviews into evidence. Every grade in an assessment rests on it.
A technical reference listing each table and field in a system, with its type, allowed values and meaning.
WhyIt's the starting point for mapping source data to model inputs. Where it's missing or stale, documentation grades fall.
The agreed business definitions of shared terms, such as what counts as an open claim or an active member, owned by named people.
WhyWhen claims, finance and actuarial define the same term three ways, a model's training labels disagree before training starts.
The person accountable for the definitions and quality of a set of data, and the first call when it's wrong.
WhyGaps without an owner don't get fixed. Naming stewards is usually one of the cheapest, earliest roadmap items.
The single, trusted version of an entity, such as a customer, policyholder or provider, assembled from the records different systems hold about it.
WhyModels that compare patterns across records need one identity per entity. Without it, one provider or claimant looks like several.
The process of deciding which records, across or within systems, refer to the same real-world person or organization.
WhyIt's how a golden record gets built, and it's often the fix that unblocks fraud and network models.
The handful of systems, fields and joins a specific AI use case depends on, as opposed to the whole data estate.
WhyFixing the narrow path gets the first use case into production in months; fixing the whole estate first rarely finishes.
The moment in a process when a model is asked to make its prediction or recommendation, such as first notice of loss or claim submission.
WhyA model can only use what exists at that moment. Everything about readiness is measured against it.
Data as it actually existed at a given moment, before later corrections, coding or adjustments were applied.
WhyTesting a model on point-in-time data is the only way to get an accuracy number that will hold in production.
Information from after the decision point creeping into the data a model is trained or tested on.
WhyIt makes demos look far more accurate than production will ever be, and it's the most common reason a funded model disappoints.
An agreed, documented threshold for how well a model must perform, measured on decision-time data, before it's built.
WhyIt turns a demo into a commitment, and makes it clear when a model is ready to go live.
The first report of a claim to an insurer, by phone, web, app or agent, and the data captured at that moment.
WhyIt's where claims triage models decide, usually on the least structured data in the whole claim lifecycle.
Improper healthcare billing, from deliberate fraud to wasteful or abusive patterns, and the programs that detect it.
WhyFWA models depend on resolving provider identities and keeping claim lines as originally billed, two common data gaps.
The National Provider Identifier for an individual or organization, and the Taxpayer Identification Number used for billing. One clinician can appear under several of each.
WhyIf they aren't resolved to one identity, the pattern a fraud model is looking for is split across what look like unrelated providers.
A change to a member's coverage that applies to past dates, often overwriting what the eligibility record said at the time.
WhyModels need eligibility as of the date of service. Systems that overwrite history lose it, which is a lineage gap.
Replacing or hiding protected health information in a dataset so it can be analyzed without exposing individuals.
WhyIt lets profiling measure sensitive fields without anyone seeing patient details.
Six weeks, from $75K fixed fee, scoped by the number of systems your use case depends on. A fraction of a typical Big Four assessment.
See exactly what's included in the assessment