AI-Assisted Data Quality Triage
The gap hunt worked. That is the problem. Three weeks ago you had a comfortable ignorance about the data feeding your invoice-exception pilot; now you have a spreadsheet with eleven findings on it, each one real, each one documented, each one capable of eating the pilot alive or of eating nothing at all, and you cannot yet tell which is which. Down the hall, a colleague who ran the same kind of audit last year is living the other ending of this story. Her audit took six weeks and produced forty-seven findings in a sixty-page report, every one of them true. Leadership read the executive summary, said "the data people have concerns," funded none of it, and let the pilot proceed unchanged, because a report that says everything says nothing. The pilot died in month four, on finding number twelve. Her "I documented that" was completely accurate and completely useless to her career. This lesson is about the discipline that separates your ending from hers: triage.
The Long List Is Where Audits Go to Die
Here is an uncomfortable truth about diligence: the failure mode of the lazy assessor is missing the gaps, but the failure mode of the diligent assessor is finding all of them and ranking none of them. The encyclopedic audit is a real and common artifact. It is thorough, defensible, alphabetized, and dead on arrival, because it hands leadership a research library when what leadership needed was a decision memo.
Think about what a steering committee can actually do with forty-seven unranked findings. They cannot fund all forty-seven; no budget on earth absorbs an unranked remediation list, because unranked means unbounded. They cannot pick the important ones themselves; if they could judge data-gap severity, they would not have commissioned an assessor. So they do the only thing an overwhelmed decision body ever does: they thank you, they file it, and they revert to the default plan, which is the plan that existed before your audit. Your six weeks of work changed nothing except the paper trail. And the paper trail, as your colleague discovered, is cold comfort. When the pilot dies on a gap you documented but did not flag as fatal, the record shows that you knew and did not say it loudly. Thoroughness without triage is how honest audits get ignored, and how honest auditors get quietly blamed anyway.
The stakes are the same ones this whole level runs on. Gartner found that 63 percent of organizations lack AI-ready data practices or are unsure of them, and predicts that through 2026, 60 percent of AI projects without AI-ready data will be abandoned. Your gap list is the raw evidence that your pilot is at risk of joining that 60 percent. But evidence does not exit you from a statistic. Decisions do. Triage is the machine that converts findings into decisions: which gaps block the pilot, which degrade it, which can wait, and which, however real, are simply not this pilot's problem.
The word triage is borrowed deliberately from emergency medicine, and the borrowed logic matters. A triage nurse with a waiting room full of patients does not treat anyone. She sorts. She sorts fast, on two questions only (how bad is it, how urgent is it), and she accepts that sorting means saying out loud that some patients will wait and some will be sent elsewhere. The sort is not a lesser form of care; it is the thing that makes care possible, because a hospital that tries to treat everyone simultaneously treats no one well. Your gap list is the waiting room. This lesson gives you the sorting instrument.
The Artifact: The Gap Severity Matrix
The deliverable of this lesson is a single named artifact: the Gap Severity Matrix. It is one table, one row per finding from your gap hunt, and it scores every finding on exactly two axes.
Axis one: impact on this pilot's core loop. Not impact on data quality in the abstract, not impact on the enterprise, not impact on some future roadmap. Impact on the specific loop this specific pilot runs: the inputs it reads, the task it performs, the output it produces, the decision that output feeds. Every finding gets one of three ratings: fatal (the pilot's core loop cannot produce trustworthy results while this gap exists), degrading (the loop runs, but measurably worse: more errors, more exceptions, more human cleanup), or cosmetic (the loop does not care; the gap offends your professional standards but not the pilot's arithmetic).
Axis two: fix economics. What it actually costs to close the gap, rated in three honest bands: days-cheap (a script, a config change, a steward's focused week), weeks-moderate (a small project with an owner and a plan), or quarters-expensive (a system change, a cross-department negotiation, a budget line). We will spend a whole section on why this axis is trickier than it looks, because the technical cost and the actual cost of a fix are routinely different by a factor of five, and the difference is almost always people.
Two axes, three ratings each, and the combinations resolve into exactly four action classes. This is the entire output of triage, and every row of your matrix lands in one of them:
| Action class | Rule | What happens |
|---|---|---|
| BLOCKER | Fatal impact, at any fix cost | The pilot does not start until this is fixed. Say so in week one, in writing. |
| PARALLEL FIX | Degrading impact, cheap or moderate fix | Fix while piloting. The degradation is measured and stated in the pilot's error budget. |
| ACCEPT AND LOG | Degrading but expensive to fix, or cosmetic | Pilot proceeds. Risk register entry. Revisit at the scale decision. |
| OUT OF SCOPE | Does not touch this pilot's loop | Documented and handed to the data program. Explicitly excluded from the pilot's budget and timeline. |
And one formatting rule that is really a governance rule: every row ends in a verb, an owner, and a date. Not "vendor master has duplicates" but "Dedupe top 200 vendors by spend: R. Okafor, data steward, by March 14." A matrix row without an owner and a date is a finding wearing a costume; it will be admired and not acted on. The verb-owner-date discipline is what makes the matrix a decision document instead of a longer, prettier version of the sixty-page report.
A finding is an observation. A triaged finding is a decision with a verb, an owner, and a date. Leadership can only fund the second kind.
Scoring Impact: Against Your Pilot, Not Against Perfection
The impact axis is where most first-time triage goes wrong, so go slowly here. The single most important sentence in this lesson is this one: impact is judged against the pilot's design, not against data perfection.
Take a concrete gap: the vendor master file contains duplicates. "Acme Corp," "ACME Corporation," and "Acme Corp." are three records for one supplier, and there are a few hundred clusters like it. Is that fatal, degrading, or cosmetic? The honest answer is: the question is unanswerable as asked, because impact is not a property of the gap. It is a property of the gap's collision with a specific pilot design. If your pilot matches incoming invoices to vendor records, the duplicates sit directly inside the core loop: every duplicate cluster is a coin-flip the matching step can lose, and the gap is fatal or close to it. If your pilot measures cycle time on exception handling and never touches vendor identity, the same duplicates are cosmetic. Same data, same gap, same spreadsheet row from your gap hunt: opposite verdicts, both correct.
This is why enterprise data-quality scores, the kind that report "vendor master: 74 percent quality," are almost useless for pilot decisions. A single score averages the gap's impact across every conceivable use, which means it is precisely wrong for any particular use. Your pilot is a particular use. The Gap Severity Matrix exists because impact must be re-derived, gap by gap, against the actual design you settled on: the fields the pilot reads, the joins it performs, the thresholds it acts on. If you cannot state the pilot's core loop in one sentence, stop and write that sentence first, because every impact rating you assign is secretly a claim about that loop.
The one-line reasoning rule
Because every impact rating is a judgment call about your specific design, every rating must carry its reasoning, and the reasoning must fit on one line. "Fatal: the auto-routing design reads the exception-reason field, and 60 percent of its values are unparseable free text." "Cosmetic: pilot never joins on cost center, so the stale hierarchy is irrelevant to the loop." One line is not a stylistic preference. It is the verification rule of this lesson: no rating ships without its one-line reasoning, because a rating without reasoning cannot be checked, disputed, or defended, and all three of those will happen. The reasoning line is also your protection: when a rating turns out wrong, a documented wrong reason is a fixable model of the pilot; an undocumented rating is just a mistake with your initials on it.
The companion rule: any rating the process owner disputes gets re-scored in the room. Not defended by email, not escalated: re-scored, live, with the person who runs the process looking at the same one-line reasoning you wrote. Half the time they know something you do not (the field you rated fatal was quietly deprecated in April). Half the time you know something they do not (the sample you pulled shows their "clean" field failing at scale). Either way the matrix gets truer, and, just as valuable, it stops being your matrix and starts being the organization's.
Fix Economics: The Technical Estimate Is the Down Payment
The second axis looks like the easy one. Estimating fix effort is a normal professional act; every operations person has scoped work before. The trap is that data fixes carry a second price tag that never appears in the technical estimate, and it is usually the larger of the two: the political cost.
Return to the free-text exception field, the one your gap hunt flagged because the pilot's auto-routing design needs structured reason codes and gets typed prose instead. The technical fix is genuinely small: add a dropdown of reason codes to the entry form, keep a free-text overflow box, migrate nothing. A competent forms team does it in two weeks. Rate it days-to-weeks-cheap and move on?
No. Because forty accounts payable (AP, the finance team that receives and pays supplier invoices) clerks type in that field every working day, and have for nine years. The field is not a schema element to them; it is how they think. Changing it means new reason-code definitions someone must write and defend, training sessions someone must run, a month of clerks selecting "Other" and typing the prose anyway, supervisors enforcing the new habit, and at least one team lead who will escalate that the dropdown "does not fit how we work" (and she will be partly right, and the codes will need a revision). The two-week form change is, in the real organization, a three-month change-management negotiation. That is not a failure of the fix; that is what the fix is. BCG's 10-20-70 rule, the finding that AI value is 10 percent algorithms, 20 percent technology and data, and 70 percent people and process, applies to remediation just as hard as it applies to deployment. Most data gaps are frozen behavior. You are not fixing a field; you are renegotiating a habit forty people rely on, and the fix-economics rating must price the negotiation, not the keystrokes.
So the axis discipline is: rate the whole fix. Days-cheap means days including the humans. Weeks-moderate means an owner could genuinely land it, adoption included, inside the pilot window. Quarters-expensive means system change, cross-team negotiation, or behavior change at scale, whatever the code diff looks like. And this axis has its own signature rule, parallel to the reasoning rule on the impact axis: no fix estimate enters the matrix without its owner's agreement. Your guess at the data steward's effort is a rumor until the data steward says the number. An estimate its owner never saw will be repudiated at the worst possible moment, in front of the steering committee, and the repudiation will taint every other row of your matrix.
The Four Classes, and the Nerve Each One Requires
The mechanics of the four classes fit in the table you already saw. What the table cannot show is that each class demands a different kind of nerve, and the nerve is the actual skill.
BLOCKER: saying "not yet" in week one
A blocker is any fatal-impact gap, regardless of fix cost. The rule is absolute on purpose: if the pilot's core loop cannot produce trustworthy results while the gap exists, then starting the pilot is not brave, it is expensive theater, because the pilot will run for months and then be undone by the thing you knew in week one. Calling the blocker early is the whole job. It is also the moment your role stops being analytical and becomes uncomfortable, because "the pilot everyone announced does not start until X is fixed" is a sentence with a blast radius. Remember the program's spine: a documented kill, or a documented delay for cause, is a win. The assessor who moves a pilot start by six weeks in week one has saved the organization from discovering the same fact in month five at ten times the price. Nobody throws a parade, but the alternative is the parade nobody wants: the post-mortem.
PARALLEL FIX: piloting with a measured limp
Degrading impact plus a cheap or moderate fix means you fix while piloting. The non-negotiable word is measured. A parallel fix is only legitimate if the degradation it is fixing has a number attached and that number is written into the pilot's error budget: "extraction accuracy runs an estimated 8 points below target until the vendor dedupe lands; results before that date are read against the adjusted bar." Without the number, a parallel fix is a euphemism for "we know it is broken and we are proceeding on vibes," and when the pilot's early results look mediocre, nobody will be able to say how much of the mediocrity was the known gap versus the tool. With the number, mediocre early results are exactly on forecast, which is a completely different steering-committee conversation.
ACCEPT AND LOG: the discipline of not fixing
Degrading-but-expensive gaps, and all cosmetic ones, get accepted: the pilot proceeds, the gap goes into the risk register with its one-line reasoning, and the entry carries a revisit date, normally the scale decision. This class exists to protect you from the perfectionist reflex. A gap that costs two quarters to fix and shaves an estimated point off accuracy is not worth delaying a pilot whose entire purpose is to find out whether the use case works at all. Accepting is not ignoring: the log entry is the difference. When the scale decision comes and the stakes multiply by ten, the logged gap gets re-triaged at the new stakes, and yesterday's acceptable limp may become tomorrow's blocker. That re-triage only happens if the log exists.
OUT OF SCOPE: scope discipline as a kindness
The fourth class is the one diligent people resist. Your gap hunt will surface real, serious problems that simply do not touch this pilot's loop: the semantic drift between Finance and Treasury definitions, the decade of inconsistent archival policy, the master-data governance vacuum. Every one of them deserves fixing. None of them deserves your pilot's budget. Write them down, with evidence, in a handoff note to the data program or whoever should own the crusade, and then refuse, politely and permanently, to let them attach themselves to the pilot. This is not cowardice; it is kindness to everyone. The pilot stays small enough to succeed, the real problems get routed to a body with the mandate to fix them, and you avoid the classic death spiral where a six-week pilot quietly becomes an eighteen-month data-transformation program that delivers neither. A pilot is a probe, not a flagship. Probes travel light.
Where AI Helps, and Where the Judgment Stays Yours
Triage is judgment work, which makes it exactly the kind of work where an AI assistant is both genuinely useful and genuinely dangerous. The division of labor has to be explicit.
AI drafts the first-pass matrix. Give the assistant your pilot's one-sentence core loop, the design notes, and the full gap sheet, and ask: "For each finding, propose an impact rating (fatal, degrading, cosmetic) on this pilot's core loop, with one line of reasoning per rating, and flag any finding where you lack the information to judge." Twenty minutes of prompting turns a raw list into a structured draft, and the drafted reasoning lines are the real value: each one is a checkable claim, and the wrong ones teach you where the AI misunderstands your pilot. Treat every rating as a hypothesis. The assistant's matrix will be neat, confident, and happy to rate gaps it does not understand; neatness is not comprehension. The flag-what-you-cannot-judge instruction matters because it gives the model permission to be uncertain, and the findings it flags are usually the ones where you needed to think hardest anyway.
AI runs the dirty-sample simulation. This is the highest-leverage move in the lesson. For any gap where impact is disputed or guessed, stop arguing and measure: take a sample of the actual dirty data, run the pilot's intended task on it with the assistant (extract these fields from these 150 real invoices, match these vendor names against this real master file), and count the failures. Two hours of work converts "I think the duplicates will hurt matching" into "on a 150-invoice sample, duplicate vendor records caused 12 mismatches, an 8-point accuracy hit." An opinion became a rate. Rates end meetings that opinions prolong, and the rate flows straight into the parallel-fix error budget. This is the same verification instinct the whole level teaches, pointed at your own ratings.
AI drafts remediation options. For each gap heading toward the owner conversation, have the assistant draft two or three fix approaches with rough effort shapes: the quick script, the process change, the system fix, with the trade-offs of each. You are not asking the AI for the estimate; you are asking it to prepare the menu so the owner conversation starts at "which of these, and what would it really take?" instead of a blank page.
What the human owns, without exception: every impact rating is finally a judgment about your pilot, and you sign it. Every fix estimate is finally a commitment by its owner, and they sign it. And the two verification rules stand over everything: no rating ships without its one-line reasoning, and any rating the process owner disputes gets re-scored in the room. The AI accelerates the drafting; it never absorbs the accountability. That is not a limitation of the tools. It is the design.
Worked Example: Eleven Findings, One Page, Twenty-Five Minutes
Back to your invoice-exception pilot, and to the eleven findings the gap hunt produced. All figures here are illustrative, but the shape is the method. The pilot's core loop, in one sentence: read incoming exception invoices, extract the key fields, classify the exception reason, and auto-route each case to the right resolver queue.
Two blockers. First, the free-text exception field: the auto-routing design classifies on the exception reason, and the reason lives in nine years of typed prose. Fatal by the one-line test: the loop's central step has no reliable input. Fix economics: quarters-expensive once the forty-clerk change negotiation is priced in, but the class does not care; fatal is fatal at any cost, so the routing design waits (or gets redesigned around classification of the prose itself, which is a design decision for the pilot team, not a data fix). Second, the evaluation-data retention hole: the systems purge resolved exception records at 90 days, and the pilot's evaluation design needs six months of history to score against. The fix is procedural and cheap (start retaining now), but nothing can retroactively create the history, so the pilot start moves six weeks while the evaluation window fills. Saying "the start moves six weeks" in week one felt terrible for about a day. It is enormously cheaper than discovering, live in month three, that the pilot cannot prove anything, which is how pilots join the 60 percent.
Three parallel fixes. The vendor-master dedupe leads: the dirty-sample simulation on 150 real invoices showed duplicate vendor records degrading extraction-and-match accuracy by an illustrative 8 points. Degrading, not fatal (the loop runs, worse), and the fix is 40 hours of data-steward time on the top 200 vendors by spend, an estimate the steward herself put her name on. Funded on the spot, precisely because the matrix showed a measured 8-point recovery for a 40-hour cost: that is the kind of arithmetic steering committees can actually approve. The two smaller parallel fixes (a date-format normalization script; a missing-PO-number backfill for one supplier segment) got the same treatment: number, owner, date, error-budget line.
Four accept-and-log. Among them: a legacy currency-code inconsistency affecting under 2 percent of volume, and a scanned-image quality problem on one regional feed that would cost a scanner-hardware refresh to fix. Each entered the risk register with its one-line reasoning and a revisit-at-scale date. The pilot proceeds; the log remembers.
Two out-of-scope. The Treasury-versus-Finance semantic drift on payment-term definitions is real, documented, and completely absent from this pilot's loop. It went into a two-page handoff note to the enterprise data program, along with a master-data ownership gap, and both were explicitly listed in the matrix as out of scope so nobody could later claim the pilot "was supposed to" fix them.
The whole matrix fits on one page: eleven rows, two axes, four classes, every row ending in a verb, an owner, and a date. The steering conversation took 25 minutes, because there was nothing to interpret: two rows said "not yet, here is why, here is the new date"; three rows said "fund 40 hours and two scripts, here is the measured payoff"; four rows said "logged, revisit at scale"; two rows said "not ours, here is who." Compare that with the sixty-page report down the hall. The difference was never diligence. Your colleague found more gaps than you did. The difference is that forty-seven true sentences without a ranking is an encyclopedia, and eleven ranked decisions is a plan, and organizations can only execute plans.
What to Do Monday Morning
You have a gap list from the last lesson. This week it becomes a Gap Severity Matrix.
- Write the pilot's core loop in one sentence: inputs, task, output, decision fed. Every impact rating you assign this week is a claim about this sentence, so it goes at the top of the matrix.
- Score impact against that sentence, with one-line reasoning per rating. Have your AI assistant draft the first pass from the gap sheet and the design notes, then challenge every rating; keep no rating whose reasoning line you cannot defend out loud.
- Run the dirty-sample simulation for your worst quality gap. Pull a real sample, run the pilot's intended task on it with the assistant, count the failures, and write the rate into the matrix. Two hours, and your most arguable rating becomes your least arguable one.
- Get one fix estimate signed by its owner. Pick the most consequential likely fix, have the AI draft the remediation options, then sit with the owner until an estimate exists that they would repeat in front of the steering committee, political cost included.
- Rewrite the list as the four action classes. Every finding lands in BLOCKER, PARALLEL FIX, ACCEPT AND LOG, or OUT OF SCOPE, and every row ends in a verb, an owner, and a date. If a blocker emerges, say so this week, in writing. That sentence, sent early, is the entire value of the assessor.
One boundary before you file the matrix, and it is the bridge to the next lesson: one class of finding never enters this triage at all. Privacy, security, and access violations are not fatal-versus-degrading questions; they are not questions. They travel in their own non-negotiable lane, which is exactly where we go next. And keep the matrix close: when you assemble the full readiness report two lessons from now, this one page is its engine room.
Key Takeaways
- Treat the unranked audit as a failure mode, not a virtue: forty-seven documented findings without a ranking reads as "the data people have concerns," funds nothing, and protects no one, including the auditor.
- Build the Gap Severity Matrix: every finding scored on impact to this pilot's core loop (fatal, degrading, cosmetic) and fix economics (days-cheap, weeks-moderate, quarters-expensive), resolving into four action classes.
- Judge impact against the pilot's actual design, never against data perfection: the same vendor-master duplicates are fatal to a matching pilot and cosmetic to a cycle-time pilot, which is why enterprise data-quality scores cannot make pilot decisions.
- Price the politics into fix economics: the two-week form change that forty AP clerks must adopt is a three-month negotiation, and per BCG's 10-20-70, the people cost usually dwarfs the technical cost.
- Call blockers in week one and in writing: a fatal gap at any fix cost means the pilot does not start, and moving the start by six weeks early is vastly cheaper than dying on the same gap in month four.
- Run the dirty-sample simulation before arguing: two hours of executing the pilot's real task on real dirty data turns an impact opinion into a failure rate, and rates end meetings that opinions prolong.
- Enforce the two verification rules: no impact rating ships without its one-line reasoning, and no fix estimate enters the matrix without its owner's signature; the AI drafts, the humans sign.
- Protect scope as a kindness: real problems that do not touch the pilot's loop go to the data program with a handoff note, never onto the pilot's budget, because probes that travel light are the ones that return answers.
Skill.re