←
AI Readiness & Process Transformation
Capable · M12 · lesson 12 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Finding the Data Gaps That Kill Pilots

15 min

Month four of the invoice-automation pilot, and the data scientist has been quiet for most of the status call. When she finally speaks, it is a simple request: "I need the rejected invoices from last year to test accuracy against. Where do we keep those?" The AP manager answers without hesitation, because he knows his systems cold: "We don't. Retention policy purges rejections after 90 days." There is a pause on the line that everyone present will remember, because in that pause a $180,000 pilot just died, and it died of something that was true, knowable, and free to discover on day one. Nobody asked. This lesson exists so that you are the person who asks, in week one, the six questions that pilots otherwise answer the expensive way in month four.

The Expensive Part Is the Lateness

In the previous lesson you built the Data Inventory Register: 23 rows for the invoice-exception process, each row a data element with its system of record, owner, format, volume, and a completeness profile you verified by hand. The register tells you what your data looks like. It does not yet tell you what to look for. Those are different skills, the way knowing your car's mileage is different from knowing that a ticking sound at idle means a specific failing part. This lesson teaches the ticking sounds.

Start with the statistic this whole chapter is built on. Gartner found in February 2025 that 63 percent of organizations either lack AI-ready data practices or are unsure whether they have them, and predicted that through 2026, 60 percent of AI projects without AI-ready data will be abandoned. Read that 60 percent the way an operations professional reads any failure rate: not as a mood, but as a population of specific, individual deaths, each with a cause. When you itemize those causes, something clarifying happens. Pilots do not die of "bad data" in general, the way people do not die of "poor health" in general. They die of specific, recurring species of gap, and each species surfaces at a specific, predictable moment in the pilot's life, usually somewhere between month three and month five, after the license is signed, the consultants are billing, and the steering committee has a slide with a green status dot on it.

That timing is the real cost. The expensive part of Gartner's 60 percent is not the abandonment; abandoning a doomed project is the correct move and a documented kill is a win. The expensive part is the lateness of the abandonment. A gap discovered in week one costs a question. The same gap discovered in month four costs everything spent in between, plus the organizational scar tissue: the next AI proposal at that company gets greeted with folded arms, and somebody's name is attached to the crater. MIT's finding that 95 percent of enterprise GenAI pilots deliver no measurable profit-and-loss return is, from this angle, substantially a chronicle of late discoveries.

Here is the good news, and it is the entire promise of this lesson: the gap species are few, their tells are visible in artifacts you already possess, and each one has a cheap, week-one hunting method. There are six species. Learn them and you can find, with the register and the samples already on your desk, what the pilot would otherwise discover with a burn rate attached. The deliverable you will build as you go is the Data Gap Taxonomy Sheet: one page, six rows, and for each species its tell, its hunting method, and the pilot-killing moment it produces if left unhunted.

Species One Through Three: Missing Fields, Stale Records, Access Walls

We will take the six species slowly, two scenes each: first the month-four failure the species causes, so you can recognize the corpse, then the week-one hunt, so you never have to see one.

Species one: the missing field

The month-four scene. A support organization pilots an AI triage model that routes incoming tickets by predicted resolution path. The design document, written with real care, assumes the model will learn from the "resolution reason" field on historical tickets. In month three the team pulls the training extract and discovers that resolution reason is blank in 70 percent of closed tickets, because agents close tickets from a keyboard shortcut that skips the field, and nobody has ever needed it for a report. The model the pilot was designed around cannot be trained, because its central input was never actually captured. The tool is not broken. The plan was fiction from the first slide.

The species, precisely: a missing field is a gap between what the pilot's design assumes is captured and what the operation actually captures. It is the most common species and the most humiliating, because it is fully visible in the register you already built. The tell is a design document whose inputs were never mapped against a completeness profile.

The week-one hunt: take the pilot's design, or even the back-of-a-napkin sketch of it, and list every input field the intended model or automation needs. Then walk that list against the register's completeness column, row by row. Any required field that is absent, or present but empty above roughly a quarter of the time, is a finding. This is an hour of work with a highlighter. AI helps at the edges: paste the design document and the register into your assistant and ask it to produce the cross-map and flag mismatches, then verify each flagged row yourself against the actual profile. The AI does the clerical cross-referencing in minutes; you certify it, because a hallucinated "field present" is exactly the failure mode you are hunting.

Species two: the stale record

The month-four scene. A procurement pilot matches invoices against the vendor master. The vendor master exists, is well structured, and is 22 years old. It contains, as your register from last lesson documented, 312 duplicate vendor entries, plus price lists that were superseded in 2023 but never retired, plus payment terms for vendors who no longer trade. The model trains happily on all of it, because data does not carry an expiry sticker. In month four the pilot's match rate is inexplicably poor, and three weeks of debugging later someone realizes the model has learned from ghosts: it is matching live invoices against dead vendors and retired prices with total confidence.

The species: stale records are data that exists and once was true. Nothing flags them; they sit in the same tables, in the same format, as the living records around them. The tell is a table with no retirement process.

The week-one hunt has two moves. First, profile the update timestamps: ask your AI assistant to chart last-modified dates for each register table and look for the shape of neglect, a long tail of rows untouched for years in a table the pilot treats as current. Second, and more revealing, ask the table's owner one question: "What is this table's retirement process? Who removes a record when it stops being true, and when did that last happen?" Listen carefully, because the silence is the answer. A table nobody curates is a table where truth and history have been allowed to mix, and a model cannot tell them apart.

Species three: the access wall

The month-four scene. The integration plan says "pull the contract terms via the vendor portal's API," an API being an application programming interface, the standard plumbing through which one system requests data from another. In month four the team discovers the vendor's SaaS product (software as a service, an application you rent and whose data lives on the vendor's servers, not yours) offers no bulk export on the current license tier, and that the licensing clause prices API access as a separate module with a five-figure annual fee and a legal review. The line item that the plan costed as "just an API call" has become a procurement negotiation with its own timeline, and the pilot idles while it runs.

The species: an access wall is data that exists, may even be clean, and cannot legally or technically be touched by anyone on the pilot. Its forms are owned-by-another-department, locked-in-a-vendor's-SaaS, and gated-by-a-contract-clause, and all three share a nasty property: they are invisible on a data inventory, because the inventory records that the data exists, not whether you can have it.

The week-one hunt is the most elegant of the six, because it produces its own evidence: for every register row the pilot needs, actually request the access the pilot would need, this week, and time the answer. Not "confirm access is possible in principle." Send the real request: the ticket to the other department's data owner, the export question to the vendor's support desk, the clause question to your contracts team. The request is the test. A three-day turnaround is a finding. A three-week approval path is a finding. A reply that begins "that would need to go through legal" is a large finding. AI helps here by drafting the full access-request list from the register in one pass, with a tracking table for send date and response date; you send them, because the timestamps are the data.

Species Four Through Six: Silent Quality Rot, History Holes, Semantic Drift

Species four: silent quality rot

The month-four scene. The exception-handling pilot is supposed to learn why invoices get flagged. The "exception reason" field is populated in 96 percent of records, which looked splendid on the inventory. Then someone finally reads the values. The top reason, by a wide margin, is "other." It turns out the reason picker defaults to "other," the field is mandatory, and clerks under volume pressure click past it, as anyone would. The second discovery is worse: the free-text comment box next to it contains the real reasons, in eleven spelling variants of "PO mismatch" and a private shorthand each clerk invented around 2021. The field is full. The field is also, for a model's purposes, mostly noise.

The species: silent quality rot is a field that is populated and wrong, and it is the inverse trap of the missing field. Completeness metrics cannot see it; only distributions can. Its classic forms are the default value everyone clicks past, free text where categories should be, and copy-paste drift, where a value gets duplicated forward from record to record until it detaches from reality.

The week-one hunt: distribution profiling, which is exactly the kind of tedious counting AI assistants do well. For every field the pilot depends on, ask for a value-frequency table. Then apply two suspicion rules. Rule one: any field where a single value dominates suspiciously (an "other" or a default sitting at the top of the chart) is presumed rotten until a human explains why the dominance is real. Rule two: any free-text field that the design treats as categorical is a finding by definition; ask the AI to cluster a sample of 200 values and show you how many real categories hide inside, then spot-check the clusters yourself against raw records.

Species five: the history hole

The month-four scene is the one this lesson opened with, so here is the general form. Every pilot that predicts, classifies, or extracts something needs historical examples of that something, both to learn from and, just as critically, to be evaluated against. Organizations, meanwhile, delete history constantly and virtuously: retention policies purge rejected invoices after 90 days, document systems overwrite the "before" version when a correction is saved, archives keep the final state and discard the journey. The pilot arrives in month four ready to measure its accuracy and finds there is nothing to measure against, because the org never kept the thing being predicted.

And here is the short, true-to-pattern failure story to keep beside it. A mid-size retailer spends five months and roughly $200,000 (illustrative figures, but the pattern is documented across the failure literature) on a demand-forecasting pilot. In month five the modelers ask for three years of promotions history to explain the demand spikes the model keeps missing. Promotions data was never retained beyond the campaign cycle. The pilot dies. In the postmortem, an email surfaces from the data team, sent before the kickoff: "Be aware our retention on campaign data is 12 months." Nobody had connected that sentence to the pilot's design. The gap was knowable in week one for the cost of one question, and the postmortem's real finding is that no one owned the asking.

The species: a history hole is an absence of the examples the pilot's training and evaluation require, usually created deliberately by a retention policy doing its job. The tell is any pilot whose target variable is a thing the operation treats as disposable once handled: rejections, errors, drafts, before-states, complaints resolved.

The week-one hunt is a single question asked of every element the pilot's evaluation needs: "Could we reconstruct last March?" Pick a real month comfortably in the past and ask whether, today, you could produce the complete set of examples from it: the rejected invoices of last March, the pre-correction versions, the promotions that ran. If the answer is no, you have found the hole, and you have found it while there is still time for the only cure that exists: start retaining now, so that evaluation data accumulates while the pilot is being built instead of being wished for after.

Species six: semantic drift

The month-four scene. The pilot's model has been trained on "closed" cases and its aggregate accuracy looks acceptable, yet users in one department insist it is wrong constantly, while another department finds it fine. Weeks of confusion later, the truth emerges: for accounts payable, "resolved" means the invoice was paid; for treasury, "resolved" means the dispute was answered, payment or not. Same field, same values, two meanings. The model learned a word that means two things and is wrong half the time in a way no aggregate metric shows, because the errors cancel out in the average and concentrate in one team's queue. There is a time-based variant too: the 2022 system migration silently changed what "priority" meant, so records before and after the cutover date disagree about the world while sharing a schema.

The species: semantic drift is the same field meaning different things across teams, systems, or eras. It is the hardest species to see in data alone, because the data is complete, populated, and internally consistent within each regime. The tell is a field whose definition nobody can state in one sentence that two departments would both sign.

The week-one hunt is two-pronged. Prong one is human: cross-team definition interviews. Take the pilot's five most load-bearing fields, ask one person in each involved team "what exactly does a value of X in this field mean, and when do you set it?", and write the answers side by side. Divergence is the finding. Prong two is statistical, and AI-friendly: profile the field's value distribution over time and look for step changes around known migration or reorganization dates; a distribution that jumps at the month of the 2022 cutover is a dated confession. AI also shines at the document side of this hunt: feed it both departments' SOPs (standard operating procedures, the written instructions for how work is done) and ask it to cross-reference definitions and list every term the two documents use differently, then confirm the discrepancies with the humans who live in them.

The Artifact: Your Data Gap Taxonomy Sheet

Now assemble the lesson's named artifact. The Data Gap Taxonomy Sheet is one page. Six rows, one per species. Four columns: the species, its tell, the week-one hunting method, and the month-four moment it produces if unhunted. Build it once as a reference, then use a working copy per pilot, where each row gains a fifth entry: what the hunt found, with evidence attached.

SpeciesTellWeek-one huntMonth-four moment if unhunted
Missing fieldDesign assumes an input the register shows as absent or largely blankMap every pilot-required field against the register's completeness profileThe training extract arrives and the central input is not in it
Stale recordTable with no retirement process; long tail of untouched rowsProfile update timestamps; ask the owner "what is the retirement process?" and log the silenceThe model learns from ghosts and matches the living against the dead
Access wallData owned elsewhere, locked in SaaS, or gated by contractSend the real access request now and time the answer; the request is the test"Just an API call" becomes a procurement negotiation
Silent quality rotOne value dominates suspiciously; free text where categories belongDistribution-profile every dependent field; cluster free-text samplesThe model trains on "other" and on eleven spellings of the truth
History holeTarget variable is something the operation discards once handledAsk "could we reconstruct last March?" for every evaluation elementAccuracy day arrives and there is nothing to measure against
Semantic driftNo single-sentence definition two teams would both sign; migrations in the field's pastCross-team definition interviews plus distribution profiling around migration datesThe model is wrong half the time in a way no aggregate metric shows

Two disciplines govern how the sheet is used, and they are what separate an audit from a vibes tour. First, the verification rule, which is this chapter's spine restated: every gap claim gets one concrete, evidenced example attached. A record ID with the blank field. A screenshot of the "other" distribution. The timestamped access request and its timestamped answer. The two interview quotes that define "resolved" differently. No exhibit, no finding; an unevidenced gap claim is an anecdote, and anecdotes lose arguments with people whose budgets you are threatening.

Every gap claim carries one evidenced exhibit: a record, a screenshot, a timed request. Without the exhibit it is an anecdote, not a finding.

Second, the division of labor with your AI assistant, consistent across all six hunts: AI does the profiling, the cross-referencing of definitions across documents, the clustering of free text, the drafting of the access-request list, in minutes instead of days. You do the sending, the interviewing, and the certifying, because every one of those AI outputs is a draft of reality, not reality, and the whole point of this audit is that unverified assumptions are what kill pilots. It would be a special kind of irony to conduct it on unverified assumptions.

The Worked Example: Three Days Against the 23-Row Register

Here is the full method run end to end, with illustrative numbers, on the storyline this chapter has been following: the invoice-exception process, its 23-row Data Inventory Register, and the proposed pilot, an AI system to auto-triage invoice exceptions, budgeted at $180,000 for the year. All figures hypothetical; the pattern is the lesson.

You block three days. Day one: the missing-field map and the distribution profiles, mostly AI-assisted, human-certified. Day two: timestamp profiling, the retirement-process interviews, and sending every access request. Day three: the reconstruction question, the definition interviews, and writing up exhibits. Here is what the six hunts return.

  • Missing fields: two, one fatal to the design. The design assumes a structured exception-reason code on every exception. The hunt shows the real reason lives in free text for 61 percent of exceptions; the structured code exists for only three exception types. Second, the "resolution owner" field the routing logic needs is blank in 38 percent of records. Exhibits: the cross-map table and ten record IDs.
  • Stale records: the vendor master. Timestamp profiling shows 41 percent of vendor rows untouched in four or more years; the owner's answer to the retirement question is "there isn't really one," faithfully quoted. The 312 known duplicates now have a mechanism, not just a count. Exhibit: the timestamp histogram and the quote.
  • Access wall: one, and it is timed. The bank-detail fields the matching step wants are legally restricted; the real access request, sent Tuesday, returns a documented three-week approval path involving the data protection officer. Exhibit: the request and reply, both timestamped.
  • Silent quality rot: two fields. "Exception reason" is 44 percent "other" (the mandatory default), and "hold code" turns out to be copy-pasted forward on batch entries. Exhibits: two distribution charts, spot-checked against 20 raw records.
  • History hole: exactly where evaluation data should be. Rejected invoices purge at 90 days; last March cannot be reconstructed. The pilot's accuracy could never have been proven. Exhibit: the retention policy clause, page and paragraph.
  • Semantic drift: one, load-bearing. "Resolved" means paid in AP and answered in treasury; two interview quotes, side by side, twelve words apart in meaning. Exhibit: the quotes, with names and dates.

Seven findings, each with an exhibit, in three days, against a register you already had. Now the punchline, and read it carefully because it is the political heart of this chapter: the pilot as originally imagined is not buildable. Auto-triage of all exceptions depends on a reason field that is 61 percent free text and 44 percent "other," evaluated against history that gets deleted quarterly. Presented alone, that sounds like the audit killed the pilot. It did not. On the same page, the audit shows what is buildable: a redesigned scope that auto-categorizes the three structured exception types (roughly 39 percent of volume, illustratively), routes free-text exceptions to humans untouched, and starts retaining rejected invoices today, so that twelve months from now an evaluation set exists and phase two becomes possible. The access wall gets its three-week clock started this week instead of in month four. The audit did not kill the pilot. It killed the doomed version of the pilot and saved the buildable one, and it converted a future $180,000 crater into a scoped project with evidence under every claim. That reframe, gap findings as redesign inputs rather than verdicts, is exactly where the next lesson picks up: triaging which gaps to fix, which to route around, and which to accept.

What to Do Monday Morning

You have a register from the last lesson and a taxonomy from this one. Combine them.

  1. Build your Data Gap Taxonomy Sheet: six rows, four columns, from the table above, in whatever tool your team will actually keep open. Twenty minutes.
  2. Run hunt one and hunt four with AI assistance: cross-map your pilot's required fields against the register's completeness profiles, and pull value-frequency distributions for every dependent field. Certify every flagged row against raw records yourself.
  3. Send the access requests today and timestamp them. One real request per register row the pilot needs, logged with send date. The clock you start Monday is the clock that would otherwise start in month four.
  4. Ask the two oral questions: "what is this table's retirement process?" to each table owner, and "could we reconstruct last March?" for each element your evaluation would need. Write the answers down verbatim; silences count as answers.
  5. Attach one evidence exhibit to every gap you find: a record ID, a screenshot, a timed request, a quote. Delete any finding you cannot exhibit; it was an anecdote.
  6. Write the one-sentence verdict: "As imagined, this pilot is / is not buildable; the buildable version is X." That sentence, with the exhibits behind it, is what you bring to the next steering conversation, and it is what the next lesson teaches you to defend.

Key Takeaways

  • Reframe Gartner's numbers as an itemized list: 63 percent of organizations lack AI-ready data practices, 60 percent of AI projects without AI-ready data will be abandoned through 2026, and the expensive part is not the abandonment but the lateness of it.
  • Hunt the six gap species by name: missing fields, stale records, access walls, silent quality rot, history holes, and semantic drift; each has a tell, a week-one hunting method, and a predictable month-four moment.
  • Map every pilot-required field against the register's completeness profile before believing any design document; a field that is absent or largely blank makes the plan fiction, not the tool faulty.
  • Test access by requesting it: send the real access request in week one and time the answer, because the request is the test and a three-week approval path found now is a finding, not a crisis.
  • Distrust completeness alone: a populated field can be rotten, so profile distributions, presume any suspiciously dominant value guilty, and treat free text posing as categories as a finding by definition.
  • Ask "could we reconstruct last March?" for every element the evaluation needs, and where the answer is no, start retaining now so evaluation data accumulates while the pilot is built.
  • Attach one concrete evidenced exhibit to every gap claim, a record, a screenshot, a timed request, or a quote; without it the claim is an anecdote and will lose the budget argument.
  • Present gap findings as redesign inputs, not verdicts: the audit's job is to kill the doomed version of the pilot early and cheaply, and to hand back the version that is actually buildable.