AI-Assisted Data Inventory and Profiling
The pilot proposal is one signature away from funding. The invoice-exception process has been inventoried, mapped, verified, baselined; the business case says $428,000 a year is on the table. Then the IT director, who has been quiet for most of the meeting, asks a question that sounds simple: "Is the data ready?" And the room does what rooms always do with that question. Someone says the data is "mostly fine." Someone else says the data warehouse project will fix everything in eighteen months. A third person repeats the sentence that has haunted every AI conversation since Gartner put numbers on it: "our data isn't ready." Here is the problem with that sentence, and it is the problem this whole chapter exists to fix: it is not false, it is meaningless. Ready for what? Which data? Nobody in the room can point to a list of the actual data elements the invoice-exception process touches, let alone say what shape any of them are in. The pilot is about to be bet on data nobody has ever looked at. This lesson teaches you the first move of a real data audit: not a quality crusade across the enterprise, but a scoped inventory and an honest profile of the specific data one specific process actually eats and produces.
The Sentence That Means Nothing
Start with the numbers hanging over the meeting, because they are the spine of this entire chapter. In February 2025, Gartner reported that 63 percent of organizations either lack AI-ready data practices or are unsure whether they have them, and predicted that through 2026, 60 percent of AI projects that proceed without AI-ready data will be abandoned. Read those together and the arithmetic is brutal: most organizations do not know where they stand, and the majority of projects launched from that ignorance will die. This is not a footnote to the failure record you studied in Level 1. MIT's finding that 95 percent of enterprise generative AI pilots deliver no measurable return has data surprises running through its autopsy like a recurring cause of death: pilots that demoed beautifully on curated samples and collapsed when production fed them the real thing.
And yet "our data isn't ready" is a useless sentence, for the same reason "our processes are inefficient" is useless. It has no object. Readiness is not a property of an organization's data in general; it is a property of specific data feeding a specific process. The vendor master might be a swamp while the purchase-order table is pristine. The invoice header fields might be clean while the exception-reason field is a landfill of free text. A company can be 63-percent-statistic unready in aggregate and still have one process whose data would support a pilot tomorrow, and the reverse is just as common: a company proud of its data warehouse discovers that the one field its pilot depends on has never been governed by anyone.
This reframing rescues you from the trap that swallows most data-readiness conversations: the enterprise data quality program. When "our data isn't ready" gets treated as a company-wide diagnosis, the prescribed cure is company-wide too: a data catalog initiative, a governance council, a master data management platform, a multi-quarter roadmap with a steering committee of its own. Those programs have their place, but they take quarters to years, and your pilot decision is needed in weeks. Worse, they answer the wrong question. The pilot does not need all the company's data to be ready. It needs the 20 or 30 data elements the target process actually touches to be understood: what they are, where they live, who owns them, and what shape they are really in.
That is a bounded question, and bounded questions have fast answers. An enterprise data-cataloging effort is measured in quarters. A process-scoped data inventory, done with AI assistance using the skills you built in the last two chapters, is measured in days. This lesson's method takes the invoice-exception storyline four working days in the worked example below. The output is the first artifact of your data audit, and everything else in this chapter, the gap taxonomy, the severity triage, the Data Readiness Report, and ultimately your Level 2 capstone, gets built on top of it.
"Our data isn't ready" is a diagnosis with no patient. Readiness belongs to specific data feeding a specific process, and the first move of a data audit is a scoped inventory, not an enterprise crusade.
The Artifact: The Process Data Register
The named artifact of this lesson is the Process Data Register: one row for every data element the target process touches, with enough columns to answer the four questions a pilot team will eventually ask about each one. What is it? Where does it live? Who owns it? What shape is it in? A data element, for our purposes, is any distinct piece of information the process reads, writes, or consults: a field, a document, a file, a report, a system record. The invoice number is an element. The vendor master table is an element. The emailed PDF of the supplier's statement is an element, even though no system admits it exists.
Here are the register's columns, and why each one earns its place.
| Column | What it records | Why it is load-bearing |
|---|---|---|
| Element name | What the element is called by the people who use it, plus the official system name if different | Half the confusion in data work is two teams using different names for one thing, or one name for two things |
| Direction | Input (the process reads it), output (the process writes it), or reference (the process consults it without changing it) | Inputs are what a pilot must be able to digest; outputs are what it must produce; reference data is where silent staleness hides |
| Source system | The system, file, or location where the element actually lives, including "Maria's spreadsheet" if that is the truth | The pilot integrates with real locations, not official ones; a wrong answer here surfaces as an integration surprise in month three |
| Format | Structured field, semi-structured export, free text, PDF, scanned image, email body | Format determines what kind of AI work is feasible and how much cleanup stands between here and a pilot |
| Owner | The named person or role accountable for the element's accuracy | "Nobody" is a common and important answer; an unowned element has no one to fix it when profiling finds problems |
| Update frequency | How often the element changes: real-time, daily, monthly, "when someone remembers" | A pilot fed monthly reference data makes decisions on stale facts for up to 30 days at a time |
| Access path | How you would actually get this data out: an export, an API, a report, a request to a named team | Data you cannot extract is data your pilot cannot use, whatever its quality |
| Profile verdict | Clean, messy, or unknown, based on evidence from profiling an actual sample | This is the column the rest of the chapter runs on, and "unknown" is an honest, legal verdict |
| Evidence citation | Where the knowledge in this row came from: SOP section, transcript timestamp, interview, profiled sample file | The register inherits the program's evidence rule; a row nobody can trace is a rumor in a table |
Two design choices deserve a pause. First, the profile verdict has exactly three values, and unknown is a verdict, not a blank. The register's job is to make ignorance visible and priced, not to hide it behind optimistic defaults. A register that shows 9 clean, 6 messy, and 8 unknown is telling the pilot team something precise: here is the map, and here are the parts of it that are still fog. Fog you can see is manageable. Fog labeled "probably fine" is how pilots join Gartner's 60 percent.
Second, every row cites where its knowledge came from, exactly as your process inventory and your baseline pack did. This is the same discipline for the same reason: the register will be challenged, by IT, by the data owner, by the vendor doing the pilot integration, and "the SOP says so, section 4.2" or "confirmed by the AP specialist in the March 12 interview" ends arguments that "we believe" cannot.
Movement One: The Inventory, or Finding Everything the Process Touches
Building the register has two movements: inventory, then profiling. The inventory answers "what data exists and where," and it has three passes, each one catching what the previous pass missed.
Pass one: walk the verified process map
You are not starting from nothing, and this is where the last chapter pays its dividend. Your verified process map, the one you built, de-hallucinated, and walked through with the people who live in the process, already encodes most of the inventory. Every step on that map consumes something and produces something, and every one of those somethings is a candidate register row. Walk the map step by step and interrogate each one: what does this step read before it can act? What does it write when it is done? What does it look up along the way? The approval step reads the exception ticket and the purchase order, consults the delegation-of-authority matrix, and writes an approval record. That single step just gave you four register rows. Do this for every step and every decision diamond, and the register's skeleton assembles itself in an afternoon.
Pass two: mine the documents with extraction prompts
The second pass turns AI loose on the material you already collected: the SOP (standard operating procedure, the official written instructions), the ticket exports, the interview transcripts from your discovery work. This is exactly the extraction pattern you learned in the prompting lessons, aimed at a new target. The prompt shape that works:
"Read the attached material. List every document, field, file, report, or system record that this material mentions the process reading, writing, or consulting. For each one, give the name used in the text, whether the process reads it or writes it, and quote the exact passage where it is mentioned. Do not infer data elements that are not explicitly mentioned; if a passage is ambiguous, list it under 'unclear' with the quote."
The citation requirement is not decoration; it is the leash. You learned in the verification chapter that AI extraction invents plausible entries when allowed to roam, and a fabricated data element in your register is exactly as poisonous as an invented SOP step. Requiring a quoted passage for every element means every claim arrives pre-attached to its evidence, and your verification pass is a matter of checking quotes rather than re-reading everything. Run the prompt separately over the SOP, the transcripts, and a sample of real tickets, then ask the AI to merge and deduplicate the lists, flagging elements that appear under different names in different sources. Expect the transcript run to surface elements the SOP never mentions. That gap is not noise. It is your first sighting of the third pass's quarry.
Pass three: hunt the shadow data layer
Every real process runs partly on data that no official system holds: the reconciliation spreadsheet on a specialist's desktop, the vendor contact list someone maintains in a personal file, the emailed PDF that is the actual source of truth for supplier bank details, the shared-drive folder where "the real notes" live. Call it the shadow data layer, and treat it with respect rather than disapproval, because it exists for a reason: the official systems failed to hold something the work needed, and a human quietly built the missing piece. The shadow layer is where processes actually run, and it is very often where pilots actually die, because no integration plan accounts for data that officially does not exist.
You will not find the shadow layer in any document, because by definition nobody wrote it down. You find it by asking the performers, in the same interviews you learned to run in the discovery chapter, questions engineered to give permission: "When the system doesn't have what you need, where do you look?" "Is there a spreadsheet or a file you'd be lost without?" "Where do you keep the notes that help you decide the tricky ones?" Ask with genuine curiosity and zero audit energy, because the fastest way to make shadow data vanish is to sound like you are hunting policy violations. Every answer is a register row, with the source system column telling the truth: "XLSX file, S. Okafor's desktop, no backup, updated ad hoc." A register that contains rows like that is not embarrassing. It is accurate, and accurate is the entire point.
Movement Two: The Profiling, or Finding Out What Shape It Is Really In
The inventory tells you what exists. Profiling tells you what condition it is in, and it is the difference between a register and a wish list. The method: for each register row that matters to the pilot, get a sample, have AI profile it, then verify the profile before a single figure leaves your desk.
Get a real sample, not a demo sample
Ask the system owner for a genuine export of 50 to 200 records: real rows, pulled the way a pilot integration would pull them, not the polished extract someone prepared for a vendor demo. Before any sample moves anywhere near an AI tool, redact it according to your organization's data policy: strip or mask PII (personally identifiable information, anything that identifies a specific person), bank details, and anything contractually confidential. Getting redaction and tool policy right is a discipline of its own, and it gets a full lesson two lessons from now; for today, the operating rule is simple: ask what the policy is before you paste, and when in doubt, mask the column. A profile of masked tax IDs still tells you 8 percent of them are missing, which is the fact you actually need.
One sampling decision matters more than all the others: sample from the full range the pilot will ingest, not just the recent past. If the pilot will process the whole vendor master, including records created in 2019 under an old numbering scheme and two acquisitions ago, then your sample must reach back into 2019. This is selection bias made plain: the clean recent export hides exactly the records that will break the pilot in production. Ask the owner explicitly: "Does this export cover the full date range and all record types the process touches, or is it filtered?" The answer changes what your profile is worth.
The profiling prompt pattern
With a redacted sample in hand, AI becomes the most patient analyst you have ever employed. Where a human eye glazes over by row 40 of a messy CSV (comma-separated values, the plain-text spreadsheet format most systems export), AI reads all 200 rows with the same attention it gave the first, and it is genuinely superb at exactly the pattern-spotting this task needs: it notices the three date formats cohabiting in one column, the vendor-name field whose conventions drift across the years ("ACME CORP," "Acme Corporation," "acme corp."), the numeric field that occasionally contains the word "pending." The prompt pattern:
"You are profiling the attached data sample. For each column, report: (1) the percentage of blank or null values; (2) every distinct format you observe, with a count for each; (3) values that fall outside the expected range or type for that column; (4) likely duplicate records, with the row numbers of each suspected pair; and (5) the five oddest records in the sample, by row number, with one line on what makes each odd. Report counts and row numbers for everything; do not summarize impressionistically."
The insistence on counts and row numbers serves two masters. Counts make the profile decision-grade: "8 percent missing" supports a severity call in the next lesson in a way "some missing" never will. Row numbers make the profile checkable, which brings us to the move that is not optional.
The non-negotiable: spot-check the profile by hand
AI can miscount. It can report 12 blanks where there are 9, garble a percentage, or hallucinate a tidy figure where the real answer is messier. So every profiling report is treated as a draft, and the verification move is mechanical: pick 2 or 3 of the reported figures and recompute them yourself on the raw sample. Count the blank tax-ID cells in the actual file. Filter the date column and count the formats with your own eyes. Pull the five "oddest records" by their row numbers and confirm they exist and are odd. If your recomputation matches, the profile has earned provisional trust. If it does not, the whole report goes back for rework, exactly like the tiered sampling rule you learned in the Verification Habit lesson. And that lesson's absolute clause still governs everything downstream: 100 percent of figures that graduate from the profile into the register, the Data Readiness Report, or any deliverable a decision will rest on get verified, no exceptions. The spot-check earns the draft a hearing; the 100-percent rule decides what gets published.
The two failure modes to guard against
Two blind spots deserve explicit guardrails, because AI will walk into both with total confidence. First, AI cannot see the data you did not sample. Its profile describes your export, not your database. If the export was filtered to recent, clean records, the profile will be glowing and worthless, and no prompt can fix that; only your sampling discipline can. Second, AI will describe a column by its name rather than its contents if you let it. Ask casually about a field called "status" and it may explain what status fields typically contain instead of reporting what this one actually holds, which might be free-text sentences typed by whoever closed the ticket. This is why the profiling prompt demands observed formats with counts: "field is named exception_reason; observed contents: free text in 61 percent of rows, one of six category codes in 39 percent" is a report about reality. A description of what the field is presumably for is a report about the schema's good intentions, and pilots do not run on intentions.
The Worked Example: Four Days on the Invoice-Exception Data
Time to run the whole method on the storyline this level has been building. The mid-size company from the previous chapters, invoice-exception process mapped, verified, and baselined at roughly $428,000 a year, now builds its Process Data Register before anyone signs the pilot. All numbers that follow are hypothetical, an illustration of the method, not research findings.
Days one and two: the inventory. The analyst walks the verified map and drafts 19 candidate rows. The extraction prompt over the SOP and interview transcripts adds elements the map walk missed, including a monthly vendor-performance report nobody had mentioned, and the transcript run surfaces two phrases that smell like shadow data. Two short interviews confirm it: an AP (accounts payable) specialist maintains a personal spreadsheet mapping cryptic supplier codes to real names, and supplier bank-detail changes arrive as emailed PDFs that live in a shared mailbox. After merging and deduplicating, the register lands at 23 data elements: 11 inputs, 5 outputs, and 7 reference elements, every row citing its source. Four rows have "nobody" in the owner column. Two have a person's desktop as their source system. The register is already telling the truth in ways the phrase "our data is in the ERP" never did. The ERP, for the record, is the enterprise resource planning system, the central software where finance transactions officially live, and 14 of the 23 elements do live there. Nine do not.
Days three and four: the profiling. The team requests redacted samples for the six elements the pilot would lean on hardest. The vendor master sample, 200 records spanning the full date range back to 2018, produces the profile that changes the meeting. AI reports, and hand-recomputation of three figures confirms: 8 percent of vendor records are missing tax IDs; the date-added column contains four distinct date formats; roughly 312 of the 5,400 vendors in the full table are probable duplicates, extrapolating the sample's duplicate rate, with the AI listing suspected pairs by row number ("ACME CORP" and "Acme Corporation" sharing an address). And then the finding that will echo through the rest of this chapter: the exception-reason field is free text in 61 percent of records. The pilot's entire premise is routing exceptions by reason. You cannot route on 61 percent free text; the categories the pilot needs do not exist yet. That one row, profiled in an afternoon, just rewrote the pilot's scope, its timeline, and its budget, before a dollar was spent.
The clock. Inventory plus profiling: four working days with AI assistance. The analyst's honest estimate of the same work done manually, reading every document, eyeballing every export line by line: six to eight weeks, which in practice means it would simply never have been done, and the pilot would have discovered the exception-reason field the expensive way. Final register: 23 rows, verdicts marked 9 clean, 8 messy, 6 unknown, every figure hand-verified, every row cited. That document is the foundation the next lesson builds on.
The Failure Story: "It's in the ERP, So It Must Be Fine"
Now the counterfactual, assembled from the pattern that data-readiness autopsies keep finding. A different company, same kind of pilot: an extraction model to pull vendor and reason data from invoice exceptions. The team is under schedule pressure, and when someone proposes a data profiling pass, the project lead waves it off with the sentence that should be printed on warning labels: "The data is in the ERP, so it must be fine." The ERP is a serious system; serious systems have clean data; the syllogism feels airtight in a status meeting. Note what it actually claims: that a storage location guarantees content quality, which is like assuming a filing cabinet's contents are accurate because the cabinet is fireproof.
The pilot proceeds. The vendor configures the extraction model against a demo dataset: a few hundred recent, well-formed records, the clean subset that always gets pulled when nobody insists on the full range. Accuracy in the demo: excellent. Champagne-adjacent emails are sent. Then production integration connects the model to the real vendor master, all of it, including the 2019 records with the old numbering scheme, the duplicate vendors from two acquisitions, the free-text fields nobody had ever read at scale. Accuracy collapses. Exceptions route to wrong queues; specialists lose trust within two weeks and revert to manual handling while the pilot runs beside them as decoration, generating the usage statistics that make dashboards green and value zero.
Month four, the autopsy. A consultant profiles the vendor master, the two-day exercise this lesson just taught, and finds the duplicates, the format drift, and the free-text reason field: every one of them present, findable, and fixable back in week one, for the cost of one sample export and an afternoon of prompting. The pilot is quietly shut down and joins Gartner's 60 percent. And here is the detail worth sitting with: in the internal retelling, the ERP keeps its reputation, because "the ERP" was never blamed; the story becomes "the AI didn't work here." The tool takes the fall for a data condition nobody looked at. That is the quiet mechanism behind a large share of the failure statistics this program keeps citing: not AI that failed, but data nobody profiled, wearing an AI costume at the funeral.
The gap between the two stories is not budget, talent, or vendor quality. It is four working days and one register. And the register's value compounds from here: the next lesson takes your messy and unknown verdicts and sorts them into a gap taxonomy (missing data, dirty data, trapped data, and their cousins), the lesson after that triages severity so you know which gaps kill and which merely annoy, and by the end of this level, the register is a load-bearing exhibit in your Data Readiness Report and your L2 capstone. Everything starts with the rows you write this week.
What to Do Monday Morning
This lesson becomes real the first time you profile a sample you were assured was fine. Here is the sequence.
- Pick your baselined process. The one you inventoried, mapped, and baselined in the previous chapters is the right target: the register slots straight into the evidence trail you have already built.
- Walk the verified map and draft the register. For every step, list what it reads, writes, and consults. Set up the nine columns and fill what you know, citing the map, SOP section, or transcript for each row.
- Run the extraction prompt over your SOP and interview transcripts, quotes required, then merge its findings into the register. Ask one performer the shadow-data questions and add what surfaces, with honest source-system entries.
- Request one sample export of about 100 records for the element the pilot would depend on most, full date range, redacted per your organization's data policy before it goes anywhere near an AI tool.
- Run the profiling prompt: percent blank per column, distinct formats with counts, out-of-range values, duplicate suspects, and the five oddest records with row numbers.
- Hand-verify three of its claims. Recompute a blank percentage, recount a format split, and pull the odd records by row number. If any check fails, send the profile back and rerun before trusting anything in it.
- Mark every row's profile verdict honestly. Clean, messy, or unknown, based on evidence you can cite. Unknown is a verdict; rows you have not profiled stay unknown, visibly, until you have. File the register next to your baseline pack: the next lesson picks it up from there.
Key Takeaways
- Treat "our data isn't ready" as a diagnosis with no patient: readiness is a property of specific data feeding a specific process, and Gartner's numbers (63 percent of organizations lack or are unsure of AI-ready data practices; 60 percent of AI projects without AI-ready data will be abandoned through 2026) are the stakes, not the analysis.
- Scope the audit to the process, not the enterprise: a company-wide data program takes quarters, while an AI-assisted, process-scoped inventory and profile takes days and answers the only question the pilot decision needs.
- Build the Process Data Register: one row per data element with name, direction, source system, format, owner, update frequency, access path, profile verdict, and an evidence citation for every row.
- Run the inventory in three passes: walk the verified process map for each step's inputs and outputs, mine the SOP, tickets, and transcripts with quote-demanding extraction prompts, and interview performers to surface the shadow data layer where the process actually runs.
- Profile real samples, not demo samples: 50 to 200 records, redacted per policy, covering the full date range the pilot will ingest, because AI cannot see the data you did not sample and the clean recent export hides the records that break production.
- Demand counts and row numbers from the profiling prompt (percent blank, formats with counts, out-of-range values, duplicates, five oddest records), then spot-check 2 or 3 figures by hand on the raw sample, and apply the 100-percent verification rule to every figure that enters a deliverable.
- Guard against the column-name trap: AI will confidently describe a field by its label rather than its contents, so require observed formats with counts, which is how a "status" field gets exposed as 61 percent free text.
- Mark unknown as a verdict, not a blank: visible fog gets priced and managed, while "probably fine, it's in the ERP" is the sentence that sends pilots into Gartner's 60 percent with the AI taking the blame.
Skill.re