Auditing the Data Foundation at Org Scale
The enterprise scorecard you built in the last lesson is on the wall, and the data dimension is wearing the worst number in the room: 2.4 out of 5, averaged across six functions of a 2,400-person company. Under that average sit two facts you already know and one you are about to learn. You know that the median wait for a data access grant is 14 days, because you pulled the ticket logs. You know that the vendor master file has had no owner since a 2022 system migration, because you went looking for the person and found a chair. What you are about to learn is what those facts cost, dataset by dataset, use case by use case, in a document precise enough that a chief financial officer can fund the repairs from it. The process-level data audit you ran at Level 2 asked whether one pilot's data could carry one pilot. The enterprise data audit asks a colder and more valuable question: which of our data assets could carry any of the next twenty use cases, and what would it cost to make the rest of them carriable?
The Question Changes at Enterprise Scale
Recall what the Level 2 audit was for. You had one candidate process, one pilot on the table, and a simple binary to defend: go or no-go. You built a Process Data Register listing the handful of data sources that one process touched, you profiled samples, you classified problems with the gap taxonomy, you triaged severity by repair economics, and you wrote a one-page report that either cleared the pilot or stopped it before it burned six figures. The unit of analysis was the process. The deliverable was a verdict.
At enterprise scale, both of those change, and the audit that does not notice the change produces an expensive pile of nothing. The unit of analysis shifts from the process to the dataset, because at portfolio scale the same datasets keep appearing under different processes. The customer master sits under the churn model, the service copilot, the collections prioritizer, and the marketing segmentation engine. Audit it once, properly, and four future audits are prepaid. Audit process by process and you will profile the same file four times, reach four slightly different conclusions, and fix it zero times.
The deliverable shifts too. A go/no-go verdict is the right output when there is one decision to make. An enterprise has twenty decisions queued, so the audit's output is not a verdict but a remediation backlog: a priced, sequenced, owner-assigned list of data repairs, each one annotated with how many pipeline use cases it unblocks. Hold that thought, because it is the hinge of this whole lesson and the reason Chapter 4.3 exists: the data program you will design there does not invent its own work. It executes the work orders this audit writes. An enterprise data audit that ends in adjectives ("quality is inconsistent," "ownership is unclear") has produced sentiment. One that ends in a backlog has produced a program.
The stakes are already quantified at the industry level. Gartner found that 63 percent of organizations either lack AI-ready data practices or are unsure whether they have them, and forecast that through 2026, 60 percent of AI projects without AI-ready data will be abandoned. Those are the enterprise diagnosis. Your audit's job is to localize it: which 63 percent, sitting in which systems, blocking which use cases, repairable at what price. The MIT autopsy of the 95 percent taught you that pilots die of missing workflow integration; the enterprise version of that lesson is that portfolios die of missing data integration, twenty pilots at a time, each one discovering the same broken vendor master six weeks after its own kickoff.
Criticality Triage: How a Dataset Earns Its Audit
Here is the move that makes an enterprise data audit finishable, and it comes first because everything else depends on it. A 2,400-person company does not have a data landscape; it has a data ocean. Count every table, extract, spreadsheet, SharePoint list, and vendor export and you will pass two thousand candidate datasets without breaking a sweat. You cannot audit two thousand datasets. You cannot audit two hundred. The organizations that try are the subject of this lesson's failure story, and their audits are still running.
So the first artifact of the enterprise audit is a filter, and the filter has a name: the Critical Dataset Register. Not every dataset. The 20 to 40 that matter, each one scored on five enterprise dimensions, each one carrying a named owner (or the documented absence of one), a consumer list, and a remediation line. The register is the named artifact of this lesson, and the rest of the audit is simply the discipline of filling it in honestly.
How does a dataset earn a slot? Not by size. Not by the prestige of the system it lives in. Not by how much the enterprise resource planning (ERP) team loves it. A dataset earns a register slot by consumption potential: the number of pipeline use cases that would consume it. You already hold the demand evidence from earlier in this level: the shadow-AI demand map that told you what employees are already trying to do, the Level 2 scorecard nominations that surfaced candidate processes, and the function interviews from the framework lesson. Walk that pipeline, use case by use case, and for each one list the datasets it would touch. Then invert the list. The datasets that appear under multiple use cases are your register. The ones that appear once can wait.
The instrument for this is deliberately humble: a dependency-count table. Down the left, candidate datasets. Across the top, pipeline use cases. A mark in each cell where a use case consumes a dataset. Sum the rows. An excerpt from our 2,400-person storyline company, with illustrative counts:
| Dataset | Pipeline use cases touching it | Dependency count |
|---|---|---|
| Vendor master | Invoice matching, procurement copilot, spend analytics, contract renewal alerts, fraud screening, supplier-risk scoring | 6 |
| Customer master | Churn model, service copilot, collections prioritizer, segmentation, quote assistant | 5 |
| Contract repository | Renewal alerts, obligation extraction, procurement copilot, legal review triage | 4 |
| Support ticket history | Service copilot, product-defect clustering, workforce forecasting | 3 |
| Marketing events list | Event-lead scoring | 1 |
Read that table the way a portfolio manager reads it. The vendor master appears under six use cases: it is infrastructure, and every hour spent auditing and repairing it is an hour invested six times over. The marketing events list appears under one: it is that use case's private problem, and it waits outside the register until demand promotes it. This is the rule worth engraving, because every stakeholder with a favorite system will test it:
Data criticality is derived from use-case demand, never from system size, system age, or IT's affection for the platform it runs on.
Expect the pushback and enjoy it, because the pushback is the triage working. The data warehouse team will argue their star schema belongs on the register because it is large and expensive. The answer is a question: which pipeline use case consumes it? If the honest answer is none yet, it waits, no matter what it cost to build. Demand-derived criticality is also your defense in the other direction: when a function head insists their niche dataset is "strategic," the dependency count either agrees with them or it does not, and the table is harder to argue with than you are.
The Five Enterprise Dimensions
Every dataset that makes the register gets scored on five dimensions, on the anchored 1-to-5 scale you built in the framework lesson, and the evidence standard from that lesson applies with full force: commissioned reports and profiled samples, not opinions in a workshop. An enterprise data audit built on testimony inherits every distortion of the org chart. One built on tickets, profiles, and retrieval tests inherits only the truth.
1. Quality: fitness for the consumer, not the producer
This is the Level 2 profiling craft, run on samples: completeness of critical fields, consistency of formats, duplication rates, referential integrity, with the L2 gap taxonomy as the cell-level classification tool when you need to name exactly what kind of broken a field is. But the enterprise adds a twist that changes scores by two full points: quality is fitness for consumption, and consumers differ. Profile at the consumer's standard, not the producer's. The vendor master that is perfectly adequate for the payments team (they key on vendor ID and bank account, and those fields are clean because money forces them clean) can be simultaneously broken for AI matching, where the model needs names, addresses, and hierarchies, and those fields are a museum of abbreviations, mergers nobody recorded, and 312 duplicate vendor records. The register therefore records quality per consumer class: one line, two scores, and a note explaining why both are true. This single practice ends the most tedious argument in enterprise data work, the one where the producing team says "our data is fine" and the consuming team says "your data is garbage" and both are right.
2. Ownership and stewardship: a name, a budget, a standard
Three questions, each with a documentary answer. Is there a named owner, in writing, in a current document? Is there budgeted steward time, meaning hours in someone's actual plan for maintaining this dataset, not goodwill? Is there a quality service-level agreement (SLA), meaning a written standard the data is held to, with consequences? Score against the anchor scale: a 5 has all three; a 3 has a name but no budget or standard; a 1 has none.
Expect to do what is best called ownership archaeology. Ask "who owns the vendor master?" and you will get a name with confidence. Check the name and you will often find not an owner but the person who complained about the data most recently, or the analyst who wrote the cleanup script in 2023, or a manager who left in the migration. Ownership by folklore is a 1 wearing a 3's clothes. The audit's job is not to fix this (that is governance work, and it gets its own treatment later in this level); the audit's job is to name the gap precisely and price the role: this dataset needs a named owner at roughly 0.2 full-time equivalents of steward time, currently funded at zero.
3. Access architecture: the tax every future pilot pays
At Level 2 you hit access walls one process at a time. At enterprise scale, access is an architecture, and its central metric is the one already glowing on your scorecard: grant latency, measured from the ticketing system, never from testimony. Our storyline company's median is 14 days from request to usable access. IT remembers it as "about a week" because IT remembers the easy tickets.
Then go deeper than latency, because latency is a symptom. Is access purpose-bound, so a grant for one use does not silently become access for everything? Is there a fast path for approved AI use cases, or does the flagship program queue behind password resets? Do the access rules survive the privacy pre-flight you learned at Level 2, or does personally identifiable information (PII) handling get improvised per request?
Access is where data programs die quietly, because nobody ever writes "access killed it" in a post-mortem. They write "the pilot lost momentum." So compute the tax and make it loud. Twenty pipeline use cases, each needing on average grants to three systems, at a 14-day median, with even modest parallelization, is hundreds of calendar-days of pure waiting distributed across the portfolio's first year. Illustratively: if 20 pilots each lose two weeks of a five-person team's momentum, that is on the order of 1,000 person-days of stall. Against that number, a dedicated access fast path costing 15 IT-days to build is not a convenience. It is the highest-return line the audit will produce, and it benefits every use case at once.
4. Lineage and freshness: would anyone know if it stopped?
Three questions per dataset. Where does this data come from, in a traceable chain rather than an anecdote? How often does it update, verified against timestamps rather than the schedule someone remembers? And the sharpest one: if it stopped updating, would a consumer know, or would models keep consuming quietly staling data until an output was embarrassing enough to investigate? This is the drift-detection question you learned at Level 3, asked one layer down, of the data itself. A dataset with no update monitoring is the silent-rot lesson in its enterprise form: nothing fails loudly, everything degrades politely, and the first alarm is a business decision made on data that died months ago. Score a 5 for documented lineage plus automated freshness monitoring with alerts; score a 1 for "we think the nightly job still runs."
5. AI-usability: reachable by our tooling, under our rules
This is the dimension the enterprise adds, and it exists because a dataset can score well on all four classic dimensions and still be unusable for AI. The question: can the organization's AI tooling actually consume this dataset, under the organization's actual rules? Four sub-checks. Format: is it structured or extractable, or is it 40,000 scanned PDFs? Tenancy: can it leave the system or region it lives in, or do residency rules pin it? PII load: how much redaction or pseudonymization infrastructure does it require before a model may see it? And contracts: do supplier or customer agreements restrict third-party processing, including sending the data to a model provider at all?
This is the Level 2 privacy and permission pre-flight, generalized into a scored dimension. Its punchline deserves italics in your mind: a clean, owned, accessible, fresh dataset that contract clauses forbid sending to any external model scores low here, and discovering that before twenty use cases have assumed the opposite is the audit's whole point. It is also where the regulatory clock touches the register: the EU AI Act's high-risk obligations, which include data-governance requirements and begin applying to Annex III systems on December 2, 2027, will ask organizations to show exactly the lineage, quality, and governance evidence this dimension collects. The register you build for the AI program doubles as the file you will want when a regulator or an enterprise customer asks how your data was governed.
The Remediation Backlog: Where the Audit Becomes a Program
Scores diagnose; the backlog treats. The rule is mechanical: every dimension score below 3, on any registered dataset, generates one backlog line. Each line carries four fields, and the fields are chosen so the backlog can be sorted, funded, and assigned without another meeting.
- The fix, stated as work, not aspiration: "assign and fund vendor-master ownership; run the deduplication using the L2 taxonomy to classify the 312 duplicates," not "improve vendor data quality."
- The effort class: days, weeks, or quarters. This is the Level 2 severity-triage economics scaled to a portfolio: you are not estimating to the hour, you are sorting repairs into cost bands honest enough to sequence.
- The beneficiary count: how many pipeline use cases this fix unblocks, read straight off the dependency-count table. This is the prioritization key and the political weapon. Fix the vendor master and six use cases move. Fix the marketing events list and one does.
- The owner-designate: the name proposed to own the fix, because a backlog line without a name is a wish.
Sort the backlog by beneficiaries per unit of effort, and something quietly important happens: the data program's roadmap writes itself, pre-justified. This is the difference between a data program that gets funded and one that gets politely deferred every quarter. A hygiene sermon ("we must invest in data quality as a foundation") loses to any revenue project in any budget fight, every time, and deserves to. A backlog line that reads "this 40-day fix unblocks six pipeline use cases carrying roughly $340,000 of projected annual value" is not a sermon, it is an investment memo one line long. That is the Level 2 money-line discipline operating at enterprise scale, and it is why the CFO, who has been the skeptic in every chapter of this program, becomes the data program's sponsor in this one. Chapter 4.3 will take this backlog and wrap a program around it: sequencing, staffing, governance, and the operating rhythm. The audit's job ends when every sub-3 score has a priced, named, beneficiary-counted line. The program's job begins there.
What Only the Enterprise Audit Can See
Two findings belong in the enterprise report that no process-level audit could have produced, because both only become visible when the whole register is on one page.
The concentration finding
Read the dependency-count column again, this time as a risk officer. Six use cases on the vendor master is an opportunity: one fix, six unblocks. It is also a single point of failure: one unowned, duplicate-ridden dataset sits under a quarter of the AI portfolio, and it currently has no steward, no SLA, and no freshness monitoring. The same reading applies to systems and vendors: if fourteen of twenty use cases assume data flowing through one integration platform, or one cloud tenancy, or one software-as-a-service (SaaS) vendor's export API, the register has just mapped the portfolio's concentration risk before a single pilot depends on it. Write the concentration finding as its own short section of the report: which datasets, systems, and vendors everything depends on, and which of them are currently orphans. You will meet this idea again at Level 5, where concentration risk gets a full lesson of its own; the register is where its evidence gets collected. Plant the seed now: every dependency count above four is both a priority and an exposure.
The dark-data honesty
One paragraph of the report, and often the most clarifying one: what does the organization believe it has that it cannot actually produce? Every enterprise carries retrieval folklore: "we have every customer interaction since 2015," "the archive has all the legacy contracts," "that history is in the old system." The test is the history-holes hunt from Level 2, scaled and made ruthless: for each critical retrieval claim, run one actual retrieval. Ask for one specific contract from 2019, one specific interaction thread from 2017, and time what happens. In our storyline company, the "complete contract archive" claim met a retrieval test and produced the contract in question after nine days, from a backup format two migrations old, missing its amendments. That is not an archive; that is a rumor with a storage bill. Use cases were being nominated on the assumption of ten years of training history that functionally does not exist. One paragraph, one test per claim, and the pipeline stops planning against phantom data.
The Register in Action, and the Audit That Never Ended
A worked example: 31 datasets out of two thousand
Here is the storyline company's audit, end to end, with all figures illustrative. The team walks the pipeline: 22 candidate use cases from the demand map, the scorecard nominations, and the function interviews. The dependency-count table starts with roughly 2,000 candidate datasets by name; the demand walk touches 118 of them; 31 have a dependency count of 2 or higher and make the register. The audit scopes to those 31, and is suddenly finishable: five dimensions, 31 datasets, evidence-based, eight weeks.
The vendor master's register line, as a scoring vignette. Quality: 2 for AI-matching consumers (the 312-duplicate history you met in earlier levels, plus name and hierarchy fields profiled at 71 percent completeness on a 500-record sample), while payments-class consumers would score it 4, and the register records both. Ownership: 1: no named owner since the 2022 migration, and the audit's archaeology shows the last three "owners" were whoever complained most recently. Notice the causal arrow the register makes visible: the duplicate pile is the symptom; ownership 1 is the disease; a dedupe without an owner is a cleanup that will need repeating in eighteen months. Access: 3: grants work but ride the 14-day median with no fast path. Lineage and freshness: 2: three feeding systems, one undocumented, no update monitoring. AI-usability: 2: bank-detail fields carry contractual and PII restrictions (the Level 2 finding, now generalized), so any AI consumption needs a redacted view that does not yet exist.
The full register produces a 22-line remediation backlog. The top of the sort:
- Vendor-master ownership plus deduplication: 40 steward-days, 6 beneficiaries, unblocks roughly $340,000 of projected annual pipeline value. The first work order of the future data program.
- Access fast path for approved AI use cases: 15 IT-days, benefits all 22 pipeline use cases, attacks the 14-day median directly.
- Customer-master redacted view for model consumption: 25 days, 5 beneficiaries.
- Contract-repository text extraction and freshness monitoring: 30 days, 4 beneficiaries.
And the audit's headline totals, the numbers the readiness heat map in lesson five will compress to a single cell: 7 of 31 critical datasets are AI-usable today. That is 23 percent, which is the organization's local, evidence-based instance of Gartner's finding that 63 percent of organizations lack AI-ready data practices: the industry statistic, localized to named files with named gaps. Remediation to bring the count from 7 to 20: approximately 210 effort-days across two quarters, sequenced by beneficiaries-per-effort, every line owner-designated. The data dimension's 2.4 is no longer a grade. It is a bill, itemized, with a payment plan.
The failure story: the audit that tried to boil the ocean
Now the other path, assembled from a pattern you will recognize the moment you see it forming. A bank's data office receives essentially the same mandate: assess the data foundation for the AI program. Their opening move is the fatal one, and it sounds impeccably rigorous: "we cannot prioritize what we have not cataloged." So they begin a full enterprise data catalog: every system, every table, every field, harvested into a metadata platform, with quality scores to follow.
Eighteen months later the catalog is 60 percent complete, which in catalog work means the easy 60 percent. The AI program has funded zero pilots, because every proposal is told to wait for the catalog that will "tell us if the data is ready." Meanwhile the business, which cannot wait, has quietly bought three SaaS tools with their own embedded data stores, none cataloged, so the territory is now growing faster than the map. The shadow expands while the map is drawn. Two data-office analysts have left; their successors are re-harvesting systems whose metadata went stale during the harvest.
Here is the uncomfortable part: the catalog was never wrong. Every entry was accurate. It was unprioritized, and at enterprise scale unprioritized is a synonym for unfinished, because enterprise completeness is a decade-long asymptote and the pipeline needed twenty datasets, not two thousand. Criticality triage exists precisely because of this arithmetic. The demand-first audit and the catalog-first audit both claim to be thorough; only one of them ever ships. Audit for the demand you have. Catalog as the program matures, funded by the credibility the first remediations bought. The bank's mistake was not ambition; it was sequencing, and sequencing is exactly what a strategist is paid to get right.
What to Do Monday Morning
You can have a draft register skeleton inside a week, using evidence you mostly already hold.
- Build the dependency-count table from your current use-case pipeline. Take every candidate use case you have (demand map, scorecard nominations, interview notes), list the datasets each would consume, invert, and sum. Two hours of honest work, and the enterprise data ocean becomes a ranked list.
- Cut the register to datasets with a dependency count of 2 or more. Expect 20 to 40 survivors. Write down what you excluded and why, because the excluded systems' sponsors will ask, and "no pipeline demand yet" is an answer that survives a steering committee.
- Commission the access-grant-time report. Ask IT for the last six months of data-access tickets with request and fulfillment timestamps, and compute the median yourself. This is your 14-day number, whatever it turns out to be, and it converts the access conversation from anecdote to arithmetic.
- Run one retrieval test on your most-cited archive claim. Pick the "we have all of that history" statement most load-bearing for your pipeline and request one specific, dated item from it. Time the retrieval. File the result in the dark-data paragraph.
- Draft the top five backlog lines with beneficiary counts. Even with rough dimension scores, you know enough to write five lines with fix, effort class, beneficiary count, and owner-designate. Sort by beneficiaries per effort. You are now holding the first page of your organization's data program, one chapter before you formally build it.
Key Takeaways
- Shift the unit of analysis from process to dataset at enterprise scale: the process audit asked whether one pilot's data could carry it; the enterprise audit asks which assets can carry any of the next twenty use cases, and prices making the rest carriable.
- Build the Critical Dataset Register by criticality triage: datasets earn their 20 to 40 slots through consumption potential read from the dependency-count table, never through system size or IT's affection.
- Score each registered dataset on five evidence-based dimensions: quality at the consumer's standard (per consumer class), ownership and stewardship, access architecture, lineage and freshness, and AI-usability under real contracts, tenancy, and PII rules.
- Measure access as an architecture, not an anecdote: pull grant latency from ticket logs, compute the portfolio-wide waiting tax, and let that number fund the fast path that benefits every future use case at once.
- Convert every sub-3 score into a remediation backlog line with fix, effort class, beneficiary count, and owner-designate; sorted by beneficiaries per effort, the backlog is the data program's pre-justified roadmap and Chapter 4.3's work orders.
- Report the two findings only the full register reveals: the concentration finding (high dependency counts are opportunities and single points of failure at once) and the dark-data honesty paragraph, verified by one real retrieval per critical claim.
- Localize the industry statistics instead of quoting them: Gartner's 63 percent becomes "7 of our 31 critical datasets are AI-usable today," which is a finding a CFO can fund against.
- Refuse the boil-the-ocean catalog: enterprise completeness is a decade, the pipeline needs twenty datasets, and an accurate but unprioritized catalog ships nothing; audit for the demand you have and catalog as the program matures.
Skill.re