←
CAP Certification
Strategic · M22 · lesson 22 of 60 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Data Quality & Master Data Management

15 min

Fabrizio Costa spent three weeks building what looked like an excellent AI-powered pricing recommendation tool for his retail chain's procurement team. The model trained cleanly, the validation metrics were strong, and his pilot users said the recommendations felt right. Then the store operations team started complaining: the recommended prices were wrong for roughly 15% of products. When Fabrizio traced the issue, the problem was not in his model. It was in the product master data. Two different ERP systems defined "unit of measure" differently for 840 product SKUs, so his model had trained on data that appeared consistent but was not. "The model was working perfectly," he said. "It was optimizing on garbage."

That is the foundational problem master data management exists to prevent. Data quality is not a technical problem with a technical solution. It is an organizational problem, about ownership, accountability, and the discipline to maintain standards over time, that technical tools support rather than solve. Organizations that treat data quality as something IT handles while the business focuses elsewhere consistently produce the kind of model Fabrizio built: technically sound, substantively wrong, and convincing enough that nobody questions it until the complaints arrive from people downstream.

What Data Quality Actually Means

Data quality is not a single property. It has six dimensions, and a dataset can score well on some while failing badly on others. That is why "is this data good?" is an unanswerable question and "is this data accurate, complete, and consistent enough for this particular use?" is a tractable one.

  • Accuracy: Does the data reflect the real world correctly? A customer record with a misspelled name or a wrong address fails on accuracy.
  • Completeness: Are there missing values where values should exist? A product record with no unit price fails on completeness.
  • Consistency: Does the same thing mean the same thing across systems? Fabrizio's unit-of-measure problem was a consistency failure.
  • Timeliness: Is the data current enough for its intended use? Inventory data that is 24 hours old may be adequate for weekly reporting but dangerously stale for real-time reorder decisions.
  • Uniqueness: Does each real-world entity appear exactly once, or are there duplicates? A customer database holding the same customer in three records under slightly different names fails on uniqueness.
  • Validity: Does the data conform to defined formats and business rules? A date field containing "N/A" is a validity failure.

For AI specifically, consistency and completeness are typically the most consequential of the six. A model trained on inconsistent data learns the wrong patterns, and it learns them confidently, because nothing in the training process flags a contradiction between two source systems as an error. A model trained on significantly incomplete data learns patterns that do not generalize, because the records that were missing are rarely missing at random. Both failures pass validation, which is exactly what makes them expensive.

Quality Is Defined Per Use Case, Not Globally

The most common mistake in a data quality programme is to pursue quality as an absolute. It is not one. Quality requirements vary by what the data is for, and the same dataset can be entirely fit for one purpose and unusable for another. Data supporting exploratory analysis, where an analyst is looking for a direction worth investigating and will sanity-check anything surprising, can tolerate gaps and inconsistencies that would be unacceptable in data feeding a regulatory report, where a single wrong figure is a filing error.

This is why the useful unit of governance is the pairing of a dataset with a use, not the dataset alone. Write the requirement down in that form: this customer table, used for this campaign, needs this level of completeness on these fields. A 95% completeness threshold on customer email addresses might be perfectly acceptable for marketing campaigns, where the cost of a gap is one unreached recipient, and entirely inadequate for compliance reporting, where the gap is a defect in the submission.

Tiering this way has a practical benefit beyond accuracy: it tells you where to spend. A quality programme that holds every dataset to the standard of its strictest consumer will exhaust its budget on records that nobody depends on. One that tiers by use case concentrates stewardship effort on the small number of data domains where the cost of being wrong is high, and accepts known, documented imperfection everywhere else. The documentation matters as much as the decision, because an undocumented tolerance is indistinguishable from an unnoticed defect the next time somebody new uses the data.

Master Data Management

Master data management, or MDM, is the practice of maintaining a single, authoritative version of the data that matters most across the enterprise. The "master" in the name refers to the master record: the version of a customer, product, supplier, or location that all other systems should treat as the truth. Everything else is a local copy that resolves back to it.

Think of MDM as a census registry. Every jurisdiction has different agencies, for tax, health, transport, and welfare, and each maintains its own records about citizens. Without a master registry, the same person might appear under four different names, two different addresses, and three different identification numbers across those agencies, and no agency would be wrong from its own point of view. With a master registry, each agency knows how to resolve its local records to a canonical source. MDM is that canonical source for enterprise data, and the discipline is less about the technology than about agreeing which system gets to be the registry.

The four most common master data domains in large organizations are:

  • Customer master data: the single authoritative record of each customer, who they are, how to reach them, and what relationship they hold with the organization.
  • Product master data: the single authoritative definition of each product, its attributes, its pricing basis, its unit of measure, and its classification.
  • Supplier master data: the single authoritative record of each supplier, their identity, their terms, and their risk profile.
  • Location master data: the single authoritative definition of each physical location, facility, territory, or jurisdiction.

Fabrizio's problem was specifically a product master data failure. Two ERP systems held different product master records for the same SKUs, with different unit-of-measure conventions, and neither system knew the other existed as a competing authority. Nothing in either system was broken; each was internally consistent. The defect only became visible at the point where a model consumed both of them at once and treated their contradictory definitions as a single vocabulary.

Data Stewardship

Data quality degrades continuously without active maintenance. New records are created inconsistently. Source systems change their definitions. Business rules evolve but data formats do not. Without someone responsible for catching and correcting this drift, quality declines on a predictable schedule, and the decline is usually invisible until something downstream depends on it.

Data stewardship is the practice of assigning explicit ownership of data quality to specific people, typically business users rather than IT staff, who have the context to understand what quality means for their domain and the authority to enforce standards. A steward for customer master data, for example, is responsible for defining what a complete customer record looks like, reviewing and resolving duplicate records on a scheduled basis, approving the addition of new fields or definitions, and escalating quality issues that require system changes to IT.

The business ownership of this role is the part organizations most often get wrong, usually by assigning stewardship to whoever administers the database. IT can build tools that detect and flag quality issues, and those tools are genuinely valuable. IT cannot know whether a flagged record represents a genuine duplicate or two different customers who happen to have similar names, whether a missing field is a defect or a legitimately unknown value, or whether a rule that fires constantly is catching a real problem or encoding a stale assumption. That judgment requires domain knowledge, and domain knowledge sits with the business.

Frameworks and Monitoring

Defining quality standards is necessary but not sufficient. You also need to measure quality consistently, so you know whether it is improving or declining and where the most significant issues are concentrated. A practical framework defines four things for each critical data domain:

  • Quality dimensions that matter: not all six are equally important for all domains. Prioritize the two or three that most affect AI model performance or business decisions.
  • Acceptable thresholds: what percentage of records must meet the standard before the data is usable for a specific purpose? This is the fitness-for-purpose decision, written down and attached to a number.
  • Measurement frequency: daily for data feeding real-time AI systems, weekly or monthly for data used in periodic reporting and models.
  • Escalation paths: when quality falls below threshold, who is notified, and what action is expected within what timeframe?

Fabrizio's post-incident response was to build a simple quality dashboard for product master data: a weekly report showing the percentage of active SKUs with complete unit-of-measure definitions, consistent across both ERP systems, with valid format values. The first version took two days to build. It found 840 inconsistent records in week one. After remediation, it became the early-warning system that prevented three similar issues in the following year. The lesson worth taking from that is not that dashboards are powerful; it is that a narrow monitor aimed at one known failure mode is cheap and pays for itself, while the comprehensive quality platform nobody has time to build never does.

Remediation and Prevention

Finding quality issues is half the work. Fixing them, and preventing recurrence, is the other half. Remediation is not a one-time cleanup; it is an ongoing process with two distinct phases that call for different skills and different budgets.

Reactive remediation addresses quality issues after they are found. Locate the error, correct the record, trace the root cause, asking whether this was a data entry error, a system integration failure, or a business rule violation, and then fix the root cause rather than just the record. Correcting the record alone guarantees you will correct it again. Proactive prevention addresses root causes before they produce errors: validation rules at data entry points, integration tests that catch inconsistencies before they propagate downstream, and automated alerts when new records fail quality checks.

The ratio between reactive and proactive work shifts as data governance matures, and that ratio is itself a useful maturity signal. Early in a programme most effort is reactive, cleaning up the backlog of existing issues. Over time, as prevention mechanisms take hold, the reactive burden decreases and the team spends more of its time on rules and tests than on corrections. Organizations that never invest in prevention stay permanently in cleanup mode, which is both more expensive and more disruptive, because every cleanup cycle interrupts whoever depends on the data while it runs.

Quality as an Operational Discipline

The most important conceptual shift in data quality governance is from thinking of it as a project to thinking of it as an ongoing operational discipline. You do not complete data quality work any more than you complete financial controls work. The conditions that produce quality problems, people entering data, systems integrating, business rules changing, are permanent features of organizational life, so the response has to be permanent too.

Fabrizio's pricing tool now runs on product master data actively stewarded by two part-time product data stewards. It is reviewed weekly. The thresholds that trigger review have been refined three times based on the problems that actually occurred rather than the ones anticipated at design time. The model's recommendation quality is measurably higher than in the pilot, and it has stayed that way for 14 months. That stability is not an accident of better technology. It is the product of an organizational commitment to treating data quality as a discipline rather than a project.

Anti-Patterns to Avoid

Data quality programmes fail in recognizable ways, and each failure has an early sign a leader can watch for.

  • Debugging the model when the data is at fault. Fabrizio lost time inside a model that was working correctly. The tell is a model whose errors cluster around particular products, regions, or record types rather than distributing evenly.
  • Treating quality as an absolute rather than a fitness-for-purpose judgment. Holding every dataset to the strictest consumer's standard exhausts the budget on data nobody depends on, and leaves no capacity for the domains that matter.
  • Assigning stewardship to IT. Tools can flag a suspected duplicate. Only someone with domain context can judge whether it is one, and a steward without that context becomes a queue rather than a decision-maker.
  • Assuming two internally consistent systems agree with each other. Neither ERP in Fabrizio's case was broken. The defect lived in the space between them, which no single system owner was looking at.
  • Fixing records instead of causes. A correction that does not change a validation rule, an integration, or a definition is a correction you have scheduled yourself to make again.
  • Running a cleanup and declaring victory. Quality declines from the day the cleanup ends unless stewardship and monitoring continue, and the second decline is harder to fund than the first.

Practice Prompts

Work these against data your own AI systems actually consume. Each should produce something you can show a colleague.

  • Score one dataset on all six dimensions. Take a table feeding a live model and rate it for accuracy, completeness, consistency, timeliness, uniqueness, and validity. The dimension you cannot measure at all is the one to instrument first.
  • Write one fitness-for-purpose statement. Pick a dataset and a specific use, and write the sentence that says which fields need what threshold for that use. Then find someone who uses the same data differently and write theirs.
  • Hunt for a competing authority. Choose one master data entity, such as product or customer, and list every system that believes it holds the definitive record. If more than one does, you have found the shape of Fabrizio's problem before it reaches a model.
  • Build the narrow monitor. Instrument the single failure mode that has already burned you once, on the model Fabrizio used, and put it on a weekly report with a named recipient.

Reflection

Think about the data feeding your most important model. If somebody asked who is accountable for its quality, would you be able to name a person, and would that person agree? Consider the last time a model produced results that looked wrong: how long did the team spend inside the model before anyone examined the inputs? And of the datasets your organization actively maintains, how many are being held to a standard that reflects what they are actually used for, rather than a standard inherited from whoever built the pipeline first?

Glossary

  • Master data. The data describing the core entities of the business, customers, products, suppliers, and locations, which multiple systems reference and which therefore needs one authoritative definition.
  • Master record. The single version of an entity that all other systems resolve to. The absence of one, rather than the presence of errors, is what produced Fabrizio's failure.
  • Consistency. The dimension measuring whether the same thing means the same thing across systems. It is the dimension AI training is least able to detect on its own.
  • Fitness for purpose. The judgment that a dataset meets the quality standard required by a specific use, rather than an absolute standard. Exploratory analysis and regulatory reporting sit at opposite ends of it.
  • Data steward. The named business owner accountable for a data domain's definitions, duplicate resolution, and escalation of issues requiring system change.
  • Reactive remediation. Correcting quality issues after discovery, including tracing and fixing the root cause rather than only the record.
  • Proactive prevention. Validation rules, integration tests, and automated alerts that stop defects being created. The balance between this and reactive work indicates governance maturity.

This chapter sits inside the wider arc on advanced data governance. Enterprise Data Architecture & Governance comes first and establishes the structures that stewardship operates within, while Data Access & Security Governance follows and covers who may reach the data once you can trust it. Data Privacy & Compliance Governance closes the arc, and depends directly on the inventory and stewardship discipline described here, since a privacy programme can only govern the data domains somebody has already agreed to own.

Closing

Fabrizio's model was never the problem, and that is the part worth remembering. He had good tooling, sound validation, and satisfied pilot users, and he still shipped recommendations that were wrong for a meaningful share of his catalogue, because two systems disagreed about what a unit of measure was and nobody owned the disagreement. The fix was not smarter modelling. It was a named owner, a defined standard, a weekly measurement, and a decision that the standard would be maintained indefinitely rather than achieved once. That is unglamorous work, and it is the work that determines whether an AI programme produces decisions anyone should act on.

Key Takeaways

  • Data quality has six dimensions: accuracy, completeness, consistency, timeliness, uniqueness, and validity. For AI, consistency and completeness tend to have the highest impact on model performance, and the highest remediation cost when ignored.
  • Quality is defined per use case, not globally. The same dataset can be fit for exploratory analysis and unfit for regulatory reporting. Write the requirement as a dataset plus a use, and document the tolerances you accept.
  • Master data management creates a single authoritative record for the entities that matter most. Without it, the same customer, product, or supplier carries multiple contradictory definitions across systems, and models learn from the contradictions.
  • Business users must own data stewardship. IT can build the monitoring. Only business users have the context to define what quality means for a domain and to judge whether a flagged issue is a real problem or an artifact.
  • Define thresholds before you need them. The question of what percentage of records must meet a standard for a given purpose belongs in the design phase, not in the post-mortem after a model has trained on substandard data.
  • Invest in prevention, not just remediation. Validation rules, integration tests, and automated alerts cost less over time than repeated manual cleanup, and the ratio of proactive to reactive work is a fair indicator of governance maturity.
  • Quality is an ongoing discipline, not a project. The conditions producing quality problems are permanent, so the stewardship, measurement, and escalation answering them have to be permanent too.

Frequently Asked Questions

How clean does data need to be before we can train a model on it? There is no universal answer, which is the point of defining quality per use case. Decide what the model will be trusted to do, work backwards to the dimensions that would break that trust, and set thresholds on those dimensions only. A recommendation tool that a human reviews before acting tolerates more imperfection than one that adjusts prices automatically, and the difference belongs in a written threshold rather than in someone's judgment on the day.

Do we need a master data management platform to do this? Not to start. What MDM requires first is an agreement about which system holds the authoritative record for each entity, and that agreement is organizational rather than technical. Fabrizio's early-warning system was a weekly report built in two days. A platform helps at scale, but buying one before the ownership question is settled tends to produce a well-engineered copy of the existing disagreement.

Our stewards keep escalating everything to IT. What went wrong? Usually the steward has responsibility without authority, or without the domain context to make a call. A steward who cannot approve a definition change, resolve a duplicate, or accept a documented tolerance is functioning as a routing queue. Fixing it means granting the decision rights explicitly and confirming that the person holds the business knowledge the role assumes.