←
AI for Government
Capable · M16 · lesson 16 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Data Governance for AI

15 min

Colleen Whitfield spent three weeks last autumn trying to get a straight answer from her own agency's data team. She is the AI program manager at a mid-sized state labor department, around 1,400 employees and a $380 million annual budget, and she had an AI pilot she had spent eight months shepherding through procurement. The pilot would use machine learning to flag unemployment claims that showed signs of potential fraud before they were paid. The model vendor needed training data: ten years of adjudicated claim records, roughly 4.2 million rows. Colleen's question was simple. Was that data clean enough to use? Three weeks later she still did not have an answer. The data team said the records were "generally complete." The vendor said they needed a formal data quality assessment before signing off. Legal said they needed to know how the data had been collected and whether the collection methods were still in use. Nobody had documentation for any of it. The pilot slipped four months while the agency tried to reconstruct a data history it had never formally recorded.

Why Data Governance Is the Foundation

Colleen's situation is common. Agencies accumulate data over decades of program operations. That data is stored in systems that have been upgraded, migrated, and patched. The people who designed the original collection processes have often retired. When an AI project arrives and asks foundational questions, such as where this data came from, who collected it, and how it has changed, the honest answer is frequently that nobody fully knows.

Every AI practitioner knows the phrase "garbage in, garbage out." A machine learning model is only as good as the data it was trained on. Poor quality data produces poor outputs. Biased data produces biased models. Incomplete data creates blind spots. Yet many agencies still treat data governance as a separate discipline from AI governance. Data teams manage quality, access, and retention on one track. AI teams build and deploy models on another. The result is predictable: data problems cascade into system problems, and nobody owns the seam between them.

This matters enormously for AI. A model trained on incomplete or biased data will produce unreliable outputs. A model trained on data that was collected in a discriminatory way will reproduce that discrimination at scale, faster and more consistently than any human process. A model whose training data cannot be documented cannot be audited, which means it cannot be defended before an Inspector General, a congressional committee, or a civil rights complaint proceeding.

Think of data governance as the building code for AI. Nobody sees it once construction is finished. But when the structure fails, the first question investigators ask is whether the code was followed. If the answer is no, if the data was not documented, quality-checked, or access-controlled, then every subsequent AI decision that rested on that data is exposed. Government agencies hold decades of program operations, millions of transactions, and extensive records. Managed well, that holding is an asset. Managed poorly, it is a liability, and the liability arrives at the worst possible moment.

The Dimensions You Must Manage

Data governance for AI is not a single activity. It spans distinct dimensions, each of which creates its own risks if neglected, and each of which needs an owner.

Quality means the data is accurate, complete, and consistent. For AI systems, completeness matters more than most analysts realize. A dataset that is missing records for rural applicants, for non-English speakers, or for applicants who filed by paper rather than online will train a model that systematically underperforms for those groups. That is not a technical failure. It is an equity obligation failure with legal exposure under Title VI and related statutes.

Provenance means you can answer where the data came from. Provenance documentation records the original source, the collection method, the date range, and any known limitations. For Colleen's agency, the problem was that ten years of records spanned three different case management systems. Each had different field definitions for the same data element. Without provenance records, the vendor could not know whether "claim status: denied" meant the same thing in 2015 as it did in 2023.

Lineage goes one step further. Lineage tracks every transformation applied to the data: cleaning steps, aggregations, joins with other datasets, feature engineering. If your agency's fraud detection model flags a claim and a constituent appeals, you need to be able to trace the model's input back through every transformation to the original record. Without lineage documentation, that trace is impossible.

Bias assessment asks whether the training data represents all relevant populations proportionally, and whether it carries historical biases that AI would propagate. Historical government datasets often reflect historical inequities. If certain zip codes were underserved by outreach programs, those residents are underrepresented in the data. A model trained on that data will learn that underrepresentation as a signal, potentially reinforcing the original inequity.

Access controls determine who can see what data and for what purpose. A data governance framework defines data classification tiers, typically public, internal, sensitive, and restricted, and enforces access permissions matching those tiers. Access to sensitive training data should require documented justification, supervisor approval, and audit logging. When a contractor's employee leaves the engagement, their access should be revoked within 24 hours, not at next quarterly review.

Retention addresses how long data is kept and what happens when the retention period expires. Many agencies have statutory retention requirements. The question for AI is more complex: must you retain the training data for as long as the model is in production? If a constituent wants to know why an AI system denied their benefit three years ago, can you reconstruct the model's inputs? Courts and Inspectors General are beginning to treat training data retention as an auditable record-keeping obligation.

Privacy protections ensure that personally identifiable information, meaning names, Social Security numbers, addresses, and medical records, is handled appropriately. Many AI projects require training data that contains PII. Governance requires reducing identifiability wherever the model can tolerate it, applying privacy-preserving techniques such as differential privacy where straightforward de-identification degrades model performance, and completing documented privacy impact assessments before training begins.

Linkage is the dimension agencies most often forget. When data from multiple sources is combined, integration itself must be governed. Joining a housing file to a benefits file can create a record that identifies individuals neither file identified alone, and it can create new correlations that neither dataset carried. Decide in advance who approves a linkage, what the joined dataset may be used for, and what new bias the join could introduce.

Documentation ties everything together. Well-documented data includes a data dictionary defining every field, notes on known quality issues, the history of schema changes, and guidance on appropriate uses and limitations. Without documentation, every new person who touches the dataset starts from scratch. Colleen's agency had four months of rework because documentation had never been created.

Data Quality Dimensions That Matter for Models

Quality is broad enough to need unpacking, because the aspects that matter for a dashboard are not the aspects that matter for a model. Accuracy comes first: if data trains a model, error in the data becomes error in the model. Completeness matters in a specific way, because the pattern of missingness matters more than the rate. Values missing at random degrade a model mildly. Values missing systematically, concentrated in one channel, one region, or one language group, teach the model that those populations look different in ways they do not.

Timeliness asks whether the data still describes the world. Stale data produces models that do not reflect current reality. Consistency asks whether similar concepts are represented similarly across datasets, because inconsistency confuses both models and the people reading their outputs. Representativeness asks whether the data covers the full population; where training data over-represents certain groups, models overfit to those groups. Validity asks whether values sit inside expected ranges, catching negative ages and impossible dates before they reach a model. Audit trails ask whether you can trace where data came from and whether changes were logged.

The implementation pattern is the same for every one of those aspects. Define written data quality standards for each dataset. Implement automated quality checks that run before data enters an AI system rather than after. Document quality issues you find and the remediation you applied. Monitor quality over time rather than testing once. Require a data quality assessment as a gate before deployment, not as a deliverable that arrives afterwards. Then establish a process for continuous improvement, because a dataset that met standard last year will not automatically meet it this year.

Provenance and Lineage in Practice

Provenance answers two questions: where did this data come from, and what has been done to it since. For AI systems that pair of questions does five distinct jobs. It explains what quality issues you should expect. It reveals when a collection process changed, which is one of the commonest causes of model drift. It makes auditing and compliance verification possible at all. It helps identify the sources of bias rather than merely detecting that bias exists. And it supports reproducibility, so a result can be validated by someone who was not there when it was produced. That is five jobs, counted from the list above.

The practical work is unglamorous. Document the sources of all training data. Track every transformation applied, including cleaning, aggregation, and feature engineering. Maintain lineage diagrams that show how datasets flow into models, so the path is legible to someone outside the data team. Version your datasets the way you version code, so a model can be tied to the exact data state it learned from. Document any change to collection or processing procedures at the time it happens. And build the ability to trace a model output back to its source records, because that is the capability an appeal or an audit will demand.

Access Control and Classification

Data is an asset and simultaneously a hazard. Access controls are how an agency keeps the asset usable without letting the hazard spread. Role-based access recognises that different people need different data: engineers building models may need record-level access, while program users need only aggregated results. Purpose limitation holds that data collected for one purpose should not be used for an unrelated purpose without additional safeguards, which for AI is the single most frequently violated principle, because training is almost always a new purpose. Audit trails record who touched sensitive data and what they did. Encryption protects sensitive data both in transit and at rest.

Reducing identifiability belongs here too. Where a model can tolerate it, data used for training or testing should be de-identified to reduce privacy risk. Note the phrasing: reduce, not eliminate. Stripping names and Social Security numbers removes direct identifiers, and it leaves quasi-identifiers such as date of birth, zip code, employer, filing channel, and rare diagnosis codes that can re-identify individuals when combined with other data an agency or an outside party already holds. Treat de-identified data as a lower-risk category, not a released category, and keep its access controls and approval processes in force.

The implementation sequence is straightforward to state and slow to finish. Define classification tiers of public, internal, sensitive, and restricted. Implement access controls that actually match those tiers, rather than matching them on paper. Require written justification for access to sensitive data and approve it on the basis of legitimate need. Audit access regularly rather than at incident time. Deactivate access when people change roles, not only when they leave the organisation. And encrypt sensitive data as a default rather than as an exception.

Retention and the Deletion Problem

Every dataset has a lifecycle. It is created, used, maintained, and eventually disposed of, and each stage needs a rule. Retention policy sets how long data is kept, and the standard should be the minimum period necessary for a legitimate purpose rather than indefinite storage by default. Archival handles data that is no longer actively used but cannot yet be destroyed. Deletion handles data that is no longer needed, and it should be secure and verified. Compliance sits over all of it, because some records must be retained for set periods by law and others must be disposed of promptly.

AI breaks this tidy lifecycle in one specific place. Data that has been absorbed into a trained model is genuinely hard to delete. Removing a record from the source database does not remove its influence from the model's parameters, and there is no simple delete operation that reaches inside a trained system. That creates a verification problem: how do you demonstrate that data has been deleted when a model has internalised it? The honest answers available today are structural rather than surgical. Retrain the model on a dataset that excludes the deleted records, or design with privacy-preserving techniques from the start so that no single record is individually recoverable in the first place.

So the retention work for an AI system has six parts. Establish a retention schedule for each dataset. Document why each period was chosen, because "we always have" is not a rationale an auditor accepts. Automate deletion where you can, since manual deletion at scale does not happen. Create a verification process that produces evidence, not assurance. Address the model-specific challenge explicitly in the system's design documents. And confirm the whole schedule against legal requirements with counsel rather than inferring them. That is six parts, counted from the list.

Implementation: The Agency Priority Sequence

Most agencies cannot implement every dimension simultaneously. The practical sequence is to start with the AI project at hand and work backward.

First, inventory the data the AI system will use. Document what exists, what format it is in, what time period it covers, and what systems it came from. This inventory is the minimum viable provenance record. It does not need to be perfect. It needs to exist.

Second, run a quality assessment. Many agencies use basic automated checks: record counts by time period, null value rates by field, value range validation. For Colleen's fraud data, this meant checking whether claim amounts were within plausible ranges, whether all required fields were populated, and whether the ratio of denied to approved claims was consistent across the ten-year period or showed suspicious discontinuities around system migration dates.

Third, conduct a bias audit before training begins. This means stratifying the training data by demographic characteristics such as geography, language, filing channel, and claim type, then checking for systematic gaps. A 12% null rate overall is manageable. A 12% null rate concentrated entirely among paper filers is a signal that the model may learn to underweight that population.

Fourth, establish access logging for training data. Before the vendor touches the dataset, configure audit logs that record every access by user, timestamp, and operation. This is both a security control and an accountability record.

Fifth, document retention requirements specific to the AI system. Work with your agency counsel to determine whether training data must be preserved, and for how long, as an administrative record under your jurisdiction's record-keeping statutes.

Data governance is not a constraint on AI adoption. It is what makes AI adoption defensible when someone with subpoena power asks how the system worked.

Worked Case: Benefit Determination Data

An agency maintains data on benefit applicants spanning 20 years. Some of the historical data is biased, reflecting older assessment methods and older practices. New AI systems will be trained on it. The governance decisions fall out along three of the dimensions above.

On quality, historical data is poor, with missing values and inconsistent categorisation, while recent data is better. The decision is to train primarily on recent data and validate separately against historical data, using the historical set to understand the bias rather than to teach the model. On bias, historical records reflect the demographics of applicants from that era while recent records reflect current demographics. The decision is to require fairness testing against recent data distributions, and to flag any model that performs worse on historical distributions as potentially carrying historical bias forward. On retention, the old data still has to be kept for compliance and audit. The decision is to segregate the two purposes: retain for compliance, train on recent data only, and never let the retention obligation become an argument for using old records in a model.

The result is a model that reflects current data and current populations, with historical bias documented rather than embedded. Note what the segregation buys and what it does not. Separating retention from training removes one pathway by which historical bias enters the model. It does not certify the model as fair, because recent data can carry current bias just as faithfully as old data carried old bias. Fairness testing still has to happen on every release.

Worked Case: Fraud Detection Data

The second case is closer to Colleen's own. An agency builds fraud detection from historical fraud cases and patterns, but some of those patterns are outdated because sophisticated fraudsters have evolved. Three quality dimensions drive the decisions.

Timeliness comes first: some training data is five or more years old and describes fraud schemes that have since changed, so the decision is to retrain regularly, quarterly in this agency's design, on recent data. Representativeness is the subtler problem. The training data was collected reactively, after fraud was detected, which means it represents only the fraud that was caught and says nothing about the fraud that succeeded. The decision is to supplement with behavioural data to probe patterns of undetected fraud, and to state that limitation openly in the model's explanations rather than let users assume full coverage. Completeness closes it out: some cases have incomplete investigation results, so only cases with complete outcomes are used.

Documentation is what makes the package defensible. The collection methods, the time periods covered, and the known limitations are all written down. The result is a model that reflects current fraud patterns, users who understand what the model cannot see, and a retraining cadence that keeps the system current. The limitation about undetected fraud is worth restating, because it is the one that gets forgotten: a fraud model trained on caught fraud measures the agency's past detection, and treating its score as a measure of actual fraud risk overstates what the data can support.

The Force Multiplier Effect

Colleen's agency eventually produced the documentation the vendor needed. The fraud detection pilot launched and has since flagged approximately $2.1 million in potentially improper payments annually. Against that annual rate, the four-month delay cost the agency an estimated $700,000 in improperly paid claims that would have been caught sooner, which is the four twelfths of the annual figure that the delay covered.

Good data governance compounds in value. Documentation created for one AI project serves the next. Bias audits conducted for one model create templates that apply to every subsequent model using the same constituent data. Governance also makes every other governance function work better: privacy impact assessments are sharper when provenance is known, risk classification is more accurate when quality is measured, and minimum practices are easier to implement on well-governed data. The governance investment does not scale with the number of AI projects. The return does.

Anti-Patterns

Data quality ignored. Building AI on poor quality data on the assumption that the model will sort it out. Models do not fix bad data; they amplify it and lend it an authoritative surface. Avoid by requiring a data quality assessment and remediation before model development starts, not as a finding that surfaces after deployment.

Provenance unknown. Building models without knowing where data came from or how it was collected. When problems emerge later, and they do, you cannot explain why. Avoid by documenting provenance thoroughly and understanding collection methods and their potential biases before training.

Access sprawl. Granting broad access to sensitive data because fine-grained access management is difficult. Everyone can see everything, so governance is easy and exposure is total. Avoid by implementing real access controls. It is harder. It is also the control that limits the blast radius when something goes wrong.

Retention confusion. Keeping data forever in case it is useful, or deleting aggressively and then needing it for an audit. Both failures come from the same root, which is having no documented policy. Avoid by setting retention periods from legal requirements and program need, and by writing down the rationale for each one.

Reuse without reassessment. Data collected for one purpose is fed to an AI system for a different purpose with no fresh assessment of whether that is appropriate. Training is nearly always a new purpose. Avoid by applying purpose limitation as a gate, and by reassessing lawful basis, notice, and consent expectations whenever the use changes.

Treating identifier removal as anonymisation. A team strips names and Social Security numbers from a training extract and describes the result as anonymised, sometimes with a claim that no individual can be identified. That claim goes beyond what the step supports. Removing direct identifiers reduces risk; it does not by itself prevent re-identification through quasi-identifiers or linkage to another dataset. Avoid by naming what was actually removed, testing re-identification risk before the data leaves its original control boundary, and keeping access controls in force on de-identified extracts.

Governing data and AI on separate tracks. The data office manages quality, access, and retention; the AI team builds models; neither owns the join. Avoid by making data governance an explicit input to AI governance, with the same named owners appearing in both processes.

Practice Prompts

Data quality assessment. Pick one AI system in your agency, in production or in planning. Assess its data across accuracy, completeness, timeliness, representativeness, and consistency. For completeness, do not stop at the overall rate; break missingness down by region, language, and filing channel. Document what you find, including the things you could not determine.

Provenance mapping. Map provenance for that same system. Where does each input come from, who collected it, under what authority, and what transformations occur between the source system and the model? Produce a lineage diagram simple enough that your general counsel can read it without a data engineer present.

Access design. Design access controls for that system's data. Assign each role the narrowest data it needs, specify how access is requested and approved, and state how you would demonstrate to an auditor that the controls were enforced rather than merely documented.

Retention policy. Draft a retention schedule for the dataset and for the trained model separately. State how long each is kept, on what legal basis, when archival applies, and how deletion would be verified. Write down what your agency would actually do if a record had to be removed from a model already in production.

Governance implementation. Sketch the data governance function your organisation would need to run all of the above. What structures, what roles, what policies, and critically, who has the authority to stop a model from training on data that fails the standard?

Reflection

Assess your organisation's data governance maturity honestly. Do you have documented data quality standards, and for the AI systems you already run, are those standards being met in practice rather than on paper? Do you know the provenance of the data in those systems well enough to trace it back to source? Are access controls implemented and enforced, and is the access that exists today actually appropriate? Do you have retention policies, and are they followed? Finally, name the single largest data governance risk in your organisation and the one change that would most reduce it.

Glossary

Data governance. The framework for managing data quality, access, security, and lifecycle across an organisation.

Data provenance. Documentation of where data came from, how it was collected, and what transformations have been applied to it.

Data lineage. Tracking of how data flows through systems and is transformed on the way to a model or a report.

Data quality. The characteristics of data, including accuracy, completeness, timeliness, and consistency, that determine its fitness for a particular use.

Retention policy. The rules governing how long data is kept before it is archived or deleted, and on what authority.

Purpose limitation. The principle that data collected for one purpose should not be used for an unrelated purpose without additional safeguards and a fresh assessment.

Data classification. The assignment of data to tiers such as public, internal, sensitive, and restricted, with access controls matched to each tier.

De-identification. The removal or masking of identifiers to reduce the risk that a record can be attributed to an individual. It lowers risk rather than eliminating it, because quasi-identifiers and linkage to other datasets can still enable re-identification.

Linkage. The combination of data from multiple sources into a joined dataset, which can create identifiability and correlations that none of the source datasets carried alone.

Data governance sits underneath the rest of the governance curriculum, so it connects outward in several directions. NIST AI RMF: The GOVERN Function and NIST AI RMF: MAP, MEASURE, MANAGE depend on it directly: MAP requires knowing what data you hold, MEASURE requires data good enough to measure with, and MANAGE requires data reliable enough to act on. Your Agency's AI Governance Structure is where the roles and authorities described here get assigned to real people.

For the quality and bias material, Data Quality and AI Performance goes deeper into measurement and remediation, and Understanding AI Bias works through how historical bias becomes model behaviour. On the privacy side, Privacy Impact Assessments for AI Systems and PII and AI: The Bright Red Lines pick up the handling obligations this lesson only frames. Minimum Risk Management Practices shows what these controls look like when they become a floor rather than an aspiration.

Closing

Colleen lost four months to a question that should have taken an afternoon. The cost was not the documentation she eventually produced; it was that the documentation had to be reconstructed from memory, migration notes, and guesswork, and that some of it could not be reconstructed at all. That is the ordinary shape of a data governance failure in government. It does not announce itself as a scandal. It shows up as a delayed pilot, a vendor waiting on a sign-off, and a legal team unwilling to approve what nobody can describe.

The next piece of the puzzle is organisational rather than technical. Data governance only happens if someone owns it, and ownership means structures, roles, and processes with authority attached. Everything in this lesson assumes a decision-maker who can say that a dataset does not meet the standard and that training will wait. Building that authority is the subject of governance structure, and it is what turns the practices described here from a reading list into an operating discipline.

Key Takeaways

  • Documentation gaps are the most common AI delay. Agencies that cannot answer basic provenance questions, such as where data came from and how it has changed, will lose months of AI project timeline reconstructing history that should have been recorded as it happened.
  • Data quality has an equity dimension in government. Missing data concentrated among specific populations, such as rural applicants, non-English speakers, and paper filers, trains models that systematically underperform for those groups, creating Title VI and equity obligation exposure. The pattern of missingness matters more than the overall rate.
  • Lineage documentation enables auditability. Every AI system deployed by a government agency must be traceable from its outputs back through every data transformation to the original records, or it cannot be defended before an Inspector General or in a civil rights complaint.
  • Access controls for training data require audit logging. Before any contractor or vendor touches sensitive training data, configure audit logs. Access should be role-based, justified, approved on legitimate need, and revoked promptly when personnel change roles.
  • Retention of training data is an emerging legal obligation, and deletion is structurally hard. Courts and oversight bodies are beginning to treat AI training data as an administrative record. Data absorbed into a trained model cannot simply be deleted from it; removal generally means retraining. Work with counsel before deleting training datasets from completed projects.
  • De-identification reduces risk rather than removing it. Stripping names and identifiers lowers exposure, but quasi-identifiers and linkage to other datasets can still enable re-identification. Keep controls in force on de-identified extracts and test re-identification risk before release.
  • Bias audits should happen before training, not after deployment. Stratifying training data by demographic characteristics before the model is built catches representational gaps early, when they can be corrected, rather than after constituent appeals begin accumulating.
  • Purpose limitation is the principle AI violates most often. Training is nearly always a new purpose for data collected for something else. Reassess appropriateness, authority, and notice every time the use changes.
  • Data governance compounds in value across projects. Documentation, quality standards, and bias audit templates created for one AI project reduce the cost of every subsequent project, and they make privacy assessments, risk classification, and minimum practices work better at the same time.

Frequently Asked Questions

Our data is imperfect and always will be. Does that mean we cannot use AI?

No. It means you have to know how it is imperfect and design around that. The failure mode is not using flawed data; it is using flawed data without knowing which flaws it has. An agency that documents a 12% null rate concentrated among paper filers can decide to fix the intake process, weight the training set, restrict the model's scope, or accept the limitation openly. An agency that never measured cannot do any of those things, and will discover the problem through an appeal.

Who should own data governance for an AI project, the data team or the AI team?

Both, with an explicit seam between them. The commonest structural failure described in this lesson is data governance and AI governance running as parallel tracks that never meet. Whatever the org chart says, the practical test is whether a single named person has the authority to say that a dataset does not meet standard and that training will wait until it does.

If we remove names and Social Security numbers, is the dataset anonymised?

It is de-identified to that extent, which is a real risk reduction and not the same as anonymisation. Direct identifiers are only one route to an individual. Date of birth, zip code, employer, filing channel, and rare case characteristics can re-identify people, particularly when the extract can be linked to another dataset. Describe precisely what you removed, test re-identification risk, and keep the extract under access control rather than treating it as releasable.

Can we delete a person's data from a model that has already been trained on it?

Not by deleting the source record. Removing a row from a database does not remove its influence from a model's learned parameters, and there is no simple delete operation that reaches inside a trained system. The available approaches are structural: retrain on a dataset that excludes the records, or design with privacy-preserving techniques from the outset. Decide which of those your agency can actually execute before you commit to a deletion promise in a notice or a contract.

How long should we retain training data?

The lesson cannot give you a period, and neither should anyone who does not know your program. The standard is the minimum necessary for a legitimate purpose, bounded by your jurisdiction's record-keeping statutes and any litigation hold. What is specific to AI is that the model itself may create a retention obligation, because reconstructing why a system produced a particular output three years ago requires the inputs it saw. Settle that question with counsel before deployment rather than at the first appeal.

What is the smallest useful first step if we have no data governance at all?

Inventory the data behind the AI project in front of you. What exists, what format, what time period, from which systems. That inventory is the minimum viable provenance record, it takes days rather than quarters, and it is the artifact Colleen's agency spent four months building under pressure because nobody had built it in advance.