Data Quality and AI Performance
Delphine Marchetti is a program analyst at a state housing assistance agency with a $220 million annual budget and about 900 employees. Eight months into her agency's AI pilot, a model designed to prioritize households for emergency rental assistance, she pulled the demographic breakdown of model accuracy. Overall accuracy was 89 percent. For Black and Latino applicants, it was 61 percent. The model had been trained on twenty-two years of housing records: records built on redlining policies, records from zip codes that were systematically under-surveyed, records that reflected which neighborhoods the agency had historically chosen to serve. The AI did not invent the disparity. It learned it and encoded it. Delphine's discovery set off six months of data remediation work that delayed the launch but almost certainly prevented a civil rights complaint that would have cost the agency far more.
The Sensor Network Problem
A weather forecast is only as good as the sensor network feeding it. Sensors clustered in wealthy neighborhoods, sensors that have not been calibrated since 1995, sensors that stopped transmitting in 2008 but were never removed from the network, all of these produce a forecast that looks precise but is wrong for the communities that need accurate prediction most. The readout says "80 percent chance of rain." The readout is confident. The readout is wrong where it matters.
Training data is your sensor network. Every gap, every outdated record, every value that was never collected from certain zip codes becomes part of what the AI learns. The model's high overall accuracy, 89 percent for Delphine's pilot, is the average across all sensors. The 61 percent accuracy for specific demographic groups is what happens when your sensors never covered those neighborhoods reliably.
This is the practical meaning of the oldest phrase in machine learning: garbage in, garbage out. No matter how sophisticated the model, poor training data produces a poor model. What the phrase leaves out is that the failure is usually not visible in the headline metric. It hides inside the average. That is why overall accuracy is dangerous when used alone. Always disaggregate performance by race, age, geography, and income bracket before you trust a single number. An AI system that performs well on average while performing poorly for a specific protected class is not a successful system. In a government context, it is a potential Title VI violation.
Six Dimensions of Data Quality
Before training any model, your agency needs a structured way to evaluate the data it plans to use. Six dimensions capture the most common failure modes.
Completeness asks whether required fields are populated, and whether entire records are missing rather than just fields. A record with no income field either gets dropped from training, shrinking the dataset unevenly, or filled with a placeholder that teaches the model a false pattern. Delphine's agency found that 31 percent of records from three zip codes were missing employment verification data. Those areas had used a paper intake process until 2017 that was never fully digitized.
Accuracy asks whether recorded values reflect reality: whether there are errors, typos, or out-of-range values. An address field populated with a shelter's address rather than a client's true residence is inaccurate. Training a model on inaccurate location data teaches it geography that does not exist.
Consistency asks whether the same concept is represented the same way across systems. If state is recorded as an abbreviation in one place and a full name in another, that is inconsistency. Delphine's agency merged two legacy databases in 2011. One used numeric codes for housing status; the other used text strings. "1" meant "housed" in one system. "Housed" meant the same in the other. The merge never resolved the difference, leaving the model with two representations of one value and no way to treat them as equal.
Uniqueness asks whether records are deduplicated. A client who applied twice under slightly different name spellings appears as two people. A model trained on duplicates over-learns that client's characteristics and produces inflated confidence scores for similar profiles.
Timeliness asks whether the data is current. Delphine's training set included records from 2001. Interest rates, rental vacancy rates, and income patterns have shifted fundamentally since then. A model calibrated on 2001 sensor readings does not know the 2024 landscape. The instrument looks functional. Its output is not.
Validity asks whether values meet defined standards and fall within expected ranges. A rental assistance amount of $47,000 per month is outside any plausible range for your program and should be flagged as an error. A date of birth that would make an applicant 140 years old is invalid. A numeric field containing letters is invalid. Automated range and type checks catch these before they corrupt a model.
How Each Defect Reaches the Model
It is worth being concrete about the mechanism, because "bad data makes bad models" is true and unhelpful. Each dimension damages a model in its own way, and knowing which way tells you whether the defect will show up as a visible error or as a silent bias. Incompleteness usually causes the second kind. Records with missing required fields are either dropped, which quietly reshapes the training population, or filled, which teaches the model a pattern that nobody observed. Neither produces an error message. The model simply becomes better at describing the people whose records were complete.
Inaccuracy is the most direct route. If a location field holds a shelter address rather than a residence, the model learns a geography that does not exist and will confidently apply it. Inconsistency splits a single concept into two, so the model treats "1" and "Housed" as unrelated signals and learns each from half the evidence it should have had. Duplication distorts the weighting: a client recorded twice contributes twice, and profiles resembling that client acquire an unearned confidence that is invisible in aggregate accuracy.
Staleness produces the failure that is hardest to catch, because the model is right about a world that no longer exists. A system trained on records from 2001 encodes the rent levels, vacancy rates, and income distributions of that year, and it will keep applying them fluently against 2024 applicants. Invalid values are the friendliest defect, because a rental assistance figure of $47,000 per month or an applicant aged 140 is obviously wrong and a range check catches it before training. The dangerous defects are the ones that look reasonable.
The pattern worth internalising is that the more plausible a defect looks, the further into the model it travels. Impossible values get stopped at the door. Systematically missing values, inconsistent encodings, and stale-but-coherent records pass every superficial inspection and become part of what the system believes. That is the case for measurement over impression, and for running the checks before training rather than diagnosing them afterwards from a production incident.
Labeling and Representativeness in Practice
Two preparation steps sit between a cleaned dataset and a trainable one, and both are more consequential than the cleaning itself. The first is labeling. A supervised model learns to reproduce the labels it was given, so the label definitions are the specification of correct behaviour, whether or not anyone wrote them down as such. If two caseworkers would label the same case differently, the model learns the average of their disagreement without any record that a disagreement occurred. Before labeling begins, write the definitions down, test them by having several people label the same sample independently, and adjudicate the cases where they diverge. Those divergent cases are usually the ones the model will handle worst.
Who labels matters as much as how. Labels produced as a byproduct of past program decisions carry those decisions' assumptions, which is the historical bias problem arriving through the back door. Labels produced fresh by staff carry those staff members' judgement, which is the labeling bias problem. Neither is avoidable, and the useful response is to record which route your labels took, so that a later reviewer can reason about what the ground truth actually represents rather than treating it as fact.
The second step is checking representativeness, and it is a different question from completeness. A dataset with no missing values can still describe a population that does not match the one your program serves, because the people who reached the intake process are not necessarily the people who were eligible for it. Compare the composition of the training set against the eligible population on the characteristics that matter for your program, document the reference source you compared against, and treat a meaningful divergence as a finding rather than a rounding issue. A model allocates its accuracy to the groups it saw most, so a training set that under-represents a community is a model that will serve that community worst, and it will do so while reporting a healthy overall score.
Where Bias Enters the Data
Quality problems and bias problems are related but not identical, and it helps to separate the stages at which bias enters. There are four, and they call for different responses.
Collection bias comes from how the data was gathered. If a survey ran online only, or if outreach reached some communities and not others, the resulting dataset systematically excludes or over-represents particular groups. Delphine's three zip codes with paper intake are a collection bias problem wearing a completeness costume.
Labeling bias comes from the humans who assigned the labels a supervised model learns from. If caseworkers were more likely to record a particular flag for particular applicants, that judgement is now the ground truth the model is optimising toward. The model will reproduce the labeller's tendencies faithfully, because from its point of view those tendencies are the definition of correct.
Historical bias is present in the events the data records rather than in the recording. Past discrimination in the program produces a dataset in which discrimination is the pattern. Cleaning the data does not remove it, because the data is accurate about an unjust history.
Measurement bias comes from how a quantity is defined and captured. If a proxy measure works better for one group than another, the model inherits that gap. A credit-related proxy, a prior-contact count, or an inspection frequency can all measure the institution's behaviour more faithfully than the applicant's.
Measuring Quality Rather Than Asserting It
Most agencies can tell you their data is "pretty good." Very few can tell you the completeness rate of a required field, and that gap between impression and measurement is where AI projects fail. Every one of the six dimensions can be turned into a number that a script produces on a schedule. Completeness becomes the populated rate per required field, reported both overall and broken out by segment. Uniqueness becomes the proportion of records that survive a deduplication pass. Timeliness becomes the age distribution of the time-sensitive fields, not the age of the newest record. Validity becomes the count of values failing type and range checks. None of this is sophisticated. It is simply written down, run repeatedly, and compared against a standard.
Accuracy and consistency are harder, because neither can be computed from the dataset alone. Accuracy requires an external reference: a sample of records checked against source documents, another system that holds the same fact, or a subject matter expert who can say whether a value is plausible for that program. Consistency requires knowing that two fields in two systems are supposed to mean the same thing, which is a documentation question before it is a data question. Delphine's numeric-code and text-string mismatch could not have been detected by any automated check that did not already know the two columns described one concept.
The segment breakdown is the part teams skip, and it is the part that matters most in government. A single overall completeness figure tells you almost nothing useful. The same figure broken out by zip code, by intake channel, by preferred language, and by application year tells you whether you have a modest uniform gap or a hole shaped exactly like a population your program is obliged to serve. Delphine's team ran the overall number first and saw nothing alarming. The breakdown by zip code is what produced the 31 percent finding and the paper intake explanation behind it.
Representativeness deserves its own measurement, separate from completeness, because a dataset can be complete and still describe the wrong population. The check is a comparison: the demographic composition of the training data against the composition of the population the program is eligible to serve. Where those diverge, the model is being taught that some groups are rarer than they are, and it will allocate its accuracy accordingly. This comparison needs an external reference for the eligible population, which usually means program enrollment data or census-derived estimates, and the choice of reference should be documented because it determines what counts as a gap.
How Government Data Encodes Historical Bias
Government datasets are not neutral recordings of reality. They are records of decisions made by institutions operating under the laws, policies, and assumptions of their era.
Housing databases contain the legacy of redlining, the federally sanctioned practice, ended in 1968 by the Fair Housing Act, of refusing mortgage guarantees in predominantly Black neighborhoods. Those neighborhoods are underrepresented in approval records and overrepresented in denial records. A model trained on that history learns that certain zip codes correlate with denial. It is not wrong about the history. It is perpetuating the history into the present.
Health records for communities with less healthcare access show lower diagnosis rates, not because those communities are healthier, but because they saw providers less often. A model trained on those records learns to underpredict conditions in those populations. Crime data reflects policing patterns. Communities with heavier police presence generate more arrest records, regardless of underlying behavior differences. A risk-scoring model trained on arrest records learns policing patterns, not risk patterns.
None of this is the fault of current data teams. The bias was introduced by historical policy. But it now sits inside your training data, structurally embedded in what the sensors recorded and in what they missed. That is why the response to historical bias is a decision rather than a cleanup. Either exclude the affected data from training, or account for it analytically and document what you did, and in both cases test the trained model against diverse groups afterwards to see whether the bias survived your remedy.
Worked Case: Benefits Eligibility Data
A benefits agency trains a model to predict whether an applicant is likely to be eligible, using five years of historical benefit decisions. Four quality problems show up immediately. Some applications have missing information because applicants did not fill in every field. Income is recorded annually in some records and monthly in others. Applicants who applied more than once appear as separate records because nothing de-duplicated them. And the historical decisions themselves may reflect biases in how decisions were made.
The remediation plan addresses each. Standardise income to a single basis, annual in this case, so the model is not reading a monthly figure as an annual one. De-duplicate the records. Analyse whether historical biases are present and then decide, explicitly and in writing, whether to exclude the affected data or account for it in the analysis. For the missing values, the conventional move is imputation, filling gaps with statistical estimates derived from other applicants.
Imputation deserves a caution, because it is the step most often mistaken for a fix. Imputing a value produces a plausible number where there was no observation. That is useful when values are missing more or less at random. It is actively harmful when missingness is systematic, as Delphine's 31 percent gap in three zip codes was, because imputation smooths over the exact signal you needed to see and hands the model a confident estimate built on nothing the agency ever collected. Before imputing anything, break missingness down by geography, language, and intake channel. If it clusters, the finding is the point, and the fix belongs upstream in intake rather than downstream in a script.
Worked Case: Compliance Monitoring Data
The second case is subtler because nothing in it looks broken. An agency monitors whether organisations comply with regulations, drawing on inspections, self-reports, and complaints. The three sources are not equally reliable: an inspection is a direct observation, a self-report is an interested party's account, and a complaint depends on someone knowing how to complain. Coverage is also uneven, because some compliance areas are monitored consistently and others are not. And there is an inspection bias, because organisations in certain regions are visited more frequently than others.
The approach is to weight data by source, treating inspection findings more heavily than self-reports, and to identify and adjust for the inspection bias so that the analysis does not mistake over-inspection for non-compliance. That adjustment is where the care is needed. A region that is inspected twice as often will generate more recorded violations even if its true violation rate is identical, and a model trained on raw counts will learn to flag that region. Weighting corrects the arithmetic; it does not tell you the true rate in the under-inspected regions, because no one measured it. Say so in the model documentation rather than letting the adjustment imply a precision the data does not contain.
Cleaning Government Data Is Slow and Expensive
Data cleaning in government is not a weekend project. It requires legal review under the Privacy Act of 1974, which governs federal agencies' handling of personal records, and HIPAA, the Health Insurance Portability and Accountability Act, for any dataset containing health information. Subject matter experts, not just automated tools, must review whether a corrected value actually reflects what happened or introduces a new error. Supervisory sign-off is required before altering records that may be subject to FOIA, the Freedom of Information Act, requests. Whatever cleaning you perform, document it, because an auditor's first question is what you changed and why.
Delphine's agency spent fourteen weeks and approximately $180,000 in contractor hours on its initial data cleaning effort. That is a real cost. The alternative, discovering the disparity after deployment, managing a civil rights complaint, and rebuilding public trust, would have cost more.
Siloed systems compound the problem. Benefits data lives in one system. Health data lives in another. Housing data lives in a third. An AI trained only on housing records cannot see that a client is also navigating a health crisis that affects their housing stability. The model is making predictions from a partial sensor network, and it does not know what it cannot see.
Paper records create a further hazard. Many agencies have decades of decisions on paper. When paper is scanned and processed through OCR, Optical Character Recognition, which converts scanned images into searchable text, the process introduces new errors. A handwritten "6" becomes "0." A torn corner drops a digit from an income figure. These OCR errors enter the dataset quietly and degrade accuracy in ways that are hard to detect without manual spot-checking.
Build a Data Quality Scorecard Before Training
The sensor network analogy points toward a concrete discipline: measure data quality before you train, not after you discover a problem in production. Build a data quality scorecard with specific, enforceable thresholds, and treat it as a gate rather than a report.
Reasonable starting thresholds for a government AI training dataset: completeness above 95 percent for required fields; no more than 5 percent duplicate records after deduplication; data staleness below 18 months for time-sensitive fields like income and address; OCR error rate below 2 percent on scanned records; and demographic representation within 10 percentage points of the program-eligible population in each category. If the dataset does not meet those thresholds, do not train. Fix the data first.
Treat those figures as starting points to be justified for your program, not as a standard handed down from outside. What makes a scorecard work is not the specific numbers but the fact that someone wrote them down before the pressure to launch arrived, and that someone has the authority to hold the line when the dataset misses. Note also what a passing scorecard does and does not tell you. It says the data met the thresholds you chose to measure. It does not say the model will be fair, because a dataset can clear every completeness and duplication check and still encode a century of housing discrimination perfectly.
A quality gate is also aimed at a moving target. Set up quarterly data quality reviews once the model is in production. Address fields go stale as people move. Income figures age as conditions shift. New intake processes introduce fields the model was not trained on. Ongoing monitoring is what separates a trustworthy system from one that quietly drifts while its dashboard still shows 89 percent accuracy overall.
Anti-Patterns
Using raw data without auditing it first. The dataset arrives, the deadline is close, and the team trains on it as delivered. Quality problems then propagate silently through every analysis and every model built on that base. Avoid by auditing before use, cleaning deliberately, and documenting every change so an auditor can reconstruct what you altered and why.
Trusting the headline accuracy number. Delphine's model reported 89 percent and was performing at 61 percent for two demographic groups. A single aggregate metric is an average across a population, and averages hide exactly the disparities that create legal exposure. Avoid by making disaggregated performance a required field in every model evaluation, not an optional appendix.
Treating imputation as a fix for missing data. Filling gaps with statistical estimates produces a complete-looking dataset and can erase the evidence of systematic exclusion. Avoid by profiling missingness by geography, language, and intake channel before imputing anything, and by fixing collection upstream when the gap clusters.
Ignoring bias in historical data. Training on historical decisions without auditing them means the model replicates and amplifies whatever the institution did before. Avoid by auditing for bias before training, deciding explicitly whether to exclude or adjust, and testing the trained model across diverse groups to see whether the bias survived.
Cleaning records without legal review. Correcting a value in a system of records is not a purely technical act. The Privacy Act, HIPAA, and FOIA obligations attach to those records. Avoid by routing cleaning plans through counsel and supervisory sign-off, and by keeping a change log that survives staff turnover.
Treating the scorecard as proof of fairness. Passing every quality threshold means the data met the criteria you measured. It does not certify the model. Avoid by keeping fairness testing as a separate gate that a clean scorecard does not substitute for.
Cleaning once and calling it done. Quality degrades continuously as people move, programs change, and intake processes are revised. Avoid by scheduling recurring reviews and by monitoring the same disaggregated metrics in production that you used to approve the model.
Practice Prompts
Audit one dataset against the six dimensions. Take a dataset your agency already uses for analysis or plans to use for AI. Score it on completeness, accuracy, consistency, uniqueness, timeliness, and validity. For each dimension, write down the specific check you ran and the number it produced, not an impression.
Profile your missingness. For the same dataset, break the null rate for every required field down by geography, language, and intake channel. Identify whether missingness is roughly even or clustered. If it clusters, name the operational cause, as Delphine's team traced their gap to a paper intake process.
Trace the bias entry points. Walk one field back to its origin and identify which of the four bias stages could have shaped it: collection, labeling, historical, or measurement. Write one sentence on what the model would learn if that bias went unaddressed.
Draft a scorecard for your program. Set thresholds for completeness, duplication, staleness, OCR error, and demographic representation. For each threshold, write the justification specific to your program, and name the person with the authority to refuse a training run when the data misses.
Design the disaggregation report. Specify exactly how model performance will be broken out before deployment: which groups, which metrics, what disparity would trigger a hold. Decide the trigger now, while nothing is at stake, rather than during a launch window.
Reflection
Think about the datasets your agency would reach for first if an AI pilot were approved tomorrow. Do you know how each was collected, and by whom? Could you produce a disaggregated accuracy report, or only an overall figure? If a constituent group performed at 61 percent while the headline read 89 percent, as Delphine's did, would anyone in your current process see it before launch, or would the first signal be an appeal or a complaint? And if your data failed a quality threshold next quarter, who has the standing to say that training waits?
Glossary
Data quality. The degree to which data is suitable for its intended use. In this lesson it covers six dimensions: completeness, accuracy, consistency, uniqueness, timeliness, and validity. Quality is a property of the fit between a dataset and a purpose, which is why a dataset can be high quality for reporting and unfit for training a model on the same subject.
Data cleaning. The process of correcting errors, removing duplicates, handling missing values, and standardising formats before analysis or training. In government it is a legal process as much as a technical one, because altering records held under the Privacy Act, HIPAA, or a FOIA obligation requires review and a documented change log.
Imputation. Filling in missing values with estimates derived from other records. It is a reasonable technique when values are missing at random, and a misleading one when missingness is systematic, because it manufactures confident values for exactly the population that was never measured and hides the gap that should have prompted a fix upstream.
Deduplication. Removing duplicate records so that one person or entity is represented once rather than several times. Duplicates matter for AI because a model trained on them over-learns the repeated characteristics and produces inflated confidence for profiles resembling whoever happened to appear most often.
Collection bias. Systematic exclusion or over-representation of groups arising from how the data was gathered, such as an online-only intake channel or outreach that reached some communities and not others. It often presents as a completeness problem, and the fix belongs in the collection process rather than in the dataset.
Labeling bias. Bias introduced by the humans who assigned the labels a supervised model learns from. Because those labels define what the model treats as correct, the model reproduces the labellers' tendencies faithfully rather than correcting them, and no amount of cleaning the input fields will surface the problem.
Measurement bias. Systematic error introduced by how a quantity is defined and captured, often favouring or disfavouring particular groups. Proxy measures are the usual culprit, because a proxy that tracks institutional behaviour, such as inspection frequency or prior contact counts, will measure the institution rather than the person.
Historical bias. Bias present in the events the data records rather than in the recording of them, reflecting past discrimination or unfair practice. Cleaning cannot remove it, because the data is accurate about an unjust history, so the response is an explicit decision to exclude the affected records or to account for them analytically.
Disaggregated accuracy. Model performance reported separately for defined subgroups rather than as a single population average. It is the difference between Delphine's 89 percent headline and the 61 percent that was true for two demographic groups, and in a government setting it should be a required field in every evaluation.
OCR. Optical Character Recognition, the conversion of scanned images into machine-readable text. It matters for agencies with decades of paper records because it introduces its own error rate quietly, turning a handwritten digit into a different one in ways that degrade accuracy without producing any obvious signal.
Related Lessons
This lesson is the measurement half of a pair. Data Governance for AI supplies the structures that make quality standards enforceable: provenance, lineage, access control, retention, and the named owner who can hold a launch. Read them together, because a scorecard with no authority behind it is a document rather than a gate.
For the bias material, Understanding AI Bias works through how bias becomes model behaviour, and Bias Detection Tools and Methods covers the techniques for finding it. Testing and Validating AI Systems and Systematic AI Output Validation pick up the disaggregated evaluation discipline that this lesson only introduces. For technical grounding on what the model does with the data you give it, Supervised vs. Unsupervised vs. Reinforcement Learning and How Transformers and LLMs Work come earlier in the sequence.
Closing
Delphine's six months of remediation looked, from the outside, like a delay. From the inside it was the difference between an agency that found its own problem and an agency that had the problem found for it. The 61 percent figure was always true. The only variable was whether anyone would look before the model started making decisions about households in the middle of a housing crisis.
Data quality is where every other AI control gets its footing. A validation regime, a fairness test, a monitoring dashboard, and a human review process all operate on the assumption that the underlying records describe the world. When they do not, each of those controls reports success on a fiction. The next stage of the curriculum follows the system forward through its full lifecycle, from development and training through testing, deployment, monitoring, and eventual retirement, and every stage of that lifecycle inherits whatever the data brought with it.
Key Takeaways
- Always disaggregate accuracy by demographic group. A high overall accuracy score can hide severe underperformance for specific populations. Break down model performance by race, age, geography, and income before approving any AI system for public-facing use.
- Six dimensions define data quality. Completeness, accuracy, consistency, uniqueness, timeliness, and validity each identify a distinct failure mode. Gaps in any one dimension can corrupt what a model learns.
- Bias enters at four stages. Collection, labeling, historical events, and measurement each introduce bias differently, and each calls for a different response. Naming the stage tells you whether the fix belongs in intake, in the label definitions, or in the decision to exclude the data.
- Government data encodes historical policy decisions. Redlining, over-policing, and unequal healthcare access are embedded in agency records. AI systems learn these patterns unless analysts actively identify and address them before training.
- Imputation is not a fix for systematic missingness. Filling gaps with statistical estimates works when values are missing at random and erases the evidence when they are not. Profile missingness by geography, language, and channel before imputing anything.
- Data cleaning in government requires legal review, not just technical tools. The Privacy Act, HIPAA, and FOIA obligations mean that altering records demands supervisor approval, counsel review, and a documented change log, and it takes weeks rather than hours.
- Siloed systems produce incomplete sensor networks. An AI trained on one agency's data cannot see what other systems know about the same client. Cross-system gaps produce models that miss the full picture and cannot report what they missed.
- Build a data quality scorecard with concrete thresholds before training. Set minimum standards, justify them for your program, and refuse to train on data that does not pass. A passing scorecard is evidence about the data, not a fairness certificate for the model.
- Data quality degrades after deployment; schedule recurring reviews. Ongoing monitoring on the same disaggregated metrics is what keeps a model's sensor network calibrated as the real world changes around it.
Frequently Asked Questions
Our overall accuracy is high. Why is that not good enough?
Because an aggregate metric is an average, and averages conceal the subgroups that create legal and ethical exposure. Delphine's pilot reported 89 percent overall while performing at 61 percent for Black and Latino applicants. Nothing in the headline number would have surfaced that. Disaggregated reporting is what converts a comfortable dashboard into an honest one, and in a government setting a model that underperforms for a protected class is a potential Title VI problem regardless of how well it does on average.
If we impute the missing values, does that solve our completeness problem?
It solves the appearance of the problem. Imputation estimates plausible values from other records, which is reasonable when values are missing at random. When missingness is systematic, as it was for the three zip codes that used paper intake, imputation manufactures confident numbers for exactly the population the agency never actually measured, and it removes the visible signal that would have prompted someone to fix intake. Profile the missingness first; the clustering is the finding.
The bias is in the historical record itself. What are we supposed to do about that?
Make an explicit, documented decision rather than an implicit one. The two defensible options are to exclude the affected data from training, or to account for it analytically and state what you did. What is not defensible is training on it silently. In either case, test the trained model across diverse groups afterwards, because a remedy applied to the data does not guarantee the bias failed to survive into the model.
Can we clean our data quickly if we bring in contractors?
Contractors can compress the technical work, not the legal work. Delphine's agency spent fourteen weeks and roughly $180,000 in contractor hours, and the constraint was rarely the tooling. Records covered by the Privacy Act, HIPAA, or a FOIA obligation need counsel review and supervisory sign-off before values change, and subject matter experts have to confirm that a correction reflects what actually happened rather than introducing a new error. Budget the review time as seriously as the processing time.
Where do the scorecard thresholds come from, and can we just adopt them?
They are offered as reasonable starting points for a government training dataset, not as a standard from an external authority. Adopt them as a first draft, then justify each one against your program: what completeness level your intake process can actually sustain, what staleness window your subject matter makes tolerable, what representation gap your eligible population makes acceptable. The value of a scorecard comes from having written the numbers down before launch pressure arrives and from naming who can enforce them.
We passed every quality threshold. Is the model fair?
Not established. A clean scorecard says the data met the criteria you chose to measure, which is a real and useful result. It says nothing about whether the historical patterns inside that clean data are ones you want a model to reproduce. A dataset can be complete, current, deduplicated, and valid while faithfully recording decades of unequal treatment. Fairness testing is a separate gate, and the scorecard does not substitute for it.
Skill.re