←
AI for Government
Capable · M37 · lesson 37 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Testing and Validating AI Systems
📖
now learning

Testing and Validating AI Systems

15 min

Bartholomew Nwachukwu-Reyes was the quality assurance lead for a state department of motor vehicles, and he had tested government software for fifteen years. His craft was deterministic: you write a test case, you specify the expected output, you run it, and it either passes or fails. So when the department deployed an AI system to verify the authenticity of submitted documents, detecting forged proof-of-residency papers, Bartholomew wrote his usual test suite, ran it, watched every case pass, and signed off. Four months later the system was quietly rejecting a disproportionate number of legitimate documents from one immigrant community, and an advocacy group had the data to prove it. Every one of Bartholomew's tests had passed. None of them had asked the questions that mattered. He had tested an AI system as if it were ordinary software, and ordinary software testing had been blind to exactly the failures that AI systems produce.

This lesson is about testing and validating AI systems in government: why it differs fundamentally from traditional software testing, and what a complete validation program actually requires. The central truth Bartholomew learned the hard way: traditional testing checks whether a system does what it was specified to do, but AI systems can do exactly what they were specified to do and still cause serious harm, because their behavior depends on data and emerges statistically. Validating AI means testing for performance, fairness, robustness, and operational readiness, and then continuing to test after launch.

Why AI Testing Is Different

Traditional software has a correct answer for every input, and testing verifies the system produces it: same test, same result, every time. AI systems are probabilistic. They are right most of the time and wrong some of the time, and "wrong" is part of normal operation rather than a bug to be eliminated. This breaks traditional testing in three ways. First, there is no single expected output to assert against; there is a distribution of behavior to evaluate. Second, a system can pass every functional test and still be unfair, brittle, or biased. Third, an AI system that tests perfectly at launch can degrade as the world changes, so testing cannot be a one-time gate before deployment. It must be continuous.

Bartholomew's document-verification system passed his functional tests because those tests asked "does it correctly classify these specific documents?" using a small, clean set he had assembled. They never asked "how does it perform across the full diversity of real documents?" or "does it fail more often for some groups?" or "is it still accurate six months after launch?" Those are the AI questions.

The government stakes make the gap consequential. An unvalidated benefits eligibility system may deny eligible citizens their benefits. An unvalidated hiring system may discriminate against protected classes. An unvalidated fraud detection system may harass innocent citizens. Agencies also carry legal obligations to ensure systems are fair, accurate and safe: the Administrative Procedure Act, civil rights laws, and agency-specific regulations all require evidence that systems work as intended, and testing is how that evidence gets produced. Accountability to the public runs alongside the legal obligation, because citizens deserve to know that systems affecting their access to services have been validated, and test documentation is part of that accountability.

The Four Dimensions of AI Testing

Accuracy and performance: how good, measured properly

Accuracy testing asks whether the system produces correct outputs, with what frequency, and on what types of case. The metrics that answer it are specific: overall accuracy, the percentage of correct decisions; precision, meaning of the cases the system flags as positive, how many really are; recall, meaning of all actual positive cases, how many the system catches; the false positive rate, cases incorrectly flagged as positive; the false negative rate, cases incorrectly flagged as negative; and how accuracy varies by case type or complexity. Reporting only overall accuracy is how a system with an unacceptable error profile passes review.

Choose the metric that matches the problem. For benefits eligibility you probably care about recall, catching all eligible cases, more than precision, since some false positives are acceptable if they route to manual review. For fraud detection you care about both: precision so you do not falsely accuse people, and recall so you catch actual fraud. For maintenance prediction you probably care about recall, since false alarms are low-cost compared with missed maintenance. The cardinal sin is testing on data that resembles the training data more than reality. Bartholomew's test set was too clean; it underrepresented the worn, photographed, multilingual documents real residents submitted. For a forgery detector, the false-rejection rate, legitimate documents wrongly flagged, matters more than overall accuracy, because that error directly denies service to real people.

Fairness: performance broken out by group

This is the dimension that would have caught Bartholomew's failure. Fairness testing measures performance separately for each relevant subgroup, disaggregating by the characteristics that matter for the decision. Had he measured the false-rejection rate by community, the disparity affecting one immigrant group would have been visible before launch, not four months after. For government AI, fairness testing is not optional polish; for rights-impacting systems it is required, and it connects directly to civil rights law and to OMB M-24-10's mandate to test for disparate impact. An aggregate accuracy number hides exactly the harm that ends up in a complaint.

The questions fairness testing answers are concrete. Does the system have equal accuracy across demographic groups? Are approval and denial rates equal across groups? Do false positive rates vary by group? Would the system face disparate impact liability in civil rights terms? And are there proxy variables that encode protected characteristics indirectly? Several formal criteria exist for judging the answers. Demographic parity means equal approval rates across groups, so 50% approval for group A and 50% for group B. Calibration means predictions are equally reliable across groups, so that when the system says "80% confidence" that means 80% in group A and in group B alike. Two further criteria compare error rates directly: equal true positive rates across groups, so the system catches 85% of eligible cases in both, and equal false positive rates across groups. No single fairness metric is perfect, which is why the choice belongs to legal and civil rights teams working with you rather than to the modelling team alone.

Robustness: behavior under stress and edge cases

Robustness testing probes how the system behaves on inputs at the edges of, or outside, its training distribution: unusual but legitimate documents, poor-quality scans, inputs it was never designed for, and deliberately adversarial inputs designed to fool it. It asks how the system handles missing data, how it handles inputs outside its training distribution, how it handles intentionally malicious inputs, what happens when the data environment changes dramatically through a pandemic, an economic shift or a policy change, and how the system fails when it fails, gracefully or catastrophically.

Concrete edge cases sharpen the exercise. An eligibility system trained on a population of 1 million suddenly receives 10 million applications: how does it perform? A hiring system trained on applications received from January to March meets completely different patterns in September: does it generalize? A fraud detection system encounters a type of fraud it has never seen: how does it respond? Citizens encounter edge cases in the ordinary course of their lives, and bad actors craft malicious inputs deliberately, so the system has to handle both. A document-verification system that has never been tested against a clever forgery, or against a legitimate but unusual document format, is untested where it matters most. Robustness testing also asks what the system does when it is uncertain: does it fail safe by flagging for human review, or fail silently by making a confident wrong call?

Operational readiness: can this be run responsibly

The fourth dimension asks whether the system is ready for deployment with appropriate governance around it, and it is the one technical teams most often skip because none of it is about the model. Can the system be monitored, meaning are we collecting the right data to see how it is doing? Are human override mechanisms in place? Can the system be audited, meaning is the decision trail clear? Are incident response procedures documented? Is rollback possible if the system fails? Are success metrics being tracked? And is there a process for updating the model as conditions change? A system that scores well on the first three dimensions and fails this one is a system nobody can steer after launch.

Designing Test Data and Test Sets

Your test data determines what you can actually know about your system: bad test data produces unreliable test results no matter how carefully you run them. Four principles govern it. Test data must be separate from the training data. It must be representative of real-world usage. It must include sufficient examples of edge cases and underrepresented groups. And it must itself be clean and accurately labelled, because mislabelled test data makes test results meaningless rather than merely noisy.

Representativeness is where most test sets fail. If the population the system will serve is all citizens applying for benefits, the test set should include demographics matching the actual application population. The characteristic error is building a test set of "typical" cases only, missing edge cases and minority populations entirely. The fix is to get demographic data on your actual population, build the test set with proportional representation, and then deliberately oversample underrepresented groups so that you have enough examples to measure performance on small groups at all.

Stratified testing is what turns that test set into findings. Rather than reporting one number, test the system separately on different population segments. A typical result reads: overall accuracy 82%, accuracy for Group A 80%, for Group B 84%, for older applicants 78%, and for applicants from rural areas 81%. That breakdown shows where the system performs worse and where fairness concerns live, and it is invisible in the aggregate figure that a single overall accuracy number would have reported.

On size, the working rules of thumb are these: at least 100 examples of each important subgroup, and a test set of at least 1,000 cases for low-stakes systems and 5,000 or more for high-stakes systems. Read those as floors rather than as certifications. Below them, a subgroup result is too noisy to interpret at all; above them, you have a number worth arguing about, not a guarantee that the number is right. A subgroup result built on 100 cases can still move substantially with a different sample, which is an argument for oversampling and for repeating the measurement, not for treating the threshold as a finish line.

How Much Validation, by System Type

System typeExamplesValidation requirementsTypical duration
Makes final decisions (high governance burden)Benefit eligibility determination, loan approval, hiring decisionsAccuracy testing on a representative population with 95% or higher confidence in results; fairness analysis across all protected classes, documenting any disparities; edge case testing demonstrating robustness to unusual inputs; human review testing showing that reviewers can effectively override the system; and a plan for continuous post-deployment fairness and accuracy monitoring2-4 months, with rigorous documentation
Flags cases for human review (medium governance burden)Fraud detection, anomaly flagging, prioritization systemsAccuracy of flagging, meaning what percentage of true cases are caught; the false positive rate, meaning how many false alarms; fairness analysis of whether flagging rates vary appropriately by demographics, which may not require perfect parity where populations legitimately carry different risk profiles; and operational validation of whether reviewers find the flags useful and what percentage of flags result in action4-8 weeks
Provides recommendations (lower governance burden)Informational chatbots, decision support toolsAccuracy on representative questions; fairness analysis so that recommendations are not biased; and basic robustness, including what happens on off-topic queries2-4 weeks

One caution on the first row. "Human review testing" demonstrates that a reviewer is able to override the system, which is a necessary property and a weak one. It says nothing about whether reviewers do override, how much time they have per case, or whether they would notice a wrong recommendation in the first place. Measure the override rate in production alongside the capability test, because an override mechanism that is never exercised looks identical, in a validation report, to one that is not needed.

Fairness Testing, Step by Step

Fairness testing is increasingly critical for government AI and it is not optional. Run it in five steps. First, define what fairness means for your system, working with civil rights and compliance teams, considering legal obligations under civil rights laws and equal protection principles, and stating explicitly what disparities you can tolerate. Second, identify protected characteristics and proxy variables. The obvious characteristics are race, gender, age and disability status. The proxies are harder: ZIP code often correlates with race, a name often signals gender or ethnicity, and the school someone attended may correlate with socioeconomic status. A system can encode a protected characteristic indirectly without ever using it.

Third, measure accuracy and rates by demographic group: calculate accuracy separately for each group, calculate approval and denial rates by group, calculate false positive and false negative rates by group, and compare across groups. Fourth, assess the disparities you find. Are the differences statistically significant or noise? Are they within acceptable tolerance? Can they be explained by legitimate factors, such as genuinely different risk profiles, or do they indicate bias? Fifth, document findings and remediation. If no disparities appear, document that the analysis was done and what it showed. If disparities exist but are acceptable, document the disparity and the justification. If they are unacceptable, plan remediation before deployment rather than after.

A worked example makes the judgment concrete. A benefits eligibility system tests at 82% accuracy overall, breaking down to 84% for White applicants, 79% for Black applicants, and 81% for Hispanic applicants. That is a disparity of 5 percentage points between White and Black applicants. Is it statistically significant? With 5,000 test cases, probably yes. Is it within acceptable tolerance? Probably not; that is a meaningful gap. What is causing it, and can it be fixed? Both need investigation, since the cause could be training data bias or could reflect genuinely different documentation rates between populations, and the remedy differs completely depending on which it is. What the analysis does not permit is deploying while the question stays open.

Degradation and Continuous Validation

Because AI degrades as the world drifts from its training data, new document formats, new forgery techniques, demographic shifts, validation must continue after deployment. This means production monitoring: tracking the system's real-world performance and fairness on an ongoing cadence, with defined thresholds that trigger investigation, retraining, or rollback. A system validated only at launch is validated against a world that no longer exists. Bartholomew's failure festered for four months precisely because nothing was watching after go-live.

Be precise about what monitoring delivers, though. Continuous monitoring catches issues that lab testing missed, but only issues on the metrics somebody chose to track, broken out by the groups somebody chose to break out, at the frequency somebody chose to run. A disparity affecting a community that is not one of your reporting categories will not appear on the dashboard however long you watch it. Deciding what to monitor is therefore a design decision with the same weight as the model choice, and it should be revisited whenever the population or the use changes.

Building a Validation Program

Bartholomew rebuilt the department's AI testing into a program with gates at every stage rather than a single sign-off. Before deployment: a representative test set reflecting real-world conditions; performance metrics appropriate to the task with explicit attention to the costlier error; fairness testing disaggregated across relevant subgroups; and robustness testing including edge cases, low-quality inputs, and adversarial attempts. At deployment: a defined human-review fallback for low-confidence cases, so the system fails safe. After deployment: continuous monitoring of performance and fairness, with named thresholds that trigger action and a named owner responsible for acting.

His concrete artifact was an AI validation report, a standard document completed before any AI system goes live and updated on each monitoring cycle. When he ran the redesigned document-verification system through this program, the subgroup fairness test surfaced the disparity immediately; the fix, retraining on a more representative document set and adding a human-review step for low-confidence rejections, brought the false-rejection rates within tolerance across communities before anyone was harmed.

Documenting Test Results for Deployment

Test documentation serves two audiences at once: internally it answers whether the system is ready, and externally it is the record of accountability to the public and to oversight bodies. A minimum document for a government AI system carries seven parts.

  • Executive summary. What was tested, the key findings, and whether the system is ready for deployment: yes, no, or conditional.
  • Test methodology. What test set was used, how large it was and what its composition was, and who performed the testing.
  • Accuracy results. Overall accuracy with confidence intervals, accuracy by demographic group, accuracy by case type, and comparison to the baseline or to human performance.
  • Fairness analysis. Demographic performance analysis, the disparities identified, an assessment of legal risk, and mitigation strategies where disparities exist.
  • Robustness testing. Which edge cases were tested, which adversarial examples were tried, and how the system behaved.
  • Limitations and caveats. What the system is not expected to do, the populations or scenarios where accuracy is lower, known failure modes, and data staleness issues.
  • Governance readiness. Whether monitoring mechanisms are in place, whether human override procedures are documented and tested, whether incident response procedures exist, and whether audit trail mechanisms are in place.

Bartholomew's report followed that shape and added one thing worth copying: it recorded the test data and why it was representative, not merely its size. It was signed by the system owner and reviewed by the governance board. Test documentation is what lets an agency answer, long after everyone involved has moved on, what exactly was validated and against what evidence, which is precisely the question that arrives with a complaint attached.

Anti-Patterns

  • Testing only on best-case data. Testing runs on clean, representative data; the system launches, meets real data that is messier and full of edge cases, and accuracy drops 15%. The system is now in crisis. Test on real data from the production environment where possible, and deliberately include realistic data quality problems in the test set.
  • Fairness testing as an afterthought. Accuracy testing is thorough, fairness testing happens in the last week before deployment as a checkbox exercise, and the results show disparities that should have been addressed months earlier. Run fairness testing in parallel with accuracy testing from the beginning, and make fairness part of the success criteria rather than a gate at the end.
  • Test results ignored. Testing reveals fairness problems or accuracy below threshold, the team acknowledges the findings, and deploys anyway on the theory that it can be fixed in production. Make test results binding: do not deploy until results meet the agreed criteria, and if you deploy despite findings, document that explicitly with named stakeholder sign-off.
  • Vanishing test data. Testing completes, the system deploys, the test data is deleted, and six months later nobody can answer what accuracy was validated. Archive test data and all test results; you will need them for audits, incident investigations, and future maintenance.
  • Over-optimization on the test set. The team finds underperformance on a specific demographic group and tweaks the model to improve that group's numbers on the test set. When deployed, the tweak does not generalize, because it fitted the sample rather than the problem. Do not tune to individual test results; if accuracy differs significantly for a group, find and address the root cause rather than the symptom.
  • Reading a passing test as a property of the system. Testing produces evidence that a system behaved a certain way on a certain sample under certain conditions. It does not make the system accurate, fair or robust, and a validation report is a description of an experiment rather than a certificate. Write conclusions that say what was measured and on what, so the next reader can see the boundary of the claim.
  • Treating a sample-size threshold as sufficiency. Hitting 100 examples in a subgroup or 5,000 cases overall clears a floor below which the numbers cannot be read at all. It does not make a subgroup estimate stable, and it does not license a confident statement about a group that only just cleared the bar. Oversample the small groups, repeat the measurement, and report the uncertainty alongside the point estimate.
  • Accepting an override mechanism in place of oversight. A validation report that records "human override available" has documented a capability, not a control. If nobody measures how often reviewers override, how long they spend per case, or whether they can tell a wrong recommendation from a right one, the human in the loop is an assumption the report is quietly relying on.

Practice Prompts

  • Test plan design. Choose a government AI system, real or hypothetical, and design a comprehensive test plan. What are the four dimensions you will test? What test set will you use, in size, composition and representativeness? Which accuracy metrics matter most for this problem, and why? What fairness metrics will you measure? What edge cases must you test? And what would your pass and fail criteria be?
  • Fairness analysis. Take a hypothetical system with these accuracy results: overall 85%, Group A 88%, Group B 82%, Group C 83%. Is there a fairness concern? What is the magnitude of the disparity? What questions would you ask to understand the root cause? And would you deploy with these results, with your reasoning either way?
  • Test documentation. Write the executive summary and limitations section for a benefits eligibility AI that achieved 81% accuracy overall, 79% accuracy for historically disadvantaged groups, showed edge case issues with cases over 10 years old, and achieved 85% accuracy on straightforward eligibility determinations. Say plainly what a reader should not conclude from those numbers.
  • Break your own test suite. Take a test plan you wrote in the first prompt and describe a real-world failure that would pass every test in it. Then write the test that would have caught it. Most validation programs are one such exercise away from finding their largest blind spot.

Reflection

Reflect on a time when you deployed something, software, a process, or a policy, that turned out to have unforeseen problems. What would advance validation have caught? What specific test would have revealed the problem before deployment, and what evidence would that test have produced? Then ask the harder version: was the problem genuinely unforeseeable, or was it foreseeable by somebody who was not in the room when the tests were designed? The second answer is more common than the first, and it points at who your test design process is missing rather than at which test you forgot to run.

Glossary

  • Test set. A collection of data examples used to evaluate an AI system, kept separate from the training data used to build it.
  • Stratified testing. Testing a system separately on different population subgroups to identify whether performance varies by demographic group or case type.
  • Precision. Of the cases the system flags as positive, the proportion that really are positive.
  • Recall. Of all actual positive cases, the proportion the system catches.
  • Fairness metric. A quantitative measure of whether an AI system treats demographic groups equitably, such as demographic parity, equalized odds, or calibration.
  • Calibration. The property that stated confidence matches actual accuracy, and that it does so equally across groups.
  • Disparate impact. A pattern of outcomes where a facially neutral practice, such as an AI system, falls disproportionately and negatively on members of a protected class.
  • Edge case. An unusual or extreme input outside the system's normal operating range, which tests the limits of its robustness.
  • Proxy variable. A feature that indirectly encodes a protected characteristic, such as ZIP code standing in for race, without the protected characteristic ever being an input.

Closing

Testing and validation are not quality control activities. They are governance mechanisms. They produce the evidence that a system is ready for deployment and safe to use, and they create accountability by documenting how a system works and what results it produced. In practice that means testing starts early, in sprint 2 rather than sprint 8; results are transparent to stakeholders; fairness testing carries the same weight as accuracy testing; results are binding, so systems do not deploy without meeting the criteria; and results are preserved for audits and future reference.

The agencies that get this right can point to documentation showing what was validated, on what data, by whom, and with what caveats, and can defend their decisions when questioned. That is a real advantage, and it is worth stating carefully: a thorough validation program does not make a system fair or accurate. It tells you what the system did on the evidence you gathered, shows you where it fails, and gives you a defensible record of the judgment you made with that knowledge. Bartholomew's tests had all passed. The change that mattered was not running more of them; it was asking the questions whose answers could have stopped a launch.

Key Takeaways

  • Functional testing is necessary but nowhere near sufficient for AI. A system can pass every functional test and still be unfair, brittle, or biased. AI validation tests behavior, not just correctness.
  • Test across four dimensions. Accuracy, fairness, robustness, and operational readiness. Accuracy alone is insufficient, and readiness is the one technical teams most often skip.
  • Choose metrics that match the problem. Recall matters most where missing a case is the costly error; precision matters most where a false accusation is. Report the error that actually harms people, such as the false-rejection rate.
  • Test data quality determines what you can know. Test data must be separate from training data, representative of real usage, inclusive of edge cases and underrepresented groups, and accurately labelled.
  • Stratify, do not aggregate. Overall accuracy hides subgroup disparities. Measuring by group is what catches the harm before it becomes a complaint, and is required for rights-impacting systems.
  • Sample-size rules of thumb are floors, not certifications. At least 100 examples per important subgroup and at least 1,000 cases overall, or 5,000 and up for high-stakes systems, get you numbers worth reading rather than numbers you can rely on.
  • Robustness testing probes the edges. Unusual inputs, poor-quality data, volume shocks, and adversarial attempts reveal where a system breaks, and good systems fail safe to human review when uncertain.
  • Validation is continuous, not a launch gate. AI degrades as the world drifts. Monitor performance and fairness in production with thresholds that trigger action and a named owner who acts, and remember monitoring only reports on what you chose to watch.
  • Make results binding and produce a validation report. A standard, signed document covering test data and its representativeness, performance, fairness, robustness, limitations, fallback, and governance readiness turns testing from a one-time sign-off into an accountable program.

Frequently Asked Questions

How much accuracy is enough to deploy? There is no universal number, and any figure quoted without a decision context should be treated as a warning sign. What matters is the error profile against the specific harm: a system at 95% accuracy that concentrates its errors in one community is worse than a system at 88% that distributes them evenly, and a fraud flagger with high recall and terrible precision generates a pile of false accusations while looking successful. Set the acceptable threshold during requirements, per metric, and per subgroup, before anyone has a result to defend.

We found a disparity between groups. Does that mean the system is biased? It means you have a finding that needs a cause. The disparity might reflect training data bias, or it might reflect a genuine difference such as different documentation rates between populations, and the remediation is completely different depending on which. What is not available is deploying while the question stays open, or documenting the disparity and moving on. Assess significance, investigate the cause, and decide with civil rights and legal colleagues whether the gap is within stated tolerance.

Can we deploy and validate in parallel to save time? Only if you are willing to accept that the citizens affected during that window are the test set. If validation later reveals a problem, the harm has already happened and your record shows the agency chose the schedule over the evidence. Where a genuine business reason requires a phased launch, limit the exposed population deliberately, monitor it far more intensively than steady state, and document the decision and its sign-off explicitly.

Our test set is too small for one subgroup. What do we do? Say so in the report, and say what it means: a subgroup that thin gives you an estimate that could move substantially with a different sample, so you cannot make a confident claim about that group either way. Oversample the group deliberately if more data can be obtained, and if it cannot, treat the group as an unresolved risk with heightened production monitoring rather than reporting a number that reads as reassurance.

How often should we re-test after deployment? Frequently enough that a degradation trend is visible before it becomes a pattern of complaints, with a cadence set in the validation report rather than by whoever remembers. Tie it to named thresholds that trigger investigation, retraining or rollback, and to a named owner who acts on them. Then re-test on the calendar even when nothing has tripped, because a drift that changes your test set's representativeness will not announce itself.

Who should perform validation, the build team or someone independent? The build team runs testing continuously through development, because they are the ones who can act on what it shows. Sign-off should not rest with them alone. The team that built the system is the team least likely to design the test that embarrasses it, which is why the validation report goes to a governance board and why "who performed the testing" is a required field in the documentation rather than a courtesy.