Systematic AI Output Validation
Diane Osei is a benefits analyst at a state health and human services department. On a Tuesday, she asked an AI assistant to summarize the eligibility rules for a new childcare subsidy and draft a one-page explainer for caseworkers. The summary was clean, confident, and well-formatted. It also cited an income threshold of $52,000 for a family of four. The real figure was $48,200. The AI had not lied on purpose; it had blended numbers from two different programs into one fluent, wrong sentence. Diane caught it only because something felt off and she checked the actual policy. Three hundred copies of that explainer were about to go to caseworkers who would have quoted the wrong number to real families.
The lesson Diane learned the hard way is the one every government professional using AI must internalize: AI output is a draft, never a fact. A large language model, the kind of AI behind tools like ChatGPT and Copilot, is built to produce text that sounds right. Sounding right and being right are different things, and in government the gap between them can mean a wrongful denial, a broken regulation, or a citizen given false information by their own government. Validation is the discipline that closes that gap, and this lesson covers two levels of it: the desk check you run on AI-drafted text, and the structured validation an agency runs on the outputs of an AI decision system.
Why Fluent Is Not the Same as True
When an AI produces a confident wrong answer, that is called a "hallucination." It is not a rare glitch; it is a built-in feature of how these tools work. The model predicts plausible-sounding text. It has no internal sense of "I am not sure about this number." That is why the dangerous errors are the confident ones, there is no tremor in the voice to warn you. Diane's $52,000 figure was delivered with exactly the same certainty as the parts that were correct.
The stakes scale with what the output touches. AI systems hallucinate, misinterpret inputs, and produce outputs that are technically correct but practically wrong. A system might confidently recommend denying benefits to someone who actually qualifies. It might recommend hiring someone it predicts will fail. It might assess a location as low-risk when conditions are actually dangerous. Government decisions built on AI output affect benefits eligibility, hiring, criminal justice risk assessment, and resource allocation, which is why validation in this setting is not a quality-control nicety but a governance requirement. Without it, bad decisions reach citizens at scale.
There is a legal dimension too. If your AI system produces discriminatory decisions, the agency faces liability. If auditors review your processes and find outputs were deployed without validation, you have violated basic governance, whatever the outcome happened to be. And there is a trust dimension that runs in both directions: citizens extend more trust to systems they know are validated before use, and that trust evaporates the moment they learn outputs went out unchecked.
The Desk Check: Five Checks on AI-Drafted Text
Five checks, in order. Run them on every AI output that matters. We will apply each one to Diane's childcare explainer. Before you start, hold one thing steady: these checks raise the odds of catching a problem. None of them makes an output correct, and a clean pass on all five means the errors you looked for were not there, not that no errors are.
1. Source verification: where did this come from?
Ask the AI to show its sources, then actually open them. If the AI cannot point to a real document, the claim is unverified by definition. Diane's mistake the first time was accepting a clean summary with no sources. The fix: she re-ran the task asking, "Quote the exact policy text and section number for each income figure." When the AI could not produce a real section for $52,000, the number revealed itself as invented. Note what did the work there. It was not the request for sources; a returned citation can be exactly as fabricated as the claim it supports. It was opening the document and failing to find the text.
2. Cross-referencing: does it match the authoritative record?
Check the AI's claim against the original source of truth, the actual statute, regulation, manual, or dataset. Never check AI against AI. For Diane, the source of truth was the published eligibility manual showing $48,200. Cross-referencing is non-negotiable for any number, name, date, dollar figure, or legal citation. These are exactly the things AI gets confidently wrong.
3. Consistency checks: does it agree with itself?
Read the output for internal contradictions. Does the dollar figure in paragraph two match the example in the table? Does the summary's conclusion follow from its own stated facts? Hallucinations often leave fingerprints as small inconsistencies. Diane's draft carried a worked example of a family earning $50,000 that the AI called eligible. Against the real ceiling of $48,200 that family does not qualify, so the example only held together if the invented $52,000 figure were true. Two numbers propping each other up, with nothing underneath either, is a characteristic pattern worth learning to recognize.
4. Expert review: would a human who knows this catch it?
For anything consequential, a qualified human reviews before release. This is not a formality. The expert brings context the AI lacks, that two programs have similar names, that a rule changed last quarter, that one community is affected differently. Diane's senior policy specialist knew both childcare programs well and would very likely have caught the mix-up. That is the honest claim: expert review catches what the reviewer happens to know. It is strong precisely where the expert's knowledge is strong, and blind everywhere else, which is why the rule of thumb is to match reviewer seniority to the stakes and to be specific about what expertise the output actually requires.
5. Stakes check: how much could a wrong answer cost?
Match the depth of your validation to the consequence of being wrong. A brainstorm for an internal meeting needs a light touch. A number going into a citizen-facing benefits explainer needs all four prior checks plus a sign-off. This step keeps validation realistic, you are not auditing every email, you are auditing every output that can hurt someone.
The principle underneath all five is worth stating on its own: the AI's confidence tells you nothing about its accuracy. Treat every number, name, date, and citation as unverified until you have seen it in the real source.
A Validation Worksheet You Can Use Today
Print this or paste it into your workflow. Run it on any AI output before it leaves your hands. The rule is asymmetric on purpose: if you cannot check a box, the output is not ready. The reverse does not hold. A completed worksheet records that someone looked at each of five things; it has never once made a false statement true, and it does not detect the error nobody thought to list.
| Check | What you do | Pass / Fail / N/A |
|---|---|---|
| 1. Source verification | Asked AI for exact sources and opened each one | ☐ |
| 2. Cross-reference | Confirmed every number, name, date, citation against the authoritative original | ☐ |
| 3. Consistency | Read for internal contradictions; figures and examples agree | ☐ |
| 4. Expert review | Qualified human signed off (required if citizen-facing or legal) | ☐ |
| 5. Stakes check | Validation depth matches the cost of being wrong | ☐ |
The highest-risk output types
Some outputs deserve every check, every time. Memorize this short list, these are where hallucinations do the most damage in government:
- Specific numbers: income thresholds, deadlines, fees, statistics, case counts.
- Legal and policy citations: statute sections, regulation numbers, court cases. AI invents these readily.
- Names and quotes: people, organizations, and "direct quotes" that were never said.
- Eligibility and entitlement statements: anything that tells a citizen what they qualify for.
- Historical or factual claims: dates, events, and "what the rule used to be."
Notice that several of Diane's danger zones overlapped inside a single sentence: a specific number, an eligibility statement, and a claim about what the rule is, all delivered in one clause. That is normal. A short AI paragraph can carry several distinct claims, each of which can be wrong independently, and each of which needs checking on its own terms rather than as part of a general impression that the paragraph reads well.
The System Check: Validating What a Decision System Produces
The desk check handles AI-drafted text. An AI system that produces a decision on an individual case needs a different set of five points, because there is more than a document to interrogate: there is an input, a rule the system was supposed to apply, and a record it should have left behind. Run these on the case, not on the prose.
- Input verification. Does the input exist in the source data? Is it complete and well-formed? Does it match expected patterns? Verify the system is actually receiving valid input, for example that a benefit application has all required fields and attachments before the AI processes it. The questions to ask: is the source data authentic, complete, and in the expected format?
- Cross-reference checking. Does the output match information from independent sources? Are there corroborating records, or inconsistencies with known facts? If the AI says an applicant is ineligible, check against prior benefit records: has this applicant ever been eligible before, and was there a policy change in between? The questions to ask: what independent sources can verify this, and do any of them contradict it?
- Logic verification. Does the output follow the stated decision rules? Are factors weighted correctly, and does the reasoning chain hold? For an income-based decision, verify that the decision reflects the actual income in the file and applies the correct income thresholds. The questions to ask: what decision rules should apply here, and did the system apply them?
- Sanity checking. Are outputs reasonable given the inputs, and are there obvious errors or outliers? A system recommending $10,000 of maintenance on a 2-year-old vehicle is an outlier that warrants investigation regardless of how the model justifies it. The questions to ask: does this make intuitive sense, and are there red flags that something is wrong?
- Documentation verification. Is the decision documented clearly, can it be explained to the affected party, and is there a sufficient audit trail? A decision record should show which factors influenced the choice, so an affected person can understand why they got the result they did. The questions to ask: can we explain this to someone affected by it, and is there a trail showing how it was made?
These five points work together because each catches a different category of problem. Input verification catches bad data going in, cross-referencing catches conflicts with the record, logic verification catches misapplied rules, sanity checking catches implausible results, and documentation catches the accountability gap. What they do not do is cover everything. They are five categories somebody enumerated, and a novel failure that falls outside all five will pass all five.
Different Output Types, Different Checks
| Output type | Examples | Validation approach | Red flags |
|---|---|---|---|
| Numerical | Benefits amount, loan interest rate, resource allocation | Compare to expected ranges, check for anomalies, verify calculations | Amounts far outside historical range; calculations that do not match stated formulas; extreme outliers |
| Categorical | Approve or deny, hire or do not hire, risk level | Check confidence scores, verify against decision rules, cross-reference known cases | Low confidence on important decisions; inconsistent treatment of similar cases; decisions that contradict policy |
| Text | System-generated recommendations, explanations, reports | Check factual accuracy, verify citations, check coherence and consistency | Factual errors; unsupported claims; incoherent reasoning; hallucinated citations |
| Probabilistic | Confidence scores, probability estimates, risk scores | Calibration testing: verify that stated confidence matches actual accuracy | Calibration issues; a system claiming certainty it does not have, or expressing uncertainty where it is actually reliable |
Layers of Validation and Who Performs Them
Different kinds of validation serve different purposes, and effective organizations run several layers rather than treating everything the same way. Self-validation is the system flagging its own uncertain outputs for review, so low-confidence decisions escalate automatically: a permit system that cannot classify an application flags it for a human expert instead of guessing. Peer validation puts two independent reviewers on high-stakes decisions, so that both a specialist and a supervisor must approve a benefit denial. Expert validation brings domain specialists to complex or specialized decisions, as when environmental scientists review an AI assessment of habitat impact.
Statistical validation is sample-based checking for high-volume decisions, such as reviewing 5% of hiring recommendations for accuracy. Continuous validation is ongoing monitoring of output quality in production, such as a monthly review of accuracy and fairness metrics. The layering principle is that high-stakes decisions get intensive validation, routine decisions get sample-based validation, and every decision gets some form of validation.
Two of those layers deserve a caution, because both are routinely oversold. Self-validation depends on the system knowing what it does not know, and a confidence score is a number the system generated, not a measurement of its own correctness; the errors it makes confidently are precisely the ones it will not flag. And a sample tells you about the sample. Reviewing 5% of recommendations reliably surfaces a large, broad quality drop and can easily miss a small one, or a serious one concentrated in a subgroup that 5% barely touches. Both layers are worth running. Neither is a floor under quality.
Documentation and Escalation
Every validation should be documented, and the record needs five things: what was validated, meaning which output, on what date, for what decision; who validated it, with reviewer name, role and expertise; what the findings were, including problems detected and questions raised; whether the output was approved, rejected or escalated; and what happened as a result, meaning whether the output was used, modified, or rejected. That last field is the one most often left blank and the one an investigator will want most, because it is the only evidence that validation changed anything.
Five conditions should trigger escalation rather than a routine sign-off: the system produces a low-confidence output and is signalling uncertainty; validation identifies a potential error; the output contradicts other sources; the output fails a sanity check; or the case is unprecedented or unusually complex. Documentation and escalation together create accountability. When issues are documented and escalated they get tracked, decision-making becomes transparent, and problems do not quietly disappear into the volume.
Building Validation Into Standard Procedure
Validation has to live in the standard operating procedure, not in the goodwill of whoever is on shift. For approval decisions such as hire, deny or approve: the system generates a recommendation with a confidence score; low-confidence outputs escalate automatically; high-confidence outputs get sample-based review, between 5% and 20% depending on the stakes; all adverse outcomes, meaning all denials, get reviewed; and documentation is completed before the decision is communicated to the affected party.
For recommendation systems: the system generates recommendations, the relevant stakeholders validate their appropriateness, concerns are documented and escalated, and the stakeholder decision is final, because AI is an input to human judgment rather than a replacement for it. For reporting systems: the AI-generated report is produced, subject matter experts review it for accuracy, stakeholders supply context and interpretation, and the final report reflects expert judgment rather than raw AI output. The key principle across all three is the same. AI output is input to human decision-making, not an autonomous decision, and validation is the process by which human judgment is actually brought to bear.
Two Worked Cases
Benefits eligibility. A federal agency's benefits eligibility AI recommends approve or deny decisions. Input verification checks that the application is complete, all required documents are attached, and the format is correct, flagging incomplete applications for an information request. Cross-reference checks the applicant's prior benefit history: have they received benefits before, under what circumstances, and have their personal circumstances recently changed? Logic verification takes a case showing monthly income of $2,400 against an eligibility threshold of $2,500 per month and asks whether the system applied the correct threshold and extracted income correctly from the documents. The sanity check looks at an applicant aged 18, employer listed as "self-employed," and a recommended benefit of $3,200 per month, and asks whether that is reasonable for an 18-year-old just starting out or whether another data source should be consulted. Documentation records that income was verified from the tax return, that the $2,500 threshold was applied, and why the decision was made, so the applicant can understand the outcome. Sampling: all denials are reviewed as high-stakes, and 10% of approvals are reviewed as an accuracy spot check.
Criminal justice risk assessment. A state criminal justice system uses AI to assess defendant risk. Input verification confirms the arrest record is complete, criminal history is accurate, and there are no data entry errors in key fields such as prior convictions. Cross-reference compares the AI assessment against the defendant's actual criminal history: does it reflect documented facts, and are any prior arrests missing? Logic verification takes a "high risk" output attributed to a prior felony conviction and asks whether that conviction is actually in the record and correctly categorized as a felony. The sanity check flags a first-time non-violent offense receiving a high-risk assessment as unusual and asks which factors drove it and whether they are documented in the defendant record. Documentation notes which prior crimes and behaviors drove the assessment so that the defendant and the judge can follow the reasoning. The validation requirement here is absolute: 100% of assessments are reviewed by human reviewers who can overturn the AI assessment, because the AI is decision support and not an autonomous decision.
Building the Habit on Your Team
One analyst running checks is good. A team where validation is the default is far better. After her near miss, Diane made three changes that cost nothing. First, every AI-assisted document now carries a footer line: "AI-drafted; figures verified against [source] on [date] by [name]." That single line prompts the checks and creates an audit trail. Second, her team keeps a shared "known traps" note, the programs that get confused, the numbers that change often. Third, citizen-facing AI output always gets a second set of eyes, no exceptions. None of this slowed the team down meaningfully; the explainer that took the AI thirty seconds and Diane five minutes to validate still beat the old two-hour manual draft.
This discipline is exactly what federal guidance expects. The government-wide AI direction issued in 2024 requires agencies to manage risks from AI that affects the public, and the NIST AI Risk Management Framework names "measure" and "manage" as core functions; that framework is voluntary and non-binding, but it supplies the vocabulary an auditor will use. Output validation is how those abstract requirements become a habit at your desk. You are not just protecting yourself; you are protecting the people who trust their government to get the numbers right.
Anti-Patterns
- Exception-only validation. Some organizations assume most outputs are fine and validate only the ones that look unusual. This is backwards: the outputs that look normal are where a systematic error hides, because a system that is wrong in a consistent way produces consistently plausible results. Regular systematic validation, not just exception-based validation, is what surfaces the pattern.
- Single-reviewer validation on high stakes. Using one person to validate a high-stakes output risks that person's blind spot becoming the agency's. Use at least two independent validators for important decisions; peer review catches what one reader misses, and disagreement between reviewers is information rather than an inconvenience.
- Undocumented validation. If validation happens but is never written down, there is no accountability, and when something goes wrong you cannot show that you validated at all. Record what was checked, by whom, what was found, and what happened as a result.
- Deploy first, validate later. Some organizations ship and then validate. If validation then reveals problems, citizens have already been harmed by the outputs you were about to check. Validation belongs before deployment, where its findings can still change the outcome.
- Findings with no defined next step. If validation detects a problem and nobody has decided what happens next, the finding does not lead to action. Define in advance who decides and what the options are: reject, investigate, escalate. An escalation path invented in the moment is an escalation path that ends in the inbox of whoever is least likely to push back.
- Treating the checklist as the guarantee. The dominant failure in validation practice is not skipping checks; it is reading a completed checklist as a statement about the output rather than about the reviewer. Five ticked boxes mean five specific questions were asked. A false number that none of the five questions was designed to catch passes cleanly, and the tick marks make it look more verified than it was.
- Trusting the system's own confidence. Routing only low-confidence outputs to humans feels efficient and leaves the most dangerous category unexamined, because a hallucination arrives with high confidence by construction. Confidence scores are useful for triage and useless as assurance; validate a share of high-confidence outputs too, or you will only ever review the errors the system already suspected.
- Letting expert review become a signature. A named reviewer who receives forty outputs a day and returns them approved within the hour is producing an audit trail, not a check. Track how often review changes an output. If that number is near zero, the review is decorative, and everyone downstream is relying on it anyway.
Practice Prompts
- Design a validation workflow. For a benefits eligibility AI in your agency: what would each of the five system validation points look like in practice? Which outputs need intensive validation? What sample-based validation would you run for routine approvals? How would you document validation? And what would trigger escalation?
- Build a validation checklist. For an AI system you know: what specific things would you check for accuracy? What sanity checks would you perform? What cross-references would you verify? What documentation would you require? And what unusual outputs would trigger escalation?
- Design a dual-reviewer process. For hiring recommendations: what would Reviewer 1, the technical reviewer, focus on? What would Reviewer 2, the domain reviewer, focus on? How would you handle disagreement between them? How would the review be documented, and what triggers escalation to leadership?
- Run the desk check on a live draft. Take an AI-drafted document you produced this month and run the five desk checks on it now. Count how many distinct factual claims the document contains, then count how many you actually verified against an authoritative original. The gap between those two numbers is your real validation rate, and it is usually a surprise.
- Try to defeat your own checklist. Write a plausible AI output that would pass all five desk checks and still be wrong in a way that harms someone. Then write the sixth check that would have caught it. This is how a checklist stays alive rather than becoming furniture.
Reflection
Pick one AI system in your agency and answer five questions about it honestly. What validation currently happens, as opposed to what the procedure says should happen? What validation gaps exist? How could you implement more systematic validation using the five-point system framework? What resources would that take? And, hardest of all, how would you measure whether your validation is actually catching problems, rather than merely occurring? A validation process that has never rejected an output is either watching a flawless system or is not looking. Develop a short improvement plan from your answers, and put the measurement question at the top of it.
Glossary
- Validation. The process of ensuring AI outputs are accurate, appropriate, and safe before use. Distinct from testing, which checks model performance on test data.
- Hallucination. A confident, fluent, and factually wrong output, produced because the model predicts plausible text rather than retrieving verified facts.
- Cross-reference. Checking AI output against independent sources to verify consistency and accuracy. Never against another AI.
- Sanity check. A common-sense check to catch obviously wrong outputs. Does this result make intuitive sense given the inputs?
- Calibration. The property of a probabilistic output where stated confidence matches actual accuracy, so that outputs labelled 80% confident are correct about 80% of the time.
- Escalation. The process for raising uncertain or problematic outputs for additional review or expert judgment.
- Audit trail. Documentation of how a decision was made, who made it, when, and what information was considered.
Related Lessons
- Quality Assurance for AI Work Products extends the desk check into a full quality process for AI-assisted work.
- Bias Detection Tools and Methods covers the fairness dimension that output-level validation alone will not surface.
- Human-in-the-Loop: Design and Implementation addresses how to design the human review this lesson depends on so it stays real.
- Testing and Validating AI Systems covers the system-level testing that happens before any of these outputs exist.
Closing
Systematic validation is the mechanism by which government checks AI outputs before they affect citizens. It is not optional, not a one-time event, and not somebody else's responsibility. It must be built into operating procedures, cover several dimensions rather than one, and be documented so that accountability exists after the fact as well as during. Organizations that validate systematically can say what they are putting into the world and can defend their processes to oversight bodies. Organizations that skip it deploy systems that eventually fail in public, and face the lawsuits, investigations, and loss of trust that follow.
Be precise about what validation buys, though, because overselling it is its own risk. Validation does not make an output true, and a thorough process is evidence of care rather than a warrant of correctness. What it does is make errors more likely to be caught before they reach a citizen, make the ones that get through traceable afterwards, and make the difference between a system nobody has examined and one somebody is accountable for. Diane's footer line did not guarantee that any figure was right. It guaranteed that a named person had looked, on a named date, against a named source, and that is what turned one careful analyst into a careful team.
Key Takeaways
- AI output is a draft, not a fact. Language models produce text that sounds right, which is not the same as being right; verify before you rely.
- Confidence is not accuracy. Hallucinations arrive with the same certainty as correct answers, so never trust the tone of an output, and never route only low-confidence outputs to review.
- Run all five desk checks: source, cross-reference, consistency, expert review, stakes. In order, every time, on anything consequential.
- Never check AI against AI. Cross-reference against the authoritative original, the statute, manual, or dataset itself, and count a source as verified only when you have opened it and found the text.
- Validate system decisions on five separate points. Input verification, cross-reference, logic verification, sanity checking, and documentation each catch a different class of failure on an individual case.
- Match the method to the output type. Numerical outputs need range checks, categorical outputs need decision-rule checks, text outputs need fact-checking, and probabilistic outputs need calibration testing.
- Use layers, and know what each layer misses. Self, peer, expert, statistical, and continuous validation cover different ground; a sample surfaces a large drop and can miss a small or concentrated one.
- Document every validation decision and define escalation in advance. Who validated, what they found, what was decided, and what happened as a result, with named triggers for raising a case rather than passing it.
- Treat AI as decision support, not autonomous decision. Human judgment validates AI output before use, and a completed checklist records that someone looked, not that the output is true.
Frequently Asked Questions
If I ask the AI for its sources, is that enough? No, and this is the single most common mistake in the whole practice. A citation is generated text like everything else, so it can be as invented as the claim it supports. The check is not asking; it is opening the document and finding the specific text. Finding a differently titled document by a similar author does not count as locating the citation, and neither does finding a plausible page on the right topic. If the exact passage is not there, treat the claim as unverified.
How much validation is enough for an internal document nobody outside will see? Match it to the consequence, not to the audience. An internal document that will become the basis for a policy decision, a briefing, or a caseworker's understanding of a rule is consequential regardless of who reads it, and internal documents have a habit of being copied outward. Use the stakes check honestly: ask what happens if this specific number is wrong and somebody acts on it.
We validated an output and it turned out to be wrong anyway. Did the process fail? Not necessarily, and treating it as a process failure is how teams end up hiding errors. Validation lowers the rate of errors reaching citizens; it does not drive it to zero. The useful question is which check should have caught it and did not, and whether that is because the check was skipped, was performed shallowly, or was never designed to catch that class of error. The third answer is the valuable one, because it tells you what to add.
Is sampling enough for a high-volume system? Sampling is the right tool for high-volume routine decisions and the wrong tool as your only tool. A random sample surfaces a large, broad quality problem and can easily miss a small one, or a severe one concentrated in a subgroup the sample barely reaches. Pair sampling with full review of all adverse outcomes, since denials are where the harm concentrates, and with monitoring that breaks results out by group rather than reporting one aggregate number.
Who should validate, the person who used the AI or someone else? Both, at different depths. The person who produced the output runs the desk check, because they know what they asked for and what the output is meant to do. Anything citizen-facing, legal, or otherwise consequential needs a second set of eyes that did not write the prompt, because the author of a request is the person least likely to notice that the answer is shaped like the request. Independence is the point, not seniority alone.
How do I tell whether our validation is real or ceremonial? Measure how often it changes something. Track the proportion of validated outputs that get corrected, rejected, or escalated. If that number is near zero across a large volume, either the system is performing extraordinarily well, which is testable, or reviewers are approving by default, which is far more common. A validation process that has never rejected anything is not evidence of quality; it is an unmeasured process.
Skill.re