Evaluating AI Outputs
Dana Okafor, a benefits eligibility specialist at a state human services agency, had a backlog of 140 appeal letters and a deadline. So she leaned on the agency's approved AI assistant to draft replies. One letter explained to a denied applicant that "per Section 8 of the state benefits code, households exceeding 185 percent of the federal poverty line are ineligible." It read beautifully. It was also wrong. The real threshold was 130 percent, and the section number did not exist. Dana caught it only because a number felt off. Three other letters had already gone out that week. Evaluating AI output is not a nice-to-have skill. For Dana, it is the line between a tool that saves her hours and one that erodes the public's trust in her office.
This lesson gives you a repeatable method for checking AI output before it leaves your hands. It works for emails, memos, summaries, and analysis. It takes a few minutes and turns "it looks fine" into a documented decision you can account for. What it does not do, and what no method does, is guarantee that the output is correct. It raises the floor under your review; it does not put a ceiling on the errors that can survive it.
Why This Is Not Optional in Government
You ask an AI system a question. It provides an answer. The answer sounds plausible. Should you trust it? Not automatically. Government decisions affect people. If you base a decision on information from an AI system and that information is wrong, someone is harmed and you are answerable for it. In a private firm a bad AI-drafted paragraph is embarrassing. In a benefits office it is a denial letter that a household will act on, and that a hearing officer will read back to you months later.
Evaluating AI outputs critically is not optional. It is fundamental to responsible use, and it needs to be systematic rather than instinctive, because instinct is exactly what fluent text is good at satisfying. What follows is a framework you can run every time, in the same order, so that the quality of your checking does not depend on how tired you are by late afternoon.
Why "It Looks Fine" Is a Trap
AI tools are built to produce fluent, confident language. That fluency is exactly what fools us. A draft with clean grammar, official tone, and a specific-sounding citation reads as authoritative even when the facts underneath are invented. Researchers call these confident fabrications hallucinations. In a government context the stakes are higher than a typo: a wrong eligibility threshold, a misquoted regulation, or a fabricated case number can deny someone a benefit, expose the agency to an appeal, or end up in the public record.
The core habit to build: the more authoritative an AI answer sounds, the more you slow down. Polish is not proof.
It helps to name where the danger actually concentrates. AI tends to be reliable when it is reshaping facts you supplied, like turning your bullet points into a paragraph. It tends to be unreliable when it is supplying facts from its own memory, like a statute number, a statistic, or a poverty threshold. Most damaging errors live in that second category. So as you read AI output, mentally sort each sentence: is this restating something I gave it, or asserting something new? The new assertions are where you spend your checking time. Dana's letter failed precisely there. Everything the tool restated about the applicant's situation was fine. Everything it asserted on its own about the law was wrong.
The VERIFY Method
VERIFY is a six-step pass you run on any AI output before you trust it. Each letter is one check. You do not need all six every time, but for anything that touches a citizen, a record, or a decision, run the full pass.
- V - Validate facts. Pull out every concrete claim: numbers, dates, names, thresholds, dollar figures. Confirm each against a real source. Some claims are cheap to settle. "The Affordable Care Act was passed in 2010" can be checked in seconds, and it is right. "Studies show that X intervention reduces Y outcome by 50%" cannot be settled at all without reading the studies. Treat the AI as a draft, not a source.
- E - Examine sources. Did the output cite anything, and is what it cited credible? "Research by [name of researcher], published in [journal], found..." is a claim you can chase. "Studies show..." with no specific citation is not a citation at all. If it names a statute, case, study, or URL, find that source yourself and check whether the AI characterized it correctly. If you cannot locate it in two minutes, assume it is invented. Finding a differently titled document by a similar author does not count as locating it.
- R - Review for bias. Ask whether the output treats people or groups unfairly, leans on stereotypes, or assumes facts not in evidence. Watch for loaded language and for one perspective presented as obviously true with no acknowledgement of alternatives. Ask directly: what perspective is this output reflecting, and which legitimate perspectives are absent? This matters most in anything touching eligibility, enforcement, or hiring.
- I - Inspect logic. Follow the steps. Are the premises sound, and do the conclusions actually follow? Consider: "People who graduate from University X are more likely to succeed in politics. Therefore, University X produces better political leaders." Is electoral success proof of being a better leader? That is a logical gap, and it survived into the sentence because the sentence was well written. Watch for confident leaps where a "therefore" hides a missing step.
- F - Find errors. Read carefully for the small stuff and the structural stuff together: factual errors, grammatical errors, internal contradictions, unsupported claims, math, units, names spelled the same way throughout, dates that line up.
- Y - Yield judgment. Make an explicit call: is this output reliable enough to use, for what purpose, and with what caveats? High-stakes decisions require higher confidence; low-stakes uses can tolerate less. You make the final call and you own it. The AI advises; you decide. Sign off only on what you would defend in front of your supervisor or an auditor.
The AI drafts. You decide. Your name is on the letter, not the model's, and a completed pass is a record of the checking you did rather than a warranty on what you missed.
VERIFY in Action: A Policy Recommendation
Suppose you ask for advice on regional unemployment and the AI answers: "To reduce unemployment in your region, implement a job training program focused on technology skills. Research shows that technology training increases employment by 40% and leads to higher wages. Cities that have implemented similar programs have seen success." It is a confident, well-formed, entirely typical answer. Run the pass.
Validate facts. The "40% increase" is specific enough to need a citation. Find the research and check whether it shows that figure for a population like yours. "Leads to higher wages" is vague to the point of being uncheckable: higher than what, by how much, in which jobs? Examine sources. The AI provided no specific citations. Get them, look them up, and check whether they support the claims made.
Review for bias. The recommendation points at technology training. Is that biased toward certain demographics? Is it accessible to people with disabilities? Is it appropriate for your region's actual economy? Inspect logic. Training leading to employment is a sound chain as far as it goes, but it quietly assumes everyone can access training and that training is appropriate for everyone. Find errors. The general claim is reasonable; what it lacks is specificity, which is a different defect and a more slippery one.
Yield judgment. The recommendation has merit, but you need more specific research before deciding. It is a good starting point for further investigation, not a basis for a policy decision. That sentence is the honest output of a VERIFY pass more often than people expect, and writing it down is not a failure of the method. It is the method working.
Dana Reruns the Letter Through VERIFY
Here is how Dana's bad letter would have failed the pass in under three minutes:
- Validate facts: "185 percent of the federal poverty line." Dana checks the agency's eligibility table. The real figure is 130 percent. Fail.
- Examine sources: "Section 8 of the state benefits code." She searches the code. There is no Section 8 covering this. Invented citation. Fail.
- Review for bias: The tone is neutral here. Pass, but she notes the letter assumed a household size it was never told.
- Inspect logic: The reasoning depended on the wrong threshold, so the conclusion collapses with it.
- Find errors: The percentage and the section number are the errors. The math downstream was built on a bad number.
- Yield judgment: Dana rewrites with the correct 130 percent figure and the real citation, then sends. The tool still saved her the first draft; the method saved her from the mistake.
Spotting Hallucinations
Sometimes AI systems generate false information with confidence. That is a hallucination, and it takes recognizable shapes: a research study that sounds real but does not exist, a citation to a law that was never passed, a description of an event that did not happen, statistics produced from thin air. What makes them dangerous is not that they are wrong. It is that they arrive with exactly the same tone, structure and specificity as the true statements around them.
Three detection habits do most of the work. Verify citations: the AI cites a study, so look it up, and if it does not exist you have found a hallucination. Check specific facts: dates, numbers and names are cheap to test, and one wrong one tells you the model was generating rather than recalling. And be especially suspicious of three shapes in particular: very specific statistics without sources, quotes attributed to famous people, and references to specific studies with specific results. Those three are where fabrication is both easiest for the model and most convincing to a reader.
A fourth habit needs a caveat attached. You can ask the model follow-up questions, and sometimes it will admit it cannot find the source, which is a useful signal. But asking a model to check itself is not verification. A denial is worth investigating. A confirmation is worth nothing at all, because the same system that fabricated the citation will fabricate the reassurance in the same voice. Only an external source settles the question.
Match Your Effort to the Stakes
You do not need a full VERIFY pass on an internal brainstorm. You absolutely need one on a denial letter. Use this simple tier to decide how hard to check, and re-tier the moment a document changes audience, because low-risk text has a habit of being pasted into high-risk documents by someone who was not in the room when you skimmed it.
| Risk tier | Examples | Check level |
|---|---|---|
| Low | Brainstorming, internal notes, reformatting your own text | Quick skim; you are relying on your own facts, not the model's |
| Medium | Internal memos, meeting summaries, first drafts for a colleague | Validate facts and inspect logic |
| High | Anything to a citizen, the public record, legal, or a decision | Full VERIFY pass, documented |
A Desk-Side VERIFY Checklist
Keep this where you work. Work through it before any high-stakes AI output leaves your hands. Understand what ticking the boxes means: it records that you looked, in a form you can show someone later. It does not certify that the output is right, and a completed checklist has never once made a false statement true.
- I listed every fact, number, date, and name and confirmed each against a real source.
- I personally located every citation, statute, case, or link the tool produced.
- I checked whether the output treats any person or group unfairly or assumes facts I never supplied.
- I confirmed the conclusion follows from the reasoning, with no hidden leaps.
- I rechecked the math, units, and internal consistency.
- I would defend this output, as written, to my supervisor and to an auditor.
- I noted in my records that AI assisted with this draft, per my agency's policy.
Worked Example: Three Percentages and No Sources
You ask an AI system: "What are the most effective interventions to reduce recidivism among people released from incarceration?" It responds with a list of approaches and these claims: employment support reduces recidivism by 35%, educational programs reduce it by 28%, cognitive behavioral therapy reduces it by 15%. Nothing about that output looks wrong. That is the problem, and it is why you run the pass rather than reading for a bad feeling.
V: Do those specific percentages actually exist in research? Find the studies and check whether they show those numbers. E: Are the studies credible, peer-reviewed, adequately sized? R: Does the list itself carry bias? Are all three approaches equally evidence-based, or are some more speculative than the uniform presentation suggests? I: The underlying logic, that helping people succeed in life reduces crime, is sound; the specificity of the percentages is what needs validation. F: One error stands out: the therapy figure is stated flatly, with no note that it applies to specific populations.
Y: The general recommendation has merit, but you would not rely on the specific percentages without reading the original research. Use it as a starting point, then research further. Notice that the pass did not turn a bad answer into a good one. It turned an unusable answer into a usable research agenda, and it stopped three unsourced percentages from entering a document that a legislator might quote.
Learning Your Tool's Failure Pattern
One habit makes the whole method faster over time: when a tool fabricates something, note what kind of error it was. Most tools fail in predictable ways. Once you know that yours invents citations while handling tone well, you know where to point your attention first, and the pass gets quicker because you are no longer searching blind.
Be careful what you conclude from that, though. A known failure pattern tells you where errors are most likely, not where they are confined. It is a priority order for your attention, not a license to stop checking the categories that have been clean so far. Tools change under you: a model gets updated, a vendor swaps something out, a new document type enters your workflow, and the pattern you learned last quarter quietly stops describing the tool you are using today. Keep sampling the areas you trust, precisely because you trust them.
Anti-Patterns to Avoid
- Assuming AI outputs are accurate. You ask, it answers with confidence, you accept. The risk is that the model hallucinated or simply erred, and confidence in the output tells you nothing about either. Fluency is a property of the writing, not of the facts.
- Skipping verification for obvious-seeming facts. "The AI says the capital of France is Paris; everyone knows that." But what if the question was "which city was incorrectly described as the capital of France in 1800?" and the model misread it? Spot-check even the obvious ones, because a right answer to the wrong question looks identical to a right answer.
- Not checking policy recommendations for bias. The AI recommends a policy, it sounds reasonable, and nobody asks who it disadvantages or what legitimate alternatives went unmentioned. The risk is implementing a biased policy with a neutral-sounding rationale attached.
- Relying on AI for high-stakes decisions. You ask "should we deny this person's benefit?" and use the model's reasoning as the basis for the decision. That decision is too important for AI to handle alone, and the reasoning is not evidence.
- Treating a completed VERIFY pass as a clearance. Six checks run is a documented review, not a certificate of accuracy. It is evidence that you looked, and it makes you far more likely to catch an error, but the honest claim afterward is "I checked these things and found these results," never "this output is correct."
- Asking the model to verify itself. Following up with "are you sure?" or "is that citation real?" and accepting the answer. The system that produced the fabrication will produce the confirmation just as fluently. Verification has to come from outside the tool.
- Letting a learned failure pattern shrink the pass. Knowing your tool invents citations but writes clean prose is useful for ordering your attention. Using it to stop checking the prose is how the next error gets through.
Practice Prompts
- Ask an approved AI tool a substantive question relevant to your work. Apply VERIFY in order. Write down, in one sentence, how confident you are in the output and exactly why.
- Ask an AI system a question about something you know well, where you can grade the answer yourself. Did VERIFY reveal problems you would not have noticed reading casually?
- Take an AI output containing a citation. Time yourself locating the source. Note what you found: the real thing, something adjacent, or nothing at all.
- Find an instance where you used information without fully verifying it. What would VERIFY have revealed, and at which letter would it have stopped you?
Reflection
- What types of AI outputs in your work are high-stakes enough to require a rigorous VERIFY pass, and who decides the tier?
- If an auditor asked you to show your review of an AI-assisted document you sent last month, what could you actually produce?
- Where in your workflow would a wrong number do the most damage before anybody noticed, and what would catch it there?
Glossary
- Hallucination: When an AI generates false information, stated with the same confidence and specificity as true information.
- Citation: A reference to a source of information. A reference you cannot locate is not a citation.
- Bias: Systematic tendency toward a particular perspective or conclusion.
- Logic: Reasoning that follows clear rules and sound premises.
- Confidence: The degree to which something seems reliable or trustworthy. A model's confident tone is a property of its writing style, not evidence about the claim.
- VERIFY: Validate facts, Examine sources, Review for bias, Inspect logic, Find errors, Yield judgment. A review sequence, run in order, on output you are about to rely on.
Related Lessons
- Your Agency's Approved AI Tools establishes which tool you are evaluating output from in the first place.
- Prompt Engineering Basics reduces how much VERIFY has to catch, by making the request specific before the output exists.
- AI for Government Tasks: Summarization and Drafting is where most of the output you will be evaluating comes from.
- AI Confidence and Hallucination goes deeper on why a fabricated citation arrives sounding exactly like a real one.
Closing
Evaluating AI outputs is not about being paranoid or distrustful. It is about being professional. Use VERIFY. Trust, but verify. Make your decisions based on information you have personally evaluated, and be precise afterward about what your evaluation actually established: which claims you checked, against what, and what you could not settle. Dana still uses the tool every day. What changed is that nothing it asserts leaves her desk with her name on it until she has been somewhere else to confirm it.
Key Takeaways
- Polish is not proof. Fluent, confident output is exactly what hides invented facts. Slow down when an answer sounds most authoritative.
- Run VERIFY. Validate facts, Examine sources, Review for bias, Inspect logic, Find errors, Yield judgment, in that order.
- Find every citation yourself. If you cannot locate a statute, case, or link in two minutes, treat it as fabricated, and do not count a near-match as a find.
- Hallucinations have shapes. Very specific statistics without sources, quotes attributed to famous people, and references to specific studies with specific results deserve the hardest look.
- Verification comes from outside the tool. Asking the model whether it is sure is not a check; a denial is a hint and a confirmation is worthless.
- Match effort to stakes. Skim low-risk drafts; run the full documented pass on anything touching a citizen, a record, or a decision, and re-tier when the audience changes.
- You yield the judgment. The AI advises and you decide. Sign off only on what you would defend to a supervisor or auditor.
- A completed pass is evidence, not a guarantee. The checklist records that you looked. It does not certify the output, so keep sampling the categories your tool has been getting right.
- Log AI assistance. Note when AI helped produce a work product, as your agency policy requires, so the record is honest.
Frequently Asked Questions
If I run the full VERIFY pass, is the output safe to send? It is as good as your checking made it, which is a real improvement and not a guarantee. The pass documents what you confirmed and what you could not. Errors can survive a competent review, so the honest statement is always about what you checked rather than about what is true.
Can I just ask the AI whether its citation is real? No. The tool that invented the citation can invent the confirmation, in the same voice, just as fast. If it volunteers that it cannot find the source, treat that as a useful hint and go look anyway. Only an external source settles it.
How many facts do I need to check in a long output? Every load-bearing one, meaning every claim a decision or a letter would rest on. For routine, low-consequence content you can spot-check a sample, but be clear with yourself that a sample tells you about the sample.
My tool has never fabricated a date, only citations. Can I stop checking dates? No. That pattern tells you where to look first, not where errors are confined, and the pattern can change when the model is updated or when you start feeding it a new kind of document. Keep sampling the areas that have been clean.
What counts as high-stakes? Anything reaching a citizen, entering the public record, touching a legal matter, or feeding a decision. When in doubt, treat it as high and run the full pass, because the cost of over-checking is minutes and the cost of under-checking lands on someone else.
Do I have to record that AI was involved? Follow your agency's policy on this, and note it where the policy requires. An accurate record of how a document was produced is part of the document being honest, and it is the first thing an auditor will ask about.
Skill.re