←
AI for Government
Aware · M19 · lesson 19 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Confidence and Hallucination
📖
now learning

AI Confidence and Hallucination

10 min

In 2023 two lawyers in New York filed a court brief that cited six prior cases supporting their argument. The cases were perfect: right court, persuasive holdings, real-sounding citations. They were also entirely fabricated by an AI chatbot. A judge sanctioned the lawyers and the story made national news. Now picture the same mistake inside a state agency: a benefits eligibility worker named Priya Raghunathan asks an AI assistant "what's the income limit for this program in 2026?" and pastes the confident answer into a denial letter. The number was invented. A family loses benefits they qualified for. Nobody was lying. The AI just sounds exactly as certain when it is wrong as when it is right. Learning to hear the difference is the most important AI safety skill a frontline government employee can have.

Here is the fact that surprises most people the first time they meet it: AI systems do not know whether they are right or wrong. They generate plausible-sounding text, and they do it with the same high confidence whether or not the text is accurate. Government work requires accurate information. If a system confidently asserts something false and you rely on it, you make a bad decision with real consequences for a real person, and you may not find out for months.

What a "hallucination" actually is

The word sounds exotic. The mechanism is mundane. A large language model, the technology behind tools like the one Priya uses, is trained to do one thing: predict the next most plausible word, given everything before it. It is, at its core, an extraordinarily sophisticated autocomplete. It has no concept of true or false. It has a concept of likely.

Watch it work. Given "The capital of France is," the model predicts "Paris" because that is the most common next word in that context. Given "A study showed that eating [blank] reduces heart disease," it predicts a plausible-sounding word based on patterns in medical writing. It might land on "vegetables," which is correct, or "blueberries," which is reasonable, or on something absurd if absurd text happened to sit near that context in its training material. The model is not consulting a fact. It is completing a shape. When it produces "Studies show that X reduces Y by 50%," it is not asserting that with knowledge, it is predicting that this is a likely sentence to appear here.

A hallucination is when the model produces something fluent, plausible and wrong. It is not a bug or a glitch. It is the system working exactly as designed, generating the most likely-sounding text, in a situation where the likely-sounding text happens not to be true. A made-up court case "looks like" a real citation. A made-up income limit "looks like" a real number. This is why the lawyers got fake cases that looked real, and why Priya could get a fake number that looked real. The AI was never consulting a law library or a policy manual.

Hallucinations happen for a few distinct reasons worth keeping separate. Sometimes the most plausible-sounding next word is simply wrong. Sometimes the model was trained on incorrect information and is faithfully reproducing it. And sometimes the model is being asked for something it has no information about at all, and generating something is what it does, so it makes one up. That last case is the dangerous one in government, because it is triggered by exactly the specific, local, current questions frontline staff need answered.

What they look like in practice

Hallucinations are easier to catch once you have seen the standard shapes. Each of the examples below looks authoritative on the page, and each is the model completing a pattern rather than reporting a fact. Read them as a rogues' gallery: the point is not to memorize these specific claims but to recognize the form, because the form is what recurs. A fabricated citation about government efficiency and a fabricated citation about zoning law are produced by the same mechanism and arrive looking equally solid.

  • Made-up citations. "Smith (2019) found that AI improves government efficiency by 40%." No such paper exists. The model invented the author, the year and the finding together, because that is the shape a supporting citation takes.
  • Invented statistics. "Studies show that 73% of government employees prefer remote work." The number is fabricated, and its precision is part of what makes it convincing.
  • False historical claims. "The Affordable Care Act was passed in 2008." It was 2010. The model was wrong and stated it plainly.
  • Nonsensical logic. "To reduce traffic, cities should require cars to be green." The sentence has the grammatical form of a policy recommendation and no sense behind it.

Why it sounds so certain

Here is the trap. Human beings use tone as a signal of reliability. When a colleague hedges, saying "I think it's around forty thousand, but check the manual," we lower our trust. When someone states a precise figure plainly, we raise it. AI assistants produce the confident, precise tone by default, because that is what most of their training text sounds like. The fluency is free. The accuracy is not, and nothing in the generation process ties one to the other.

So the usual human shortcut, "they sound sure, so they probably know," fails completely with AI. The confidence carries zero information about correctness, and a true answer and a hallucination arrive in the same calm, authoritative voice. This is the confidence-accuracy gap, and it runs in every direction. A system can be wrong but confident, insisting the capital of France is Rome. It can be right but uncertain, hedging that maybe the capital of France is Paris. It can be right and confident, which usually happens when it was trained clearly on the material. You cannot use its confidence as a measure of its accuracy, in any of those cases.

The gap widens with complexity. Ask about something genuinely contested, such as the effectiveness of job training programs in reducing unemployment, and the system may invent specific statistics with complete confidence, present them as established fact, sound authoritative throughout, and be entirely wrong. Nothing in the output distinguishes that answer from a well-grounded one. The AI sounds exactly as certain when it is wrong as when it is right, and confidence is not evidence.

Where hallucinations cluster

Hallucinations are not random. They concentrate in predictable places, and knowing where turns vague anxiety into a targeted checklist you can actually apply at speed. Priya now treats the categories below as red zones: answers she never accepts without verifying, regardless of how the response reads. The value of a list like this is that it works before you know whether an answer is wrong, which is the only moment when knowing is any use to you.

  • Specific facts the AI "remembers": dates, dollar figures, statute and case numbers, phone numbers, names, percentages. The more precise the claim and the more it comes from the model's own memory, the higher the risk. A specific number with no source attached, of the form "X reduces Y by 47%," deserves suspicion on sight.
  • Recent events: anything after the model's training cutoff. It may confidently describe a policy that changed last month using last year's rules.
  • Niche or local specifics: your agency's exact procedures, budget or policies, a small program's eligibility rules, county-level details. The internet has little text about these, so the model improvises. "The best approach to increase civic participation in your city is" and "your agency should implement this specific policy" are both the model generating plausible text about something it has no information on.
  • Anything you would want to cite: quotes, sources, statistics, legal references. If the model offers a citation unprompted, assume it may be fabricated until you find the real one.
  • Math and counting: step-by-step arithmetic, totals and tallies are frequently wrong despite looking precise.
  • Claims that are almost right. Some hallucinations sit close to the truth without landing on it. "The Affordable Care Act was passed in 2009" is near 2010 and still wrong, and near-misses survive a skim in a way that obvious errors do not.

One more signal is worth watching for, because it does not depend on you knowing the subject matter. If the response contradicts itself internally, asserting one thing early on and the opposite further down, that suggests the system is generating text without a consistent underlying model of the topic. You do not need to know which half is wrong to know the answer is not trustworthy.

The fix: grounding beats trusting

The most reliable defense is not "be skeptical," because skepticism fades under deadline pressure. It is a workflow change: give the AI the source, so it organizes a fact instead of inventing one. Priya's reframed habit comes from the difference between the two ways she could have asked her question.

The dangerous way is "what's the 2026 income limit for this program?", where the AI answers from memory and may hallucinate. The safe way is to paste the current eligibility policy page and ask "according to the text above, what is the 2026 income limit? Quote the exact line." Now the AI is reading rather than remembering, and the quoted line is checkable against the document sitting in front of you. If the answer is not in the text, instruct it to say "not stated" rather than guess, and good prompts ask for that explicitly.

Grounding moves the risk, and it moves a great deal of it, but it does not remove it. A model given a document can still misread it, blend the pasted text with something it remembers, or summarize a conditional rule into an unconditional one. What grounding buys you is that the claim becomes checkable in seconds, because the source is right there, which is exactly why asking for the exact line matters more than asking for the answer.

Other habits that reduce the risk

You cannot eliminate hallucination risk, but you can reduce it. What follows are the standard mitigations, each paired with the limit that stops it from being a substitute for checking. Learning the limits matters as much as learning the habits, because a mitigation you over-trust is worse than none at all: it produces the feeling of having verified something without the fact of it, and that feeling is what carries an unchecked number into a document that leaves the building.

  • Ask for sources. "What sources support this claim?" Sometimes the system will admit it does not have any. Sometimes it will produce a citation that is as fabricated as the claim, so the request is a prompt to go looking, not a verification step in itself.
  • Ask what is uncertain. "How confident are you in this? What is uncertain here?" This sometimes surfaces genuine caveats. Remember what the rest of this lesson establishes, though: a self-reported confidence level is another generated sentence, not a measurement of anything, so treat it as a hint about where to look rather than as a score.
  • Verify everything important. Do not trust factual claims without checking them, and do not let the checking be optional on the items that matter most.
  • Be most skeptical of specific claims. Specific statistics, dates and names are more likely to be hallucinated than general statements.
  • Watch the pattern. If a tool consistently hallucinates on certain topics, be extra skeptical of its output on similar topics in future.
  • Compare outputs. Ask the same question of different AI systems and see whether they agree. Disagreement is a useful signal that something is being made up. Agreement is not the reverse signal, because two systems trained on overlapping material can be confidently wrong in the same way.

Skepticism without paralysis

The goal is not to distrust AI into uselessness. If you verify every word, you have lost the time savings entirely. The skill is calibrated trust, matching how hard you check to how much the answer matters. Priya uses a simple triage.

Stakes of the outputExamplesVerification level
Low: disposable or privateBrainstorming, rephrasing my own draft, a rough outline I will rewrite.Light. Read for sense; errors cost nothing.
Medium: internal, reversibleA summary for my own use, a first-draft email I will edit.Spot-check facts before relying on them.
High: leaves the building or affects a personNumbers in a constituent letter, eligibility decisions, anything cited externally, legal or medical content.Verify every fact against an authoritative source. No exceptions.

The single dividing line: does this output affect a real person or leave the agency? If yes, it is high-stakes, and confidence in the AI's tone is irrelevant. You check the source. Notice that the triage is about the consequence of the output rather than the difficulty of the question, because a hallucinated answer to an easy question does exactly as much damage in a denial letter as a hallucinated answer to a hard one.

A worked example: the number that looked sourced

You ask an AI: "What's the average cost of implementing a municipal bike-sharing program?" It responds: "According to research by the International Transportation Association (2021), the average cost of implementing a municipal bike-sharing program is approximately $1.2 million for a city of 500,000 people. This includes infrastructure, bikes, and first-year operational costs." That answer is authoritative, specific, scoped to a population size, and broken out by category. It is exactly the kind of sentence that ends up in a briefing memo without further thought.

Take it apart. The International Transportation Association might not exist, or might exist and never have published this research. The figure is specific and carries an apparent source, which is precisely what makes it persuasive, and you still cannot verify it from the response itself. The confidence is high and tells you nothing. Note the trap in the structure: the citation makes the number feel checked, when in fact the citation is one more generated string that arrived by the same process as the number.

So do the work. Search for the named association and the 2021 research and find out whether it exists. Search independently for actual bike-sharing cost data and see what you get. Ask the tool for more sources, understanding that this is a lead-generation step. Cross-check what the AI said against what you found on your own. And only rely on the figure after you have verified it, which in practice often means replacing it with a number you found yourself and can point to.

A usable artifact: the hallucination red-flag scan

Run this scan on any AI answer before you act on it. It takes about ten seconds and it sorts answers by how much verification they need; it is a triage tool, not a verification step, and nothing on this list confirms that an answer is correct.

  1. Did it state a specific fact from memory? Date, dollar figure, statute, name, number, quote, citation. Verify against a real source.
  2. Is it about something recent, local, or niche? High hallucination risk, so confirm independently.
  3. Did I give it a source, or is it improvising? No source means no trust. Re-ask with the document pasted in.
  4. Will this affect a person or leave the agency? Verify everything, regardless of how certain it sounds.
  5. Did it offer a citation I did not ask for? Find the original. If you cannot, it may not exist.
  6. Does the answer contradict itself anywhere? Internal inconsistency means the answer is unreliable even where it happens to be right.

Anti-Patterns to Avoid

  • Trusting confident-sounding output without verification. The AI sounds certain, so you assume it is right, and you rely on hallucinated information for a decision that matters. This is the failure the entire lesson exists to prevent, and it happens under time pressure rather than out of ignorance.
  • Sharing AI-generated citations without checking them. The tool cites a study in your report and the report goes out unchecked. Your credibility takes the damage when someone looks up the citation and finds nothing there.
  • Asking about topics where hallucination is near-certain. You ask about your own agency's policies, budget or procedures, and get plausible-sounding invention back, because the model has no information about your agency. Then you act on false information about your own organization.
  • Treating a self-reported confidence as a measurement. "I am highly confident in this answer" is generated by the same next-word process as the answer. Asking the model how sure it is can surface useful caveats and it never produces a reliability score, so a hedge is worth reading and a reassurance is worth nothing.
  • Reading a returned citation as verification. Asking for sources is a good habit whose output still needs checking. A fabricated claim can arrive with a fabricated source attached, as the bike-sharing example shows, and the presence of a citation makes the claim feel checked without anyone having checked it.
  • Treating cross-model agreement as confirmation. Two systems agreeing tells you they share a pattern, not that the pattern is true, since overlapping training material produces overlapping errors. Disagreement is informative; agreement is not a clearance.
  • Assuming a grounded prompt cannot go wrong. Pasting the source removes most of the invention risk and leaves misreading, over-generalizing and blending in remembered material. Ask for the exact line and read the line, rather than trusting the summary because a document was attached.
  • Running the red-flag scan and stopping there. The scan tells you which answers need verifying. Clearing it is not the same as having verified anything, and on anything that reaches a citizen the verification still has to happen.

Practice Prompts

  • Ask an approved AI tool a specific factual question in your subject area. Does it provide sources? Can you verify the facts it gives you against an authoritative source?
  • Deliberately ask it to do something it is likely to hallucinate on, such as describing your agency's specific procedures. What does it produce, and how confident does it sound while producing it?
  • Take something you know extremely well, a city you have lived in, an organization you have worked in, a topic you have studied, and ask the tool about it in detail. Where does it get things right, and where does it drift? This exercise trains the ear for the difference.
  • Take one question you asked from memory and re-ask it grounded: paste the authoritative document and ask for the exact line. Compare the two answers.
  • Take an AI answer that included a citation and try to find the original. Time how long it takes, and note whether you found it.

Reflection

  • In your work, what types of claims would be most dangerous if hallucinated, and who absorbs the harm when one gets through?
  • Which of the outputs you produced with AI last month would land in the high-stakes row of the triage table? Did you verify them at that level?
  • Where in your workflow does deadline pressure make verification most likely to get skipped, and what would have to change for it not to be skipped there?
  • If a constituent challenged a number in a letter you sent, could you say today where that number came from?

Glossary

  • Hallucination. When an AI generates false information, often presented with confidence.
  • Confidence-accuracy gap. The lack of correlation between how confident an AI sounds and whether it is actually right.
  • Plausible-sounding. Text that reads as reasonable and likely but may be false.
  • Token prediction. The process by which language models generate text word by word based on patterns.
  • Large language model. The technology behind general-purpose AI assistants, trained to predict the next most plausible word given everything before it.
  • Grounding. Supplying the source material in the prompt so the system organizes a fact from the text in front of it rather than recalling one from training.

Closing

Hallucination is not a flaw that will be patched away in the next release. It is fundamental to how these systems work, because predicting likely text and reporting true text are different jobs and only one of them is being done. So learn to recognize where it clusters, ground your questions in real sources, verify the claims that matter, and never let confident-sounding text talk you into trusting information you have not checked. Priya's habit is now the whole lesson in one move: before the number goes in the letter, she can point to the line it came from.

Key Takeaways

  • A hallucination is fluent, plausible, wrong text. The model predicts likely words rather than true facts, and it has no built-in sense of correctness to consult.
  • Confidence carries no information. AI sounds equally certain whether right or wrong, so the human "they sound sure" shortcut fails completely, and a self-reported confidence level is just more generated text.
  • Risk clusters predictably. Specific remembered facts, recent events, local or niche details, unprompted citations, math, and near-miss claims are the danger zones.
  • Grounding beats skepticism. Paste the source and ask for the exact line, which makes the answer checkable in seconds even though it does not make misreading impossible.
  • Tell the AI to admit ignorance. Instruct it to say "not stated" rather than guess when an answer is not in the source you supplied.
  • Calibrate verification to stakes. Light for disposable work, full verification against an authoritative source for anything that affects a person or leaves the agency.
  • Treat citations and cross-model agreement as leads, not proof. A returned source can be fabricated, and two systems agreeing can share the same error, so both point you toward checking rather than replacing it.

Frequently Asked Questions

Will hallucination be fixed in a future version? Treat it as permanent for planning purposes. It is fundamental to how these systems work, because the model generates the most likely-sounding text and likelihood is not truth. Build the verification habit into your workflow rather than waiting for the problem to go away.

If I ask the AI how confident it is, can I trust the answer? Not as a number. Its stated confidence is generated by the same next-word process as everything else, and the whole point of the confidence-accuracy gap is that the two are not correlated. Asking what is uncertain can still be useful, because it sometimes surfaces genuine caveats worth chasing, but a reassuring answer is not evidence of anything.

The response included a citation. Doesn't that mean it is grounded? No. Fabricated citations are one of the most common hallucination shapes, and a made-up study with a plausible author, year and finding is exactly what the pattern-completion process produces. If the model offered a citation you did not ask for, assume it may be invented until you find the original.

Two different AI tools gave me the same answer. Is it confirmed? No. Disagreement between systems is a useful warning sign that something is being made up, but agreement is not the mirror image of that signal, because systems trained on overlapping material make overlapping errors. Confirm against an authoritative source rather than against a second model.

How do I check without losing all the time AI saved me? Use the triage. Low-stakes disposable work needs only a read for sense. Internal reversible work needs a spot check. Anything that affects a person or leaves the agency gets every fact verified against an authoritative source, no exceptions. Most output is not in that top row, which is what makes the top row affordable.