←
AI for Government
Aware · M28 · lesson 28 of 31 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The Human in the Loop
📖
now learning

The Human in the Loop

10 min

Maria Ortiz processes unemployment claims for a state workforce agency. Last winter her office switched on a new fraud-detection tool that flagged "suspicious" claims for review. The tool came with a green Approve button and a red Deny button, and a quota: 80 claims an hour, which leaves 45 seconds a case. She was, on paper, the human reviewing every decision. In practice, she was a rubber stamp. When a local reporter found that the tool had wrongly frozen benefits for 1,100 seasonal construction workers, the agency's defense was that "a human reviewed each case." That defense did not survive the first hearing.

Maria is the human in the loop. Her story is why the phrase, used carelessly, has become one of the most dangerous comfort blankets in government AI. It rests on a dangerous assumption: that if an AI system is accurate enough, it can make decisions without human input. That assumption is wrong. Even highly accurate systems should not make important decisions alone. This lesson is about the difference between the kind of human oversight that protects people and the kind that just protects a slide in a briefing deck.

What "the loop" actually means

An automated decision process is a loop: data comes in, the system scores or sorts it, an action follows. "Human in the loop" means a person sits inside that cycle with real authority to change the outcome before it affects someone. The opposite is "human out of the loop," where the machine acts on its own. Between those two sits the pattern that causes most of the trouble in practice, and the one most agencies have without knowing they have it.

That third pattern is "human on the loop." Here a person can watch and intervene, but the system keeps moving whether they engage or not. Maria was technically in the loop. The 45-second clock made her function like someone merely on the loop, swept along by a process she could observe but not realistically stop. The label said one thing; the reality was another. The question is never "is there a human?" The question is "could this human realistically have said no?"

Why humans must stay in government decisions

In the private sector, a bad automated decision usually costs money. In government, it can take away a benefit, a license, a child placement, an immigration date, or freedom. Government decisions are not just about accuracy. They are about judgment, context, values, and accountability, and AI is genuinely good at exactly one item on that list. Machines are excellent at recognising patterns in data. Everything else on the list is a human function, and pretending otherwise is how agencies end up defending the indefensible at a hearing.

Context. The AI sees data. The human sees the person and their circumstances. Values. How do we balance competing goods? AI does not do this. Humans do. Exceptions. There is always a person whose case does not fit the pattern and requires human compassion. A flood victim files an address that does not match records because their house no longer exists. A veteran's paperwork uses an old legal name. Models trained on typical cases handle the typical well and the unusual badly. Accountability. When a decision goes wrong, who is responsible? Not the algorithm. The human.

Due process sits underneath all four. People affected by a government decision have a constitutional and statutory right to a fair process and, often, a real person who can explain and reconsider. An algorithm cannot be cross-examined at a hearing. When the public, the legislature, or a court asks who made a decision, "the system decided" is not an answer a government can give. Someone has to be able to stand up, describe the reasoning, and own the outcome, and that person needs to have genuinely made the call.

Automation bias: the quiet failure

Maria did not stop caring. She fell into automation bias: the tendency of people to rely on automated decisions even when they should not, especially when they are tired, busy, or unsure. When a screen shows a confident "92% fraud risk," disagreeing feels like sticking your neck out. Agreeing feels safe. This is well documented, it happens to conscientious people, and it is not solved by telling staff to be more careful.

It shows up in four recognisable forms. "The AI said so": a reviewer is presented with a recommendation, assumes it is right, and does not carefully review it. Trusting accuracy metrics: "the system is 94% accurate, therefore it is probably right in this case." Deferring judgment: "I am not sure, so I will go with what the AI says." Reduced scrutiny: a decision made by a human gets careful review, while a decision made by an AI gets less careful review, or none at all.

The second form deserves a moment, because it is a reasoning error dressed as diligence. An accuracy figure is an estimate of past performance on the cases that were tested. It is not a promise about the case in front of you, and it says nothing at all about whether this case resembles the tested ones. The unusual applicant sitting in front of Maria is precisely the applicant least likely to be represented in that test set. The correct use of an accuracy figure is to decide how much scrutiny the system deserves overall, never to decide an individual case.

Automation bias is dangerous because of what it does to error correction. AI systems make errors and humans make errors, but human errors are sometimes caught by other humans who ask questions. Automated errors are less often caught. If a human just accepts the AI without question, the error becomes the final decision. Three conditions make this worse, and all three are common in public agencies:

  • High volume. The more cases per hour, the less time to think, the more you default to the suggestion.
  • No friction to override. If approving the AI's call is one click and overriding it requires a written justification, people approve.
  • Punishment for disagreeing. If overrides get second-guessed by managers but agreements never do, staff learn to agree.

Meaningful oversight means designing against automation bias on purpose, not just hoping staff stay sharp. Every condition above is a design choice someone made, which means every one of them is a design choice someone can unmake.

What makes a review meaningful

Meaningful human review is genuine oversight by a person who can override the AI. It is a mitigation, not a cure: it reduces the rate at which a system's errors become final decisions, and it gives the affected person somewhere to go. It does not make the underlying system correct, and an agency that installs a reviewer and stops measuring the model has swapped one blind spot for another. Four conditions have to hold at once before a review deserves the word.

  • The human has authority. They can actually override the AI if they think it is wrong.
  • The human has information. They have access to the same data the AI had, plus context the AI did not have.
  • The human has responsibility. If they override the AI and the decision goes wrong, they are responsible.
  • The human has incentive to review carefully. They are not just rubber-stamping; they have motivation to get it right.

You can hear the difference in the language reviewers use. Meaningful review sounds like this: "Here is what the AI recommended and why. Here is what I think. We differ on this point. Given this situation, I override the AI and recommend the alternative." Or: "The AI recommends this, but this person's circumstances fit a pattern the AI might not recognise, so I am escalating to a supervisor." Or: "The AI's reasoning seems sound, but it is based on aggregate data and this person might be an exception, so I am requesting additional information before deciding."

Performative review sounds different, and once you have heard it you cannot unhear it. "The AI said this. I agree. Approved," with no actual review. "The AI said this. I checked and do not see any obvious errors. Approved," which is minimal review. "The AI said this. I cannot understand why, but it is probably right. Approved," which is an abdication of responsibility. The third one is the most dangerous because the reviewer has openly stated they do not understand the decision they are about to sign.

Telling them apart in your own workflow

Performative oversight exists to be pointed at. Meaningful oversight exists to change outcomes. The line between them is visible in the design, not the org chart, which is fortunate: it means you can audit it without interviewing anyone about their intentions.

SignalPerformative (a rubber stamp)Meaningful (real control)
Time per caseSeconds; a quota forces speedEnough time to read the actual file
What the reviewer seesOnly the score and a buttonThe reasons, the source data, and the confidence
Override effortHard; requires extra justificationAs easy as agreeing
Override rateNear 0% (red flag)A real, tracked, non-trivial rate
Reviewer expertiseAnyone with a loginTrained in the program and the tool's limits
Consequence of decisionReviewer feels noneReviewer owns and signs the outcome

One number on this table is a powerful early warning: the override rate. If a tool that flags 5,000 cases produces almost no overrides, you do not have a great model. You have a rubber stamp, and you should treat the override rate near zero as a defect to investigate, not a success to celebrate. The pattern matters as much as the level. If reviewers override frequently on certain types of cases, investigate those too, because the AI may be systematically wrong about that category and your reviewers have found it before your monitoring did.

A worked case: the same decision, reviewed two ways

Your agency uses an AI system to determine benefit eligibility. The system reviews applications and makes a determination. Take one application and run it down both paths, because the org chart, the software, and the job title are identical in each. Only the behaviour differs, and the difference decides whether a citizen gets a decision or gets processed.

The meaningful path. The AI reviews an application and recommends denial. The recommendation goes to a caseworker. The caseworker reviews it and writes: "The AI flagged lack of recent employment history. But this person has been caring for a seriously ill family member, which I know from the notes, and employment might not be feasible. The AI does not account for this. I am overriding and approving." The caseworker's decision stands. If the person later complains, there is a human who made the decision and can defend it.

The performative path. The AI reviews the same application and recommends denial. The recommendation goes to the same caseworker. The caseworker skims the AI's reason, thinks "looks right," and signs off. The person is denied without any meaningful human review. When the person appeals, the human says: "The AI determined eligibility. I just confirmed it." One of these is meaningful human review. One is theater, and the second one leaves the agency with no defensible account of its own decision.

Designing so meaningful review is possible

Reviewers cannot fight the workflow they are given. If you have any influence over how a tool is deployed, six design choices decide whether real review is even available to the person doing it. None of them require a better model, and most of them cost less than the incident they prevent.

  • Make AI decisions explainable. Humans cannot review what they do not understand. Opacity is not a technical detail; it is a decision to make review impossible.
  • Provide context to reviewers. Give them the data the AI saw, plus the relevant context it did not see.
  • Train reviewers. They need to understand the system well enough to evaluate its recommendations critically, including where it is known to be weak.
  • Empower reviewers. Make it easy to override, and make override authority explicit rather than something people have to guess at.
  • Monitor override patterns. Frequent overrides on a case type are a signal about the model, not about the reviewer.
  • Build feedback loops. When an override turns out to be wrong, that should feed back into retraining, so the system learns from the correction instead of repeating the error.

Applying it: the Meaningful Oversight Test

Before any AI tool that affects the public goes live in your office, walk it through this five-question test. It draws on the federal AI risk management guidance that asks agencies to ensure humans can understand, oversee, and intervene in automated decisions. Score each question 0 (no), 1 (partly), or 2 (yes). A tool scoring under 7 is not ready for a member of the public.

  1. Time. Does the reviewer have enough time per case to actually read it, not just glance at a score?
  2. Reasons. Does the reviewer see why the system reached its recommendation, in plain language, plus the underlying data?
  3. Symmetry. Is overriding the AI exactly as easy as agreeing with it?
  4. Authority. Can the reviewer change the outcome without manager pre-approval, and does the final record show their name?
  5. Feedback. Are override rates and outcomes tracked, reviewed monthly, and used to fix the tool?

Be careful about what a passing score means. Scoring above the cutoff tells you that five known failure modes have been addressed in the design. It does not tell you that oversight is working, because the score measures the workflow you built and not the behaviour that emerges inside it. Treat it as a gate before launch and re-run it against observed practice afterwards, using the override rate and a look at what reviewers actually write. A tool can pass the test on paper and drift into rubber-stamping within a quarter.

Run Maria's tool through this and it scores about 2 out of 10. The fix was not buying a better model. The agency cut the quota in half, put the flag reasons and the claimant's full history on the review screen, made "deny" require the same single click as "approve," and started reviewing the override rate every month. Wrongful freezes fell and reviewers reported they finally felt like they were doing their job instead of feeding a machine. Note what the agency did not claim: it did not claim the tool was now correct. It claimed reviewers could now catch it when it was not.

When taking the human out is acceptable

Human review is not free, and not every decision needs it. The honest principle is to match the level of human control to the stakes. Spam filtering on a public inbox can run fully automated; a wrong call costs a re-sent email. Routing a pothole report to the right department can be automated with light monitoring. But any decision that grants, denies, reduces, or revokes a right, a benefit, or a status, or that flags someone to law enforcement, needs a meaningful human in the loop. The bright line is harm: the closer a decision comes to taking something from a person, the more real the human control must be.

This is not a loophole in the rule that important decisions need a human. It is the definition of "important" doing its work. The failure mode to watch is scope creep, where a tool approved for low-stakes routing quietly starts informing a consequential decision because it was already there and already trusted. When the stakes of a tool's output change, the oversight requirement changes with it, and somebody has to notice.

Anti-patterns to watch for

  • "Human review" that is rubber-stamping. Humans "review" AI decisions but check an approved box without actually reviewing. Risk: the system functions as if it were fully automated while the agency tells the public there is human oversight.
  • Calling meaningful review a solution. Treating a reviewer as the fix that makes a flawed system safe. Risk: monitoring of the model stops, because the human is assumed to be catching everything, and nobody measures whether they are.
  • Treating the oversight score as a clearance. A tool passes the five-question test and is declared safe. Risk: the score describes the workflow as designed, not as practised, and nobody re-runs it once real quotas and real fatigue arrive.
  • Responsibility diffusion. "The AI recommended it, but the human signed off, so it is not clear who is responsible." Risk: no one feels accountable and errors are never addressed.
  • Inadequate reviewer training. Humans are supposed to review AI decisions but do not understand the system or its limitations. Risk: they cannot do meaningful review, so they accept the recommendation.
  • No authority to override. A reviewer thinks the AI made a mistake but cannot overturn it. Risk: known errors go uncorrected, and the reviewer learns to stop noticing.
  • Viewing humans as a bottleneck. "AI is faster and more efficient if we remove the human from the loop." Risk: you remove the checks and balances and increase the risk of systematic errors and harm at scale.

Practice prompts

  • In your agency, when AI systems are used to make decisions, who reviews those decisions, and what does that review actually involve minute by minute?
  • Have you ever disagreed with an AI system's recommendation? What did you do, and what did it cost you to do it?
  • What would need to be true for human review of AI decisions to be meaningful in your context? Name the specific change that is missing.
  • Take one tool your team uses and score it on the five-question test. Where does it lose points, and who owns each of those five answers?

Reflection

Pick one AI system your agency uses and answer five questions honestly. Is there human review of its decisions? Is that review meaningful or performative? Do humans have authority to override? Are reviewers trained on the system? Who is accountable for the decisions it produces? If the review is performative, work out what would need to change to make it meaningful, and be specific about whether the obstacle is time, information, symmetry, authority, or feedback. Then ask who in your organisation can change that thing, and what it would take to get it on their list.

Glossary

  • Automation bias. The tendency to rely on automated decisions even when they should not be trusted.
  • Meaningful human review. Genuine oversight where a human can override an automated decision.
  • Performative review. Review that goes through the motions but does not actually influence decisions.
  • Override. A decision by a human to reverse or change an AI recommendation.
  • Responsibility. Accountability for decisions and outcomes.
  • Human on the loop. A pattern in which a person can watch and intervene, but the system keeps moving whether they engage or not.

Closing

The goal is not to remove AI from government. The goal is to use AI well, with humans exercising judgment, taking responsibility, and catching the errors that AI systems make. That is how you get the benefits of the technology while keeping the human values that make a government decision legitimate: dignity, context, and accountability. It is the hard way, because it costs time per case, screen space for reasons, and the managerial nerve to let staff say no. It is also the way that survives a hearing, and the only version of "a human reviewed it" that means anything to the person on the other end.

Key Takeaways

  • A human in the loop only counts if they could realistically say no. A reviewer with no time, no reasons, and no authority is a rubber stamp, and "a human reviewed it" will not hold up.
  • Meaningful review requires authority, information, responsibility, and incentive. All four at once, not someone checking a box.
  • Meaningful review is a mitigation, not a cure. It stops errors becoming final decisions; it does not make the underlying system correct, so keep measuring the model too.
  • Automation bias is the default, not the exception. Tired, busy people trust confident machine scores, and an accuracy figure describes past tested cases rather than the case in front of you.
  • Watch the override rate and its pattern. Near zero across thousands of cases is a warning sign; frequent overrides on one case type say the model is wrong about that type.
  • Make overriding as easy as agreeing. If saying no costs more clicks and more risk than saying yes, staff will say yes, and the AI effectively decides.
  • Transparency enables review. If the system's reasoning is opaque, meaningful review is impossible no matter who is assigned to it.
  • Match human control to harm. Low-stakes routing can run automated; any decision that takes a benefit, right, or status away from a person demands meaningful human review.

Frequently Asked Questions

Our reviewers agree with the AI almost every time. Is that good news?

Treat it as a defect to investigate rather than a success to celebrate. If a tool flags thousands of cases and produces almost no overrides, the most likely explanation is that reviewing is harder than agreeing. Look at time per case, whether reviewers see the reasons behind each recommendation, and whether overriding requires paperwork that agreeing does not.

The system is 94% accurate. Why should a caseworker second-guess it?

Because that figure is an estimate of past performance on the cases that were tested, not a probability about the application on the screen. It tells you nothing about whether this case resembles the tested ones, and the people who do not resemble them are exactly the people a human is there to catch. Use accuracy to decide how much scrutiny a system deserves overall, never to decide an individual case.

Does every automated decision need a human reviewer?

No, and pretending otherwise wastes the review capacity you need for decisions that matter. Match control to stakes: spam filtering and routing a pothole report can run automated, while anything that grants, denies, reduces, or revokes a right, benefit, or status needs a meaningful human. Watch for scope creep, where a low-stakes tool quietly starts informing a consequential decision.

What if a reviewer overrides the AI and gets it wrong?

Then they are responsible, and that is the point rather than a flaw. Responsibility is one of the four conditions that makes review meaningful. A reviewer who bears no consequence has no reason to think hard, and a reviewer who is punished only for overriding and never for agreeing has been taught which answer is safe. The feedback loop should also carry that wrong override back into retraining.

Our tool passed the five-question oversight test. Are we covered?

You have addressed five known failure modes in the design of the workflow. That is a gate before launch, not a clearance. The test scores what you built, not what people do inside it once quotas and fatigue arrive, so re-run it against observed practice and read what your reviewers actually write in their decisions.