Human-in-the-Loop: Design and Implementation
Catriona Hargrove manages the benefits processing division at a county human services agency: 34 case managers, approximately 12,000 active cases, and a growing backlog that her director had flagged as a performance risk two budget cycles in a row. When the agency's IT department proposed an AI-assisted eligibility screening tool, Catriona was cautiously supportive. The tool would ingest application data and produce a preliminary eligibility determination for each case, Eligible, Ineligible, or Needs Review, before the case manager opened the file. Her concern was not the technology. Her concern was what "human-in-the-loop" would actually mean in practice, for 34 case managers with an average caseload of 353 files. She had seen "human oversight" reduced to a checkbox in other agencies. She was not going to let that happen in hers. The design process took four months, longer than anyone expected, and produced a system that Catriona now considers the most important governance decision her division has made in a decade.
Why Most Government AI Cannot Be Fully Autonomous
Government decisions routinely affect fundamental rights: access to benefits, freedom of movement, housing security, the ability to work in a licensed profession. These decisions carry due process obligations under the Fifth and Fourteenth Amendments, statutory requirements under program legislation, and equity obligations under civil rights law. When a human makes one of these decisions and gets it wrong, there is a process for correction: an appeal, an administrative hearing, a supervisory review. When an AI system makes these decisions autonomously and gets them wrong at scale, the correction mechanisms are often absent, the affected populations are frequently among the most vulnerable, and the discovery that something went wrong may take months or years.
The curriculum states the principle broadly: government decisions often affect fundamental rights, and those decisions should not be fully automated. Read literally that sweeps in a great deal of routine administrative work, but it errs on the side of protection, so treat it as the default and require a documented, reviewed exception rather than assuming a category is out of scope. The burden sits with whoever wants to remove the human, not with whoever wants to keep one.
Human-in-the-loop, abbreviated HITL, is the framework for keeping human judgment in the decision path when AI assists with these determinations. HITL does not mean humans rubber-stamp AI decisions. It means humans retain genuine authority and exercise genuine judgment at the decision points that matter most. The stated purpose is a division of labor: AI provides analysis, humans make decisions, and the design balances speed and consistency against accountability and human judgment.
What the Combination Does and Does Not Buy You
The case for HITL usually runs like this: humans are slow and inconsistent, AI is fast and consistent, so the combination is better than either alone. That is true only under a condition the sentence leaves out. The combination is better than either alone when the human review is real. When it is not, the pairing is worse than either component on its own, because an automated decision now travels with a human signature on it, and that signature suppresses exactly the scrutiny that would otherwise catch the error.
This is the single most important correction to make before designing anything. Meaningful human review is a mitigation, not a cure. It reduces the rate at which an AI error becomes a final, consequential decision. It does not drive that rate to zero, and it does not convert an unreliable system into a reliable one. A design document that lists "human review" in the control column and then treats the residual risk as handled has not managed the risk; it has renamed it.
The pilot analogy holds. An autopilot handles routine conditions faster and more consistently than a human can. The pilot monitors for conditions the autopilot is not equipped to handle, decides when those conditions arise, and maintains the skills to take control at any moment. An autopilot that the pilot does not understand and cannot override is a safety hazard, not a feature. The same logic applies to government AI, and the same question applies: can the person in the seat actually fly the aircraft today, or only in the job description?
The Four Design Patterns
HITL systems in government typically implement one of four patterns, often in combination. The four are named in the curriculum and each has a canonical example.
The recommendation pattern presents the AI's assessment to a human who then makes an independent decision. The AI provides analysis; the human decides whether to accept, modify, or reject it. The textbook example is a lending decision: the system recommends approval, and a human reviews the file and makes the final call. For Catriona's eligibility tool, this is the primary pattern for cases the tool flags as Needs Review. The case manager sees the preliminary determination and the factors that drove it, then makes the final eligibility decision, and that decision is recorded regardless of whether it matches the recommendation.
The triage pattern routes cases to the appropriate level of review based on complexity, not just on AI confidence. Simple, clear-cut cases, a new application meeting all criteria with complete documentation, go to a junior case manager with AI-provided analysis. Complex cases, prior appeals, conflicting documentation, edge-case eligibility criteria, route to a senior case manager or supervisor. The AI determines which queue the case enters; humans work the queues. Catriona's implementation uses this pattern for initial routing: about 40 percent of cases route to a standard queue, 45 percent to a review queue, and 15 percent to a supervisor queue.
The override pattern lets the AI process routine determinations while giving human reviewers the authority and the interface to intervene, with a recorded justification, when they identify a problem. A system approves an application, and a human can still deny it if review surfaces an issue. This pattern is appropriate only where the error rate is very low, the consequences of error are limited, and an appeal mechanism exists. Benefits renewal for unchanged cases is a plausible application. Initial eligibility determination for new applicants is not.
The escalation pattern uses confidence thresholds to decide what gets human review. The curriculum's illustration routes cases above 90 percent confidence to automated handling and cases below 70 percent to a human. Notice that this leaves the band between the two undefined, and an implementation has to write a rule for it rather than let cases fall through. Catriona's division uses a single threshold instead: any case where the tool's confidence in its preliminary determination is below 80 percent enters the review queue, whatever the determination says.
Treat every one of those percentages the same way: they are thresholds an agency sets for itself, records in advance, and revisits, not legal tests and not properties of the technology. Two agencies running the same tool can defensibly choose different numbers. What is not defensible is choosing a number after the fact to explain a workload you already have.
What a Confidence Score Is Not
Confidence thresholds only work if everyone in the workflow understands what the number means. A confidence score is the system's own estimate, produced by the same machinery that produced the determination. It is not a probability that this determination is correct, and a high score is not evidence that this particular case was handled well. The same caution applies to a headline accuracy figure. A stated accuracy rate summarizes how the system performed on the cases someone tested it against in the past. It says nothing binding about the file open on the screen right now, which may differ from every case in that test set.
The operational consequence is that "the model was 90 percent confident, so I approved it" is not a review, and it is not a defensible entry in a case record. Confidence belongs in the workflow as a routing signal, deciding which queue a case enters, and it belongs in the reviewer's context as one input among several. It does not belong as a substitute for looking at the evidence.
Alert Design: Triggers and Thresholds
HITL systems are only as good as their alert design. Alerts should trigger human attention when the stakes are highest, not only when the AI is most uncertain. Those two conditions are related but not identical, and a system tuned purely on uncertainty will route confident errors straight through.
The curriculum lists six conditions that should trigger human review: low-confidence decisions; unusual or unprecedented cases; high-stakes decisions such as large benefit awards, job offers, and parole; cases that fail sanity checks; demographic outliers; and decisions that contradict other information the agency holds. Catriona's division added its own high-stakes triggers on top: decisions affecting a constituent's access to housing, food, or medical care; cases involving individuals flagged as vulnerable in the system, including domestic violence history, active child welfare involvement, or documented disability; first-time determinations for any applicant; and cases where the application data contains apparent inconsistencies the tool flagged but did not resolve.
Demographic outlier triggers deserve a word of their own. They detect when a case pattern is rare in the training data, which means the system's confidence rests on limited precedent. Those cases need human review precisely because reliability on novel pattern types is lower than reliability on common ones, and the confidence score does not always reflect that gap.
Alert quality matters as much as alert quantity. A good alert is specific and actionable rather than generic, explains why it triggered, provides the context the reviewer needs, and states a clear path forward. An alert that says "Review recommended" satisfies none of those. An alert that says "Eligibility determination: Ineligible. Primary factor: Income exceeds threshold by $127/month. Confidence: 74%. Flag: Applicant income figure differs from prior-year income by 38%, possible reporting error or income change. Recommend case manager verify income documentation before finalizing" satisfies all four, because it gives the reviewer a specific task and a specific reason for it.
The Review Workflow in Four Parts
An effective review process has four components, and skipping any one of them produces a system that looks supervised without being supervised.
The alert system needs clear trigger conditions, delivery to the right person rather than a shared inbox nobody owns, and enough context and data attached that the reviewer can act without hunting for the file. The review interface shows the AI recommendation, the evidence and factors behind it, and the demographic context of the case; allows an easy override; and captures the human decision. Ease matters here more than it sounds. An override buried behind a long justification form and a chain of confirmation screens is an override that will not happen at the end of a long day.
Decision capture records four things for every reviewed case: what decision was made, why if it was an override, how long the review took, and whether any issues were identified. Review duration is the underrated field. It is the closest thing to a direct measurement of whether reviews are happening, and it is the metric a supervisor can compare across reviewers and across case types without waiting for an appeal to surface a problem.
The feedback loop is where captured decisions flow back to improve the system. Be precise about what that sentence promises. Overrides sitting in a database change nothing. A feedback loop improves the system only when someone reads the override reasons, decides what they imply, retrains or reconfigures, and revalidates the result before it goes back into production. Absent those steps, "the system learns from human overrides" describes an intention, not a control, and it should not be counted as one in a risk register.
Reading Override Rates
How often humans override the AI is the most useful single indicator of whether the loop is real, provided it is read correctly. The curriculum gives three bands and one cross-cut.
| Observed override rate | What it suggests | What to do |
|---|---|---|
| High, above 30 percent | Staff do not trust the system's recommendations | Investigate why. Improve the system or lower the confidence threshold at which cases route to automation. |
| Low, below 2 percent | The system may be over-trusted | Verify that humans are actually reviewing. Do not read the number as evidence of model quality. |
| Middle band, 5 to 15 percent | Humans are occasionally catching things the system missed | Consistent with healthy oversight, but not proof of it. Confirm with sampling audits. |
| Split by demographic group | Uneven override rates across groups | Treat as a potential bias indicator and investigate, whether the unevenness favors or disfavors a group. |
Two cautions attach to that table. First, the bands are numbers an agency adopts and records in advance so that later readings mean something; they are not legal thresholds and nothing turns on them by itself. Second, an override rate inside the healthy band is consistent with genuine review, not a demonstration of it. A team can produce a middling override rate by overriding the easy cases and waving through the hard ones. The rate tells you where to look; the sampling audit tells you what is actually happening.
An override rate at or near zero is a finding. Catriona's rule is that a case manager showing a zero percent override rate over three months triggers a supervisor conversation. It is not a signal of excellent AI performance and it should never be reported upward as one.
Preventing Automation Bias
The most persistent challenge in HITL implementation is not technical. It is behavioral. Automation bias, the tendency to over-trust and under-scrutinize automated recommendations, is well documented in research on human decision-making in supervised automation. Case managers who process twenty AI-recommended denials a day will, over time, begin approving them without the scrutiny they applied in week one. The curriculum names four failure modes: reviewers get busy and skip reviews, reviewers develop automation bias, reviewers do not understand the recommendations they are approving, and reviewer quality varies across the team.
Six mitigations answer those four. Sampling audits randomly check whether humans actually reviewed. Training on AI recommendations teaches reviewers what to look for. Clear guidelines define when an override is appropriate, so that overriding is a documented professional act rather than an act of nerve. Escalation options let hard cases move higher in the organization instead of being resolved by whoever happens to hold them. Workload management keeps reviewers from being handed more cases than anyone could genuinely review. Performance monitoring tracks reviewer quality and consistency over time.
Catriona implemented three of these concretely. Case managers must document their decision rationale for every case where their final determination differs from the preliminary one, which forces an independent judgment rather than passive acceptance. Supervisors run monthly sampling audits, pulling a random 5 percent of cases where the case manager agreed with the tool and reviewing whether the documentation reflects genuine independent review or formulaic sign-off language. And the division tracks override rates by case manager and by case category, treating outliers in either direction as questions rather than verdicts.
Note what the audit looks at. It samples the agreements, not the disagreements. Disagreements document themselves. Agreement is where a rubber stamp hides, and it is the only place an audit can find one.
What Makes a Review Real
Pulling the sources together, human review that deserves the name has four qualities, and a design should be able to point at where each one lives.
It is genuine: the reviewer forms an independent judgment and has the authority to reach a different answer, with the override made easy to exercise and normal to use rather than exceptional and effortful. It is timely: it happens before the decision takes effect, not after a constituent has already been denied. It is competent: the reviewer has been trained on how the system works, what it is good at, and where it fails, and can see the reasoning behind a recommendation rather than just its output. It is appealable: the affected person can contest the result and obtain a further human look from someone not involved the first time.
A review process missing any one of the four is not a human alternative to automated decision-making. It is automated decision-making with a countersignature, and describing it otherwise in a system inventory or an impact assessment misstates the agency's actual control posture.
Appeal Mechanisms Close the Loop
Any government decision that AI assists in making must have a clear, accessible appeal mechanism that does not require the constituent to navigate the AI system's logic to exercise their rights. The appeal pathway must be prominently disclosed at the point of decision notice, must provide a timely hearing before a human reviewer who was not involved in the original determination, and must include a full explanation, in plain language, of the factors that led to the original decision.
The curriculum sets out five components of the appeal process itself. The affected person receives a clear explanation of how the system contributed to the decision. They can request human review by senior staff. They can supply additional information the original determination did not have. They receive a written explanation of the appeal outcome. And there is a path to further escalation if the appeal does not resolve the matter. Appeals and their outcomes should be tracked and monitored as a body of data, not filed and forgotten.
That tracking is what makes appeals useful twice. Cases where an AI-assisted determination was overturned identify specific decision types where the recommendations are unreliable, and that signal is more trustworthy than any benchmark a vendor supplies, because it comes from your own caseload and your own population. That data should flow back into the system improvement process, subject to the same caution as the internal feedback loop: it improves the system only when someone acts on it.
Anti-Patterns
- Humans become rubber stamps. The review step exists in the workflow diagram and in nobody's actual behavior. Countermeasure: monitor override rates and audit the cases where the human agreed with the system.
- Calling meaningful review a solution. Listing "human review" as the control and treating the risk as closed. Meaningful review reduces the rate at which errors become final decisions; it does not eliminate them and it does not fix the underlying system.
- Calling a rubber stamp a human alternative. Describing a countersignature step as human oversight in an inventory, impact assessment, or public notice, when the reviewer is not genuine, timely, competent, and backed by an appeal.
- No training for reviewers. Staff are asked to evaluate recommendations from a system nobody has explained to them. Countermeasure: invest in training on how the system works and where it fails.
- No appeal mechanism. The decision lands on the constituent with no route to challenge it. Countermeasure: provide a clear, disclosed process for challenge.
- Reviewers not empowered to override. The override exists but is slow, discouraged, or career-risky. Countermeasure: make override easy to exercise and normal to use.
- The system trusted implicitly. Outputs are accepted because they come from the system. Countermeasure: build in structured skepticism and regular auditing.
- Reading a confidence score as a probability of correctness. "It was 90 percent confident" is not a review finding and does not transfer to the case in front of you.
- Counting captured overrides as system improvement. Data in a table is not learning. The loop closes only when someone analyzes, retrains or reconfigures, and revalidates.
- Celebrating a zero override rate. Reporting a near-zero override rate upward as evidence the tool is performing. It is a finding that requires investigation, not a result.
Practice Prompts
- Design a human-in-the-loop workflow for benefits eligibility. Specify which of the four patterns applies to each case type, where the alert fires, what the reviewer sees, and what gets captured at decision time.
- Develop alert rules for a fraud detection system. Start from the six trigger conditions, decide which apply, and write each alert so that it is specific, explains its trigger, carries context, and states a next action.
- Create a training program for reviewers of an AI system in your agency. Cover what the system does, its known failure modes, when an override is appropriate, and how to document one.
- Design an appeal process for AI-influenced hiring decisions. Include the explanation of the system's contribution, the request for senior human review, the route for additional information, the written outcome, and the further escalation path.
- Take one system your agency already runs and write down the override rate band you would consider healthy, before you look at the data. Then look. Record both, and record the date you wrote the first number.
Reflection
Think about a decision process in your own work where a recommendation arrives from somewhere else, from a system, a scoring tool, a checklist, or a more senior colleague, and you are formally the decision-maker. How often do you reach a different answer? If the honest number is close to never, what would it take for you to notice that the recommendation was wrong? Would you have the information, the time, and the standing to say so? That is the question every case manager in a HITL workflow faces every day, and the answer is a property of the design around them, not of their character.
Now consider the reverse. If your override rate were high, would anyone ask why, or would the number simply be read as friction to be engineered away? Agencies that only investigate one tail of that distribution end up optimizing toward agreement, which is the outcome the whole design was meant to prevent.
Glossary
- Human-in-the-loop (HITL). A design in which humans retain authority and exercise judgment in the decision path while AI provides analysis and recommendations.
- Recommendation pattern. The AI produces an assessment and a human makes an independent final decision on whether to accept, modify, or reject it.
- Triage pattern. The AI routes cases to the appropriate level of human review based on complexity, and humans work the resulting queues.
- Override pattern. The AI processes routine determinations while a human retains documented authority to intervene and reverse a result with justification.
- Escalation pattern. Confidence thresholds decide which cases proceed through streamlined handling and which are flagged for mandatory human review.
- Confidence score. The system's own estimate of its certainty, produced by the same machinery as the determination. A routing signal, not a probability that the specific determination is correct.
- Automation bias. The documented tendency of people supervising automated systems to over-trust and under-scrutinize the recommendations those systems produce.
- Override rate. The proportion of reviewed cases in which the human decision differs from the AI recommendation, tracked overall, by reviewer, by case category, and by demographic group.
- Sampling audit. A periodic random check of cases, especially cases where the human agreed with the system, to test whether documented review reflects genuine independent judgment.
- Demographic outlier trigger. An alert condition that fires when a case pattern is rare in the training data, so the system's recommendation rests on limited precedent.
- Appeal. The disclosed route by which an affected person contests a determination and obtains review by a human not involved in the original decision.
Related Lessons
- Systematic AI Output Validation covers the checking discipline that a reviewer applies once a case is in front of them.
- Bias Detection Tools and Methods supplies the techniques behind demographic outlier triggers and the group-by-group reading of override rates.
- Quality Assurance for AI Work Products extends sampling audits from decisions to work products generally.
- Continuous Monitoring Fundamentals is the next step: watching the deployed system over time rather than case by case.
- Minimum Risk Management Practices places human review and override among the required safeguards for federal AI systems.
- The Human in the Loop introduces the concept at a foundational level for readers who want the shorter treatment first.
Closing
Catriona's four months were not spent choosing a vendor or tuning a threshold. They were spent deciding, case type by case type, who decides, what they see when they decide, what gets recorded, and what happens when they disagree. That is what a HITL design actually is. The technology question, which pattern, which threshold, which interface, is downstream of the accountability question, and an agency that answers them in the other order ends up with a workflow diagram that describes a control it does not have.
The test she applies now is simple enough to say in a hallway. Point at the person who decides. Show that they can reach a different answer, that they did so recently, that they understood what they were looking at, and that the constituent can challenge the result in front of someone new. If any of those four cannot be demonstrated, the loop is open and the system is running without the safeguard everyone believes it has.
Key Takeaways
- HITL in government is a due process requirement, not just a design preference. Decisions affecting fundamental rights carry constitutional and statutory obligations that require genuine human judgment in the decision path, not ceremonial approval of automated outputs.
- Meaningful human review is a mitigation, not a cure. It reduces the rate at which AI errors become final decisions. It does not eliminate them, and it does not make an unreliable system reliable.
- The four patterns serve different use cases. Recommendation, triage, override, and escalation each have appropriate applications, and most deployments combine several depending on case type, consequence level, and confidence.
- Confidence thresholds are numbers you set and record, not legal tests. Write them down in advance, define what happens in any band between them, and revisit them deliberately rather than adjusting to fit a workload.
- A confidence score is not a probability that this case is right. Neither is a headline accuracy figure, which describes past performance on cases someone tested, not the file open now.
- Alerts must tell reviewers what to do, not just that review is needed. An effective alert is specific, explains its trigger, carries context, and states a clear next action.
- The workflow has four parts and all four are load-bearing. Alert system, review interface, decision capture, and feedback loop. Capturing review duration is the cheapest early warning that reviews are not happening.
- Automation bias is the dominant operational risk. Sampling audits of the cases where humans agreed, documentation requirements, workload limits, training, and reviewer performance monitoring are the countermeasures.
- A near-zero override rate is a warning signal, not a success metric. Investigate low override rates. Investigate uneven ones across demographic groups in either direction.
- Appeal mechanisms are a required component. Explanation of the system's contribution, senior human review, a route to add information, a written outcome, and a further escalation path, all disclosed in plain language at the point of decision.
- Appeal outcomes are your best improvement signal, once someone acts on them. Overturned determinations show where recommendations are systematically unreliable, but only a completed analyze, retrain, revalidate cycle turns that data into a change.
Frequently Asked Questions
Is human-in-the-loop the same thing as human oversight? Not necessarily. Oversight is a broad term that covers governance bodies, audits, and after-the-fact review. HITL is narrower and more demanding: a human sits inside the decision path, before the determination takes effect, with the authority to reach a different answer.
If we have a human reviewing every case, is the AI risk handled? No. Review is a mitigation that lowers the rate at which errors become final decisions. It has its own failure modes, chiefly automation bias, and its effectiveness has to be measured rather than assumed. A design that treats review as closure has stopped managing the risk.
What override rate should we target? The curriculum describes rates above 30 percent as a sign the system is not trusted, rates below 2 percent as a sign it may be over-trusted, and a band of roughly 5 to 15 percent as consistent with healthy oversight. Those are figures your agency adopts and records for itself, and none of them proves anything on its own without a sampling audit behind it.
Our override rate is zero and our vendor says that proves the model is excellent. Is it? A zero rate is a finding to investigate, not a result to report. It is equally consistent with an excellent model and with a review step that has become ceremonial, and the only way to tell the difference is to audit the cases where the human agreed.
Which decisions can be fully automated? The curriculum's position is that government decisions affecting fundamental rights should not be. That is a broad statement, and treating it as the default is the safer reading: require a documented, reviewed exception to remove the human rather than assuming a category is out of scope.
How do we stop reviewers from just clicking approve? Require a documented rationale wherever the human decision differs from the recommendation, sample the agreements monthly and read them for formulaic language, cap reviewer workload, track review duration, and train reviewers on the system's known failure modes so they know what they are looking for.
Does the system get better automatically when reviewers override it? No. Captured overrides are inputs to an improvement process, not the process itself. Someone has to read the reasons, decide what they mean, retrain or reconfigure, and revalidate before anything changes in production.
Who should hear an appeal? A human who was not involved in the original determination, working from a plain-language explanation of the factors behind the decision, with the ability to accept additional information, issue a written outcome, and refer the matter further if it remains unresolved.
Skill.re