←
CAP Certification
Visionary · M18 · lesson 18 of 55 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Designing Human-AI Collaboration Frameworks

15 min

Tae-Yang Lindqvist designed what she thought was a clean human-AI system for her insurance company's claims adjustment team. The AI would triage incoming claims, flag the high-complexity ones for human review, and auto-process the straightforward ones. In the first month, adjusters started overriding the AI on simple claims to check its work, doubling handling time on cases that were supposed to be fully automated. In the second month, they began rubber-stamping the AI's recommendations on complex claims without reading the rationale, precisely the cases where human judgment was most needed. Six months in, the worst of both worlds was operating: humans adding friction to easy decisions and removing themselves from hard ones.

Tae-Yang's experience is a classic failure of framework design, and it is worth being precise about what actually went wrong, because the obvious diagnosis is the wrong one. The system was not badly built and the adjusters were not badly behaved. The design defined what the AI would do. It did not define what the humans would do, when they were responsible for a decision versus when they were reviewing one, or what "high complexity" meant in terms an adjuster could apply the same way twice. People filled those gaps with their own judgment, as people always will, and their judgment pulled in opposite directions on the two claim types. This lesson shows you how to design human-AI collaboration that actually produces the outcomes each party is better suited to deliver.

The Core Design Question

Before deciding what AI will do in any collaboration, answer one foundational question: for which decisions do the costs of human error exceed the costs of AI error, and for which is the reverse true? This is not a technical question. It is a business risk question, and answering it requires understanding that the two kinds of error have completely different shapes.

AI makes errors that are systematic, hard to predict in advance, and capable of occurring at high volume simultaneously. When the model is wrong, it is wrong the same way across every case sharing the relevant feature, and it is wrong at machine speed. Humans make errors that are individual, cognitively predictable, and resource-constrained. Human error rates rise with fatigue, time pressure and volume, but each mistake is bounded to the case in front of the person making it. Those two profiles fail in different directions: a systematic AI error in a claims triage system could auto-approve 10,000 fraudulent claims in one night, while a human adjuster making the same mistake affects one claim. The reverse cost is just as real. A human reviewing every claim introduces capacity constraints that create 5-day processing delays, where AI processing every claim reduces that to 4 hours.

Designing the framework means deciding, explicitly and in writing, where each type of error is tolerable and where it is not. The word to hold onto is explicitly. Every organisation makes this allocation whether or not it discusses it; the difference is that an undiscussed allocation gets made by whoever configured the system, on grounds nobody recorded, and it cannot be revisited when the business context changes. A written allocation can be reviewed when volumes grow, when the fraud landscape shifts, or when the model is retrained.

Four Collaboration Modes

Human-AI collaboration does not have to be binary, and the most common design mistake is treating it as though it were. Most well-designed frameworks use different modes for different decision types within the same workflow. The four below cover the great majority of cases, and the differences between them are differences in where decision authority sits, not just in how much the AI helps.

ModeWho decidesUse when
AI decides, human auditsThe AI, with humans sampling decisions after the factVolume is high, individual decisions are low-stakes, and audit-based quality control is sufficient
AI recommends, human decidesThe human, on the AI's recommendationIndividual decisions carry moderate to high stakes and AI can reduce but not replace judgment
Human initiates, AI augmentsThe human, throughoutJudgment, relationship context or creativity is central to the task
Human decides, AI monitorsThe human, with the AI watching the aggregateHuman authority is non-negotiable but systematic oversight is still wanted

Mode 1: AI Decides, Human Audits

The AI makes the decision, and humans audit a sample of decisions after the fact, typically 5 to 10 percent, selected either by risk level or at random. Use this where volume is high, individual decisions are low-stakes, and audit-based quality control is genuinely sufficient. Auto-approving small standard claims below a defined threshold, with a monthly audit review, is the archetypal case. The design work in this mode is almost entirely in the sampling strategy: a purely random sample tells you about the average case, while a risk-weighted sample tells you about the cases most likely to hurt you, and most workflows need some of each.

Mode 2: AI Recommends, Human Decides

The AI produces a recommendation with a confidence score and a rationale, and the human makes the final decision. Use this where individual decisions have moderate to high stakes and the AI can reduce but not replace human judgment. In a claims context, the AI flags a claim as high-risk and suggests denial, showing the factors that triggered the flag; the adjuster reads the case and makes the determination. This is the mode most vulnerable to the rubber-stamping failure, because it looks like human control while being trivially easy to perform without exercising any.

Mode 3: Human Initiates, AI Augments

The human does the work and the AI provides supporting information, suggestions or quality checks, with the human remaining in full control throughout. Use this where judgment, relationship context or creativity is central to the task. A claims negotiator managing a phone call with a claimant while the AI surfaces relevant policy details and prior case outcomes in real time is a clean example: the AI is changing what the human knows, not what the human decides.

Mode 4: Human Decides, AI Monitors

Humans make all the decisions and the AI monitors the decision stream for anomalies, drift, or patterns that individuals cannot see at scale. Use this where human authority is non-negotiable but you still want systematic oversight. Clinical physicians make all treatment decisions while the AI monitors aggregate prescribing patterns for statistical outliers that might indicate errors or potential overprescribing. Note what the AI is watching here: not the individual decision, which remains entirely the clinician's, but the shape of the decisions in aggregate, which no individual clinician is positioned to observe.

Most real workflows need two or three of these modes operating simultaneously on different decision types. That is not a sign of a messy design; it is what a design that has taken the core question seriously looks like, because a single workflow almost never has a uniform error-cost profile across every decision it contains.

Defining Handoffs Explicitly

The failure in Tae-Yang's framework was not that the modes were wrong. Mode 1 for simple claims and Mode 2 for complex claims was the right design. The failure was that the handoff criteria between them were vague, so the boundary that was supposed to be a rule became a matter of individual interpretation.

"High complexity" is not a handoff criterion. It is a category. A handoff criterion has to be specific enough that both the AI system and the human operator can apply it consistently without making judgment calls about the criterion itself. The test is simple: if two competent people can look at the same case and route it differently while both following the rule, you have written a category rather than a criterion. Concrete handoff criteria for an insurance claims triage system look like this.

  • Any claim exceeding $25,000 in estimated liability routes automatically to Mode 2.
  • Any claim where the AI confidence score is below 0.75 routes automatically to Mode 2.
  • Any claim involving a claimant with three or more prior claims in the past 24 months routes automatically to Mode 2.
  • Any claim flagged by the fraud detection model routes to Mode 2 with the fraud flag visible to the adjuster.
  • All others stay in Mode 1, with a 5 percent random audit sample.

These criteria can be tested, audited and adjusted. You can ask what proportion of volume each rule diverts, whether the threshold is set too low to be useful or too high to be safe, and what happened to the cases that just missed it. "High complexity" supports none of those questions, which is why a framework built on it cannot be improved: there is nothing concrete enough to change. Writing criteria this way also makes the design reviewable by people who are not close to the model, including risk, compliance and the operational team who will live with the routing.

Designing for Appropriate Human Engagement

The other half of Tae-Yang's problem, adjusters rubber-stamping complex claims, is a design failure about what it means to "review" a recommendation. If the interface shows a recommendation and a button, most humans will click the button, and they will do so faster as they grow more confident in the system. Human review is only meaningful when the design creates genuine engagement with the reasoning. Three interface principles carry most of the weight in human-in-the-loop decisions.

Show the reasoning, not just the recommendation. The AI's recommendation should arrive with the specific factors that drove it, in plain language. "Recommended: Deny. Reason: Claim exceeds policy coverage for the listed vehicle by $8,400" is reviewable, because the reviewer can check the policy and the figure. "Recommended: Deny" is not reviewable at all; it can only be accepted or refused on instinct. The distinction matters more as accuracy improves, because a reviewer who cannot see reasoning learns to trust the label, and trust in the label is precisely what removes the human from the loop.

Create friction proportional to stakes. For high-stakes decisions, require the human to select a reason for their choice from a structured list rather than a free text override. This creates a brief moment of deliberation at the point of decision and generates audit data that a free text box does not. For routine decisions, minimal friction is appropriate, and adding more is not a safety measure but a tax. Friction is a design resource with a real cost, so spend it where the error would hurt.

Make non-agreement visible without penalty. If overriding the AI recommendation requires a supervisor notification or a special process, humans will avoid overriding even when they should, and the override rate will fall to something that looks like agreement and is actually avoidance. Design override processes that are fast and blame-free, and track override rates as a quality signal rather than a performance problem. A team whose override rate has dropped to near zero is telling you something about the incentives around the interface, not about the model's accuracy.

Anti-Patterns

  • Designing the AI side only. Specifying in detail what the system will do and leaving the human role implicit. Tae-Yang's framework failed here, and the gap was filled by individual improvisation in two opposite directions.
  • Category-as-criterion. Routing on words like "complex", "unusual" or "sensitive" that two competent people will apply differently, producing a boundary that cannot be tested or improved.
  • One mode for the whole workflow. Choosing a single collaboration mode for a process containing decisions with wildly different error-cost profiles, then wondering why it is either too slow or too risky.
  • Review as a button. Presenting a recommendation with no reasoning and calling the resulting click human oversight.
  • Uniform friction. Applying the same confirmation steps to routine and high-stakes decisions, which trains people to click through the ones that mattered.
  • Punishing disagreement. Making overrides slow, visible or career-relevant, then reading the falling override rate as improving model performance.
  • Set and forget. Writing the allocation once and never revisiting it as volumes, risk exposure and model behaviour change.

Practice Prompts

  • Take one workflow you are responsible for and answer the core design question for each decision it contains: which error costs more here, the systematic kind or the individual kind?
  • Map each decision in that workflow to one of the four modes, and note where you find yourself wanting a mode that does not exist.
  • Find a routing rule in your organisation that uses a category word. Rewrite it as a criterion specific enough that two people would route the same case identically.
  • Look at the interface your reviewers actually use and ask whether a recommendation arrives with its reasoning attached. If not, write the sentence it should carry.
  • Find out what your override rate is, and whether anyone knows. Then find out what happens, socially and procedurally, to someone who overrides.
  • For a Mode 1 process, examine the sampling strategy and determine whether the sample is random, risk-weighted, or unexamined.

Reflection

Consider a system you work with where humans are formally in the loop. Are they exercising judgment, or performing it? The distinction is usually visible in the timing: if reviews take about as long as reading the recommendation itself, the reasoning is not being engaged. Consider also what your framework assumes about the people in it. Designs fail when they assume constant attention, uniform interpretation of vague words, or a willingness to disagree with a system that management has publicly endorsed. Tae-Yang's adjusters behaved rationally at every step, which is exactly why the framework, not the adjusters, needed changing. Finally, ask who in your organisation could answer the core design question for your highest-volume workflow. If the answer is nobody, that allocation has already been made by default.

Glossary

  • Collaboration mode. A defined pattern of decision authority between human and AI within a workflow, specifying who decides and who checks.
  • Handoff criterion. A rule specific enough that both the system and the operator route a case the same way without interpreting the rule itself.
  • Systematic error. An error repeated identically across every case sharing a feature, the characteristic failure shape of an automated system.
  • Individual error. An error bounded to a single case, the characteristic failure shape of a human decision-maker under load.
  • Confidence score. The model's own estimate of how reliable a given output is, usable as a routing input when its behaviour is understood.
  • Rubber-stamping. Approving recommendations without engaging the reasoning, producing the appearance of oversight without its substance.
  • Override rate. The proportion of AI recommendations that human reviewers reject, best read as a signal about the framework rather than about individuals.
  • Proportional friction. Deliberately matching the effort a decision requires to the cost of getting it wrong.

This lesson sits alongside Human-Centered AI Work Design, which approaches the same territory from the perspective of the work itself and the people doing it, rather than from the allocation of decisions. Human Oversight & Accountability takes up what happens after the framework is designed: who answers for a decision once authority has been distributed between a person and a system, and what evidence of oversight actually demonstrates. High-Risk AI & Enhanced Oversight covers the cases where the allocation is not entirely yours to make, because the decision category itself carries obligations that constrain which collaboration modes are available to you.

Closing

The framework Tae-Yang originally built was not missing a feature. It was missing a set of decisions that nobody had realised were decisions: where each type of error was tolerable, what the boundary between modes actually was, and what a reviewer was supposed to do differently from a rubber stamp. Designing human-AI collaboration well means making those choices deliberately and writing them down in terms specific enough to test. The reward is a system that can be improved rather than merely defended, because every element of it, the allocation, the criteria, the sampling rate, the override process, is concrete enough to measure and adjust. The alternative is not the absence of a framework. It is a framework designed by accident, one interface decision at a time, by people who did not know that is what they were doing.

Key Takeaways

  • Start by asking which error type costs more in your context, systematic AI error or individual human error. The answer determines your allocation of decisions.
  • Four collaboration modes cover most use cases: AI decides and human audits; AI recommends and human decides; human initiates and AI augments; human decides and AI monitors.
  • Most workflows need multiple modes operating simultaneously on different decision types within the same process, because error-cost profiles are rarely uniform.
  • Handoff criteria must be specific enough to apply without judgment calls about the criteria themselves. "High complexity" is a category, not a criterion.
  • Human review is only meaningful when the design requires engagement with reasoning, not just approval of a recommendation.
  • Friction should be proportional to stakes. Too little on high-stakes decisions produces rubber-stamping; too much on low-stakes decisions adds cost without benefit.
  • Override processes must be fast and blame-free, with override rates tracked as a quality signal rather than a performance deficit.
  • Write the allocation down and revisit it, because an undocumented framework cannot be reviewed when volumes, risks or model behaviour change.

Frequently Asked Questions

How do I know whether a decision belongs in Mode 1 or Mode 2? Work from the error costs rather than from the difficulty of the decision. If a wrong answer on this decision type is cheap to absorb individually and the volume is high, audit-based control is proportionate. If a single wrong answer is expensive or hard to reverse, the decision needs a human in front of it, however easy the case looks.

What if my organisation has no confidence score to route on? Use the criteria you do have. Value thresholds, claimant or customer history, and flags from other systems are all concrete and testable. A confidence score is a convenient routing input, not a prerequisite, and routing purely on model confidence is fragile in any case because confidence behaviour changes when a model is retrained.

Is a high override rate a sign the framework is failing? Not by itself. A high rate may mean the model is poorly calibrated for a segment, or that the handoff criteria are sending the wrong cases to review, or that reviewers are engaging properly with borderline decisions. What it should never be treated as is a performance problem belonging to the reviewers, since that is the fastest way to make the number stop telling you anything.