←
AI Readiness & Process Transformation
Proficient · M22 · lesson 22 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

The Handoff Map: AI to Human and Back

15 min

Two handoffs, seventeen years apart. In 2009, a plant scheduler faxes a change order to the night shift and page two never leaves the machine; the line runs 4,000 units to the old spec, and the rework bill lands at $61,000 and one very quiet morning meeting. In 2026, an accounts payable reviewer named Priya opens her queue and finds what the new AI tool has left her: a paragraph. "This invoice from Meridian Supply appears to be a duplicate exception with a possible amount discrepancy; the vendor terms may have changed and the category is likely freight-related; recommend review." No fields. No source. No score. Nothing that tells her what was checked, what was guessed, or where to look. She now has two choices, and both of them are failure: redo the whole task herself, or trust a paragraph that cannot answer a single question. The fax machine at least had the decency to fail visibly. This lesson is about the invisible version: the handoff between a machine and a human, the newest and least specified interface in your process, and the boring, total fix that makes it work.

A New Species of Handoff

Every process professional already knows the oldest truth in operations: workflows do not rot in the middle of tasks, they rot at the seams. The ticket thrown over the wall between sales and fulfillment. The email with the missing attachment. The spreadsheet version nobody can identify. The "I thought you had it" that surfaces three weeks later as a customer escalation. Handoffs are where context evaporates, where ownership blurs, and where the work that was 90 percent done becomes 60 percent done the moment it changes hands. Entire methodologies, from swimlane mapping to RACI charts (responsible, accountable, consulted, informed, the ownership grid you already live in), exist mostly to discipline handoffs between humans.

AI-augmented workflows add a new species to this old genus: the machine-to-human handoff and its mirror, the human-to-machine handoff. In the invoice-exception redesign this chapter has been building, they are everywhere. The AI drafts a categorization and hands it to a reviewer. The reviewer corrects it and hands the correction back. The routine lane escalates a novel case to a specialist. Each of those seams is a handoff, and each one fails in ways the human-to-human version never did, because the parties are asymmetric in a specific, dangerous way: one party cannot ask clarifying questions, and the other party assumes it did.

Think about what actually rescues most human handoffs. Not the process documentation. The follow-up question. "Which version?" "Did you check the PO?" "Is this the same Meridian issue as last month?" Humans repair bad handoffs conversationally, dozens of times a day, without noticing. A machine on either end of the seam removes that repair mechanism entirely. The AI will not call Priya to ask whether "freight-related" was what she meant by her correction. Priya cannot ask the AI what it compared the invoice against. Whatever crosses the seam is all that crosses the seam. If the handoff format does not carry the context, the context is gone, permanently, silently, at volume.

And an unspecified AI handoff does not fail randomly. It fails into exactly two wastes, and you have already met both of them in this level.

The first waste is re-doing. Two lessons ago, in the cow-path story, a legal team bought a contract-summary AI and attorneys ended up reading the summary and then the full contract anyway, because nothing defined when the summary could be relied on, for what clause types, verified how. At the time we treated that as a redesign failure, and it was. But look at it mechanically now and it is a handoff failure: the summary arrived carrying no signal of what had been checked. A receiver who cannot tell what the sender verified has only one safe move, which is to verify everything, which means doing the task twice. The organization paid for the AI's work and then paid a human to repeat it. That is negative automation: a tool that adds a step instead of removing one, purely because its output format hid its own coverage.

The second waste is rubber-stamping. The previous lesson opened with Dana approving item 213 of 240 in 11 seconds, and we diagnosed the accountability sink: a gate designed to absorb blame rather than errors. But ask what fed that sink. Dana's queue handed her 240 undifferentiated items a day with no signal of where to look, no confidence score, no flag, no pointer to the risky field. When every item looks identical and the volume is impossible, the rational human stops looking. Rubber-stamping is not a character flaw in the reviewer; it is the predictable output of a handoff that transmits work without transmitting attention. A bad handoff does not just waste effort. It quietly disables the control you built in the last lesson.

Re-doing and rubber-stamping are opposite responses to the same missing information, and they bracket every unspecified AI handoff: the human either trusts nothing or trusts everything. Both are ruinous. The 95 percent of GenAI pilots that MIT found delivering no measurable return failed, in the report's own autopsy, on workflow integration and learning loops, and the handoff seam is precisely where both live or die. McKinsey's finding that high performers are about three times more likely to fundamentally redesign workflows points at the same seam from the other side: redesigning a workflow largely means redesigning its handoffs. The fix is not clever. It is boring and total.

The Handoff Contract: Specify Every Seam Like an Interface

Software engineers solved this problem for machine-to-machine seams decades ago, and their solution has a name you can borrow: the interface. When two systems exchange data, nobody relies on vibes. There is a specification: exactly these fields, in exactly this format, with exactly these error codes, and either side can be tested against it. The seam is boring because it is specified. Your AI-to-human seams deserve the same discipline, and this lesson's named artifact provides it.

The Handoff Contract is a one-page specification, one per AI-human boundary in your redesigned process, that answers five questions in writing:

  • WHAT passes. The payload: the exact fields, in the exact format, that cross the seam. Named, typed, ordered. Never "the output."
  • With WHAT SIGNAL. The metadata that tells the receiver how to treat the payload: confidence score, evidence citations, risk flags, anomaly markers.
  • WHAT THE RECEIVER DOES with each signal level. The routing and review behavior, keyed to the signal, so two reviewers treat the same signal the same way.
  • WHAT RETURNS. The structured path back: corrections, reason codes, escalations, in a format the system (and the improvement cycle) can consume.
  • WHAT GETS LOGGED. The audit line each handoff event writes: who received what, with what signal, and what they did with it.

Notice what the contract is not. It is not a model specification, a prompt library, or a vendor document. It is a process artifact, owned by the process owner, sitting beside the Gate Spec you wrote in the previous lesson. The Gate Spec defines the checkpoint; the Handoff Contract defines what arrives at the checkpoint and what leaves it. A gate with a four-minute review standard and no handoff contract is a promise without a delivery mechanism, because whether four minutes is physically possible depends entirely on what crosses the seam and how it is arranged. Gates and handoffs are co-designed or they are both fiction.

Specify every AI-human handoff like an interface: what passes, with what signal, what the receiver does with it, what returns, and what gets logged. An unspecified handoff always resolves to re-doing or rubber-stamping.

The five questions sound quick. They are not, and they should not be. The next three sections take the contract's four working elements slowly, because each one hides a design decision that determines whether your pilot lands in McKinsey's redesigned minority or MIT's unmeasurable majority.

Element One: The Payload, Structured, Never Prose

Return to the paragraph Priya received in the opening scene. Its problem is not that it is wrong. Its problem is that it is prose. To act on it, Priya must extract from that paragraph the vendor, the suspected issue, the amount in question, and the category guess, and then go find the source documents herself. Extraction and lookup are precisely the work the AI supposedly did. A prose handoff makes the human re-perform the extraction on top of the review, which is negative automation delivered with perfect grammar.

The rule: the payload is fielded, not narrated. For the invoice-exception routine lane, that means the handoff carries named fields: exception category, invoice amount, PO amount, variance, vendor name, vendor match status, and a pointer to the evidence for each claim. A field can be checked in a glance. A paragraph must be read, parsed, and mentally re-fielded, and every reviewer will re-field it slightly differently, which means your review standard is no longer standard.

The design question that generates the field list is one you already know how to ask, because it is a process question, not a technical one: "What does the receiver's next action need, in the order they need it?" Not "what does the model produce," which is the vendor's default and produces payloads organized for the model's convenience. Walk the reviewer's checklist from the Gate Spec: verify amount against invoice image, verify vendor against PO record, confirm category against the three criteria in SOP-AP-07 (a standard operating procedure, the written instruction that defines how a task is done). The payload should serve those checks in that order: amount claim beside amount evidence, vendor claim beside vendor evidence, category claim beside the criteria test. The receiver's checklist is the payload's table of contents.

The Screen Is the Handoff Made Visible

Here is the consequence most teams miss until month three: the reviewer's screen is not decoration around the handoff. It is the handoff, made visible. If the contract says the amount claim travels with a pointer to the invoice line, but the reviewer's screen shows the claim in one system and the invoice image three clicks away in another, then the contract is being violated at the last inch, where it costs the most. Every click between claim and evidence is time subtracted from judgment, and the arithmetic is unforgiving: the previous lesson budgeted four minutes per gated item, about 290 items a month at the invoice gate. A layout that adds ninety seconds of hunting per item silently inflates that budget by nearly 40 percent, roughly seven reviewer hours a month, and the review either overruns its staffing or degrades into the 11-second click that killed Dana's gate. The four-minute standard is only achievable if the screen was designed for it: claim and evidence side by side, checklist order top to bottom, flags visible without scrolling. When you write a Handoff Contract, sketch the screen on the same page. If the screen cannot be laid out to honor the contract, the contract is wrong or the tooling is.

Element Two: The Confidence Signal, Routing, Not Reassurance

The confidence score is the most misunderstood element of any AI handoff, so let us define it in operator terms before designing with it. When a model reports "92 percent confident," it is not issuing a guarantee, and it is not measuring truth. It is producing a self-assessment, a number that expresses how strongly the model's internal patterns matched this case. Treat it the way you would treat a seasoned employee saying "I'm pretty sure": genuinely useful information, worth acting on, and absolutely not the same thing as being right. The correct use of confidence is routing: deciding which lane an item travels, not whether its contents are true.

Calibration: A Number You Can Actually Test

A confidence signal is only worth routing on if it is calibrated, and calibration has a plain-language definition: when the tool says 95 percent, it should be right about 95 percent of the time. A tool whose "95 percent confident" outputs are wrong 20 percent of the time is not slightly off; it has a calibration problem that will misroute hundreds of items a month into your lightest-touch lane. And here is the good news that most teams never hear: calibration is not a philosophical property. It is measurable, by you, in your pilot, with a spreadsheet.

The check works like this. Sample outputs from each confidence band, say 50 items per band. Have a qualified human judge each one right or wrong against ground truth. Then compare claimed confidence to actual accuracy, band by band: the tool claimed 90 to 100 percent on these fifty, and it was actually right on 46 of them, so 92 percent, healthy; it claimed 70 to 89 on these fifty and was right on 31, so 62 percent, broken. That two-column comparison, claimed versus actual, is the calibration check, and it belongs in week two of your pilot plan as a named task with an owner, right beside the baseline measurements you learned to take in Level 1. It also belongs in your vendor due diligence, as a direct descendant of the Level 1 question sheet: do not accept "our model is well calibrated" from a deck. Ask "show me calibration data on our sample," and treat a refusal or a blank look as data. Every vendor benchmark is a claim to verify, never a guarantee, and calibration is the cheapest verification you will ever run.

Bands Map to Routes, and Sometimes the Route Is No Handoff

Once the signal is validated, the design rule is simple: confidence bands map to routes. Not to feelings, not to font colors, to routes. The invoice-exception lane uses three:

BandRouteWhat the receiver does
High (validated at or above threshold)Sampled laneItem proceeds; a 10 percent random sample routes to the gate for full review, so the escape rate stays measured
MediumFull gate reviewItem routes to the gate; reviewer runs the complete four-minute checklist against the evidence pointers
LowHuman does the task from scratchAI output is suppressed entirely; the specialist receives the case and the source documents, not the draft

The low band deserves a pause, because it contains this lesson's most counterintuitive rule: at low confidence, the best handoff is no handoff. The instinct is to show the human the AI's draft anyway, "as a starting point." Resist it, because of a one-line psychological fact: humans anchor on whatever they see first, even when they know it is unreliable. A reviewer shown a bad draft does not start from scratch; they start from the draft, negotiating with its wrong category and its wrong amount, and they end up slower than a clean start and closer to the AI's error than their own judgment would have landed. Anchoring is why the low-confidence route suppresses the output entirely. Sometimes the most valuable line in a Handoff Contract is the one that says: below this line, nothing crosses.

One boundary from the previous lesson still holds: confidence routing never overrides risk routing. Items above $10,000, first-time vendors, and anything touching a regulated ledger go to the gate regardless of the score, because confidence measures the model's uncertainty and risk measures your exposure, and only one of those shows up in an audit. Gartner's projection that over 40 percent of agentic AI projects will be canceled by the end of 2027 is, at bottom, a forecast about exactly this class of control being skipped: autonomy granted on the model's self-assessment with no independent risk lane. Your contract keeps both.

Elements Three and Four: The Evidence Pointer and the Return Path

The Evidence Pointer: Cite Your Source, Now as a Wire Format

In Level 2 you learned the personal discipline of cite-your-source: never accept an AI claim without the passage it came from. The Handoff Contract industrializes that habit into a format rule: every claim in the payload carries a pointer to its evidence. The extracted amount points to the invoice line it was read from. The vendor match points to the PO field it was compared against. The category points to the clause of SOP-AP-07 whose criteria it satisfied. Not "sources available on request." Pointers, in the payload, resolvable in one click or zero.

Evidence pointers are what make the four-minute review standard arithmetically possible at all. Without them, "verify the amount" means searching the invoice, the PO system, and the vendor record: verifying the universe. With them, it means comparing two numbers that are already side by side: verifying the citation. The reviewer's job collapses from re-research to spot-check, which is the entire economic case for the hybrid step. And evidence pointers buy you a second, subtler asset: teachable overrides. When a reviewer corrects the AI, the pointer shows exactly why the model was wrong, this invoice format puts credits in the debit column, this vendor abbreviates its own name. That "why" is what turns an override from a one-off fix into a labeled lesson, which is precisely what the improvement cycle downstream will feed on.

The Return Path: Where the Learning Loop Lives or Dies

Everything so far has traveled AI-to-human. The fourth element travels back, and it is the one your pilot's future depends on. MIT's autopsy of the 95 percent named the missing learning loop as a central killer: tools that never improve because nobody captures what they got wrong. You installed the capture point in the previous lesson, the override log with reason codes. The return path is that log's wire format: the human's response crosses the seam as structured data, never as an untracked edit. A reviewer who silently fixes the category in the ledger has corrected the invoice and taught the system nothing; the same error returns next Tuesday, and the Tuesday after, which is the exact decay curve that emptied the stalled pilots' user bases. A reviewer who returns CORRECT-CATEGORY with the right value and a reason code has produced a labeled training example worth real money. A handoff with no structured return is a one-way street to a model that never improves.

Then comes the return path's unglamorous plumbing question, the one that separates a learning loop from a wish: where, physically, do the corrections go? Name the destination. A log table the team clusters monthly? The vendor's feedback endpoint? A retraining dataset with a review cadence? "We send feedback to the vendor" without a named destination, a named consumer, and a named cadence is a wish with a progress bar. And the vendor's side of this plumbing belongs in the commercial contract, as another Level 1 due-diligence descendant: ask, in writing, "what happens to our corrections?" Do they tune our instance? On what schedule? Are they visible in any changelog? Do they train shared models (a data-governance question your legal team will want answered anyway)? A vendor who cannot answer is telling you that your return path terminates in a void, and you should price the deal, and the pilot's improvement curve, accordingly.

The fifth question, what gets logged, is the shortest to specify and the one your auditor will read first: every handoff event writes a line. Item ID, payload version, confidence, route taken, receiver, action, return code, timestamp. Accountability stays human, and the log is what makes that sentence checkable rather than aspirational.

The Contract in Full: The Routine Lane, the Reverse Handoff, and a Failure Story

Here is the Handoff Contract for the invoice-exception routine lane, all five questions answered, using the running numbers this chapter has built (all figures illustrative: roughly 1,150 exceptions a month at about $31 each, about 290 items a month reaching the gate). It fits on one page, and it took the AP team lead and the process owner one working session to draft.

Handoff Contract: Routine-Lane Exception Handoff (v1.0). Boundary: categorization AI to gate reviewer. Owner: AP team lead.

  • 1. Payload (fielded, in checklist order): exception_category (one of the three SOP-AP-07 types); invoice_amount; po_amount; variance ($ and %); vendor_name; vendor_match (exact / fuzzy / none); invoice_id and po_id. No free-text summary field exists in v1.0, deliberately.
  • 2. Signal: confidence score (0 to 100) with band label; risk flags (OVER-10K, FIRST-TIME-VENDOR, REGULATED-LEDGER); anomaly flag for any field the extractor marked low-certainty.
  • 3. Receiver behavior by signal: high band routes to sampled lane, 10 percent random pull to full review; medium band routes to full gate review, four-minute checklist per Gate Spec; low band suppresses AI output and routes the raw case to the specialist queue. Any risk flag forces full review regardless of band. Band validity is conditional on the calibration check: 50 sampled outputs per band in pilot week two, claimed-versus-actual plotted; any band off by more than 5 points triggers threshold review before the sampled lane opens.
  • 4. Evidence pointers, per field: invoice_amount cites the invoice line image region; po_amount cites the PO record field; vendor_match cites both name fields compared; exception_category cites the SOP-AP-07 clause whose test it passed. Reviewer screen renders each claim with its evidence adjacent, checklist order, zero clicks.
  • 5. Return codes: APPROVE; CORRECT-AMOUNT (with corrected value); CORRECT-CATEGORY (with corrected type); REJECT-NOT-EXCEPTION (item should never have entered the lane); ESCALATE-NOVEL (no defined type fits; routes to specialist). The correct-* codes descend directly from the Gate Spec's WRONG-AMT and WRONG-CAT override codes, so the override log and the return path are one dataset, clustered monthly, feeding the iteration backlog. Corrections also post to the vendor's feedback endpoint; the vendor's written commitment (monthly instance tuning, changelog entry per cycle) is clause 7.3 of the order form.
  • 6. Log line per event: item ID, payload version, confidence, band, flags, route, reviewer ID, return code, correction values, timestamp. Retained per the audit policy; sampled in the monthly gate health review alongside override-rate and time-per-item.

Now the reverse handoff, because the map in this lesson's title runs both directions. When a low-band item or an ESCALATE-NOVEL lands with the human specialist, that is also a handoff, and it deserves its own short contract, because the failure mode of an unspecified escalation is the specialist starting cold: re-pulling the invoice, re-finding the PO, re-discovering what made the case odd. The escalation payload therefore carries a context packet: the source documents, whatever fields the extractor did read (marked as unverified, displayed dimmed, never as conclusions, to keep the anchoring rule honest), the SOP criteria that were tested and failed, and links to the three most similar past cases with their resolutions. The specialist's resolution then returns through the same coded path, because a novel case a specialist solves is the single most valuable training example the process produces, and letting it evaporate into an email thread is burning the tuition after paying it.

Failure Story: The Handoff That Hid the Error

A procurement team (illustrative, assembled from real patterns) pilots an AI that categorizes supplier invoices. Reviewers complain about queue noise in week one, and the team makes a reasonable-sounding call: instead of per-item handoffs, the AI will email a weekly summary of its categorizations. Less noise, cleaner inbox, and the summary even includes an accuracy estimate. Everyone approves.

Here is what that decision actually did: it changed the handoff's granularity from per-item to aggregate, and granularity determines what the receiver can act on. You cannot approve, correct, or escalate a paragraph that says "412 invoices categorized this week, estimated accuracy 96 percent." So nobody does. Review quietly ceases; the summary is skimmed and archived; the accuracy line stays comfortingly high. In month five, a controller doing an unrelated reconciliation finds it: one vendor's credit memos have been booked as debits, systematically, every single one, for nine weeks, roughly $180,000 of misstatements to unwind across two closed periods. Aggregate accuracy never flinched, because that vendor's lane was a few percent of volume; 96 percent overall accuracy is fully compatible with one lane being 100 percent wrong. The postmortem's finding, and write this sentence down because it is the whole lesson in one line, was not "the model erred." It was "the handoff design hid the error." A per-item handoff with evidence pointers would have surfaced the first flipped credit memo inside a day, as one reviewer's CORRECT-CATEGORY with a pointer to the sign of the amount. The wrong granularity made a fixable error invisible for nine weeks. Handoff design is not packaging. It is what determines what your organization can see.

One bridge before Monday. You have now designed the gates (previous lesson) and the seams between AI and humans (this lesson). The assembly looks solid, which is exactly when a process professional gets suspicious. The next lesson takes the completed redesign and does to it what manufacturing does to a new line before a single unit ships: FMEA, failure mode and effects analysis, the discipline of enumerating every way the assembly can fail, scoring each one, and fixing the worst on paper, before launch does the enumeration for you.

What to Do Monday Morning

  1. Write the Handoff Contract for your pilot's main AI-human boundary. One page, five questions: what passes, with what signal, what the receiver does per signal level, what returns, what gets logged. Draft it with the receiver in the room; the payload's field order is their checklist order.
  2. Run the calibration check. Sample 50 outputs per confidence band, judge each against ground truth, and plot claimed versus actual accuracy per band. Any band off by more than a few points means your routing thresholds are guesses; fix the thresholds before the light-touch lane opens.
  3. Redesign one reviewer screen so evidence sits beside claim. Take the highest-volume review view and eliminate every click between an AI claim and its source. Time five reviews before and after; the delta is your business case for fixing the rest.
  4. Define the return reason codes. Extend your Gate Spec's override codes into a complete return set (approve, correct-with-value per field, reject, escalate-novel) and route them to one named destination with one named monthly consumer.
  5. Get the vendor's written answer on where corrections go. Ask "what happens to our corrections, on what schedule, visible where?" and put the answer in the contract. If the honest answer is "nowhere," you have learned the pilot's improvement curve is flat, and you have learned it for free.

Key Takeaways

  • Treat AI-human handoffs as a new species of the oldest failure in operations: one party cannot ask clarifying questions and the other assumes it did, so whatever the format fails to carry is lost silently, at volume.
  • Recognize the two signature wastes of an unspecified handoff: re-doing (the human repeats the AI's work because the payload hides what was checked) and rubber-stamping (the human waves everything through because the payload carries no signal of where to look).
  • Write a Handoff Contract for every AI-human boundary: what passes, with what signal, what the receiver does per signal level, what returns, and what gets logged, co-designed with the Gate Spec it feeds.
  • Structure every payload as fields in the receiver's checklist order, never prose, and design the reviewer's screen as the handoff made visible: evidence beside claim, zero clicks, or the four-minute review standard is fiction.
  • Use confidence as a routing signal, never a guarantee: validate it with the 50-per-band calibration check in pilot week two, map bands to routes, and demand calibration data on your own sample from any vendor.
  • Suppress AI output entirely on the low-confidence route, because anchoring makes a human shown a bad draft slower and less accurate than a clean start; sometimes the best handoff is no handoff.
  • Attach an evidence pointer to every claim in the payload, turning review from verifying the universe into verifying the citation, and turning overrides into teachable, labeled examples.
  • Build the return path as structured codes flowing to a named destination with a named consumer, because the return path is where MIT's missing learning loop physically lives, and match handoff granularity to what the receiver can act on: aggregates cannot be reviewed, and the wrong granularity can hide a 100-percent-wrong lane inside a healthy average.