Structured Outputs Your Systems Accept
Watch the desk for four minutes and you will see the whole failure. The accounts payable clerk has two windows open. On the left, the AI tool's output: a genuinely good paragraph, fluent and correct, explaining that invoice INV-88412 from vendor V-2041 shows a price mismatch of $310 against purchase order PO-77120, likely caused by an outdated rate card, recommended action: route to procurement review. On the right, the ERP, the enterprise resource planning system where the exception record actually lives. And between the two windows, the clerk: reading the paragraph, finding the amount, typing the amount, finding the vendor, typing the vendor, finding the category, choosing it from a dropdown, forty seconds of transcription per exception, eleven hundred times a month. The AI extracted the data from the invoice. The clerk is now extracting the data from the AI. When MIT's researchers wrote that 95 percent of enterprise GenAI pilots fail for lack of workflow integration, this desk is what the finding looks like from eighteen inches away. Nobody wrote the thirty lines of specification that would have made the left window's contents flow into the right window on their own, and this lesson exists so that you become the person who writes them.
The Retyping Tell
Every stalled pilot has a tell, one observable behavior that gives away the diagnosis before you read a single report. For integration failure, the tell is retyping. If you can find one human anywhere in the workflow whose job includes reading AI output and manually re-entering its contents into another system, you have found the exact spot where the pilot stopped being integration and became homework. The tool generates; a person transcribes; the system of record never met the AI at all.
The previous lessons in this level built up to this moment deliberately. You redesigned the invoice-exception process instead of paving the cow path. You mapped which steps are AI-ready and which stay human. You ran the FMEA, the failure mode and effects analysis, and found where outputs go wrong. You designed the human gate with its four-minute review standard, and you wrote the Handoff Contract that specifies the payload crossing every AI-to-human seam: named fields, typed, ordered, never "the output." This lesson is the final, unglamorous plank in that floor: the mechanics of making the AI's output arrive in a shape that your systems, and your reviewers, can accept without a human transcription layer in between.
Here is the fact that changes how you negotiate with every vendor and every internal builder for the rest of your career: the model can emit almost any shape you demand. A modern language model will produce flowing prose, or a rigid table, or a machine-readable record with exactly the six fields you name, in exactly the order you name them, with exactly the allowed values you list. It does not prefer prose. It defaults to prose because prose is what happens when nobody specifies anything else. The difference between a pilot whose outputs flow into the ERP and a pilot whose outputs get retyped is usually about thirty lines of output specification that nobody wrote, and writing those thirty lines is a specification skill, not an engineering one. You do not need to code. You need to name fields, list allowed values, and state what happens when a value cannot be found. That is process work. It is your work.
McKinsey's research says the organizations getting real earnings impact from AI are the ones that fundamentally redesign workflows, roughly three times more often than everyone else. Redesign sounds grand in a steering committee deck. At the desk, redesign is this: deciding, in writing, the shape of what the machine hands to the next step. Skip it and you have adopted a tool without transforming anything, which is MIT's autopsy finding restated with a keyboard.
Why prose fails systems: the extraction the workflow un-does
Think about what the AI actually did in that opening scene. An invoice arrived as a semi-structured document: a PDF, a scan, a vendor's idiosyncratic layout. Buried in it were the facts that matter: the amount, the vendor, the purchase order reference, the nature of the mismatch. The model's genuinely valuable act was extraction: pulling those facts out of the noise. And then, in the very last step, it poured the extracted facts back into noise, wrapping them in a paragraph. A paragraph containing an amount, a vendor, and a category is not data. It is a container that a human, or a brittle text-parsing script that breaks the first time the phrasing shifts, must open to get the data back out. The AI did extraction, and the output format un-did it. The clerk's forty seconds of retyping is the un-doing, performed by hand, at volume, forever.
Systems are literal. The ERP's exception record has a field called amount that accepts a number, a field called vendor_id that accepts an identifier from the vendor master, a field called exception_code that accepts one of a fixed list. It cannot accept "roughly $310 over the agreed rate, likely due to an outdated rate card." Someone must translate. The entire question of this lesson is who: a human at forty seconds a time, or a specification you write once.
Why prose fails reviewers: the search work the schema should have done
Now walk the same paragraph up to the human gate. The gate you designed earlier in this level runs on a four-minute review standard: verify the amount against the invoice image, verify the vendor, check the category against the criteria, scan the evidence. Four minutes is physically possible only if the fields sit in fixed positions, claim beside evidence, every single time. Hand that reviewer free prose instead and watch what happens to the clock: they are reading the whole paragraph hunting for the amount, then re-reading it hunting for the category, doing search work that the output shape should have done before the item ever reached them. The four-minute standard quietly becomes seven, throughput drops, the queue backs up, and the gate either gets more staff or becomes a rubber stamp. This is the same lesson this chapter keeps proving from different angles: gates, handoffs, and schemas are one design, said three ways. The Gate Spec defines what the reviewer checks, the Handoff Contract defines what crosses the seam, and the output schema defines the shape it crosses in. Change any one without the others and all three degrade together.
So prose fails both audiences. Systems need structure because they are literal. Reviewers need structure because attention is a budget and search burns it. Prose is for humans reading at leisure, and nobody in your workflow is reading at leisure.
Prose is for humans. Systems and reviewers need structure. If you do not specify the shape of the output, a clerk will specify it for you, forty seconds at a time.
The Three Shapes an Output Can Take
You need exactly three shapes in your vocabulary, and you need to know which consumer each one serves. None of them requires you to write code. All of them require you to make decisions.
Shape one: JSON, when a system consumes the output
JSON (JavaScript Object Notation, though nobody expands it in practice) is nothing more intimidating than labeled boxes: a set of field names, each holding a value, wrapped in a machine-readable envelope of braces and quotes. Every modern business system can read it. Here is the invoice-exception output as six labeled boxes:
{
"invoice_id": "INV-88412",
"vendor_id": "V-2041",
"amount_claimed": 4180.00,
"exception_code": "PRICE_MISMATCH",
"confidence": 0.87,
"evidence_ref": "PO-77120 line 4; contract C-309 p.12"
}
Read it slowly once and the fear evaporates. Field name, colon, value. Text values in quotes, numbers bare. That is the entire grammar you need. And here is the part that matters for your role: you do not write JSON, you approve the field list. The decisions inside that block, that there are six fields and not nine, that the field is called exception_code and not category, that PRICE_MISMATCH is one of a fixed set of codes, that confidence is a number between 0 and 1, are process decisions. They came from your Handoff Contract. The vendor's engineer types the braces; you own what goes between them.
Shape two: the table, when humans consume in bulk
When the consumer is a human looking at many records at once, the right shape is rows and columns. The weekly override-cluster report from your gate design is the canonical example: the reviewer's overrides, grouped into patterns, delivered every Friday not as a narrative memo but as a table:
| Cluster | Overrides this week | Share | Dominant pattern | Owner |
|---|---|---|---|---|
| Wrong category | 19 | 58% | Vendor V-2041 invoice format | Process owner |
| Amount mismatch | 8 | 24% | Currency rounding on imports | AP lead |
| Stale vendor data | 6 | 18% | Master data lag over 30 days | Data steward |
The table's power is comparability: the eye runs down a column and spots the outlier in seconds. A narrative version of the same content forces the reader to hold numbers in memory across paragraphs. If a human will scan, compare, or sort the output, demand rows.
Shape three: the fixed template, when the output is a document
Some AI outputs genuinely are documents: the decision memo recommending write-off versus recovery on a disputed invoice, for instance. Even there, free generation is the wrong demand. You specify a skeleton with mandated sections in a mandated order: Situation (three sentences maximum), Amount at stake (one figure, sourced), Options considered (exactly the options in your SOP, the standard operating procedure, no inventions), Recommendation (one of the allowed dispositions), Evidence (pointers, not paraphrase). Generation constrained to a fixed shape gives reviewers a superpower: they can diff. Two memos with identical skeletons can be compared section by section; two free-form essays cannot. The template is a schema wearing a suit.
The Output Schema Card: Five Rules of the Craft
Now the lesson's heart, and its named artifact. The Output Schema Card is a one-page specification, one per AI step in your process, that states: the exact fields, their types, their allowed values, which are required and which optional, the null rule, the validation checks that run at the boundary, and the rejection path for malformed output. It is the thirty lines nobody wrote in the failed pilots. Writing a good one is a craft with five rules, and each rule exists because of a specific way outputs go wrong.
Rule 1: Allowed values over free text, wherever a decision branches
Any field that a downstream decision branches on must be a closed list, never free text. The exception category drives routing: PRICE_MISMATCH goes to procurement, QUANTITY_MISMATCH goes to receiving, and so on through your six defined codes. If the schema lets the model write the category in its own words, you will get "pricing discrepancy" on Monday, "rate variance" on Tuesday, and "cost mismatch, probably" on Wednesday, three synonyms your routing logic treats as three unknown categories. Savor the irony, because it is instructive: back in Level 2 you profiled this very process and found the dispute-reason field was free text in 61 percent of records, which is precisely why the data was not AI-ready. Buy an AI tool, skip the schema, and the tool reintroduces the 61 percent problem at machine speed, free-texting the exact field you bought it to standardize. Closed lists are how you stop the disease from reinfecting the cure. Name every allowed value on the card. Six codes means six codes, spelled exactly.
Rule 2: The null rule: what the model says when it does not know
Every schema card must answer one question explicitly: when the model cannot find a value, what does it emit? The only acceptable answer is an explicit marker: null, or NOT_FOUND, stated in the field, on purpose. Never a guess, and never silence. You met this distinction in the FMEA lesson: a blank is a kindness, because blanks are easy to detect and route, while a missed field returned as hallucinated filler, a plausible invented vendor ID, a confident wrong amount, is among the most dangerous failure modes in the catalog because it is invisible to every downstream check that only looks for emptiness. The null rule takes that FMEA insight and enforces it by schema: the card says, in writing, "if the vendor cannot be identified from the document, emit vendor_id: NOT_FOUND, and the item routes to the human lane." A model that is permitted to say "I could not find it" in a structured way is a model you have relieved of the pressure to invent. This single line on the card prevents more silent damage than any other.
Rule 3: Confidence and evidence are fields, not decoration
Your Handoff Contract already established that every payload crossing an AI-to-human seam carries a signal (a confidence score with defined thresholds) and pointers to evidence. The schema card is where those commitments become columns: confidence as a number with a defined scale and defined routing thresholds, evidence_ref as a required field pointing to the source (the purchase-order line, the contract page), never a paraphrase of it. When confidence and evidence live in the schema, the gate's routing rule becomes mechanical: below 0.70, human lane, no discussion. When they live in prose ("I am fairly sure about this one"), routing becomes vibes. The program's verify-every-figure rule needs somewhere to point; evidence fields are where it points.
Rule 4: Validation at the boundary, with a rejection path
A schema that nothing enforces is a suggestion. Rule 4 says every output is machine-checked at the boundary, before it touches a system or a reviewer: type checks (amount is a number), range checks (amount is positive and under the process ceiling), allowed-value checks (exception_code is one of the six), and arithmetic checks wherever an identity exists. The reconciliation catch from your FMEA, line items must sum to the invoice total, is exactly such a rule: a one-line check that automatically catches the transposed digit no human reliably spots. This is the verify-every-figure discipline, automated for everything checkable, so human attention is spent only on what only humans can check.
And when validation fails, the output does not limp onward. It goes to the rejection path, which the card must define: retry the AI step once (transient formatting failures often self-correct), and if the second output also fails, route the item to the human lane with the raw output attached so the human sees what the machine attempted. Two prohibitions, absolute: never silently repair (a system that "fixes" a malformed amount is now inventing data), and never silently drop (a vanished exception is an unpaid vendor and an unexplained aging report). The rejection path is a Gate Spec entry like any other: it has a queue, an owner, and a service standard.
Rule 5: Version the schema, and let versions travel
Fields change. Someone will add a seventh exception code, split amount into amount and currency, rename a field. When that happens, every consumer of the old shape, the ERP staging table, the reviewer's screen layout, the weekly report, breaks quietly, which is the worst way to break. So the card carries a version number, the version travels inside every payload (schema_version: "1.3" as a field like any other), and a schema change triggers the same change process your living SOP uses: propose, review with every downstream consumer named in the RACI, communicate, dated changelog. This is the version-control discipline from the SOP lesson closing its loop: the SOP governs how humans do the step, the schema governs what the machine emits from it, and both are controlled documents because both have dependents.
The Retyping Tax, and the Refund That Got Away
A worked example: computing the tax (illustrative numbers)
Here is the arithmetic that turns this lesson into a steering-committee slide. All numbers are hypothetical, built on the invoice-exception process this level has used throughout: 1,150 exceptions a month.
Pre-schema pilot week. The AI emits prose. For every exception, a clerk retypes five fields into the ERP: about 40 seconds per exception. At 1,150 exceptions a month, that is 46,000 seconds, roughly 12.8 hours a month of pure transcription waste: skilled attention spent moving data between two windows on the same monitor. But the hours are the smaller finding. The reconciliation checks also surface a 1.2 percent retype-error rate: roughly 14 exceptions a month where the ERP record is wrong even though the AI's output was right, because a human transposed digits or clicked the adjacent dropdown entry. Sit with that sentence, because it is the single most persuasive integration statistic you can put in front of a committee: the workflow was adding defects to correct output. The pilot's accuracy debate was moot; the un-integrated last step was manufacturing errors the model never made.
Post-schema. The Output Schema Card goes to the vendor; three weeks later, outputs land as validated JSON in the ERP staging table. Validation now rejects 3.1 percent of outputs (about 36 a month) to the human lane, mostly the null rule firing on illegible scans. Notice the framing discipline: those rejections are not a defect statistic, they are the system working as designed, uncertainty routed to judgment instead of guessed over. Retyping: eliminated. The 12.8 recovered hours are redeployed to the verification gate that the FMEA analysis flagged as understaffed. Write the conclusion the way you would say it to the CFO: the schema paid for the gate. Thirty lines of specification bought back a third of a working week every month and spent it on the control the risk analysis wanted. Gartner's warning that 60 percent of AI projects without AI-ready data will be abandoned through 2026 has a mirror image on the output side: outputs that are not system-ready meet the same fate, one retyped field at a time.
A failure story: the one-line check that did not exist
Now the other path, composited from a pattern that recurs across service operations. A customer-service team deploys an AI that drafts responses to complaint emails, in free prose, including, when relevant, a promised refund amount. No schema. No validation. No allowed range on the refund figure, because the refund figure is not a field, it is just words inside a paragraph. One busy afternoon the model drafts a goodwill refund of $1,890 on a case where policy caps goodwill at $189: a typo-shaped hallucination, one digit of confident nonsense. The agent, forty tickets deep, sends it unread. The customer accepts within the hour, in writing.
The postmortem is the instructive part. The finding was not "the model erred." Models err; that is a design input, the whole premise of your FMEA. The finding was that no boundary existed anywhere in the workflow where a refund ten times the policy cap could be caught. Not in the output (no field to check), not at a gate (no validation to run), not in the send flow (no hold rule). A one-line range check, refund_amount must not exceed policy cap for the case type, else route to supervisor, was the entire missing control. It would have taken a morning to specify. Gartner's consistent message on GenAI risk is that the projects that survive are the ones that pair the capability with designed controls, and here is that message at its smallest and sharpest: structure is cheap. Its absence is not.
Specifying the Shape Without Writing a Line of Code
The conversation script for vendors and IT
You now hold everything needed to run the acquisition conversation from the specifying side of the table, which is where a process professional belongs. The script has two sentences, and both do heavy lifting:
- "Here is the field list, the allowed values, the null rule, and ten sample exceptions with their correct outputs." You hand over the Output Schema Card and an acceptance set: ten real (anonymized) inputs from your process, each paired with the output a competent human would produce. This is acceptance-test thinking, the same discipline you applied to retrieval quality in the RAG lesson (RAG, retrieval-augmented generation, the grounding technique from the previous lesson), reapplied to output shape. You are not asking the vendor whether their tool is good. You are defining what good means, on your data, in advance.
- "Show me your tool emitting this schema on our samples, and show me what it emits when the invoice is illegible." The first half tests the happy path. The second half tests the null rule, and it is the more revealing question in the room. A vendor whose tool confidently extracts fields from an unreadable scan has just demonstrated the hallucinated-filler failure mode live, at their own demo. A vendor whose tool returns NOT_FOUND and routes to the human lane has just demonstrated that they have met this problem before. You will learn more from the illegible sample than from the nine clean ones.
Notice what is absent from the script: any code, any model talk, any architecture. The vendor's engineers translate your card into their configuration. Your leverage is that you arrived with the specification and the acceptance set, which converts the demo from theater into a test you wrote.
The in-house version: structure at chat scale
Structure discipline scales all the way down to a team of one working in a chat tool with no integration budget. Three moves reproduce the whole architecture in miniature. First, the prompt clause that demands the shape: end your prompt with "Return only a table with these exact columns: invoice_id, vendor_id, amount, exception_code (one of: PRICE_MISMATCH, QUANTITY_MISMATCH, DUPLICATE, MISSING_PO, TERMS_MISMATCH, OTHER), confidence, evidence. If a value cannot be found, write NOT_FOUND. No prose before or after the table." That paragraph is an Output Schema Card administered orally. Second, the human eyeball as validator: before using the output, run the boundary checks yourself, thirty seconds, same list: types plausible, values on the allowed list, amounts in range, NOT_FOUND where the source was thin. Third, the paste-ready format: demand the shape your destination accepts, columns matching your spreadsheet's columns in order, so the transfer is one paste, not a retype. The 12.8-hour tax and the 1.2 percent retype-error rate operate at every scale; so does the cure.
One bridge before Monday. Your outputs now arrive shaped, validated, and versioned, with malformed items routed instead of repaired. But validation only catches what machines can check. The next lesson takes the harder half: designing human verification into the flow itself, with sampling rates and error budgets, so that checking the AI is a designed system rather than a hope.
What to Do Monday Morning
- Write the Output Schema Card for your pilot's main AI step. One page: field list, type for each field, allowed values for every field a decision branches on, required versus optional, and the null rule stated in one explicit sentence. Steal the six-field invoice-exception example as your skeleton.
- Collect ten sample inputs with their correct outputs. Real cases, anonymized, including at least one ugly one (an illegible scan, a missing reference). This is your acceptance set; date it and file it with the card.
- Run the vendor conversation. Send the card and the samples with the two-sentence script: emit this schema on our samples, and show us the output on the illegible one. If you are in-house at chat scale, convert the card into the prompt clause instead and test it on the same ten samples yourself.
- Compute your current retyping tax. Time five real transcriptions, multiply by monthly volume, and ask reconciliation for the retype-error rate if one exists. Put both numbers on one slide; this is the business case that funds the integration work.
- Add the rejection path to your Gate Spec. One entry: validation failure retries once, then routes to the human lane with raw output attached, named owner, service standard. Never silently repaired, never silently dropped.
- Version the card. Mark it v1.0, add schema_version to the field list, and register the card in the same change process as your living SOP, with every downstream consumer named.
Key Takeaways
- Treat retyping as the diagnostic tell of integration failure: a human re-entering AI output into a system of record is MIT's 95 percent finding made visible at a single desk, and it marks the exact spot where an output shape was never designed.
- Demand shape, because the model can emit any shape you specify: JSON for systems, tables for humans reading in bulk, fixed templates for documents reviewers need to diff; the shape is a specification decision, not an engineering one.
- Build one Output Schema Card per AI step: exact fields, types, allowed values, required versus optional, the null rule, the boundary validations, and the rejection path, roughly thirty lines that separate integration from homework.
- Close every decision-bearing field with allowed values, because free-text categories reintroduce the Level 2 61 percent data problem at machine speed through the very tool bought to fix it.
- Write the null rule explicitly: an unfindable value is emitted as NOT_FOUND and routed to a human, never guessed, never silent, enforcing the FMEA's blank-versus-hallucinated-filler distinction by schema.
- Validate at the boundary and define the rejection path: type, range, allowed-value, and arithmetic checks before any output touches a system or reviewer; failures retry once then route to a human with raw output attached, never silently repaired or dropped.
- Compute and present the retyping tax: in the illustrative pilot, 40 seconds across five fields at 1,150 exceptions a month equals 12.8 wasted hours plus a 1.2 percent retype-error rate, defects the workflow added to correct output, and the recovered hours staffed the verification gate.
- Specify without engineering: hand vendors the card, the null rule, and ten samples with correct outputs, ask to see the schema emitted on your data and on an illegible input, and scale the same discipline down to a prompt clause when the tool is chat.
Skill.re