←
CAP Certification
Capable · M48 · lesson 48 of 54 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Structured Output Engineering: JSON, XML, Tables

15 min

Mateus Reyes is an operations analyst at a regional freight company. He had been using AI to extract key information from carrier contracts, things like payment terms, liability clauses and rate schedules, and pasting the results into a spreadsheet. It worked, but the copy-paste step took 20 minutes per contract and introduced errors. He asked whether AI could output the data directly in the right format. His first attempt returned something that looked like JSON but had a missing comma on line 14. His second attempt came back as a nicely formatted table, but in markdown, which broke when he piped it into the next system. The problem was not that AI cannot produce structured output. The problem was that he did not yet know how to engineer the prompt to get reliable, machine-ready structure every time.

Structured output engineering means prompting AI to produce data in a specific, parseable format rather than flowing prose. It is one of the highest-value technical skills for practitioners who want to connect AI to real workflows, because it is the point where an AI experiment stops being a demonstration somebody watches and becomes a component other software can depend on. Everything downstream of that point, the database write, the spreadsheet import, the API call to the next system, assumes the shape of what arrives. This lesson covers the three formats you will meet most often, the prompt patterns that make each one reliable, and the validation discipline that keeps a pipeline standing when the model occasionally gets it wrong.

Why Structure Matters

Prose output is great for humans. Computers need structure. When an AI summarizes a contract in paragraphs, a human can read it and extract what they need, filling in gaps with judgement and context as they go. But if you want that information to flow automatically into a database, a spreadsheet or another application, paragraphs are a problem. Software cannot infer. It needs output that is machine-readable: consistent fields, predictable format, no ambiguity about where one value ends and the next begins. The whole discipline of structured output engineering is about removing the interpretive work that a human reader does for free and that a parser cannot do at all.

The three most common structured formats you will work with are:

  • JSON (JavaScript Object Notation), the lingua franca of modern APIs and web applications. A nested, key-value format readable by almost every programming language and integration tool.
  • XML (Extensible Markup Language), older than JSON and still dominant in enterprise systems, healthcare (HL7) and government data exchanges. More verbose but more explicit about data types.
  • Tables, a row-and-column structure in either HTML or markdown, suitable for human-readable reports and direct import to spreadsheets.

The choice between them is not a matter of taste. It follows from who or what consumes the output at the other end, and picking the wrong one is exactly the mistake Mateus made when he accepted a markdown table for a job that needed machine-parseable data.

The JSON Pattern

JSON is the format you will use most often when connecting AI output to downstream systems, so it is worth learning the prompt pattern properly rather than by trial and error. The difference between a prompt that works occasionally and one that works consistently is almost entirely a matter of how much you leave to the model's discretion. Every decision you do not make explicitly is a decision the model will make for you, and it will not necessarily make it the same way twice.

A weak prompt looks like this: "Extract the key terms from this contract and give me the data in JSON." This fails because "key terms" is ambiguous and the model will decide for itself what fields to include. Results will differ between runs and between documents, which means that even when any individual output looks fine, you cannot build anything on top of it.

A strong prompt looks like this: "Extract the following fields from the contract text below and return them as a single JSON object with exactly these keys. Do not include any text before or after the JSON object. If a field is not found in the contract, use null as the value. Required fields: payment_terms_days (integer), liability_cap_usd (number), notice_period_days (integer), renewal_automatic (boolean), governing_law_state (string)."

What makes this work:

  1. Field names are explicit and use consistent naming conventions (snake_case).
  2. Data types are specified for each field (integer, string, boolean).
  3. Handling of missing data is defined (null).
  4. The model is told to return nothing but the JSON object.

That fourth instruction earns its place more often than people expect. Left to itself a model likes to introduce its output conversationally, and a helpful sentence of preamble in front of an otherwise perfect object is enough to make a strict parser reject the whole response. Telling the model that the object is the entire response, not the centrepiece of one, removes a whole category of failure that has nothing to do with whether the extraction itself was correct.

Even with a strong prompt, AI models occasionally produce invalid JSON: a trailing comma, an unclosed bracket, a quote character inside a string value. This is the failure Mateus hit on his first attempt, and it is worth understanding that it is a property of generated text rather than a bug you can prompt your way out of entirely. For production use, always wrap your JSON parsing in error handling that catches malformed output and flags it for review, so that a bad response becomes a queued exception rather than a silent corruption of whatever sits downstream.

Schema Definition: The Advanced Approach

For critical workflows, move beyond field lists in natural language to explicit schema definitions. A JSON Schema is a standardized way to describe the expected structure of a JSON object, including required fields, data types and constraints. It says the same things a good natural-language field list says, but in a form that is unambiguous and that other tools in your stack can also read, which means the description of the data stops living only in the prompt.

Including a simplified schema in your prompt gives the model a much clearer target. "Produce output conforming to this schema: {"type": "object", "required": ["payment_terms_days", "governing_law_state"], "properties": {"payment_terms_days": {"type": "integer", "minimum": 0}, "governing_law_state": {"type": "string", "maxLength": 50}}}"

Notice what the schema carries that a sentence would struggle to: which fields are merely expected and which are genuinely required, and constraints on the values themselves rather than just their types. A minimum of zero on a day count rules out a negative number that would otherwise pass a type check and quietly break a date calculation later. A maximum length on a state string rules out the case where the model decides to explain the governing law rather than name it.

Many modern AI platforms now support structured outputs or function calling natively, modes where you supply a schema and the platform guarantees the output conforms to it. These modes use constrained decoding to mathematically ensure the output is valid JSON matching your schema, which eliminates the malformed-output problem entirely rather than merely making it rarer. The distinction matters when you are deciding how much error handling to build: a well-prompted model is usually right, while a constrained decoder is structurally right. If your use case is critical enough to build on, use native structured output modes where they are available to you.

Working with Tables

Tables are the right format when the output will be read by a human, imported into a spreadsheet, or included in a document or report. Two formats are common, HTML tables and markdown tables, and the choice between them again comes down to the consumer rather than to preference.

For spreadsheet import, ask for a markdown table. Most spreadsheet tools can import markdown tables with minimal friction. Specify column headers explicitly, as in: "Produce a markdown table with these exact column headers: Carrier Name | Base Rate per Mile | Fuel Surcharge | Minimum Volume Commitment | Contract End Date. Populate one row per carrier mentioned in the document." Naming the headers exactly does the same work that naming JSON keys does. It fixes the shape of the output so that the column your import expects in a given position is the column that actually arrives there.

For reports and web display, HTML tables are more reliable for rendering. Request HTML table output and specify whether you want a header row (<thead>) and whether rows should include any CSS class names your system expects. These details are easy to skip and annoying to retrofit, since a table that renders unstyled in a finished report has to be regenerated rather than patched.

A consistent trap with either format is that AI models will sometimes produce tables with extra whitespace in cell values that causes import errors. The value looks correct on screen and fails on ingest, which makes it one of the harder problems to diagnose from the output alone. Specify "no leading or trailing whitespace in cell values" if this is a problem in your system.

XML: When and How

You will encounter XML most often when integrating with legacy enterprise systems, healthcare data standards or government APIs, which is to say in exactly the settings where the receiving system is least likely to be flexible about what it accepts. The prompting approach mirrors JSON: define the schema explicitly, name the tags, specify attributes, and tell the model to return nothing outside the XML structure.

"Return the extracted data as XML conforming to this structure: <ContractSummary><PaymentTermsDays type="integer"></PaymentTermsDays><GoverningState type="string"></GoverningState></ContractSummary> Include an XML declaration at the top. Use null as element content when data is not present."

XML validation is stricter than JSON in most parsers. A single malformed tag or unescaped special character (&, <, >) will cause the entire document to fail parsing, not just the element containing it. That all-or-nothing behaviour is the practical difference between the two formats when things go wrong, and it is why extracted text that might legitimately contain an ampersand or an angle bracket deserves attention before it reaches the model rather than after. Always validate against a schema and catch errors.

Choosing the Right Format

The decision is made by the consumer of the output, not by the person writing the prompt. Working backwards from what receives the data will pick the format for you in almost every case.

FormatUse it when the consumer isWhat to watch
JSONAnother system: an API, a database, a web application, any machine-to-machine integrationTrailing commas, unclosed brackets, quote characters inside string values
XMLA legacy enterprise system, a healthcare data standard such as HL7, or a government data exchangeStricter parsing than JSON; one malformed tag or unescaped special character fails the whole document
Markdown tableA spreadsheet, via importLeading or trailing whitespace in cell values; column headers that drift between runs
HTML tableA rendered report or a web pageWhether a header row and any expected CSS class names are actually present

The Extraction-First Pattern

One of the most useful patterns for structured output work is separating reasoning from formatting. Instead of asking the model to reason about the document and produce structured output in a single step, break it into two. Step one: "Identify and quote the relevant sections of the document for each of these fields: [list]." Step two, as a separate prompt feeding in step one's output: "Given these extracted quotes, fill in the following JSON object: [schema]."

The reason this works is that the two tasks compete for attention when they are combined. Finding the right clause in a long contract is a reading problem, and producing syntactically valid nested output is a formatting problem, and a model asked to do both at once tends to do each of them slightly worse. Separating them also changes what a failure looks like. When the output is wrong, you can see whether the model quoted the wrong clause or quoted the right clause and then filled the object badly, which turns a vague "the extraction is unreliable" complaint into a specific fix in one of two prompts.

This two-step approach catches more information, reduces hallucination and makes it easy to spot where the model went wrong when output is incorrect. For Mateus's contract extraction, the two-step approach increased accuracy from about 82% to 96% across a sample of 50 contracts.

Anti-Patterns

The failures in this area are consistent enough to be worth naming, because each of them looks reasonable at the moment you commit it and only reveals itself once something downstream depends on the output.

  • Asking for "the key fields" and letting the model choose. Ambiguity in the request becomes variance in the output. If you have not named the fields, you have delegated your schema to a system that will redesign it on the next document.
  • Specifying names but not types. A day count that arrives as a quoted string on one run and a bare integer on the next will pass a shallow check and break a calculation later.
  • Leaving missing data undefined. If you have not said what absence looks like, you will get some mixture of null, empty strings, the text "not found", and omitted keys, and every consumer of the data will need to handle all four.
  • Treating a well-formed sample as proof. One good response tells you the prompt can work, not that it will. The malformed cases are the ones you have to design for.
  • Parsing without error handling. A pipeline that assumes valid JSON will not fail loudly on a trailing comma; it will fail somewhere further along, in a place that gives no hint about the real cause.
  • Picking the format for the author rather than the consumer. Markdown looks tidy in a chat window and is the wrong answer when the recipient is an API.

Practice Prompts

Work these against a document you actually use, not a sample, since the edge cases that matter are the ones your own material contains.

  • Take a structured-output prompt you already use and rewrite it to name every field, give every field a type, and define explicitly what the model should return when a field is absent. Run both versions against the same set of documents and compare the outputs field by field.
  • Convert your field list into a JSON Schema that includes a constraint which is not a type, such as a minimum on a numeric field or a maximum length on a string. Note which invalid outputs the constraint would have caught.
  • Split a single-step extraction prompt into the extraction-first pattern: one prompt that quotes the relevant passages, a second that fills the object from those quotes. Diagnose one wrong result and identify which of the two steps produced it.
  • Deliberately feed your prompt a document that is missing several of the required fields, and confirm that the output represents absence the way you specified rather than inventing plausible values.
  • Ask for the same extraction as a markdown table and as JSON, then attempt to import each into its intended destination. Record every manual correction the import required.

Reflection

Consider a workflow where you currently move AI output by hand. What format does the receiving system actually want, and is that the format you have been asking for? If the answer is no, the manual reformatting step you have been treating as unavoidable is a prompt-design problem rather than a data problem.

Then consider what happens today when the model returns something malformed. Does anyone find out, and how? A workflow where invalid output is caught, flagged and reviewed is in a different category of reliability from one where it passes silently into a spreadsheet that somebody trusts, and the gap between those two states is usually a few lines of error handling rather than a better prompt.

Glossary

  • JSON (JavaScript Object Notation): a nested, key-value data format readable by almost every programming language and integration tool, and the common currency of modern APIs and web applications.
  • XML (Extensible Markup Language): an older tag-based format still dominant in enterprise systems, healthcare standards such as HL7, and government data exchanges. More verbose than JSON but more explicit about data types.
  • JSON Schema: a standardized way to describe the expected structure of a JSON object, including required fields, data types and constraints on values.
  • Structured outputs / function calling: platform modes in which you supply a schema and the platform guarantees the output conforms to it.
  • Constrained decoding: the mechanism behind those modes, which restricts generation so that the output is mathematically guaranteed to be valid against the supplied schema.
  • snake_case: a naming convention in which words in a field name are lowercase and separated by underscores, used here to keep field names consistent across a schema.
  • Extraction-first pattern: a two-step approach that separates finding the information from formatting it, using one prompt to quote relevant passages and a second to populate the structure.

This lesson sits alongside several others in the practitioner track. Structuring Data for AI approaches the same boundary from the input side, covering how data should be shaped before it reaches a model. Tool Integration & APIs takes the output you have learned to generate here and covers the mechanics of getting it into another system. Prompt Libraries & Version Control is the natural follow-up once you have a structured-output prompt worth keeping, since a prompt that encodes a schema is an asset that needs to be versioned like one. Multi-Modal Prompting: Text, Images, and Documents extends the same extraction discipline to sources that are not plain text.

Closing

Mateus's problem was never that AI could not produce structured data. It was that he was asking for structure without specifying it, and then treating each malformed result as bad luck rather than as a missing instruction. The work of structured output engineering is unglamorous and largely consists of saying out loud things that feel obvious: what the fields are called, what type each one is, what absence looks like, and that nothing else should appear in the response. Do that consistently, use native schema-constrained modes where the workflow is critical, split reasoning from formatting when accuracy matters, and validate everything before it crosses into a system that trusts it. The result is not a cleverer prompt. It is a pipeline that other people can rely on without knowing there is a model inside it.

Key Takeaways

  • Structured output requires explicit specification. Ambiguous prompts produce inconsistent formats; define field names, data types and missing-value handling explicitly for every structured output task.
  • Use native structured output modes for production. Platforms that support JSON Schema-constrained generation eliminate malformed output mathematically, so use them for any workflow where parsing errors have real consequences.
  • Choose your format based on the consumer. JSON for machine-to-machine integration, markdown tables for spreadsheet import, HTML tables for rendered reports, XML for legacy enterprise and regulated-data systems.
  • The extraction-first pattern improves accuracy. Separating the step that finds the information from the step that formats it typically increases structured output accuracy and makes errors easier to diagnose.
  • Validate every output before downstream use. Even well-prompted models occasionally produce syntactically invalid JSON or XML, so build error handling into any production pipeline that consumes structured AI output.
  • Specify edge cases explicitly. Instruct the model on how to handle missing fields, multi-value fields, special characters and whitespace, the cases where AI output most commonly breaks downstream parsers.

Frequently Asked Questions

Why does my prompt work on one document and fail on the next? Almost always because the prompt left something to the model's discretion that the two documents exercise differently. A contract that contains every field you asked for gives the model no opportunity to improvise, while one missing several of them forces a decision you did not specify. Name the fields, name the types, and define what absence looks like, and the variation between documents largely disappears.

Should I ask for JSON or use a native structured output mode? Use the native mode when it is available and the workflow is critical, because constrained decoding guarantees conformance to your schema rather than making conformance likely. A carefully written prompt is a good substitute when native support is not available, but it is a substitute, and it needs the error handling that a guaranteed mode does not.

Do I still need validation if I am using a schema-constrained mode? You need less of it for syntax and just as much for meaning. A constrained mode guarantees that the output parses and matches the shape you declared. It does not guarantee that the value in a field is the correct value from the document, which is a separate question that only sampling and review can answer.

My model keeps adding a sentence before the JSON. How do I stop it? Tell it explicitly not to include any text before or after the object, and treat that as a standard clause in every structured output prompt rather than a fix applied after you notice the problem. If it persists, the extraction-first pattern helps, since the second prompt receives quotes rather than a document and has less to explain.

When is a table the right answer rather than JSON? When a person is the consumer. A markdown table for spreadsheet import and an HTML table for a rendered report both go to a human eventually. Anything that feeds another system directly should be JSON, or XML where the receiving system requires it.