←
AI Agent Builders & Citizen Developers
Capable · M15 · lesson 15 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Structured Output: Forcing JSON You Can Actually Use
📖
now learning

Structured Output: Forcing JSON You Can Actually Use

15 min

The most common failure in an operator-built document pipeline is the most innocent-sounding one: the model returns prose where you wanted JSON. The CRM webhook 400s. The downstream Code node throws "Unexpected token in JSON at position 0." Your error queue fills with what looks like a perfect summary of the RFP — except it begins "Here is the extracted specification:" and ends with a sentence about how the model "hopes this helps." That single newline of friendly preamble is the difference between a workflow that shipped to production and a workflow that does not. This lesson is how to use OpenAI Structured Outputs and Anthropic's tool_use response format to guarantee JSON that your CRM accepts — and how to build a fallback retry path that handles the rare case when the model returns prose anyway.

Why Strict Structured Output Changed Everything in 2024-2026

Until late 2024, the operator-builder's standard technique for getting JSON out of an LLM was a prompt that said "respond with valid JSON only, no preamble, no explanation." This worked 92-97% of the time. The 3-8% failure tail is where pipelines died. The model would add a Markdown code fence around the JSON. It would prefix with "Here's the JSON:" or suffix with "Let me know if you need anything else." It would drop a required field. It would put a trailing comma somewhere that strict parsers reject. It would emit the field "value": null when the schema demanded "value": 0.

Operators worked around this with regex extraction, JSON5 parsers, retry loops, and "self-correction" prompts. All of these were brittle. The fundamental problem was that JSON was a soft target — the model was trained to produce text that matched the prompt's intent, not to produce text that satisfied a formal schema.

OpenAI's Structured Outputs (August 2024) and Anthropic's tool_use response format (mid-2024, refined through 2025) solved this differently. Both vendors now offer an API mode in which the model's token-level decoding is constrained by a JSON schema. The model literally cannot emit a token that would violate the schema. There is no Markdown code fence. There is no preamble. There is no trailing comma. The output is parseable JSON conforming to your schema or the API errors out before sending you a response.

This is not a slight improvement. It is a change in what kind of pipelines you can ship. Before strict structured output, every JSON-consuming workflow needed a defensive parser and a fallback. After, the defensive parser is a belt for the suspenders — most pipelines no longer hit the failure mode at all.

Strict structured output is the most important feature shipped to LLM APIs since function calling. Treat it as the default. The free-form JSON prompt pattern is a 2023 technique that operators still use out of habit and should retire.

OpenAI Structured Outputs: The Mechanics

OpenAI's Structured Outputs is enabled by setting response_format on a chat-completions request to a JSON schema object:

{
  "model": "gpt-5",
  "messages": [...],
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "rfp_spec",
      "strict": true,
      "schema": { ...your JSON Schema here... }
    }
  }
}

The "strict": true flag is the magic. With strict mode, OpenAI's API constrains decoding so the output exactly matches your schema. The schema must follow specific rules:

  • Every property must appear in the required array. No optional fields. Use "type": ["string", "null"] for fields that may be absent.
  • additionalProperties: false must be set on every object.
  • Schemas have a maximum nesting depth of 5 levels and a maximum of 100 properties total.
  • Some JSON Schema keywords are not supported in strict mode (regex patterns, conditional schemas, recursive references).

These rules feel restrictive at first. They are not — they reflect what is actually decodable with token-level constraint. The reward is that you stop writing JSON-repair code.

A working schema for RFP extraction

Given the RFP spec we built in Lesson 1 of this chapter, here is the strict-mode-compatible schema:

{
  "type": "object",
  "additionalProperties": false,
  "required": ["issuer", "opportunity", "products_in_scope", "evaluation_criteria", "confidence"],
  "properties": {
    "issuer": {
      "type": "object", "additionalProperties": false,
      "required": ["company_name", "domain", "industry"],
      "properties": {
        "company_name": { "type": "string" },
        "domain": { "type": ["string", "null"] },
        "industry": { "type": ["string", "null"] }
      }
    },
    "opportunity": {
      "type": "object", "additionalProperties": false,
      "required": ["estimated_value_usd_min", "estimated_value_usd_max", "decision_date", "is_renewal", "is_competitive"],
      "properties": {
        "estimated_value_usd_min": { "type": ["number", "null"] },
        "estimated_value_usd_max": { "type": ["number", "null"] },
        "decision_date": { "type": ["string", "null"], "description": "ISO-8601 date or null" },
        "is_renewal": { "type": "boolean" },
        "is_competitive": { "type": "boolean" }
      }
    },
    "products_in_scope": { "type": "array", "items": { "type": "string" } },
    "evaluation_criteria": {
      "type": "array",
      "items": {
        "type": "object", "additionalProperties": false,
        "required": ["criterion", "weight_pct"],
        "properties": {
          "criterion": { "type": "string" },
          "weight_pct": { "type": ["number", "null"] }
        }
      }
    },
    "confidence": {
      "type": "object", "additionalProperties": false,
      "required": ["overall", "missing_fields"],
      "properties": {
        "overall": { "type": "number" },
        "missing_fields": { "type": "array", "items": { "type": "string" } }
      }
    }
  }
}

Note the flatten of estimated_value_usd_range from a two-element array to two explicit min/max fields. Strict mode does not support tuple schemas; explicit-field is cleaner and CRM-friendly anyway.

Anthropic Tool_Use Response Format: The Mechanics

Anthropic exposes structured output through the tools mechanism. You define a tool with an input_schema, ask the model to call that tool, and the model's response is the tool-use block with arguments matching your schema. Anthropic's approach treats structured output as a special case of tool calling, which is elegant if slightly counter-intuitive at first.

A working invocation:

{
  "model": "claude-sonnet-4-5",
  "messages": [...],
  "tools": [{
    "name": "extract_rfp_spec",
    "description": "Extract the structured RFP specification from the provided document.",
    "input_schema": { ...your JSON Schema here... }
  }],
  "tool_choice": { "type": "tool", "name": "extract_rfp_spec" }
}

The tool_choice with explicit name forces the model to call that specific tool, eliminating the chance the model emits a text response instead. The response will contain a content array with one tool_use block whose input property is your structured output.

Anthropic's tool_use is slightly more permissive than OpenAI's strict mode in what JSON Schema features it accepts. Optional fields are allowed (no required-everything rule). additionalProperties: false is recommended but not enforced. Recursive schemas and patterns are accepted. The trade-off is that Anthropic's schema enforcement is "best effort" rather than the token-level constraint of OpenAI's strict mode. In practice, with Claude Sonnet 4.5 we measure 99.4% schema-conformance on a 500-call benchmark in March 2026, versus 99.96% for OpenAI strict mode on GPT-5.

Why use one vs the other

  • OpenAI strict mode wins when your schema fits within the strict-mode constraints (required-everything, no patterns, no recursion) and you want the token-level guarantee. Cost is identical to normal calls. Latency is approximately equal.
  • Anthropic tool_use wins when your schema needs flexibility (optional fields, patterns, recursive nesting) and you can accept the ~0.5% schema-deviation rate (which your fallback path catches anyway). It also wins when you have already standardized on Claude for other reasons (cost, quality on long documents, prompt-cache friendliness).

Most operator-built pipelines do not switch providers between calls. Pick one, build your schema for that provider's constraints, and move on.

The Fallback Retry Path: When the Model Still Returns Prose

Even with strict structured output, you will see prose responses. Three causes:

  1. The API itself failed to honor strict mode due to a transient bug. We have seen this <0.04% of the time on OpenAI in 2025; never reported on Anthropic tool_use with explicit tool_choice in 2025-2026.
  2. The schema was violated in a way the validator allowed (e.g., the model emitted an empty string where the application logic needs non-empty). Schema-valid but business-invalid.
  3. The model called a different tool, or refused. With Anthropic this is possible if tool_choice is set to "auto" rather than explicit-name.

The fallback path handles all three. Architecturally it has three layers.

Layer one: schema validation

After receiving the response, parse and validate against your schema using a strict validator. AJV (Node), pydantic (Python), and the workflow tool's JSON validation modules all work. If validation fails, capture the validation error and the raw response — you need both for the retry.

Layer two: business-rule validation

Schema validation does not catch every problem. Add a second-stage validator for business rules:

  • Required-but-empty: a string field is present but empty. Reject.
  • Out-of-range numbers: a confidence score outside [0, 1]. Reject.
  • Date parsing: "decision_date" is a string but is not parseable as ISO-8601. Reject.
  • Cross-field consistency: estimated_value_usd_max < estimated_value_usd_min. Reject.
  • Domain validity: domain field does not look like a domain. Reject.

These are the "schema-valid but business-invalid" cases. Each one becomes a specific error type the retry path can address.

Layer three: the corrective retry

On validation failure, retry with a prompt that explicitly cites the error and asks for a corrected response:

The previous response had the following validation errors:

{{validation_errors_json}}

Please correct these errors and provide the complete, valid response according to the schema. Do not change values that were correct; only fix the cited errors.

On the retry, pass the same schema and the prior conversation history. The retry succeeds 92-96% of the time in our measurements. If the second attempt also fails, escalate to a human queue. Do not retry indefinitely — multi-retry storms are the cost anti-pattern from the previous chapter.

The full state machine

  1. Call LLM with strict structured output / tool_use.
  2. Parse response. If parse error, increment retry counter; if counter < 2, retry with error message; else escalate.
  3. Validate schema. If validation error, increment retry counter; if counter < 2, retry with errors; else escalate.
  4. Validate business rules. If business-rule violation, increment retry counter; if counter < 2, retry with errors; else escalate.
  5. On success, pass the validated object to the next workflow step.

In production this state machine succeeds on first attempt 98.5%+, on first or second attempt 99.9%+, and escalates the remaining 0.1% to a human. Those are numbers worth shipping.

Prompt Design With Strict Output

Strict mode does not eliminate the need for good prompts. It eliminates the need for "respond in JSON only" hectoring at the end of every prompt. Three prompt design moves matter even with strict output enabled.

Use the schema's descriptions to guide the model

Every field in your JSON schema can have a description property. The model reads these. They are the right place to put extraction rules:

  • "description": "ISO-8601 date (YYYY-MM-DD). Set to null if the document does not specify a decision date."
  • "description": "Overall confidence in this extraction, 0 to 1. 0.9+ means every required field was clearly stated in the document; 0.6-0.9 means most fields were stated; below 0.6 means the document was ambiguous or partially unreadable."
  • "description": "List of field paths (dot-notation) that the model could not derive from the document. Empty array if all fields were derivable."

Field descriptions consistently outperform same-text instructions in the system prompt because they sit immediately next to the field at decoding time. Use them.

Provide explicit enum values where the universe is closed

If a field has a known set of valid values, encode it in the schema with "enum": [...]. Examples: industry sectors, certification names (SOC2 / ISO27001 / HIPAA / FedRAMP / PCI-DSS), submission format (email / portal / sftp). Enum-constrained fields cannot be wrong-spelled. The model literally cannot emit "SOC 2" if the enum has "SOC2."

Use system prompt for instructions, user prompt for the document

The system prompt should be stable across calls so it benefits from prompt caching (Anthropic cache_control or OpenAI's automatic prefix cache). Put your extraction philosophy, your "never invent values" rule, your null-and-missing_fields convention there. The user prompt is the variable content — the extracted document text — and any per-document context (sender domain, page count). This structure maximizes cache hits and minimizes per-call cost.

Real Failure Modes From Production

Failure: empty string where business logic needed non-empty

A model with strict-mode output emitted "company_name": "" on a document where the issuer name was OCR-garbled. Schema-valid (string type). Business-invalid (empty company name). The Salesforce upsert fails because Account.Name is required. The fix: business-rule validator catches this, retry asks "the company_name field is empty; either populate it from the document or set the entire response's confidence.overall to below 0.6 if the issuer cannot be determined."

Failure: enum drift

The model emitted "industry": "Banking and Financial Services" when the schema's enum was ["financial_services", "healthcare", "retail", "tech", "manufacturing", "other"]. In OpenAI strict mode this is impossible. In Anthropic tool_use this happens roughly 0.3% of the time when enum values are not idiomatic. The fix: use enum values that match what the model would naturally produce ("Financial Services" rather than "financial_services") or add a description with the rule.

Failure: array of one

The model emitted "products_in_scope": ["Identity and Access Management, Privileged Access Management, Single Sign-On"] as a single comma-joined string in an array of length 1, instead of three separate array items. Schema-valid (array of strings, length >= 0). Business-invalid because downstream code expects three items. The fix: add field description with explicit splitting rule, and a business-rule validator that flags suspiciously long single-string array items.

Failure: out-of-range confidence

The model emitted "confidence.overall": 95 (interpreting "95% confidence" as the natural numeric form) when the schema expected [0, 1]. Schema "number" type is too permissive; add "minimum": 0, "maximum": 1 to constrain. In strict mode this constraint is honored at decode time.

Failure: date string that is not parseable

The model emitted "decision_date": "early November 2026" for a document that was vague about the deadline. Schema-valid (string). Application-invalid (not ISO-8601). The fix: business-rule validator parses the string, retries with "the decision_date field must be ISO-8601 (YYYY-MM-DD) or null. If the document does not specify an exact date, set to null and add 'opportunity.decision_date' to confidence.missing_fields."

Testing Your Structured Output Pipeline

A pipeline that depends on structured output needs a test fixture. Build one with three layers.

Layer one: synthetic schema-conformance tests

Pre-generated example outputs that exercise every schema branch — every enum value, every nullable field set to null, every nullable field set to a value, every array length from 0 to 10+. Run these through your validator to confirm validation rules behave correctly. This catches schema bugs before the model sees them.

Layer two: golden-document tests

A set of 20-50 real documents with their expected extracted JSON written by hand. Run the model on each document. Compare extracted JSON to expected JSON. Compute exact-match accuracy and field-level accuracy. This is your accuracy bar. Re-run the golden tests on every model swap (Sonnet 4.5 → Sonnet 4.7, Haiku → Sonnet, etc.) to catch regressions before they reach production.

Layer three: adversarial documents

Documents specifically designed to trip up extraction — ambiguous deadlines, multi-issuer RFPs, documents in non-English, redacted sections, half-OCR-garbled scans. The model should either extract correctly or set confidence.overall below 0.6 and list missing_fields. The wrong answer is a confident-but-wrong extraction. Track this on every release.

Prompt Caching With Structured Output

Anthropic's cache_control works with tool_use. The tool definition (including the input_schema) goes in the request and benefits from caching the same way the system prompt does. If your schema is large (1,000+ tokens, common for nested RFP schemas), this is a meaningful cost saver. Place the tool definition in the request before the variable user content and add cache_control: {"type": "ephemeral"} to the last cacheable block.

OpenAI's automatic prefix cache also helps. The schema appears at the start of the system content. The variable user content (document text) appears in the user message. Cache prefix matches as long as the schema and system prompt are stable across calls.

Measured cache savings on a real RFP pipeline with a 1,800-token schema and ~1,500-token system prompt: 47% reduction in input-token cost across a 1,000-call test, when calls are spaced under 4 minutes apart on average.

When to Use Multiple LLM Calls vs One Big Structured Call

The temptation is to extract every field in one call. For small schemas (under ~30 fields) and short documents (under ~10K tokens), this is correct. For larger pipelines, split the call.

The "extract-then-enrich" pattern

First call: extract the core fields with a smaller schema. Issuer, opportunity basics, products in scope, deadline. Second call: take the core fields as input and extract the detailed nested fields (evaluation criteria, key questions, certifications). The second call has more context and a tighter focus, which improves field-level accuracy on the detailed fields.

Trade-off: more total tokens, two calls instead of one. The cost may be higher; the accuracy is consistently better for nested schemas above ~50 fields. The decision should be made on the eval set.

The "extract-then-validate" pattern

First call: extract with a permissive schema and explicit confidence per field. Second call: a smaller model (Haiku 4.5) takes the extraction and the document and answers "are these extracted fields accurate? List any that should be revised." This is the LLM-as-judge pattern applied to extraction. It catches the silent-confident-wrong failures that human review otherwise has to handle.

Key Takeaways

  • Strict structured output (OpenAI Structured Outputs with "strict": true, Anthropic tool_use with explicit tool_choice) eliminates the most common LLM-pipeline failure: prose where JSON was expected. Treat strict mode as the default in 2026.
  • OpenAI strict mode constraints: every property in required, additionalProperties: false on every object, max nesting 5, max 100 properties, no patterns/recursion/conditionals. Returns schema-valid JSON at the token level.
  • Anthropic tool_use is more permissive on schema features (optional fields, patterns, recursion). 99.4% schema-conformance measured on Claude Sonnet 4.5 in March 2026, vs 99.96% on OpenAI GPT-5 strict mode.
  • Build a three-layer fallback retry path: parse + schema validation + business-rule validation. On failure, retry with an explicit error message citing the violation. Max 2 retries before escalation. Production success: 98.5% first try, 99.9% within two tries, 0.1% to human.
  • Business-rule validators catch the schema-valid-but-business-invalid cases: empty strings where non-empty is required, dates that don't parse, confidence out of [0,1], cross-field consistency violations. These are real failures from production.
  • Prompt design with strict output: put extraction rules in JSON schema field description properties (the model reads them), use enum for closed sets, keep the system prompt stable for prompt-cache benefit.
  • Use prompt caching with structured output. Tool definitions cache the same way system prompts do. Measured savings: 47% input cost reduction with a 1,800-token schema and tight call spacing.
  • Test with three layers: synthetic schema-conformance tests, 20-50 golden documents with hand-written expected JSON, and adversarial documents that should produce low-confidence extractions rather than confident-wrong ones.
  • Multi-call patterns for large schemas: extract-then-enrich (core fields then detailed) for nested schemas above ~50 fields; extract-then-validate (LLM-as-judge) for catching silent-confident-wrong failures.
  • The free-form "respond in JSON only" prompt pattern is a 2023 technique. Operators who still use it ship pipelines with 3-8% failure tails. Move to strict mode and recover those failures.