←
AI Agent Builders & Citizen Developers
Capable · M10 · lesson 10 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
PDF and Table Extraction Without Hand-Rolled Code
📖
now learning

PDF and Table Extraction Without Hand-Rolled Code

15 min

There is exactly one PDF on every operator's desk that does not parse correctly. Sometimes it is the customer's vendor security questionnaire with the merged-cell header row that splits a 47-row matrix into 89 phantom rows. Sometimes it is the scanned addendum where every page is an image of a fax of a signed page. Sometimes it is a 240MB pricing matrix where the tables run sideways across two-page spreads. The operator who knows how to pick the right extractor for that one PDF — and who knows the exact moment OCR cost beats LLM cost — ships pipelines that do not break in week three. This lesson is the head-to-head bake-off: LlamaParse, Unstructured, Reducto, and platform-native extractors on a messy real PDF, with the arithmetic on when each one wins and when you should escalate the page image to a vision LLM.

Why You Do Not Write Your Own Extractor in 2026

It is tempting. PyPDF2 is free, pdfplumber is free, Tesseract OCR is free, and a Python developer can wire them together in an afternoon. We have shipped both. The hand-rolled stack works for 70% of PDFs out of the box. The remaining 30% — the ones that actually matter to your inbound triage pipeline — are where you spend the next nine months patching edge cases. Multi-column layouts. Rotated tables. Embedded vector graphics. Mixed font encodings. Scanned pages with skew. Headers that repeat mid-table. Footnotes that bleed into the next page's text layer. Each of these is a weekend of your time, and there are forty of them.

The 2026 managed extractors — LlamaParse, Unstructured, Reducto, plus the platform-native extractors in n8n, Make, and Zapier — exist because four well-funded teams spent the last 24 months solving those forty edge cases. The list price for processing a single PDF page through any of them is between $0.003 and $0.04. The cost of one engineering hour debugging your hand-rolled stack is, on average, $200. The break-even is roughly 6,667 pages per engineering hour. If your pipeline processes more than 6,667 pages per quarter, the managed extractor is strictly cheaper than maintaining your own. Almost every production inbound-document pipeline crosses this line in month one.

The free extractor is not free. It costs you the time to debug the 30% of PDFs it cannot handle. The managed extractor is the answer for anyone whose throughput exceeds 6,667 pages per quarter — and that is everyone running this kind of pipeline at scale.

What "managed extractor" actually means

A managed extractor is a hosted service that accepts a PDF (or image, or DOCX, or PPTX) and returns one or more of: extracted text, structured Markdown, JSON-shaped tables, page-level metadata (orientation, language, character count, OCR confidence), and optionally bounding boxes for figures and tables. The differentiation across vendors is mostly in three areas: table fidelity (how well do complex tables come out), OCR quality (how well do scanned pages render), and chunking semantics (whether the output preserves logical document structure for downstream LLM use).

The Four Extractors in May 2026

LlamaParse (LlamaIndex)

The default choice for most operator-built pipelines in 2026. LlamaParse offers three modes — "fast" ($0.003/page for plain extraction), "balanced" ($0.005/page with basic table support), and "premium" ($0.015/page with vision-LLM-backed table reconstruction and complex layout handling). The API is one POST with a file upload and a results-polling URL. Output is Markdown by default, with optional JSON-tabular for tables. The premium mode uses Claude Sonnet 4.5 under the hood for layout-sensitive pages and gives the strongest table fidelity outside of Reducto.

Strengths: simple API, predictable pricing, robust on long documents (no per-document size cap up to 1GB), good Markdown semantics that downstream LLMs handle well. Weaknesses: premium mode latency is 8-25 seconds per document depending on length; complex nested tables sometimes degrade in the "balanced" tier; the OCR mode for scanned pages is competent but not best-in-class.

Unstructured (Unstructured.io)

The most flexible extractor, especially for self-hosted deployments. Unstructured exposes both a hosted API ($0.01/page on the "hi-res" strategy that includes table extraction) and an open-source library that runs on your own infrastructure (free; you pay for compute and OCR). The output is a list of "Element" objects (Title, NarrativeText, Table, Image, ListItem, etc.) with rich metadata. Tables come out as both HTML and a structured cell representation.

Strengths: best-in-class for diverse document types (HTML, Markdown, DOCX, PPTX, EML, HTML email, plus PDFs), open-source option for compliance-sensitive deployments, element-level granularity is excellent for RAG pipelines. Weaknesses: the hosted API is rate-limited tightly on free tier; complex tables come out structurally clean but cell-content fidelity sometimes lags Reducto and LlamaParse premium; self-hosted "hi-res" requires a GPU for performant OCR.

Reducto

The specialist. Reducto was built specifically for table-heavy documents — financial statements, pricing matrices, lab reports, vendor questionnaires. Pricing in May 2026 is $0.04/page on tables-included extractions; cheaper tiers for text-only pages. The output includes a "tables" array with each cell typed and positioned, plus a confidence score per cell.

Strengths: highest table fidelity in the market by an observable margin, especially for merged cells, multi-row headers, and tables that span page breaks. Strong handling of rotated and landscape-oriented tables. Per-cell confidence scores enable downstream validation. Weaknesses: most expensive option; overkill for text-heavy documents where the tables are simple or absent; smaller community and fewer integrations than LlamaParse.

Platform-native extractors

n8n's "Extract from PDF" node, Make's PDF.co bundle, and Zapier's Formatter+Files modules all expose competent text extraction directly from their workflow canvas. Pricing is bundled into the platform subscription rather than per-page. The quality bar is "fine for text, weak for tables." None of the platform-native options handle multi-column layouts well; none reconstruct complex tables; none expose per-page confidence.

Strengths: zero additional cost, zero additional auth, immediate availability in the workflow canvas, sufficient for simple text-extraction use cases. Weaknesses: not production-grade for table-heavy or scan-heavy workloads; failure modes are silent (you get back garbled text instead of an explicit error); no escalation path for OCR-blank pages.

The Bake-Off: One Messy Real PDF

To make this concrete, we ran the same RFP PDF through all four extractors. The document is a 38-page vendor security questionnaire from a Fortune-500 financial services firm. It has: 6 pages of prose (cover letter, scope, instructions), 22 pages of multi-part questionnaire with a mix of yes/no toggles and free-text fields, 4 pages of pricing matrix with merged headers and multi-row groupings, 3 pages that are scans (the issuer's standard NDA), and 3 pages of network architecture diagrams that are largely images. It is, in other words, a realistic real-world RFP. Here is what came out of each extractor.

LlamaParse premium results

Total processing time: 14 seconds. Total cost: $0.57 (38 pages × $0.015). The prose pages came through clean as Markdown. The 22 questionnaire pages came through with the question-and-answer structure intact, including the yes/no toggles rendered as bracket-pair indicators. The pricing matrix came out as a Markdown table that was 90% correct: 47 of 51 rows were structurally clean; 4 rows had merged-cell content split incorrectly across columns. The 3 scanned NDA pages came through with OCR but had 3-4 character-recognition errors per page. The diagram pages returned a description ("Network architecture diagram showing three-tier topology with...") rather than the diagram itself, which is acceptable for downstream LLM use but loses fidelity.

Unstructured (hi-res) results

Total processing time: 22 seconds. Total cost: $0.38. The element-based output gave us 384 Element objects across the document. Prose came out as a mix of Title and NarrativeText elements with strong structure. The questionnaire structure was less coherent than LlamaParse's Markdown — questions and answers were sometimes split into separate elements requiring downstream re-association. The pricing matrix came as Table elements with HTML representation; row count was correct at 51 but two merged-cell groups in the header lost their merge semantics. Scanned NDA pages came through with comparable OCR quality to LlamaParse. Diagrams returned as Image elements with metadata but no content description.

Reducto results

Total processing time: 18 seconds. Total cost: $1.52 (38 pages × $0.04). The pricing matrix was perfect: 51 rows, all merged cells correctly preserved, per-cell confidence scores ranging 0.91-1.00. The prose pages came through fine but with fewer structural cues than LlamaParse — Reducto's optimization for tables shows in the prose output, which feels less Markdown-friendly. Scanned NDA pages came through with OCR confidence per word, and the recognition errors were down to 1-2 per page versus LlamaParse's 3-4. Diagrams returned bounding boxes and a brief description.

Platform-native (n8n Extract from PDF) results

Total processing time: 3 seconds. Total cost: $0 (bundled). The output was a single flat text string of approximately 22,000 characters. Prose came through readable. Questionnaire structure was completely lost — Q&A pairs were jumbled and the toggles were absent. The pricing matrix was a free-text dump with no row alignment. Scanned NDA pages returned blank. Diagrams returned blank.

Side-by-side scoring

If we score each extractor on five dimensions (prose fidelity, table fidelity, OCR quality, downstream-LLM friendliness, cost) on a 1-5 scale:

  • LlamaParse premium: Prose 5, Tables 4, OCR 4, LLM-friendliness 5, Cost 4. Total: 22/25.
  • Unstructured hi-res: Prose 4, Tables 4, OCR 4, LLM-friendliness 3, Cost 4. Total: 19/25.
  • Reducto: Prose 4, Tables 5, OCR 5, LLM-friendliness 4, Cost 2. Total: 20/25.
  • Platform-native: Prose 3, Tables 1, OCR 0, LLM-friendliness 2, Cost 5. Total: 11/25.

The headline: LlamaParse premium is the best general-purpose default. Reducto is the right choice when tables dominate. Unstructured is right when you need self-hosting or non-PDF input variety. Platform-native is right when you only need rough text and the bar is low.

Picking the Cheapest That Hits Accuracy

The operator-builder's question is never "which is best" but "which is cheapest at my accuracy bar." Define the accuracy bar first, then walk down the price ladder until you hit something below the bar.

Define the accuracy bar in measurable terms

Vague accuracy targets ("good enough") are worthless. Pick three measurable metrics that map to downstream business outcomes:

  • Table row recall: percentage of rows in source tables that appear in extracted output. For a pricing matrix, 100% is mandatory. For a less critical appendix, 95% may suffice.
  • Question-answer pair preservation: percentage of Q&A pairs in source questionnaires that remain associated after extraction. For a security questionnaire, 95%+ is required.
  • OCR character error rate: percentage of characters incorrectly recognized on scanned pages. For legal text (NDAs, contracts), <0.5% is required. For non-load-bearing pages, 2-3% is acceptable.

Build a 20-document evaluation set that reflects your real production distribution. Run all four extractors. Compute the three metrics. Now you have data instead of vibes.

The price ladder

Walk from cheapest to most expensive:

  1. Try platform-native first. If it hits the bar on text-only pages, use it for text and escalate tables/scans elsewhere.
  2. Try LlamaParse "balanced" ($0.005/page). If it hits the bar on tables and OCR, you are done.
  3. Try LlamaParse "premium" ($0.015/page) or Unstructured "hi-res" ($0.01/page). Most pipelines stop here.
  4. Try Reducto ($0.04/page) if and only if tables are central to your accuracy bar and the cheaper options fall short.
  5. Escalate per-page to a vision LLM only on the pages where the extractor flags low confidence or returns sub-threshold character count.

The "hybrid" pattern — text-only pages through LlamaParse balanced, table-heavy pages routed to Reducto — saves real money on table-light documents and pays for itself in compliance-sensitive ones. You do this with a two-pass workflow: first pass uses the cheap extractor and inspects per-page output; pages flagged for re-extraction (low char count, table detected, OCR confidence low) get routed to the expensive extractor.

When OCR Cost Beats LLM Cost

Here is the moment that surprises most operators. There is a price point where it is cheaper to send the rendered page image to a vision LLM and ask it to transcribe than it is to run premium OCR. The math depends on three numbers: page size in tokens after vision encoding, LLM input price, and OCR price.

The arithmetic

A typical PDF page at 150 DPI renders to roughly 1,400-1,800 vision tokens at Claude Sonnet 4.5's encoding density. At $3/M input tokens, that is $0.0042-$0.0054 per page just for input. Output tokens for a transcription are usually 400-800 tokens at $15/M, so $0.006-$0.012 per page output. Total per-page LLM transcription cost: $0.010-$0.017.

Compare to extractor prices:

  • LlamaParse premium: $0.015/page — LLM transcription is competitive ($0.010-$0.017).
  • Unstructured hi-res: $0.01/page — LLM transcription is more expensive (unless you need vision-LLM judgment on what is on the page, in which case it is a different feature).
  • Reducto: $0.04/page — LLM transcription is dramatically cheaper.
  • Platform-native: $0/marginal — LLM transcription is the most expensive option on a per-page basis, but the platform-native option fails on scans entirely, so the comparison is moot.

When is the LLM-transcription path actually right?

Three scenarios. First, when you only need to escalate a handful of pages per document (the per-page premium pricing is acceptable when volume is small). Second, when the page contains content that requires interpretation as well as transcription — diagrams that need description, mixed text-and-image content where the relationship matters. Third, when the document is small enough that the entire-document path through vision LLM is competitive with sending the PDF to a managed extractor.

For a 38-page PDF, end-to-end vision-LLM transcription at $0.013/page averages out to $0.49 per document. LlamaParse premium at $0.015/page is $0.57 per document — comparable. Reducto at $0.04/page is $1.52 per document — three times more. For a table-heavy document the Reducto premium is justified by table quality. For a prose-heavy or diagram-heavy document, the vision-LLM path may win on both cost and fidelity.

The "page image to vision LLM" escalation pattern

Most pipelines do this only on flagged pages. The flag triggers: per-page character count below 50, OCR confidence below 0.8 (where the extractor reports it), or a "Table" element with cell count below the expected count for that page region. The escalated page renders the PDF page to an image (1024x1024 or thereabouts), wraps it in a vision-LLM API call with a transcription prompt, and merges the result back into the document text stream.

n8n's "PDF to Image" node, Make's PDF.co Render module, and Zapier's Image Conversion bundle all expose this conversion. The downstream LLM call is a standard chat-completion with a vision-capable model — Claude Sonnet 4.5, GPT-5, or Gemini 2.5. The prompt:

You are transcribing one page of a PDF document. Output the page content as Markdown. For tables, use Markdown table syntax. For figures and diagrams, write a one-paragraph description prefixed with "[FIGURE]". Preserve heading hierarchy. Do not invent content not visible on the page. If a region is unreadable, mark it as "[ILLEGIBLE]" with a brief description of what is partially visible. Output only the Markdown, no preamble.

This prompt is portable across models. The "balanced" extractor pipelines that use this pattern hit accuracy bars comparable to Reducto at a fraction of the all-pages cost.

Chunking for Downstream LLM Use

Extraction is not the whole story. The output has to feed into a downstream LLM call. Bad chunking eats accuracy at the next step.

Why chunking matters even for "long-context" models

Claude Sonnet 4.5 supports 200K input tokens. GPT-5 supports 1M. So why chunk a 38-page PDF that fits comfortably in either? Three reasons. First, attention degrades on long inputs — even within the context window, models pay less attention to content in the middle of very long inputs (the "needle in a haystack" effect). Second, cost is linear in input tokens; sending the whole document for every micro-question wastes money when one chunk is sufficient. Third, retrieval-augmented patterns (RAG) require chunked storage; if you want to query "what did this RFP say about HIPAA compliance" later, you need chunks.

Chunking strategies

  1. Fixed-size token windows (e.g., 500 tokens per chunk, 50-token overlap). Simple, robust, but cuts mid-sentence and mid-table. Use only when document structure is unreliable.
  2. Structural chunking — chunk by Markdown heading (H2 boundaries) or by Element type (Title-bounded sections for Unstructured). Preserves meaning. Recommended default.
  3. Semantic chunking — chunk by topical similarity using embeddings. Best fidelity, highest cost, used in RAG pipelines where the additional setup is justified.

For inbound-document triage where the goal is a one-shot structured-output extraction, structural chunking is usually unnecessary because you pass the whole document to the LLM in one call. For multi-stage workflows where you ask many questions of the same document, semantic chunking pays off.

Real Numbers, Real Budgets

Pipeline A: SaaS company, 60 RFPs per quarter, average 25 pages

Pages per quarter: 1,500. LlamaParse premium at $0.015: $22.50/quarter, $7.50/month. Cheapest path through the price ladder. Total extraction cost per RFP: $0.38. Compared to the LLM extraction cost of $0.33 per RFP at Sonnet 4.5 pricing — extraction is the more expensive step on this workload.

Pipeline B: Financial services compliance team, 200 documents per month, average 80 pages, table-heavy

Pages per month: 16,000. Reducto at $0.04/page: $640/month. Stiff but justified by the table fidelity — the compliance team's downstream tool requires exact pricing-matrix preservation. Compared to LlamaParse premium at $240/month — Reducto's $400/month premium buys the difference between 90% and 100% table fidelity, which on this workload prevents two compliance escalations per month at an average remediation cost of $1,800 each. ROI: 9x.

Pipeline C: Insurance claims processing, 4,000 claim packets per month, average 6 pages, mixed scan and digital

Pages per month: 24,000. Hybrid pattern: LlamaParse balanced at $0.005/page for digital, vision-LLM escalation at $0.013/page for the ~25% of pages that are scans. Total: 18,000 × $0.005 + 6,000 × $0.013 = $90 + $78 = $168/month. Compared to all-pages Reducto: $960/month. Hybrid saves $792/month. The hybrid pattern is dominant in scan-heavy pipelines.

The Extractor Failure Runbook

Failure: extractor returns 200 OK with empty pages

Symptom: the API succeeds, the page count is correct, but several pages have content length zero. Cause: page is a scan and the extractor's free tier did not run OCR, or page is an image with no text layer. Fix: per-page character-count check after extraction; escalate empty pages to vision-LLM transcription.

Failure: extractor returns 200 OK with garbled text

Symptom: pages have content but the content is character-soup ("Th3 v3nd0r mu$t pr0v1de..."). Cause: font encoding mismatch on a non-standard PDF generator. Fix: re-extract with the premium tier (which uses vision-LLM-backed extraction and bypasses the font-encoding path) or escalate to direct vision-LLM transcription.

Failure: tables come out with shifted columns

Symptom: every row has the right number of cells but the cell contents are shifted one column right or left. Cause: header row had merged cells the extractor failed to detect. Fix: re-extract with Reducto, or add a downstream LLM validation step that compares column headers to expected schema.

Failure: extractor never returns

Symptom: the polling URL returns "processing" for 5+ minutes. Cause: file size near the per-document limit, or the document has a corrupted xref table. Fix: implement a 90-second timeout on the polling, after which the workflow either retries with a cheaper tier or routes to manual triage with the original PDF attached.

When to Add a Vision-LLM Pass as the Default

For most pipelines, the vision-LLM escalation is a per-page exception path. There is a class of workflows where it becomes the primary extraction strategy: documents where the visual layout itself carries meaning that text extraction loses. Examples include architectural drawings with annotations, marketing one-pagers with text laid out over images, lab reports with handwritten margin notes, and invoices where the visual arrangement of line items is structurally significant. For these documents, every page goes through a vision LLM with a "transcribe and describe" prompt, and the cost-per-page is higher but the fidelity is dramatically better than text-extraction alone.

The trade-off is latency. Vision-LLM transcription at 1024x1024 input runs 3-8 seconds per page on Sonnet 4.5; an extractor processes the same page in 0.5-2 seconds. For real-time use cases the extractor is still right; for asynchronous workflows the vision-LLM-first pattern is increasingly competitive.

The Decision Tree: Picking Your Extractor in 90 Seconds

  1. Will your documents have complex tables that matter to downstream decisions? If yes, default to Reducto for those documents.
  2. Will your documents include diverse formats (HTML, email, DOCX, PPTX) beyond PDFs? If yes, default to Unstructured.
  3. Will your pipeline run on documents that contain visual content requiring interpretation? If yes, use vision-LLM transcription as primary or per-page escalation.
  4. Is regulatory or compliance sensitivity such that you must self-host? If yes, use Unstructured open-source library.
  5. If none of the above special cases apply, use LlamaParse premium as the default with a vision-LLM escalation path for low-char-count pages.

Most operator-built inbound pipelines end with: LlamaParse premium for 90-95% of pages, vision-LLM escalation for the remaining 5-10%, and structural chunking on Markdown headings. This stack is the cheapest one that hits the accuracy bar for the median RFP-triage workload.

Key Takeaways

  • Do not hand-roll an extractor in 2026 unless your pipeline processes fewer than 6,667 pages per quarter. Managed extractors (LlamaParse, Unstructured, Reducto, platform-native) priced $0.003-$0.04/page are strictly cheaper than the engineering hours required to maintain a hand-rolled stack.
  • Four extractors in May 2026: LlamaParse (general-purpose default, $0.003-$0.015/page across tiers), Unstructured (flexible, element-based, $0.01/page hosted or self-hosted), Reducto (table specialist, $0.04/page, highest table fidelity), platform-native (free, text-only, weak on tables and OCR).
  • Pick the cheapest that hits your accuracy bar. Define the bar in measurable terms: table row recall, Q&A pair preservation, OCR character error rate. Build a 20-document eval set. Walk the price ladder from cheapest to most expensive.
  • On a 38-page bake-off, scores were LlamaParse premium 22/25, Reducto 20/25, Unstructured hi-res 19/25, platform-native 11/25. LlamaParse premium is the best general-purpose default; Reducto wins on table fidelity; Unstructured wins on format diversity; platform-native is text-only sufficient.
  • Vision-LLM transcription costs $0.010-$0.017 per page at Sonnet 4.5 vision pricing. It beats Reducto's $0.04/page on prose-heavy documents and complements LlamaParse premium on the 5-10% of pages flagged for low character count or low OCR confidence.
  • Hybrid extraction (cheap extractor for clean digital pages, vision-LLM escalation for scans) dominates scan-heavy pipelines. Insurance claims example: $168/month hybrid vs $960/month all-Reducto, $792/month savings.
  • Chunking matters even with long-context models: attention degrades mid-document, cost is linear in input tokens, RAG patterns require chunks. Structural chunking (by Markdown heading or Element type) is the default; semantic chunking via embeddings for RAG pipelines.
  • Failure modes: extractor returns 200-OK-but-empty (escalate to vision-LLM), 200-OK-but-garbled (re-extract premium or vision-LLM), shifted-column tables (re-extract Reducto or LLM-validate against schema), extractor-never-returns (90-second timeout, retry or manual triage).
  • Decision tree: complex tables → Reducto; diverse formats → Unstructured; visual content carrying meaning → vision-LLM primary; compliance/self-host → Unstructured open-source; otherwise → LlamaParse premium plus vision-LLM escalation for low-char-count pages.