←
AI Agent Builders & Citizen Developers
Capable · M20 · lesson 20 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The RFP-and-Inbound-Email Triage Workflow
📖
now learning

The RFP-and-Inbound-Email Triage Workflow

15 min

An RFP lands in a shared inbox at 4:47pm on a Thursday. Forty-three pages of attached PDFs, a vendor questionnaire in Excel, two reference URLs, and a deadline in eleven days. By Monday morning, three different people on the team have read the same first six pages, two have read different sections of the technical questionnaire, and nobody has filed a Salesforce opportunity yet. This lesson is the end-to-end pipeline that turns that inbound chaos into a structured spec, an opportunity record, and attachments filed in the right Salesforce folder — automatically, before anyone reads page one. We will cover the trigger, the extraction, the structured spec, the CRM write, and the four failure modes that kill the pipeline in week three: orphan attachments, OCR-blank pages, scanned PDFs with embedded images, and the malformed-MIME problem that swallows your most important RFP of the quarter.

Why RFP Triage Is the Canonical Inbound Workflow

Operators who are picking their first inbound-document agent almost always pick the wrong one. They pick the "summarize all customer support tickets" project because it has the largest call volume. Volume is the worst possible starting criterion. The right starting criterion is economic per-event value: how much does a single missed or mis-triaged inbound event cost the business? RFPs and large inbound emails win this race by a wide margin. A single missed RFP at a B2B SaaS company is a $40k-$400k pipeline event. A single missed support ticket is, on average, less than $50 of operational drag. The math on attention-per-dollar overwhelmingly favors building the inbound-document pipeline first.

There is a second reason. RFP triage exercises every primitive that every other inbound-document workflow uses: email parsing, attachment handling, PDF extraction, OCR, structured-output coercion, CRM writes, and human-in-the-loop review. Ship the RFP pipeline and you have shipped a reusable template for "vendor security questionnaires," "compliance attestation requests," "RFI responses," and "due-diligence packet processing" — five workflows for the price of one.

If you only build one inbound-document agent in 2026, build the one that catches your biggest revenue events. Volume is a distraction. Build for value.

What "shipping" looks like end-to-end

Concretely, a shipped RFP-triage workflow does the following in roughly the order below, with the elapsed wall-clock target of 90 seconds from email arrival to Salesforce record:

  1. Gmail "watch" or Pub/Sub trigger fires the workflow when an email matching the RFP rule set lands in the shared inbox.
  2. The workflow downloads the email body and every attachment, hashing each attachment to detect duplicates within a 30-day window.
  3. PDFs and Office documents pass through an extraction node (LlamaParse, Unstructured, Reducto, or platform-native — Lesson 2 of this chapter covers the trade-offs) that returns Markdown plus tables.
  4. The LLM call extracts a structured RFP spec (issuer, products, deadline, evaluation criteria, key questions, contract value range) using a JSON schema with OpenAI Structured Outputs or Anthropic tool_use response format (Lesson 3 covers this).
  5. The workflow creates a Salesforce Opportunity (or upserts on existing), attaches the documents to the related ContentDocument, and posts to a Slack channel for human review with a "claim" button.
  6. An exception path catches every failure mode (attachment download failed, OCR returned blanks, LLM extraction below confidence threshold) and routes those to a manual-review queue with the original email link.

The point of laying it out this way is that every step is independently testable and every step has a known failure mode. Operators who skip the failure-mode plumbing ship pipelines that work on Tuesday and silently rot by Friday.

The Trigger: Gmail Pub/Sub Watch and the RFP Rule Set

Gmail offers three trigger styles for shared inboxes in 2026: a polling search-query node (most workflow tools default to this), a Gmail Pub/Sub push notification (Google's recommended approach for low-latency triggering), and an inbound-MX route through SendGrid or AWS SES. For most operator-built workflows, the polling search-query approach is the right starting point. It is robust, it is debuggable, and it does not require provisioning a Google Cloud Pub/Sub topic.

The polling search query

In n8n, the Gmail Trigger node polls a search query every 1-5 minutes. The query is the unsung hero of the whole pipeline. A bad query yields false positives that triple your costs and irritate your humans. A good query yields one RFP per hit. Examples that work in practice:

  • label:inbound-rfps is:unread newer_than:2d has:attachment — works when an upstream Gmail filter labels suspected RFPs based on sender domain, subject keywords, or routing address.
  • to:[email protected] is:unread newer_than:2d — works when you have a dedicated catch-all address that vendors and customers know to use.
  • (subject:RFP OR subject:RFI OR subject:proposal OR subject:"request for") is:unread has:attachment newer_than:1d — works when you do not control the routing address and have to pattern-match on subject lines.

The pattern-match query is the most error-prone. A "request for" search matches "request for status update on the SSO integration" and you end up triaging your own internal threads. The cheap fix is a second-stage classifier — a tiny Haiku 4.5 call costing $0.0004 per email that asks "is this email an inbound RFP or proposal request? Answer yes or no in one token." This filter sits between the trigger and the expensive extraction step. It throws away 60-80% of false positives at a per-call cost so low it does not appear on the dashboard.

What to extract from the email itself before touching attachments

Before any attachment is downloaded, capture: sender email and domain, sender name, subject line, body plaintext and HTML, all attachment MIME types and sizes, the Message-ID header (for threading), and the Date header. The sender domain alone often determines routing — an email from @statefarm.com is going to a different evaluator than an email from @acme-consulting.io. Don't waste an LLM call to classify what a domain check trivially answers.

Downloading Attachments Without Losing the Important One

Here is the failure mode that consumes the most engineering hours in the first month of an RFP pipeline: orphan attachments. The email arrives. The body downloads cleanly. One of three attachments — usually the biggest, usually the actual RFP — fails to download because of a transient Gmail API error, a 25MB attachment-size limit on the workflow tool, or a malformed MIME boundary that the parser silently swallows. Your pipeline ships a Salesforce opportunity with the wrong document attached and your sales engineer thinks the customer sent a 4-page summary instead of the 43-page RFP.

The defensive download pattern

Three rules. First, enumerate all attachments before downloading any. Most workflow tools expose a metadata array; capture it. Compare the count against the actual downloads later — if they do not match, raise an exception. Second, hash each downloaded attachment (SHA-256 of the bytes) and store the hash with the metadata. This gives you a deduplication key and a tamper-evident record. Third, store the attachment in a known-stable location (S3, GCS, or Google Drive) before passing the URL to the extraction node. Workflow-tool temporary storage is not reliable on retries.

Attachment size limits, the silent killer

Zapier's default file-handling limit is 150MB per file but only 5MB for files passed between Zaps without explicit storage steps. Make's "binary data" handling caps at 5GB total per scenario run but with execution-time constraints. n8n self-hosted has no hard limit but throws "JavaScript heap out of memory" errors on files above a few hundred MB. A 200MB RFP-response packet with embedded videos and reference imagery will break your default workflow. The fix is to stream large attachments directly to object storage with a presigned URL and pass only the URL to subsequent nodes, never the bytes.

A team in March 2026 lost a $180k RFP because their workflow downloaded the 8MB cover letter and silently failed on the 240MB technical addendum. The sales engineer responded with the wrong scope of work. The customer noticed. The deal went to a competitor. The post-mortem was a single Slack message: "we need attachment-size telemetry." Build it day one.

PDF Extraction: The Three Failure Modes That Burn You

Lesson 2 of this chapter goes deep on tooling choice. For now, assume you have picked an extraction provider (LlamaParse is a sensible default; Reducto for complex tables; Unstructured if you self-host; platform-native if your platform exposes a competent extractor). The question is what to do when the extractor returns garbage. There are three flavors of garbage that the operator-builder needs to recognize on sight.

Failure one: the blank-page scan

The PDF looks like a PDF. It opens in Acrobat. It has text. But when your extractor runs OCR on page 7, the result is an empty string. Why? Because page 7 is a scan of an image — usually a network diagram, a logo-heavy cover page, or a hand-marked-up architecture sketch — that the OCR engine could not interpret as text. Your downstream LLM sees the gap and proceeds as if the page did not exist.

The defensive move is to check character count per page. After extraction, compute the character count for each page. If any page returns below a threshold (say, 50 characters for a page that should reasonably contain prose, or 0 characters for any page), flag the document for human review or escalate to a vision-capable LLM call. Claude Sonnet 4.5 and GPT-5 both accept page images directly. The escalation pattern: extractor returns text + image for each page, you send the image of any low-character page to the vision model with a prompt like "transcribe and describe what is on this page." Cost: roughly $0.01-$0.04 per page depending on resolution. Worth every penny for the page that contains the deal-breaking compliance requirement.

Failure two: the embedded-image-with-text

Worse than blank pages: pages that look like they have text but the "text" is actually an image of text. Common in RFPs where the issuer has pasted screenshots of their own slide deck. A naive extractor returns the page's text layer (often empty) and skips the image. The vision-LLM escalation handles this case identically to failure one. The character-count threshold catches both.

Failure three: table-shaped chaos

The RFP includes a 3-page pricing matrix. LlamaParse returns the table as Markdown with the columns roughly intact but with rows split across page breaks, footers merged into cells, and the header repeated mid-table. Reducto handles this better because it was built for table-heavy documents. Unstructured returns elements as a list of "Table" objects but the cell structure is wobbly on complex headers.

The defensive move when tables matter: do not rely on the extractor alone. Have the downstream LLM call validate the table by re-rendering it as JSON and checking that the row count matches the document's expected row count (you can extract the expected count from the surrounding prose: "the following table lists 47 evaluation criteria..."). If the LLM's JSON has 41 rows when the prose claims 47, raise an exception and escalate. We cover this validation pattern in detail in Lesson 3.

The LLM Call That Produces the RFP Spec

The structured spec is the heart of the workflow. Get this right and downstream Salesforce writes are trivial. Get it wrong and you are debugging null fields in production at midnight. The schema we recommend after shipping eleven of these pipelines:

The RFP spec schema

{
  "issuer": { "company_name": "string", "domain": "string", "industry": "string", "headcount_estimate": "string" },
  "opportunity": { "estimated_value_usd_range": ["min", "max"], "decision_date": "ISO-8601 date or null", "incumbent_vendor": "string or null", "is_renewal": "boolean", "is_competitive": "boolean" },
  "products_in_scope": ["string"],
  "deal_breakers": ["string"],
  "evaluation_criteria": [ { "criterion": "string", "weight_pct": "number or null" } ],
  "key_questions": [ { "section": "string", "question": "string", "answer_owner_team": "string" } ],
  "required_certifications": ["SOC2", "ISO27001", "HIPAA", "FedRAMP", "..."],
  "submission_format": { "type": "email/portal/sftp", "address_or_url": "string", "required_attachments": ["string"] },
  "confidence": { "overall": "0-1 float", "missing_fields": ["string"] }
}

Two things to note. First, every field has an explicit null-or-absent representation. The model is not allowed to invent. If a value is not in the document, the model is instructed in the prompt to set it to null and add the field name to confidence.missing_fields. This sounds obvious but the alternative — silent hallucinations that look like real values — is the most common production failure in inbound-document agents. Second, the schema explicitly includes a confidence object. Downstream routing uses it: high confidence routes straight to Salesforce, lower confidence routes to Slack for review first.

The prompt template, abridged

System prompt (kept stable for prompt-caching benefit, around 1,800 tokens in real deployments):

You are an RFP intake analyst. You read inbound RFP and proposal documents and extract a structured specification. You never invent values. If a field is not explicitly present in the document, set it to null and add the field name to confidence.missing_fields. You output only the JSON object matching the provided schema, with no surrounding prose. The schema is enforced; you must populate every required field. For dates, prefer ISO-8601. For currencies, normalize to USD using the rate noted in the prompt or null if not derivable. For company names, use the exact legal name where available; otherwise the trading name as written in the document.

User prompt (varies per RFP):

The following is the extracted text of an inbound RFP. The sender domain is {{sender_domain}}. The subject line was: "{{subject}}". The document has {{page_count}} pages.

--- BEGIN DOCUMENT ---
{{extracted_markdown}}
--- END DOCUMENT ---

Extract the RFP specification according to the provided JSON schema. If the document does not contain enough information for a field, set the field to null and list it in confidence.missing_fields. Compute an overall confidence between 0 and 1 reflecting your certainty in the extraction.

Model choice: Claude Sonnet 4.5 is the workhorse for this extraction task in May 2026. GPT-5 is competitive. Haiku 4.5 fails on complex multi-page documents because it loses fidelity on the deal-breakers and evaluation-criteria fields. We measured 94% extraction accuracy on Sonnet 4.5 versus 81% on Haiku 4.5 on a 50-RFP benchmark in March 2026, with the gap concentrated on long documents above 30 pages.

Writing to Salesforce and Keeping Attachments Linked

Salesforce is the dominant CRM target for RFP pipelines and exposes a clean REST API. The pattern is straightforward: upsert an Opportunity record keyed by some idempotent identifier (the email Message-ID is a reasonable choice), then create ContentVersion records for each attachment and link them via ContentDocumentLink.

The opportunity upsert

Use the Salesforce Upsert action keyed by an external ID field on Opportunity. Create a custom field like RFP_Message_Id__c with external-ID + unique constraints. The upsert pattern means re-runs of the workflow on the same email do not create duplicates. The fields to populate from the extracted spec:

  • Name: "{issuer.company_name} — {products_in_scope[0]}" (sales-friendly opportunity name)
  • Account: lookup by issuer domain via the Account.Website matcher; if no match, create a new Account first
  • Amount: midpoint of opportunity.estimated_value_usd_range if both bounds are present
  • CloseDate: opportunity.decision_date if present; otherwise default to deadline + 30 days
  • StageName: "RFP Received" (a custom stage your sales ops team has wired into the funnel)
  • RFP_Message_Id__c: the email Message-ID
  • Description: a 200-word summary that the LLM generates as a separate field in the spec, formatted for hover-card display

Attachment linking

Each attachment becomes a ContentVersion (the file content) plus a ContentDocumentLink (the join to the Opportunity). The two-step pattern:

  1. POST to /services/data/v60.0/sobjects/ContentVersion with the file bytes (base64-encoded for the JSON path, or multipart for direct upload). Salesforce returns the ContentVersion Id; the ContentDocumentId is on the response.
  2. POST to /services/data/v60.0/sobjects/ContentDocumentLink with the ContentDocumentId, the LinkedEntityId (the Opportunity Id), and ShareType "V" (Viewer) for read-only attachment behavior.

The gotcha: Salesforce caps a single ContentVersion upload at 2GB but the REST API requires the file fit in a single request, and the practical limit for reliability is around 100MB. For larger files use the multi-part upload pattern (POST a ContentVersion with no body, then PATCH the VersionData in chunks). Most workflow tools have an "upload to Salesforce file" action that handles this.

The Slack notification

The final step posts to a #rfp-intake Slack channel with a card containing: issuer company, estimated value range, deadline, top 5 evaluation criteria, a button "Claim this RFP," and a link to the created Salesforce Opportunity. Slack interactivity (Block Kit buttons) means a sales engineer can claim the RFP from Slack and the workflow back-fills the Owner field on the Opportunity record. Round-trip time from email arrival to Slack notification: 60-90 seconds in our production deployments.

Confidence Routing: The Human-in-the-Loop Gate

The single most important architectural decision in an RFP pipeline is when to involve humans. Too eager and your humans become a bottleneck; too lazy and bad data lands in Salesforce. The pattern we have shipped successfully is the confidence-threshold router, a cousin to the pattern covered in the previous chapter's lesson on LLM-output routing.

The three-bucket gate

  1. Confidence ≥ 0.85 and missing_fields list is empty: Write to Salesforce directly. Post the Slack card as a confirmation, not a request for review.
  2. Confidence 0.60–0.85 OR missing_fields list non-empty: Write a draft Opportunity in stage "RFP Review Needed" with a clear flag. Post to Slack as a review request with the missing fields listed.
  3. Confidence < 0.60: Do not write to Salesforce. Drop into a manual-triage queue with the original email link, the extracted Markdown, and the LLM's stated reason for low confidence.

In practice on a 200-RFP-per-quarter pipeline, around 70% of RFPs sail through bucket one, 25% land in bucket two, and 5% fall to bucket three. The bucket-two reviews take a sales engineer 90 seconds each. The bucket-three triage takes 5-15 minutes. Total human time spent on the inbound pipeline: about 90 minutes per quarter. Before the pipeline: about 6 hours per week.

Why the empty-missing-fields check matters

A model can be 0.92 confident in an extraction that is missing the deadline. Confidence alone is not enough. The presence of any value in missing_fields means at least one field the model knows about is unanswerable from the document — usually because a section was OCR-garbled or the issuer left a placeholder. Treat the missing-fields list as a separate hard gate, independent of the confidence score.

The Failure-Mode Runbook

What kills these pipelines in week three is not the happy path. It is the long tail of weird inputs. Here is the runbook every operator-builder should write before shipping, with concrete responses.

Failure: Gmail trigger silently stops firing

Symptom: no new RFPs arrive in Salesforce for 24+ hours and the workflow run history is empty. Cause: Gmail OAuth token expired, or the search query returned 0 hits and the polling history shows successful empty runs. Fix: implement a "heartbeat" workflow that posts to Slack every 6 hours saying "RFP pipeline checked Gmail at {{time}}, found {{count}} hits." Silent absence becomes obvious silence.

Failure: extraction succeeds but the LLM returns an empty spec

Symptom: Opportunity created in Salesforce with name "TBD" and no fields populated. Cause: the LLM was given OCR-blank text and dutifully extracted nothing. Fix: gate the LLM call on a minimum character count from the extractor (we use 500 characters as a floor). If extraction returned less, escalate to the vision-model path.

Failure: duplicate Opportunities created on resends

Symptom: vendor resends the RFP with "Re: " prefix and a new Message-ID; workflow creates a second Opportunity. Cause: Message-ID-based upsert key does not match. Fix: secondary deduplication by attachment SHA-256 hash within a 30-day window. If 80%+ of attachment hashes match an existing Opportunity, link to that Opportunity instead of creating a new one.

Failure: Salesforce write fails with a validation rule error

Symptom: workflow run shows a 400 response from Salesforce with "FIELD_CUSTOM_VALIDATION_EXCEPTION." Cause: your Salesforce admin added a validation rule that requires "Lead Source" to be set on Opportunities and your workflow does not populate it. Fix: catch the error, log the validation message, default-fill the offending field, retry once. Long-term fix: ask the Salesforce admin to alert you before deploying validation rules that touch the Opportunity object.

Failure: confidential RFP gets posted to a too-public Slack channel

Symptom: NDA-bound RFP details appear in #general because someone misconfigured the workflow notification target. Cause: hard-coded channel name without environment awareness. Fix: route via a config object, gate on a "confidential" boolean in the spec (the LLM can flag this from "Confidential" markings in the document), and post sensitive RFPs to a private #rfp-intake-confidential channel with restricted membership.

Real Numbers from Three Deployments

Deployment A: enterprise SaaS, 60 RFPs/quarter

Pre-pipeline: 6 hours/week of inbound triage, average 14 hours wall-clock from email to Salesforce record, two RFPs missed entirely in Q4 2025 due to inbox overflow. Post-pipeline (live since January 2026): 90 minutes/quarter of triage time, 75-second average wall-clock from email to Salesforce, zero missed RFPs. Cost: $42/month in LLM spend, $0.71 per RFP processed.

Deployment B: cybersecurity vendor, 110 RFPs/quarter

Higher volume drove cost up but also drove value capture up. Pre-pipeline: a dedicated half-FTE sales operations role for triage. Post-pipeline: that role redeployed to RFP response drafting (a higher-leverage task). Cost: $103/month in LLM spend, $0.94 per RFP. Annualized labor savings: ~$60k. ROI: visible in week three.

Deployment C: management consulting firm, 35 RFPs/quarter

Lower volume but extremely high-value RFPs (six- and seven-figure engagements). Pre-pipeline: partners read every RFP themselves, delays of 3-5 days were common, prioritization was inconsistent. Post-pipeline: partners get a Slack card within 90 seconds with the engagement-value range and recommended next action. Most striking metric: time-from-RFP-to-partner-decision dropped from 4.1 days median to 6 hours median, and the firm's RFP-win rate moved from 18% to 27% in the first six months. Cost: $28/month in LLM spend.

What to Build Next on Top of This Pipeline

Once the RFP intake pipeline is stable, three natural extensions emerge:

  • Auto-draft response shells. The same extracted spec drives a second workflow that pulls relevant content from a content-library (case studies, security questionnaires, reference architectures) and assembles a response shell. The sales engineer edits a draft instead of writing from scratch. Time-to-response often drops from days to hours.
  • Competitor and incumbent intelligence. The extracted incumbent_vendor field triggers a separate enrichment step that pulls public information about the incumbent and surfaces it on the Slack card. Closing the gap from "I know there's an incumbent" to "I know who the incumbent is and what they probably charge" changes how sales engineers prepare.
  • Win/loss feedback loop. When the Opportunity is marked Closed Won or Closed Lost in Salesforce, a follow-up workflow pulls the original extracted spec and the actual outcome into a learning database. Over 12+ months, this becomes the training data for a "should we bid?" scoring model that runs on every new RFP.

None of these extensions require new primitives. They reuse the trigger, the extraction, the structured-output coercion, and the Salesforce write. That reuse is why building the RFP pipeline first pays back across the whole document-intake portfolio.

Key Takeaways

  • Build for value, not volume. A single missed RFP costs $40k-$400k in pipeline; a missed support ticket costs $50 in operational drag. RFP triage is the highest-economic-value inbound-document workflow most operators can ship.
  • Six steps end-to-end: Gmail trigger → defensive attachment download → extraction (LlamaParse/Unstructured/Reducto/native) → LLM structured-output extraction with explicit confidence → Salesforce Opportunity upsert plus ContentDocument link → Slack human-in-the-loop card. Target wall-clock: 90 seconds.
  • Orphan attachments are the silent killer. Enumerate before downloading, hash for dedup, store in S3/GCS before passing to the extractor. A team lost a $180k RFP because the 240MB technical addendum failed to download silently.
  • Three PDF extraction failure modes: blank-page scans, embedded image-of-text pages, table-shaped chaos. Defense: per-page character-count thresholds plus vision-LLM escalation for low-character pages. Cost: $0.01-$0.04 per escalated page.
  • The structured-output spec must include an explicit confidence object with overall and missing_fields. The model is instructed to set null and list any field it could not derive from the document — never to invent.
  • Confidence-routing gate: ≥0.85 with empty missing_fields writes direct to Salesforce; 0.60-0.85 or any missing fields routes to Slack review; below 0.60 routes to manual triage. Typical distribution: 70%/25%/5%.
  • Salesforce write pattern: Upsert Opportunity keyed on a custom external-ID field set to the email Message-ID, then create ContentVersion plus ContentDocumentLink for each attachment.
  • Heartbeat workflow every 6 hours catches the silent-trigger-failure mode. Silent absence becomes obvious silence.
  • Real deployments: $0.71-$0.94 per RFP processed, 90-second median wall-clock, 70-90 minutes of human triage per quarter replacing 6 hours/week. Win-rate lifts of 9 percentage points observed in consulting firms moving from days-to-decision to hours-to-decision.
  • Build on top of the pipeline: auto-draft response shells, incumbent/competitor enrichment, win/loss feedback loop. All reuse the same primitives.