←
AI Agent Builders & Citizen Developers
Capable · M2 · lesson 2 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Lead Enrichment with Clay or Apollo Plus an LLM
📖
now learning

AI Lead Enrichment with Clay or Apollo Plus an LLM

15 min

A RevOps lead opens HubSpot on a Tuesday morning. There are 847 leads from last week's webinar sitting in "New" status. Each one is a row: first name, last name, work email, "How did you hear about us?" Nothing else. No company size. No industry. No ICP fit signal. The sales team will look at the list, pick the ten with familiar logos, and let the rest age until the auto-nurture sequence catches them. This lesson is how to wire Clay or Apollo plus an LLM so all 847 get a sales-ready summary written into the right CRM field within five minutes of form submission — with the LLM's confidence logged so you know which ones to trust.

What Modern Enrichment Actually Looks Like in May 2026

In 2022, "enrichment" meant Clearbit. You pointed at an email, you got back a JSON blob with company size, industry, and a logo URL. You stuffed it in the contact record. Done. That world is gone. The 2026 equivalent is a 10-15 step enrichment chain in Clay or a multi-source waterfall in Apollo, each step pulling from a different source — LinkedIn for employee count, BuiltWith for tech stack, the company website for the "About" page, Crunchbase for funding history — and an LLM at the end that reads everything and writes the four sentences a sales rep actually wants to read before the discovery call.

The unit of work is not "look up the company." It is "build a sales-ready brief that takes 30 seconds to read and answers: who are they, why might they care about us, what's the timing signal, and how confident are we." Clay's spreadsheet-like interface and the Sculptor natural-language column builder make this composable — you describe what you want a column to contain in English and Clay writes the steps. Apollo's Plays interface plus their Apollo AI scoring does the same thing through a different UX. Both produce, at the end, the same artifact: a populated CRM record with a paragraph summary, a fit score, and ideally a "next best action."

The four things the brief has to answer

If your enrichment output cannot answer all four of these in two sentences each, the SDR will not read it, and you will have spent your Clay credits to no purpose. Print this list on a sticky note and tape it to the side of your monitor:

  1. Who is this company? Industry, size, geography, stage. One sentence.
  2. Why might they care about us? Pain signal — recent hire, tech-stack change, funding round, hiring spree, expansion. One sentence with a date.
  3. How is this contact relevant? Title, seniority, function. Buyer/champion/influencer/blocker. One sentence.
  4. Confidence? A number 0-100 and what's driving the uncertainty. One sentence.

The confidence number is the operator's discipline. It is the single field that converts "AI did some enrichment" into "AI did some enrichment and we know how much to trust it." Without confidence, the SDR has to re-verify every claim. With confidence, they spend their re-verification budget on the 12 leads that scored below 60.

The Clay Pipeline: From Form to HubSpot in Five Minutes

Let's build the actual pipeline. The trigger is a webhook from HubSpot when a new lead is created with source = "Webinar Registration." The destination is the same HubSpot record, with three fields populated: ai_lead_summary, ai_fit_score, ai_confidence. The middle is a Clay table with 13 columns that fire in sequence.

Step 1: The trigger and the input row

HubSpot's webhook is configured under Settings → Integrations → Webhooks, scoped to the "Contacts created" event with a filter on the lead source property. Clay receives the payload at a table-specific webhook URL (each Clay table has one under "Sources → Add source → Webhook"). The first row in Clay is now populated with: email, first_name, last_name, company_name (if the form captured it), hubspot_contact_id. The hubspot_contact_id matters — you'll need it to write back.

Step 2-3: Find the company domain and enrich the company

Column 2 is "Find company domain from email." This is a built-in Clay column that strips the email domain and validates it. Free emails (gmail, yahoo) get flagged for a fallback path — usually "look up the company name they typed in the form." Column 3 is the company enrichment waterfall: Clay tries People Data Labs, then Apollo, then ZoomInfo, then a LinkedIn scrape, until one returns data. You get back employee count, industry, headquarters, founding year, LinkedIn URL, description.

The waterfall is the single most valuable feature in Clay. One source is wrong 30% of the time. Three sources stacked are wrong less than 4% of the time. The cost is more credits per row, but a $0.40 row that ships to a $30,000 ACV deal is the best ROI in the building.

Step 4-5: Person enrichment and seniority classification

Column 4 is "Enrich person from email" — Clay's people waterfall returns LinkedIn URL, current title, current company, tenure, prior companies, location. Column 5 is an LLM call (the "OpenAI" or "Claude" column in Clay) that takes the title string and returns a normalized seniority — IC, Manager, Director, VP, C-level. This matters because titles are messy ("Senior Manager of Strategic Operations" vs. "Strategic Ops Lead" vs. "Head of Strategic Operations") and your downstream routing rule needs a clean enum.

The prompt for column 5 is two sentences: "You are normalizing a job title. Output ONLY one of: IC, Manager, Director, VP, C-level. Title: {{title}}." Short, instructed, no preamble. Clay's column will return the raw text and a follow-up column will validate it's one of the five values before passing it on.

Step 6-8: Timing signals — funding, hires, tech stack

Column 6 looks up recent funding rounds (Crunchbase or the Clay funding-database integration). Column 7 hits a "recent job posts" data source for the company's open roles — a company hiring 5 SDRs is in a very different posture than one hiring 5 software engineers. Column 8 is BuiltWith for tech stack — particularly useful if your product integrates with specific tools.

These three are the "why now" signals. A company is always there. A company that just raised Series B and is hiring sales leadership is there and ready. The brief has to surface that.

Step 9: News and recent activity (the LLM-driven web search)

Column 9 is a Sculptor-built column that hits Claude with web search enabled. The prompt: "Search for news, announcements, or press about {{company_name}} from the last 90 days. Return up to three items as a JSON array with fields: date, headline, source, summary. If nothing relevant, return []." This is the "what changed lately" surface, and it's the one most likely to surface the actual hook for the SDR's opener.

Step 10: The fit-scoring LLM call

Now everything goes into a Claude Sonnet 4.5 column with a structured prompt. The system prompt describes your ICP — for example: "Our ICP is B2B SaaS companies, 50-500 employees, North America or Western Europe, with a sales team of 10+, that use Salesforce or HubSpot. Bonus signals: recent Series A/B/C, hiring sales ops, currently using a competitor tool we displace."

The user prompt assembles the row data: {{company_enrichment}}, {{person_enrichment}}, {{funding_data}}, {{job_posts}}, {{tech_stack}}, {{news}}. The output schema is JSON:

{"fit_score": 0-100, "tier": "A|B|C|D", "rationale": "...", "confidence": 0-100, "confidence_drivers": "..."}

The confidence field is the critical one. The prompt tells the model: "Confidence reflects how much data you had to make this judgment. If company_enrichment or news returned empty, confidence should be below 60. If three or more of the data sources were rich, confidence can be above 80. Never above 95."

Step 11: The brief-writing LLM call

Column 11 is the final LLM call. It takes everything plus the fit-score JSON from column 10 and writes the four-sentence brief the SDR will actually read. The prompt is explicit about format:

You are writing a sales-ready brief for an SDR. Output exactly four sentences in this order: (1) Who is the company. (2) Why they might care about us (cite a specific signal with a date). (3) Who the contact is and their likely buying role. (4) Confidence and what's driving it. Keep it under 80 words. No marketing language. No hedging filler ("interestingly," "notably"). Plain prose.

The output goes into ai_lead_summary. Score goes into ai_fit_score. Confidence goes into ai_confidence.

Step 12-13: Write back to HubSpot and log the LLM confidence

Column 12 is "Write to HubSpot" — Clay's native HubSpot connector takes the hubspot_contact_id and updates the three properties. Column 13 logs the row to a Google Sheets "enrichment audit log" with timestamp, row hash, cost, latency, confidence. That log is the artifact you use a month from now when someone asks "how is the enrichment performing?"

The Apollo Version of the Same Pipeline

If your shop is on Apollo instead of Clay, the moving parts are similar but the surface is different. Apollo's Plays builder lets you chain enrichment, but the LLM steps are leaner — you have Apollo AI for scoring and writing, plus the ability to call out to OpenAI or Claude via webhook for custom prompts.

The Apollo waterfall sources person data from their database (over 275M contacts as of May 2026), plus email verification, plus their LinkedIn integration. For B2B operators on the Apollo stack, the simpler version is:

  1. HubSpot webhook fires Apollo's "Enrich Contact" endpoint.
  2. Apollo's native enrichment fills firmographic and person fields.
  3. Apollo's "Engagement Scorer" runs against your ICP (configured once in Apollo settings).
  4. An HTTP webhook from Apollo's Play sends the enriched payload to Claude with the brief-writing prompt.
  5. The brief lands back in Apollo and syncs to HubSpot via Apollo's native HubSpot integration.

Apollo gives you fewer custom steps than Clay but a smoother native CRM experience. Clay wins when you need 10-15 custom enrichment sources and granular control over each column. Apollo wins when you already have Apollo as your prospecting database and want enrichment to "just work" with one less vendor.

Cost arithmetic for both stacks

Clay costs vary by table size and credit consumption. As of May 2026, a 13-step enrichment per row consumes roughly 8-15 Clay credits depending on the waterfall hits. At Clay's Pro plan ($349/month for 10,000 credits, May 2026 pricing), that's ~$0.35-$0.50 per fully-enriched row. Apollo's Organization plan ($149/user/month for enrichment + Apollo AI) includes enrichment credits up to a soft cap, plus you pay LLM tokens separately if you're piping to OpenAI/Anthropic.

For 847 weekly leads, Clay sits around $300-$420/month in credits. Apollo sits around $0 incremental (you pay seats either way) plus ~$15-30/month in Claude tokens. The differential is real if you're cost-conscious, but the operator-time math usually favors Clay for sophisticated pipelines and Apollo for simple ones.

Writing to HubSpot and Salesforce: The Fields That Matter

The brief is useless if it lands in the wrong field. Three fields earn their keep:

HubSpot setup

Create three custom contact properties under Settings → Properties → Contact:

  • ai_lead_summary — Long text, shown in the right-hand sidebar of the contact record. Place high in the sidebar (drag up under "About") so it's visible without scrolling.
  • ai_fit_score — Number, shown as a badge. Use a workflow to set lifecycle_stage based on score thresholds (e.g., 80+ → MQL, 60-79 → Lead Nurture, below 60 → Marketing).
  • ai_confidence — Number, shown next to the summary. Build a saved view "Low Confidence Leads" filtered to ai_confidence < 60 for the data team to spot-check weekly.

The HubSpot UI in 2026 supports inline rendering of the summary in the contact card. Make sure the property is set to "Show on contact record" and ordered high. SDRs read what is visible. They do not scroll.

Salesforce setup

For Salesforce-based teams, the same three fields go on Lead (and on Contact, if you convert): AI_Lead_Summary__c (Long Text Area, 1,000 char), AI_Fit_Score__c (Number), AI_Confidence__c (Number). Add them to the page layout under a "AI Enrichment" section at the top of the layout. Build a Lightning report "Low Confidence Leads" with filter AI_Confidence__c < 60 for the QA pass.

One Salesforce-specific note: if you have lead-to-account matching enabled, populate the matched account ID in a hidden AI_Matched_Account__c field for the routing rule downstream. The Apollo and Clay integrations both expose the matched account ID. Don't lose it.

The logging table you'll thank yourself for

Whether you're on Clay or Apollo, log every enrichment to a separate table (Google Sheets, Airtable, or Supabase). One row per enrichment with: timestamp, lead_id, source_system, fit_score, confidence, llm_cost_cents, llm_latency_ms, source_count (how many waterfalls hit), and a free-text "brief_preview" with the first 200 characters of the summary.

A month from now, when the SDR director says "the AI scoring isn't working," this table is what you query. You can plot confidence distribution. You can correlate confidence with downstream conversion. You can find the rows where the model was confident but the lead was bad — those are your fine-tuning examples for the next iteration.

Logging the LLM Confidence and Why It Matters

Confidence is the lead operator's superpower. It transforms an AI enrichment pipeline from "magic black box" to "instrument with calibration."

The five-sentence guidance on confidence

The instruction in your prompt has to be specific enough that the model doesn't just write "95" every time. The prompt template that produces well-calibrated confidence in our testing as of May 2026:

Output a confidence score 0-100. Treat 50 as "moderate evidence, some inference required." Score above 80 only when at least three data sources independently confirm the key claims. Score below 40 when more than one data source was empty or contradicted another. Add a one-sentence confidence_drivers that names the strongest signal you had AND the strongest absence. Never output above 95. The model is uncertain about edge cases; reserve 95+ for cases where you'd bet $1,000.

The "never above 95" rule is the difference between an LLM that hedges into the 70s and an LLM that anchors everything at 99. The first is useful. The second is decoration.

What confidence drift looks like

If your weekly mean confidence drifts from 72 to 88 over two months, something has changed. Either your input data got cleaner (waterfall now hits more sources), or the model got more confident (a prompt change, a model upgrade), or your audit isn't being run. A drift up of more than 8 points in 30 days warrants a re-eval. A drift down might mean a data source went stale.

Build a Sunday-night automation: pull last week's confidence distribution from your audit log, render a histogram, post it to the #revops Slack channel. Five minutes of operator-time on a Sunday saves an "AI enrichment seems off" panic on a Thursday.

The audit ratio that earns finance trust

For every 100 enrichments, sample 10 and have a human grade them: "Was the summary accurate? Was the score reasonable? Was the confidence reasonable given the data?" Track the agreement rate weekly. When you can tell finance "the AI agrees with a human reviewer 91% of the time on score and 87% on confidence," you've converted a guess into an audited process. That's the artifact that lets you defend a per-lead enrichment cost of $0.40 to a CFO who wants to know what they're getting.

The Failure Mode and How to Catch It

The most common failure of this entire pipeline is silent. The waterfall returns empty on a free-email lead (someone from gmail with no listed company). The LLM still writes a brief, because the model is helpful by nature. The brief is generic, plausible, and totally fabricated.

The defense: explicit "I don't have enough" outputs

Modify your brief-writing prompt with a hard rule: "If the company_enrichment field is empty AND no other data source provided a company name, output exactly: {\"summary\": \"INSUFFICIENT_DATA\", \"fit_score\": null, \"confidence\": 0}. Do not invent a company. Do not infer from the email domain alone unless the domain matches a known corporate domain in the input data."

Then in HubSpot/Salesforce, build a workflow that detects ai_lead_summary = "INSUFFICIENT_DATA" and routes those leads to a "needs human review" queue instead of the normal SDR flow. You will catch 60-200 of these a month depending on your form discipline, and the SDRs will thank you for not feeding them garbage.

The 100-lead validation test

Before you ship the pipeline to production, run it on 100 leads from last month's pile. Spot-check ten. Are the summaries accurate? Are the scores reasonable? Are the confidence numbers spread out (not all clustered at 90)? If 70 of 100 score above 80, you've miscalibrated the rubric — you're scoring on availability, not on fit. If 90 of 100 say confidence 95, your prompt is broken. Iterate. Don't ship the broken version because the demo looks pretty.

One team in March 2026 ran this test, found that 78 of 100 leads were scored A-tier, audited the input data, and discovered their ICP definition in the system prompt was too broad. They tightened the definition to require employee count above 50 (not above 10), reran on the same 100 leads, and got a 22-A/41-B/29-C/8-D distribution. That's a usable distribution. The first one was not.

The Five-Minute Rule

The pipeline has to complete in under five minutes from form submission, or the SDR sees the lead before it's enriched and the whole point evaporates. Optimize the waterfall to fail fast: if a data source hasn't returned in 8 seconds, move on. Run the three independent enrichment steps in parallel (Clay supports this; Apollo's Plays do too). Use Claude Sonnet 4.5 with streaming for the brief-writing step — the first sentence appears before the fourth is generated.

A well-tuned 13-step Clay pipeline in May 2026 ships brief-to-HubSpot in 90 seconds median, 4 minutes p95. The slow ones are the ones with a slow LinkedIn step or an LLM call that's hitting Anthropic's overflow region. Both are tolerable. Both stay under five minutes if you don't add unnecessary serial steps.

Key Takeaways

  • Modern enrichment in May 2026 is a 10-15 step Clay chain or a multi-source Apollo waterfall, ending with an LLM that writes a four-sentence sales-ready brief into HubSpot or Salesforce.
  • The brief must answer four questions: who is the company, why might they care now, who is the contact, and what's the confidence. Print the list on a sticky note.
  • Clay's Sculptor natural-language column builder lets you describe a step in English and have Clay write the pipeline; Apollo's Plays do the same through a different UX. Both produce the same artifact.
  • Use a waterfall of three data sources stacked: one source is wrong 30% of the time, three stacked are wrong less than 4% of the time. The credit cost is worth it on any deal above $5K ACV.
  • Three fields in your CRM: ai_lead_summary, ai_fit_score, ai_confidence. Show them at the top of the contact record. SDRs read what's visible.
  • Log every enrichment to a separate audit table with cost, latency, confidence, source count. This table is your defense against "is the AI working?" questions.
  • The "never above 95" rule on confidence is what separates a calibrated LLM from a decorative one. Without ceiling instructions, models anchor at 99.
  • The most common failure is silent: empty waterfall → fabricated brief. Add an explicit INSUFFICIENT_DATA output path and route those to human review.
  • Run the 100-lead validation test before shipping. If 78% score A-tier, your rubric is wrong; tighten the ICP definition and re-run.
  • The five-minute SLA from form submission to enriched brief is the operational bar. Anything slower and the SDR sees the raw lead first, and the pipeline is decoration.