←
AI Agent Builders & Citizen Developers
Capable · M9 · lesson 9 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
LLM-Powered Lead Scoring vs. Rule-Based Scoring
📖
now learning

LLM-Powered Lead Scoring vs. Rule-Based Scoring

15 min

There is a meeting that happens once a quarter in every RevOps org. The Marketing director says "we should use AI to score leads." The Sales director says "our current scoring works." The CFO says "what does it cost and what will it return." And the operator in the room — that is you — has thirty seconds to either commit to a real comparison or commit to a year of dueling spreadsheets. This lesson is how to run that comparison: a 200-lead holdout set, an LLM scoring node, the existing rule-based system, head-to-head on the metric finance trusts. The version that wins ships. Not the trendy one.

Why "The Comparison," Not "The Replacement"

Most teams blow this decision in the first ten minutes. They read a vendor blog post about how LLM scoring "outperforms traditional rules by 31%," they shop a vendor demo, and they sign a contract. Six months later, the AI scoring system is rating lukewarm prospects as "hot" and the SDRs have learned to ignore the score entirely. The rule-based system, meanwhile, sits unused, even though it was working fine.

The operator's job is not to install the trendy thing. The operator's job is to ship whichever scoring system reduces the false-positive rate that wastes SDR time. Sometimes that's an LLM. Sometimes it's a tuned rule-based system. Sometimes it's a hybrid where the LLM scores qualitative signals (recent news, hiring intent, tone of the form-fill) and rules handle the firmographic gate (employee count, geography, industry blocklist). You don't know which until you run the comparison.

The metric that matters: false-positive rate

Lots of vendors quote "accuracy." Accuracy is a trap on imbalanced datasets, which all lead-scoring datasets are. If only 8% of your leads convert and your model predicts "no convert" for everything, you have 92% accuracy and zero useful predictions. Useless.

The metric finance trusts is the false-positive rate: of all the leads your system rated as "hot" (top tier, send to SDR), what percentage turned out to not be worth the SDR's time? Every false positive is roughly 25 minutes of an SDR's day wasted on a discovery call that goes nowhere. At a fully-loaded SDR cost of $90,000/year, that's about $5.40 per false-positive call. A scoring system that produces 200 false positives per month costs you $1,080/month in pure SDR time, plus the opportunity cost of the calls they would have made on real leads.

Track also the false-negative rate — leads you rated low but that converted anyway. False negatives are missed revenue. The relationship between the two is a precision/recall trade-off; you tune your threshold to balance them. But the headline metric you put in front of the CFO is false-positive rate, because it converts directly to dollars wasted.

The 200-Lead Holdout Set

The comparison rests on a labeled holdout set. Without labels, you have no truth; without truth, you cannot measure either system. The lesson here is how to assemble the holdout, label it honestly, and run both scoring systems against it.

Sourcing the 200 leads

Pull 200 leads from the past 90 days that have closed outcomes — either won, lost, disqualified at SQL, or aged out (90+ days untouched without escalation). The mix should approximate your real lead flow: if 25% of your real leads come from webinars, 25% of your holdout should be from webinars. Don't cherry-pick. The holdout is only useful if it mirrors the distribution you'll see in production.

Don't go below 200. The math is unforgiving — with 100 leads and an 8% conversion rate, you have 8 positives to learn from. Your confidence intervals on false-positive rate will be wide enough to drive a truck through. 200 with a 12% positive rate gives you 24 positives, which is the floor for a comparison that won't get torn apart in the next CFO meeting.

The label schema

Each lead in the holdout gets a label assigned by a human (you or a senior SDR) blind to the scoring outputs. Three values:

  • True positive (TP): This lead converted to opportunity, or is still in pipeline at MQL-quality after a discovery call. SDR time was well spent.
  • False positive (FP): This lead got an SDR discovery call and turned out to be poor fit, wrong stage, no budget, or unreachable. SDR time was wasted.
  • True negative (TN): This lead was not contacted (low score) and the outcome confirms that was correct — they aged out, unsubscribed, or were a clear non-fit.

The fourth case — false negative (lead was rated low but should have been high) — is harder to label retrospectively because if you didn't contact the lead, you don't know what would have happened. The proxy: leads that scored low but eventually self-converted (came back in via demo request, or were sourced again by an SDR cold and converted). Count those as FNs. The count will be small; that's a feature, not a bug.

The labeling discipline

Two people label independently. Compare. Where you disagree, discuss and resolve. Lock the labels. Do not adjust them after seeing scores. The holdout exists because human judgment is the ground truth; if you let scores change your labels, you have a feedback loop that produces high agreement and zero insight.

The Rule-Based Baseline

Before you run the LLM, write down what your current rule-based system does. If it's HubSpot's native lead scoring, export the rules. If it's a custom system, document the formula. Most rule-based systems in 2026 are a weighted sum:

Score = (Company Size Score × 0.3) + (Industry Score × 0.2) + (Title Seniority Score × 0.2) + (Form Source Score × 0.15) + (Recent Activity Score × 0.15)

Threshold for "hot" (send to SDR): 70+

Run this rule set against all 200 holdout leads. Record the score for each. Determine the binary classification at threshold 70: "hot" or "not hot." Now you have a baseline confusion matrix:

  • Rule score 70+ AND TP: True positive of the rule system
  • Rule score 70+ AND FP: False positive of the rule system
  • Rule score below 70 AND TP/FN: False negative
  • Rule score below 70 AND TN: True negative

From that matrix you compute the false-positive rate for the rule system. Let's call it FP_rules. In a typical mid-market RevOps setup, this number sits at 35-50%. That means roughly four in ten leads the SDRs are told are "hot" turn out to be cold once they pick up the phone.

Building the LLM Scoring Node

The LLM scoring node lives wherever your workflow lives — n8n, Make, Zapier, Clay, or directly in your CRM via a custom integration. The mechanics are similar across all of them.

The system prompt

The system prompt encodes your ICP, your scoring rubric, and your output schema. It is the operator's job to write this well. A typical structure:

You are a B2B lead scorer for [Company Name], a [product description]. Our ICP is [employee count range], [industry list], [geography], with [tech stack signals]. Score leads on a 0-100 scale based on:
1. ICP fit (40%): employee count, industry, geography, stage
2. Buying signals (30%): recent funding, hiring, expansion, tech stack changes in last 90 days
3. Contact relevance (20%): seniority, function, decision-making authority
4. Engagement signal (10%): form source, page visits, content downloaded

Output JSON: {"score": 0-100, "tier": "A|B|C|D", "rationale": "two sentences", "confidence": 0-100}

Confidence must reflect data availability. Never above 95. If two or more input fields are empty, confidence cannot exceed 60.

The user prompt is the lead data — every enriched field from the prior lesson's enrichment pipeline. The model returns the JSON.

The model choice

For this scoring task, the right starting point is Claude Sonnet 4.5 ($3/M input, $15/M output) or GPT-5 ($5/M input, $20/M output). Haiku 4.5 ($0.25/M input) is tempting for cost but produces less reliable rationales — fine for the score itself, weaker for the explanation. Since rationale is what the SDR reads, the bigger model earns its cost.

Per-lead cost at Sonnet 4.5: roughly 2,000 input tokens (system prompt + enriched lead data) + 200 output tokens. That's $0.006 input + $0.003 output = ~$0.009 per scoring call. At 8,000 leads/month, that's $72 in LLM costs for scoring. Trivial against the SDR time saved if the false-positive rate moves favorably.

Run it on the 200 holdout leads

Run the LLM scoring against the same 200-lead holdout. Capture: score, tier, rationale, confidence per lead. Apply the same threshold logic ("hot" = score 70+) and compute the LLM's confusion matrix. Call its false-positive rate FP_llm.

The Head-to-Head Arithmetic

Here's where most teams skip the actual numbers and end up regretting it. Do the math.

A worked example: a mid-market SaaS RevOps run

In April 2026, a RevOps team at a mid-market sales-tech company ran this comparison on 200 holdout leads with a 12% positive rate (24 TPs in the set). Results:

Rule-based system at threshold 70:

  • Predicted "hot": 56 leads
  • Of those, 14 were true positives, 42 were false positives
  • False-positive rate: 42 / 56 = 75%
  • True-positive recall (of all 24 real positives, how many caught): 14 / 24 = 58%

LLM scoring (Claude Sonnet 4.5) at threshold 70:

  • Predicted "hot": 38 leads
  • Of those, 19 were true positives, 19 were false positives
  • False-positive rate: 19 / 38 = 50%
  • True-positive recall: 19 / 24 = 79%

The LLM caught more true positives and had a lower false-positive rate. In this specific case, the LLM was the clear winner. The recommendation went to leadership with the spreadsheet attached: "200-lead holdout, blind-labeled, LLM scoring reduces false positives from 75% to 50% and catches 21% more real opportunities."

The CFO converted to dollars. At 8,000 monthly leads with this performance, the rule system surfaces ~2,240 SDR-eligible leads of which ~1,680 are false positives (75%). The LLM system surfaces ~1,520 SDR-eligible leads of which ~760 are false positives (50%). At 25 minutes per false positive and $5.40 per call cost, that's $9,072/month for rules vs $4,104/month for LLM. Savings: roughly $5,000/month. LLM cost: $72/month. Net: ship it.

The case where rules win

That's not always the result. In a different March 2026 comparison at a mature enterprise sales team where the rule system had been tuned for three years and the ICP was extremely narrow (only one industry, only one geography, only public companies above $500M revenue), the rule system actually beat a generic LLM scoring setup. The rules had 38% false-positive rate; the out-of-the-box LLM scoring had 47%. Why? The rules were doing a hard firmographic gate ("must be public company in North America with revenue above $500M"). The LLM was reading the firmographics correctly but also scoring on weak qualitative signals that introduced noise.

The fix in that case was a hybrid: rules handle the firmographic gate, LLM scores the qualitative signals only when the firmographic gate passes. The hybrid got the false-positive rate down to 28%. Best of both. The operator who ran the comparison didn't know that was the answer when they started; they discovered it by running both head-to-head, seeing the result, and asking why.

Thresholds and Tuning

The threshold of 70 is arbitrary. The right threshold for your business is a decision, not a fact. Tune it.

The ROC-style sweep

For each scoring system (rules and LLM), compute the false-positive and false-negative rate at thresholds 50, 60, 65, 70, 75, 80, 85. Plot them. Pick the threshold where the false-positive rate is at or below your tolerance and the false-negative rate isn't catastrophic.

A typical tolerance for B2B SaaS: false-positive rate under 40% (SDRs accept this), false-negative rate under 25% (sales VP accepts this). Inside those bounds, you have flexibility. Outside them, you're either flooding SDRs with junk or missing too much revenue.

Two-tier thresholding (the operator's secret)

Instead of a single threshold, use two. Anything 80+ goes auto-routed to SDR. Anything below 50 goes to nurture only. The 50-80 band — the uncertain middle — goes to a daily review queue where a senior SDR or the marketing ops team spot-checks them, escalates the obvious good ones, downgrades the obvious bad ones. The queue is small (maybe 30-40 leads/day at mid-market scale), but the leverage is enormous because the uncertain middle is where most of the cheaply-correctable false positives live.

Confidence as a Routing Input

The LLM's confidence field (the one you forced into the output schema) is not decoration. It is the second axis of routing.

The confidence/score matrix

Cross score against confidence:

  • High score + high confidence (≥80 score, ≥80 confidence): Auto-route to SDR. The model says hot and is sure.
  • High score + low confidence (≥80 score, <60 confidence): Manual review before routing. The model says hot but admits uncertainty — usually because the enrichment data was sparse. Cheap to verify.
  • Low score + high confidence (<50 score, ≥80 confidence): Route to nurture. The model is confident this isn't a fit. Trust it and save the SDR cycles.
  • Low score + low confidence (<50 score, <60 confidence): Re-enrich first. The model couldn't see enough to judge. Trigger a re-run of the enrichment pipeline with deeper sources before any routing decision.

This matrix is the difference between "AI tells me what to do" and "AI tells me what it knows and how sure it is, and I route accordingly." The first one collapses on edge cases. The second one degrades gracefully.

The Monthly Re-Run and the Drift Question

Ship the winning version. Then run the comparison again on a fresh 200-lead holdout 30 days later. And again 30 days after that. This is the discipline that separates the operator who installed a tool from the operator who runs a system.

What you're watching for

Three drifts matter:

  1. Input distribution drift: Your lead source mix changes (a new ad campaign brings in different prospects), your form changes (a new field gets added), your enrichment sources change (a vendor's data quality degrades). Your scoring model — rule or LLM — was tuned for the old distribution. The new one needs re-evaluation.
  2. Model drift: For the LLM, Anthropic or OpenAI ships a new model version. You upgrade because you should. The new model scores differently. Re-run the comparison.
  3. Outcome drift: Your conversion rate at the funnel shifts. A 12% positive rate becomes 8%. The threshold of 70 was tuned to 12%. At 8%, the threshold should probably shift up.

The 30-minute monthly cost review from Lesson 2.2.4 has a sister discipline: the 60-minute monthly scoring review. Pull a fresh 50-lead sample, label, run both systems, compute current false-positive rates. Plot them next to last month's. Drift catches itself if you look.

What Finance Wants on the One-Pager

The one-pager you bring to the CFO when you ship the winning system should be three lines:

  1. "On a 200-lead blind-labeled holdout from the past 90 days, [system X] reduces false-positive rate from [Y%] to [Z%]."
  2. "This translates to [N] hours per month of SDR time recovered, valued at $[$$$$$]."
  3. "Monthly cost of [system X]: $[$$]. Net monthly benefit: $[$$$$$]. ROI: [X]x. Re-run quarterly."

That's what gets the approval. Not a vendor blog post. Not a demo screenshot. Three lines of arithmetic from your own data. Finance approves what finance can audit, and the only thing they can audit is your numbers from your leads.

The Five Mistakes That Tank This Comparison

Operators who run this badly hit the same five mistakes. Avoid them.

Mistake 1: Letting the LLM see the labels

If your scoring prompt has been "tuned" using the same leads you're going to test on, the LLM has essentially seen the answers. The result is bogus. Use a separate set for prompt iteration and reserve the 200-lead holdout for the final comparison.

Mistake 2: Re-labeling after seeing scores

You glance at the LLM output and think "actually that one is a yes, I'll change my label." Don't. The labels were set by human judgment for a reason. Adjusting them post-hoc is how teams talk themselves into the AI being better than it actually is.

Mistake 3: Comparing against an untuned rule baseline

If your rule system has been ignored for two years and the LLM is fresh, you're comparing the LLM to a strawman. Spend an afternoon tuning the rule system on a separate 100-lead set before the comparison. Give it a fair fight. If the LLM still wins, the result is real. If the tuned rules win, you've saved yourself $72/month in LLM costs and the operational overhead of an LLM dependency.

Mistake 4: Ignoring the recall side

A scoring system that says "no" to everyone has a perfect false-positive rate. It is also useless. Track recall (true-positive rate) alongside false-positive rate. Both have to be in the acceptable range.

Mistake 5: Skipping the confidence dimension

If your LLM scoring outputs only a score and no confidence, you're throwing away half the information. Force the schema. Use confidence in routing. Re-run when confidence drops as a leading indicator.

Key Takeaways

  • The operator's job is to ship the scoring system that reduces false-positive rate. Sometimes that's an LLM, sometimes tuned rules, often a hybrid where rules gate the firmographics and the LLM scores qualitative signals.
  • Run a 200-lead blind-labeled holdout from the past 90 days. Two humans label independently. Lock the labels before running either scoring system.
  • Track false-positive rate as the primary metric — that's the one finance can convert to dollars (25 minutes × $5.40 per wasted SDR call). Track false-negative rate as the constraint.
  • A typical mid-market run shows the LLM winning: 50% false-positive rate vs. the rule system's 75%, with the LLM also catching 21% more true positives. But this is empirical, not theoretical — run YOUR comparison on YOUR data.
  • The case where rules win: extremely narrow ICP with a hard firmographic gate. The fix: hybrid scoring with rules for firmographics, LLM for qualitative.
  • Two-tier thresholding beats single threshold: 80+ auto-routes to SDR, below 50 goes to nurture, 50-80 lands in a small daily review queue.
  • Confidence is the second axis of routing. The confidence/score matrix tells you when to route, when to re-enrich, when to verify, when to skip.
  • Re-run the comparison monthly on a fresh holdout. Catch input drift, model drift, and outcome drift before they accumulate.
  • The CFO one-pager is three lines: holdout result, SDR time recovered in dollars, monthly cost vs. monthly benefit with ROI. Three lines of arithmetic from your own data is what gets approval.
  • Five mistakes to avoid: letting the LLM see labels, re-labeling after scoring, comparing against an untuned rule baseline, ignoring recall, skipping confidence. Avoid all five and the comparison earns its keep.