←
AI Agent Builders & Citizen Developers
Capable · M14 · lesson 14 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Routing on LLM Output: The 'Confidence Threshold' Pattern
📖
now learning

Routing on LLM Output: The 'Confidence Threshold' Pattern

15 min

An operator at a Series B SaaS company shipped an LLM-powered routing workflow on a Tuesday in February 2026. By Thursday, the human queue had 1,400 leads in it. Every single lead from two days of inbound had been routed to "uncertain — needs human review." The auto-assignment branch had fired zero times. The Slack message from the head of sales was three words long. This lesson is how to design the confidence-threshold pattern so it actually routes — high-confidence to auto-assignment, low-confidence to humans — and the failure mode that lands 100% in the wrong bucket.

What the Confidence-Threshold Pattern Actually Is

You have an LLM node in your workflow that produces a structured output — a score, a classification, a routing decision, plus a confidence number. You have two downstream paths: auto-assignment (the agent does something without a human in the loop) and human queue (a person reviews before action). The confidence-threshold pattern is the simplest possible routing rule: if confidence is above some threshold T, take the auto path; otherwise take the human path.

That's it. The entire pattern is one if statement. The interesting work is everywhere else — in writing a prompt that produces calibrated confidence, in choosing T, in documenting the override path, in measuring whether the routing is actually correct, and in catching the failure mode where the LLM produces unusable confidence numbers and everything ends up on one side of the branch.

Why two branches, not three or four

Could you have a "medium confidence" tier that gets a different treatment? Yes, and the previous lesson described it. For most operators shipping their first AI routing workflow, two branches is the right starting complexity. A two-branch workflow is auditable: every lead either auto-routed or went to a human. A three-branch workflow has more decision rules to maintain and more edge cases when confidence numbers cluster near a boundary. Start with two. Add the middle tier once you've measured where it would actually help.

The Failure Mode: The 100% "Uncertain" Disaster

The Tuesday-to-Thursday story above is the most common failure of this pattern, and it has a precise cause: the prompt did not ask the LLM for a numeric confidence score, so the LLM returned strings like "fairly confident" and "not sure" and "probably qualified." The downstream filter — which was looking for confidence >= 80 — couldn't parse any of those, defaulted to "low confidence," and routed everything to the human queue.

The fix is mechanical and immediate. Three changes to the prompt:

  1. Force the schema. The prompt must say: "Output JSON only. Schema: {\"decision\": \"...\", \"confidence\": INTEGER 0-100, \"rationale\": \"...\"}. The confidence field must be an integer. Never use words like 'high' or 'low.'"
  2. Anchor the scale. "Confidence 90+ means: enough data to bet $1,000 on this decision. Confidence 70-89 means: clear signal with one or two minor uncertainties. Confidence 50-69 means: workable inference, multiple gaps. Below 50: do not auto-route. Output 'INSUFFICIENT_DATA' if you cannot judge."
  3. Test the output format BEFORE you wire the routing. Run the prompt on 20 sample inputs. Look at the JSON. Confirm every output has an integer confidence. If even one has a string, fix the prompt.

The Series B team's fix in February 2026 was exactly this. They added explicit numeric instructions, ran 50 test leads, and the routing started working. Of 50 leads, 31 routed to auto-assignment, 14 to human review, 5 returned INSUFFICIENT_DATA. That distribution was usable. The original distribution — 100% to human review — was not.

The diagnostic question every operator should ask

If your routing workflow looks broken and 95%+ of items are going to one branch, ask: what does the LLM actually return for confidence? Open the execution log. Look at the raw output. If you see strings, the prompt is the problem. If you see all numbers near 99, the prompt has no ceiling rule. If you see all numbers near 50, the prompt is asking the model to score on a dimension it has no data for. Each diagnosis has a different fix; the symptom of "everything routes to one branch" maps to a specific data issue every time.

Choosing the Threshold T

The threshold is the dial. It is also the place where operators are most tempted to choose by feel rather than measurement.

The threshold by sample

Pull a sample of 50-100 outputs from your LLM node. Get the confidence distribution. Plot it. You'll see something like: a mode around 75-85, a tail down to 30, occasional 95s. Now overlay outcomes — for the ones you can verify (a sample of past data where you know the right answer), at each confidence bin, what percentage was the LLM correct?

If at confidence 80+ the LLM is correct 92% of the time, that's a defensible auto-route threshold for many use cases. If at confidence 60-79 it's correct 78% of the time, that's the band that goes to human review. Below 60, route to "needs more data" or to nurture. The actual threshold for your business depends on how costly a wrong auto-route is.

The cost-of-error tuning

The threshold should reflect: what does a wrong auto-route cost compared to a wrong human-review-queued? If a wrong auto-route fires off an email to the prospect (recoverable, slight reputation hit), threshold can be aggressive — 70 is fine. If a wrong auto-route creates a Salesforce opportunity and assigns it to a rep, threshold should be conservative — 85+. If a wrong auto-route triggers a refund or a contract clause, threshold should be very conservative — 95+ and probably with a "double-check" step.

Write this down in your routing-decision document. Future-you will thank you when someone says "why is the threshold 85 here and 70 there?" because the answer ("cost of wrong auto-route differs") will already be on the page.

The Override Path: The Operator's Document

The override path is the second-most-important thing about the pattern, after calibrated confidence. It is the answer to: "what if the LLM is wrong AND confident?"

Three override mechanisms

Every confidence-threshold routing workflow needs at least one. The best ones have all three.

  1. Manual reroute by the assignee. The auto-assigned SDR can move a lead back to the human queue with one click ("this isn't a fit, route for review"). The button writes to an audit log so you can see how often the auto-assignment is being overridden. If the override rate is above 15%, your threshold is too aggressive.
  2. Periodic spot-check. 5% of auto-routed items get flagged for human review post-hoc. A senior SDR or marketing ops person reviews them weekly and labels: was the auto-route correct? This is the calibration data that lets you tune the threshold over time.
  3. Hard override on flag. Certain conditions force a route to human regardless of LLM confidence. Examples: lead from a strategic account (your top 50 logos), lead from a closed-won customer's email domain (they're back!), lead with a specific keyword in the form ("urgent," "RFP," "legal"). These hard rules sit in front of the LLM routing as an overlay.

The override audit

Build a saved view in your CRM or workflow tool called "Routing overrides — last 30 days." It shows every lead where the auto-route was reverted, plus every spot-check that flagged an error. Review it monthly. The patterns are diagnostic: overrides cluster around free-email leads → enrichment is failing on that segment; overrides cluster around a specific industry → ICP definition is wrong; overrides spike after a model upgrade → confidence calibration shifted.

Documenting the Threshold (The Operator's Three-Question Doc)

Every operator-shipped LLM routing workflow should have a one-page document. Three questions. Three answers. Three sentences each.

Question 1: What confidence threshold did we choose and why?

"We chose 80 because at our holdout sample (n=200, March 2026), confidence 80+ correlated with 91% correct auto-routes. A wrong auto-route in this workflow costs the SDR roughly 22 minutes of recovery time (apologize to the prospect, re-route, update notes). At an SDR cost of $5.40/call, that's $4.80 per wrong auto-route. We tolerate up to 10% wrong auto-route rate before reviewing the threshold."

Question 2: What is the override path?

"Three mechanisms: (a) Auto-assignee reroute button writes to audit log; if override rate >15% over rolling 14 days, threshold review triggered. (b) 5% post-hoc spot-check by [name] in marketing ops, weekly. (c) Hard-override rules: strategic-account list (top 50), closed-won re-engagement, RFP/legal/urgent keywords in form. These hard rules execute BEFORE the LLM routing node."

Question 3: When does the threshold get re-evaluated?

"Monthly on a fresh 100-lead sample. Plus immediately if override rate exceeds 15% on the rolling 14-day window. Plus on model upgrade (Anthropic or OpenAI new version) — re-run the calibration before re-enabling auto-route."

Three answers. One page. Done. This document is what makes the workflow defensible when someone asks "why is the AI doing that?" four months from now.

Building the Routing Node in the Workflow Tool

The mechanics differ slightly across platforms. The pattern is identical.

In n8n

After your LLM node (Anthropic or OpenAI), drop an "IF" node. Condition: {{$json.confidence}} >= 80. True branch goes to the auto-assignment chain (Salesforce/HubSpot update, Slack notification, calendar invite). False branch goes to a "Create Task" node in your CRM that routes to the human review queue. The IF node is the entire routing logic. Two branches. One condition.

Add: a third branch that catches the INSUFFICIENT_DATA case explicitly. Use a "Switch" node instead of IF if you want three branches. The third branch routes to a re-enrichment workflow with deeper sources, then loops back.

In Make

The "Router" module has multiple paths. Each path gets a filter. Filter A: confidence ≥ 80 AND decision ≠ "INSUFFICIENT_DATA" → auto-assignment chain. Filter B: confidence < 80 AND decision ≠ "INSUFFICIENT_DATA" → human queue. Filter C: decision = "INSUFFICIENT_DATA" → re-enrichment chain. Three explicit paths. Make handles the routing.

In Zapier

Use Paths (Zapier's branching primitive). Each Path is a separate sub-workflow with a filter on the input. Same logic as n8n's IF or Make's Router. Note: Zapier Paths are billed as separate tasks per branch evaluated, so the cost math is slightly worse than n8n self-hosted; that's covered in the Level 2 self-host lesson.

In Clay

Clay's "if/then" column type creates conditional column outputs. Wire confidence > 80 to "Update HubSpot — auto assign," <= 80 to "Update HubSpot — assign to review queue," and equal-to-null/INSUFFICIENT_DATA to "Tag for re-enrichment." Clay's strength is that all three paths execute in one row, so the whole pipeline stays in one auditable table.

Logging and Observability (The Part Operators Skip and Regret)

Every routing decision should write a row to an audit log. The columns:

  • timestamp — when the routing decision was made
  • lead_id or item_id — what was routed
  • llm_decision — the model's classification or score
  • llm_confidence — the numeric confidence
  • routed_to — "auto" or "human" or "re-enrich"
  • threshold_used — what the threshold was at decision time (lets you reconstruct after threshold changes)
  • override_status — initially null; updated if someone overrides
  • override_reason — free text if applicable

Store this in a Google Sheet, Airtable, or Supabase table. A month from now, when the question is "is the routing actually working?" — this table is the answer. You can compute: routing distribution, override rate, confidence distribution by routed-to, time-to-action by route.

The Sunday review

One operator habit that compounds: open the audit log every Sunday evening. Spend 10 minutes. Look at the override rate this week vs. last week. Look at the confidence distribution. Look at any INSUFFICIENT_DATA spikes. If nothing changed, close the tab. If something changed, you have Monday morning to act before the head of sales walks into your office.

The Second Failure Mode: Confidence Anchored at 99

The opposite of the Tuesday-Thursday disaster. The prompt asks for numeric confidence, the LLM happily complies, and every output has confidence 95+ because the model is trained to be helpful and certain. Now the auto-route fires for every lead, including the ones that should have gone to human review.

The diagnostic

If your confidence distribution shows mean 94, median 95, p10 of 92, your prompt has no ceiling. The model is anchoring. The lesson from the earlier enrichment chapter applies: add the explicit "never above 95" rule and the explicit "confidence 80+ requires three independent data sources confirming the claim" rule. Re-run. The distribution should spread.

The validation

A healthy confidence distribution from a well-calibrated prompt looks bimodal-ish: a cluster of high-confidence (80-92) for the leads with rich data, a cluster of mid-confidence (50-70) for sparser data, occasional INSUFFICIENT_DATA outputs, and almost nothing in the 95-100 range. If you see a distribution centered at 99 with a tiny tail to 90, your prompt is broken.

Real Numbers from a Shipped Workflow

A mid-market RevOps team in April 2026 shipped this pattern for inbound lead routing. The numbers, after the first 30 days:

  • Total inbound leads routed: 3,247
  • Auto-routed (confidence ≥ 80): 2,194 (67.6%)
  • Human-routed (confidence < 80, ≥ 50): 832 (25.6%)
  • INSUFFICIENT_DATA / re-enrich: 221 (6.8%)
  • Auto-route override rate: 9.4% (within tolerance)
  • Spot-check error rate (5% sample, weekly): 8.1% wrong
  • Mean time-to-first-action for auto-routed leads: 4 minutes
  • Mean time-to-first-action for human-routed leads: 47 minutes
  • SDR time saved vs. previous all-manual routing: ~18 hours/week

The threshold was set at 80, chosen because the holdout test showed 91% correct at that level. The override rate at 9.4% is below the 15% trigger for threshold review. The spot-check confirms the override rate's signal. The 4-minute time-to-action for auto-routed leads is the real competitive advantage — leads engaged within 5 minutes of form submission are 5-9x more likely to convert than leads engaged in the first hour, by every B2B sales study from 2010 onward.

The deeper win: routing is the lever, not enrichment

The team spent the prior quarter building the enrichment pipeline (Lesson 2.3.1). That pipeline produced richer data. The richer data improved the LLM's confidence calibration. The improved calibration made the threshold-routing pattern reliable. The reliable routing saved 18 hours/week of SDR time.

The lesson chain matters. You cannot route confidently on bad data. The order — enrich, then score, then route — is the order this chapter teaches because it's the order the work has to happen in.

The Five Anti-Patterns to Avoid

Five mistakes will make this routing pattern fail. Avoid all five.

1. No numeric confidence in the output schema

You skipped the "force JSON with integer confidence" step. The LLM returns "high confidence" or "moderate." Your downstream filter can't parse it. Everything routes to one bucket. Fix: explicit schema in the prompt, integer required, test 20 outputs before wiring.

2. No ceiling rule on confidence

The model anchors at 99 and the auto-route fires for everything. Fix: "never above 95" plus "confidence above 80 requires three independent confirmations" plus an explicit rubric for what 70 vs. 85 vs. 95 actually mean.

3. No override path documented

When the auto-route gets it wrong (and it will), no one knows how to revert it. The SDR works the bad lead and resents the system. Fix: button to reroute, audit log, override rate threshold for re-tuning, hard-override rules for known edge cases.

4. No periodic spot-check

The pattern ships, looks fine in week one, drifts in week three, and no one notices until week six when a strategic account complains. Fix: 5% spot-check, weekly, owned by a named person, results captured in the audit log.

5. No documented threshold rationale

Six months later someone asks "why is the threshold 80?" and no one knows. The doc gets dropped, the threshold gets changed on a hunch, the system performance regresses. Fix: the three-question doc. One page. Updated when the threshold changes.

Key Takeaways

  • The confidence-threshold pattern is one if statement: confidence above T → auto-route; below T → human queue. All the interesting work is in the prompt, the threshold choice, and the override path.
  • The most common failure: 100% of items route to "uncertain" because the prompt didn't force a numeric confidence score. Fix: force JSON schema with integer confidence, anchor the scale (90+ = bet $1,000), test the output format on 20 samples before wiring the routing.
  • The second failure: confidence anchored at 99 because the model is trained to be certain. Fix: "never above 95" rule and explicit "80+ requires three confirmations" rubric.
  • Choose the threshold by measurement, not feel. Sample 50-100 outputs, plot confidence vs. correctness, pick T where correctness is at your tolerance (typically 90%+ for auto-routing).
  • The threshold should reflect cost-of-wrong-auto-route. Recoverable mistakes → aggressive threshold (70). Costly mistakes (Salesforce opportunity creation) → conservative (85+). Contract or refund triggers → very conservative (95+).
  • Three override mechanisms: assignee reroute button (with audit log), 5% post-hoc spot-check, hard-override rules for strategic accounts and keywords. The best workflows have all three.
  • The three-question doc: what threshold and why, what is the override path, when does it get re-evaluated. One page. Update when the threshold changes.
  • Log every routing decision: timestamp, item_id, llm_decision, llm_confidence, routed_to, threshold_used, override_status. Sunday-night review on the audit log catches drift before Monday-morning complaints.
  • Real numbers from April 2026: 67.6% auto-routed, 25.6% human queue, 6.8% re-enrich. Override rate 9.4%. Mean time-to-action for auto-routed: 4 minutes — that's the competitive advantage worth the engineering.
  • Five anti-patterns to avoid: no numeric confidence, no ceiling, no override path, no spot-check, no documented threshold rationale. All five are tempting. All five are fatal.