←
AI Agent Builders & Citizen Developers
Capable · M19 · lesson 19 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The 'Refuse if Unsure' Pattern
📖
now learning

The 'Refuse if Unsure' Pattern

15 min

The agent that always answers is the agent that quietly costs the most. It is the one that auto-tags a billing dispute as a sales question, drafts a confident reply citing a product feature that does not exist, or marks a security incident as routine. The fix is the smallest prompt-engineering change with the highest production payoff: teach the model to say "I don't know" out loud, in a way the workflow can detect, and route those cases to a human queue. This lesson covers how to write the refusal instruction, how to structure the output schema so refusals are detectable, and how to wire the human handoff in n8n, Make, and Zapier without melting the ops team's inbox.

The Agent That Always Answers

A workflow we audited in March 2026 at a 70-person customer-success team had been running for three months. The agent took inbound customer emails and routed them: billing, technical, cancellation, upsell, partnership, spam. Volume: ~600 emails/day. The team was proud of it — 96% of emails routed automatically, only 4% landed in a "manual review" queue.

We pulled a week of decisions and re-labeled them. The agent's actual accuracy: 83%. The 13% gap between "routed without flagging" and "actually correct" was the cost. A cancellation flagged as upsell went to a sales rep who treated the customer as a prospect and made them angrier. A security incident flagged as technical sat in a Tier 1 queue for 18 hours before someone realized it was an active credential leak. A partnership inquiry tagged as spam went to deletion; the prospect followed up two weeks later asking why nobody had responded.

The agent had not "always known the answer." The agent had always produced an answer. Without instruction to refuse, the model picked the most likely category every time, including for ambiguous inputs that a human reviewer would have flagged as "I need more context." The cost of that confidence was invisible in the agent's logs (which only counted "did the model return a category") and visible only in the downstream customer experience.

An agent that never refuses has zero false negatives in its logs and zero accountability for its false positives. It looks great on dashboards and burns out the customers it gets wrong.

The fix is teaching the model to say "insufficient context" — explicitly, in a structured way the workflow can detect — and routing those cases to a human queue. Done correctly, this lifts effective accuracy from "83% silent" to "92% routed + 8% explicitly escalated." The 8% is the agent doing its second job: knowing what it does not know.

The Three-Component Refusal Pattern

Three things have to be in the system prompt for refusal to work reliably.

Component one: a declared escape value in the schema

The output schema must include a sentinel value the model can emit when uncertain. For a classifier, this is typically a category like insufficient_context or needs_human. For an extractor, it might be a null with a reason field. For a scorer, it might be a special value like -1 or a flag like "confidence": "low".

The value must be enumerated in the schema alongside the normal values, not as an afterthought. If your classifier accepts billing, technical, cancellation, upsell, partnership, spam, your schema also accepts insufficient_context. The model treats it as a normal first-class output, not as a meta-comment.

Component two: explicit instruction on when to use it

The model needs to know what "unsure" means. Vague instructions ("use insufficient_context if you're not sure") don't produce reliable behavior. Specific criteria do. For our six-category email classifier:

Use insufficient_context when ANY of the following is true:
- The email mentions topics from more than two of the six categories (e.g., a billing question that's also a cancellation).
- The email is shorter than 20 words and the topic is ambiguous.
- The email contains a screenshot or attachment reference but no body text.
- The email references a previous conversation thread you don't have access to.
- You would assign a confidence below 70% to your primary classification.

The last criterion is the most general. The first four are specific cases where we saw the agent guess wrong in production. Both kinds of instruction help: specific cases handle the patterns we know about; the general criterion catches new patterns we haven't seen yet.

Component three: a rationale field that explains the refusal

When the model emits insufficient_context, it should also write 1-3 sentences explaining why. This serves three purposes: (1) it helps the human reviewer in the routing queue understand the case quickly, (2) it gives you a way to audit whether the refusals are well-calibrated (lots of "user said 'help'" refusals means the model is over-refusing), and (3) it slows the model down — generating a rationale forces a deliberate decision instead of a reflex.

Sample output for a refused case:

{
  "category": "insufficient_context",
  "rationale": "Email mentions both a billing dispute ($340 charge) and a request to cancel. Cannot determine whether the user wants the charge refunded and to stay, or to cancel and dispute the final charge. Human review needed.",
  "primary_candidate": "billing",
  "secondary_candidate": "cancellation"
}

The primary_candidate and secondary_candidate fields are optional but useful: they let the human reviewer see what the model was leaning toward without forcing it to commit. This is the "show your work" pattern, and it speeds up human review measurably.

Calibrating the Refusal Rate

The danger of teaching refusal is that the model over-refuses. We've seen agents that, after a too-aggressive refusal instruction, refused 40% of inputs — including cases a human would have classified easily. That's worse than not refusing at all because now you have a human queue twice as large as it should be and the team stops trusting the agent.

The target band: 5-15% refusal rate for a well-tuned workflow. Below 5%, the agent is probably still over-confident on edge cases. Above 15%, it's over-refusing and the human queue is overloaded.

Calibration steps:

  1. Measure the baseline refusal rate. Run 200 historical records through the refusal-instrumented prompt. Count how many came back as insufficient_context.
  2. Audit the refusals. Have an operator label each refusal as: "correctly refused" (human would have flagged it too), "should have classified" (clearly belongs to a category), or "correctly refused but easy" (refused unnecessarily).
  3. Adjust the criteria. If "should have classified" is high, your refusal criteria are too aggressive — relax them. If "correctly refused" is dominant, you're well-calibrated. If "correctly refused but easy" is high, the model is being lazy on hard-but-decidable cases — sharpen the criteria.
  4. Re-run. Iterate until the refusal rate is in the 5-15% band AND the audit shows the refusals are mostly "correctly refused."

Three common over-refusal patterns

Pattern one: the "I'm not 100% sure" refuser. The model treats any non-certain case as a refusal. Fix: change the confidence threshold from "below 70%" to "below 50%," or remove the confidence criterion entirely and rely on specific case criteria.

Pattern two: the "any ambiguity" refuser. The model refuses any input that mentions multiple topics, even when one is clearly primary. Fix: change "mentions topics from more than two categories" to "primary topic is from more than one category" with examples showing how to pick the primary.

Pattern three: the "I don't have full context" refuser. The model refuses any input that references previous conversations, even when the current message is self-contained. Fix: scope the criterion to "the topic cannot be determined without the previous conversation," not just "previous conversation is referenced."

Wiring the Human Handoff

The refusal pattern is only half the work. The other half is the workflow plumbing: when the model emits insufficient_context, what happens next? Three approaches, in order of robustness.

Approach one: dedicated review queue

The workflow has two branches downstream of the LLM node. If category != insufficient_context, the normal routing happens. If category == insufficient_context, the record goes to a "needs review" queue (a Slack channel, a Linear sub-status, a dedicated Airtable view, or a Trello list).

The reviewer sees the email, the model's rationale, and the primary/secondary candidates. They make a decision in seconds. The decision feeds back: either as a manual route (the human picks a category and the workflow continues) or as a labeled training example (the next quarterly refresh adds this case to the few-shot examples).

For 600 emails/day with a 10% refusal rate, that's 60 reviews/day. A trained reviewer does each in 30-45 seconds. Total team load: 30-45 minutes/day. That's the actual cost of the refusal pattern, and it's significantly less than the cost of misroutes the team was eating before.

Approach two: Slack-based approval inline

The workflow sends a Slack DM to the reviewer with the email content, model's rationale, and three buttons: the top two candidate categories, and "Other (specify)." The reviewer clicks. The workflow continues with the human's choice.

This is more interactive but adds 30 seconds of friction (Slack notification delay, click latency, channel-switching cost). Works well at smaller volumes (under 50 refusals/day) where each one warrants attention. Becomes overwhelming at high volume.

Approach three: confidence-tiered routing

For workflows where partial automation is acceptable, treat refusals as "auto-route with low confidence" rather than "stop and ask." The workflow routes to the primary candidate but flags the record with a "review next business day" tag. A daily batch review checks the flagged records. If the model was right, no action. If wrong, the team adjusts the routing and updates the few-shot examples.

This is the least friction approach but accepts a small percentage of wrong routes during the batch interval. Suitable for workflows where speed matters more than precision on edge cases (marketing-ops, lead enrichment) and not for workflows where precision matters more than speed (security, billing, compliance).

Implementation in n8n, Make, and Zapier

n8n

After the LLM node, add an "IF" node that checks {{ $json.category }} === "insufficient_context". The true branch goes to a "Slack" node posting to the review channel, with the full email body, rationale, and candidate categories in the message. The false branch continues to the normal routing.

For the review queue's feedback loop, add an n8n webhook node that the Slack interactive buttons call back to. The webhook updates the original record's category and triggers the rest of the workflow as if the model had emitted that category.

Make

Use a Router module after the LLM call. Add two routes: one for insufficient_context (goes to the review queue), one for everything else (goes to normal routing). Make's Filter conditions on the routes make this trivially declarative.

For the feedback loop, Make's Webhook trigger receives the human decision and writes back to the source system (HubSpot, Salesforce, Zendesk, etc.).

Zapier

Use Zapier's "Paths" feature. Path A: category equals insufficient_context, send to Slack channel for review. Path B: any other category, continue normal routing. For the feedback loop, the Slack action's reply triggers a second Zap that updates the record.

Zapier's main constraint is task cost — each Path counts as a task. A 600-email/day workflow with 60 refusals/day will incur ~660 tasks daily, which adds up. We covered when this becomes indefensible in the Chapter 2.1 lesson on Zapier's $60K threshold.

Real Stories Where Refusal Saved the Bill

Story one: the billing-vs-cancellation classifier

A FinTech company's billing-ops team auto-classified support emails as billing, cancellation, or other. Before refusal: 91% accuracy on holdout. In production, customer NPS dropped 8 points over two months. Audit revealed the agent was routing "I want to dispute this $340 charge AND cancel" emails as billing (because the dispute was salient), which sent them to the disputes team. The disputes team processed the refund. Then the customer's auto-renew hit a month later, and the customer went ballistic on social media: "you charged me again, I cancelled."

The fix: add insufficient_context as an output value, instruct the model to refuse when an email mentions both billing and cancellation. Refusal rate: 7%. Those 7% went to a human reviewer who made the explicit dual-action decision (refund + cancel) or contacted the customer for clarity. NPS recovered in six weeks. Customer-effort score on cancellation flow improved 14 points.

Story two: the contract-clause extractor

A legal-ops team extracted eight key clauses from MSAs. Before refusal: model would happily produce a clause extraction even when the contract didn't contain that clause type, hallucinating a plausible-sounding clause from the surrounding context. Downstream review caught most of them, but two slipped through into client deliverables.

The fix: each clause field could be null with a reason. Instruction: "If the contract does not contain a clause of type X, return null for that field and explain why in the rationale. Do not fabricate." Hallucination dropped from ~3% of fields to ~0.2%. The remaining 0.2% were cases where the contract had unusual phrasing the model misread; those went to a manual review queue.

Story three: the security-incident triager

A SecOps team used an agent to triage security alerts: noise, investigate, page_oncall. Before refusal: the agent classified an active credential-leak alert as investigate (not page_oncall) because the alert text was unusual and didn't match any of the team's training examples. The leak ran for 18 hours before someone noticed in the daily review.

The fix: add insufficient_context as the safe-default. Instruction: "When the alert pattern does not closely match any of the trained examples or contains terms not seen in training, classify as insufficient_context and page the on-call." The refusal rate went from 0% to 4%. Of those 4%, about a quarter turned out to be real escalations the agent would have under-classified. Three months in, the team's mean-time-to-respond on novel alerts dropped 65%.

The Anti-Patterns

Anti-pattern one: refusal as a hidden behavior

The model is instructed to say "I'm not sure" in prose, but the output schema doesn't enumerate that as a valid category. Downstream parsing breaks; the workflow either crashes or treats "I'm not sure" as a normal category and routes garbage. Fix: enumerate the refusal value in the schema as a first-class output.

Anti-pattern two: no rationale on refusal

The model emits insufficient_context with no explanation. Human reviewers can't tell why the model refused, so they can't quickly decide. Review time triples. Fix: require a 1-3 sentence rationale and the top candidates.

Anti-pattern three: refusal as a dumping ground

The team treats "needs review" as "deal with whenever." Refusals pile up; customers wait; complaints arrive. Fix: set an SLA on the review queue (typically 1-4 business hours for customer-facing workflows; 30 minutes for security). Monitor queue age; alert if any item exceeds the SLA.

Anti-pattern four: silent over-refusal

The refusal rate climbs to 25% over time as the model encounters new patterns. The review queue overloads. The team starts ignoring the queue. Fix: monitor refusal rate as a daily metric. Alert if it exceeds 15% for two consecutive days; trigger a prompt review.

Anti-pattern five: no feedback loop

Human reviewers make decisions but those decisions never feed back into the prompt's few-shot examples. The model keeps refusing the same kinds of cases for months. Fix: quarterly refresh of few-shot examples includes 2-3 newly-correctly-decided cases from the human queue. The model learns by absorption.

When Not to Use Refusal

Three cases where refusal is overkill or harmful:

  • Output space is binary and the cost of wrong is low. A "is this email spam, yes/no" classifier doesn't benefit much from a third "I don't know" option. The cost of false positive (real email marked spam) is low, false negative (spam not marked) is also low, and the third option just creates a queue with no real value.
  • Downstream is fully reversible. If a wrong decision can be undone trivially (e.g., a draft email that a human reviews before sending), the refusal pattern adds friction without preventing harm. The human review step is already the safety net.
  • Refusal would always result in the same human action. If every refusal goes to the same person who does the same thing (e.g., "send to John, John classifies"), you might as well route everything to John and skip the model. The point of the model is to reduce John's load; refusal is the model's way of keeping the hard cases on John's plate while taking the easy ones.

The Build Routine

  1. Enumerate the refusal value in the schema. Add insufficient_context (or your task-specific equivalent) as a valid output value alongside the normal categories.
  2. Write specific refusal criteria. 4-6 specific cases where the model should refuse, plus one general criterion ("confidence below 70%" or similar).
  3. Require a rationale field. 1-3 sentences explaining why the model refused.
  4. Optionally add primary/secondary candidates. Helps human review without forcing the model to commit.
  5. Build the review queue. Slack channel, Linear sub-status, Airtable view, or Trello list. Set an SLA.
  6. Wire the workflow branch. n8n IF, Make Router, Zapier Path. insufficient_context → review queue. Everything else → normal routing.
  7. Calibrate the refusal rate. Run 200 historical examples. Audit. Tune criteria until the refusal rate is 5-15% and the audit shows refusals are mostly well-justified.
  8. Monitor refusal rate daily. Alert if it exceeds 15% for two consecutive days.
  9. Feed reviewed cases back into few-shot. Quarterly refresh adds 2-3 newly-correctly-decided cases.

Key Takeaways

  • An agent that never refuses has zero false negatives in its logs and zero accountability for false positives. Refusal is the agent's second job: knowing what it does not know.
  • The three-component refusal pattern: declared escape value in the output schema, explicit criteria for when to use it, and a required rationale field for human reviewers.
  • Target refusal rate is 5-15%. Below 5%, the agent is over-confident on edge cases. Above 15%, it's over-refusing and the human queue overloads.
  • Three over-refusal patterns to watch: the "not 100% sure" refuser, the "any ambiguity" refuser, and the "missing context" refuser. Tune criteria to fix each.
  • Three handoff approaches: dedicated review queue (most robust), Slack inline approval (lower volume), and confidence-tiered routing (least friction, accepts batch interval cost).
  • Five anti-patterns: refusal as hidden behavior (no schema value), no rationale, refusal as dumping ground (no SLA), silent over-refusal drift, and no feedback loop into few-shot examples.
  • Skip refusal when output is binary with low cost of wrong, when downstream is fully reversible, or when every refusal would result in the same human action.
  • Real production wins: billing-vs-cancellation classifier saved 8 NPS points; contract-clause extractor cut hallucination from 3% to 0.2%; security-incident triager cut MTTR on novel alerts by 65%.