Few-Shot Prompting with Real Records
Three to five anonymized examples in your prompt will move classifier accuracy from 78% to 94% on the same model. We've watched it happen across nine workflows. But examples carry a hidden tax: if you pull them from real tickets, real emails, real Slack threads, you are now shipping customer data into every prompt — sometimes into a cached system block that lives in a third-party provider's cache. This lesson is how to do few-shot well, how to measure the accuracy delta on a holdout set, and how to anonymize examples so PII doesn't leak into the model context (or into your provider's training data on a vendor whose terms allow that).
What Few-Shot Actually Does
Few-shot prompting is two to five labeled input-output pairs placed in the system prompt, just above the model's actual task input. The model uses them for what researchers call in-context learning: pattern induction from the examples to the new input. The examples teach three things simultaneously.
One: the input-output mapping. Given inputs that look like X, the desired outputs look like Y. For a classifier, this means "given ticket text that sounds like this, emit this label." For an extractor, "given a contract paragraph like this, return this JSON shape."
Two: the output format. The examples are the most reliable way to lock format because they show, not tell. If your three examples all output single-line JSON, the model emits single-line JSON. If they output multi-line JSON with indentation, the model emits that. Schema descriptions help; examples teach.
Three: the tone and verbosity of any prose field. If your examples have 40-word rationales, the model writes 40-word rationales. If your examples have terse one-sentence outputs, you get terse one-sentence outputs. This is the most underappreciated lever in workflow prompt design — you can shape register without ever using an adjective like "concise" or "professional."
Few-shot is teaching by showing. The model learns the mapping, the format, and the tone in one pass. Adjectives like "be concise" tell. Examples show. Showing wins.
The Accuracy Delta on a Real Holdout Set
The pattern we measure across most operator workflows: zero-shot achieves 70-85% accuracy on the task; 3-shot moves it to 85-92%; 5-shot to 90-95%; 10-shot adds marginal gain at significant token cost. Past 5 examples, returns diminish fast.
Here's a real measurement from an account-health classifier we built with a customer-success team in February 2026. The task: given 30 days of usage data + 2 recent CSM call notes + open tickets, classify the account as green, yellow, red, or insufficient_context.
Holdout set: 200 historical accounts the team had already manually classified. We froze the labels before any prompt iteration.
- Zero-shot (no examples). 78% accuracy. 22% errors split: 14% wrong classification, 5% format failures (model returned a paragraph instead of a label), 3% hedging that broke parsing.
- 3-shot. 89% accuracy. Format failures dropped to zero. Wrong classification dropped to 9%. Hedging dropped to 2%.
- 5-shot. 93% accuracy. Wrong classification dropped to 6%. Hedging eliminated.
- 10-shot. 94% accuracy. Wrong classification at 5%. Cost: input tokens doubled.
The team picked 5-shot. The marginal gain from 5 to 10 was not worth the doubling of input cost. And we noticed something quieter: at 5-shot the model's tail behavior (the 7% wrong) had a pattern — it was over-classifying as yellow when accounts were genuinely green. That kind of structural error you can fix by adjusting the examples. At 10-shot, the residual errors were more random and harder to fix.
How to build the holdout set
The holdout set is what separates "I think this is better" from "this is better by 11 percentage points." Without it, every prompt iteration is vibes. Three rules:
- Use historical data, not synthetic. Pull 50-200 real records that the workflow would have processed if it had been running last month. Synthetic data drifts from production distribution.
- Label before iterating. Have the operations team label the holdout set before you touch the prompt. Otherwise you'll subconsciously tune the prompt to match what you remembered, not what was actually correct.
- Don't use holdout examples as few-shot examples. Holdout is for measurement. Few-shot is from a separate pool (drawn from a different time window if possible). Mixing them produces inflated accuracy that won't survive contact with production.
For most workflows, 50 holdout records is enough to detect a 5+ percentage point delta. For high-stakes workflows (medical triage, financial decisions), 200 is more defensible. Below 30, the noise band is wider than typical accuracy gains and your measurements are inconclusive.
The PII Leak Problem
Here is the dirty secret of few-shot prompting in business workflows: the most useful examples are real customer interactions. The customer who literally said "I'm canceling because your support is terrible and refund me $4,200 by Friday" — that's a perfect refund-classifier example. The customer's actual name, email, account ID, and dollar amount are right there in the example.
If you paste that example into your system prompt verbatim, you are now shipping that customer's data into every prompt to the model provider. With most providers (Anthropic, OpenAI as of 2025, Google for paid Gemini tiers) the API tier does not train on your data. But the example still lives in the prompt cache, gets logged in your observability tooling (PromptLayer, Langfuse, Braintrust), and is reachable by anyone with access to that tooling. If you have GDPR or HIPAA obligations, you have just spread customer data across systems that may not be in your data processing agreements.
Three things can go wrong:
- The example contains direct identifiers (name, email, phone, address, account number). Trivially identifiable.
- The example contains indirect identifiers that, in combination, identify a person (job title + company + city + specific complaint). Harder to spot. We've seen "VP of Product at Acme Logistics in Denver, frustrated with the Q2 outage" — three of those data points together identify a specific person to anyone with LinkedIn.
- The example contains quasi-sensitive content (medical, financial, legal). Even anonymized, if the content reveals a condition or a financial situation, you may be in regulated territory.
Anonymization is not finding-and-replacing names. It is ensuring that no individual person can be re-identified from the example. The bar is harder than you think, and the cost of getting it wrong is reputational, regulatory, or both.
Anonymization Patterns That Work
Three patterns we use in practice. Each one trades fidelity for safety. Pick the level that matches your data sensitivity.
Pattern one: token replacement
Replace direct identifiers with role placeholders. "Sarah Chen" becomes "[CUSTOMER]." "acme.com" becomes "[COMPANY_DOMAIN]." "$4,200" becomes "[AMOUNT]" or "$X,XXX" if the magnitude matters for classification. "March 14, 2026" becomes "[DATE]."
Strengths: keeps the example readable and preserves the linguistic structure that helps the model learn. Weaknesses: indirect identifiers may still leak through (the specific complaint, the specific feature mentioned).
Use when: examples come from a single tenant, the data is low-sensitivity (support tickets without medical/financial detail), and your team has GDPR/CCPA basics in place.
Pattern two: synthesis from a template
Take 20 real examples, identify the structural patterns (length, sentiment, key entities, typical phrasing), and write synthetic examples that hit those patterns without using any real content. The synthetic examples look real — they sound like your customers — but no specific real customer is identifiable.
This is harder than it sounds. Operators often write synthetic examples that are too clean (no typos, no rambling, no irrelevant tangents) and the model learns a false pattern. A better technique: take the 20 real examples, write a prompt that asks Claude or GPT to generate 50 synthetic examples in the same style, then have a human review and pick the 5 best.
Strengths: zero PII risk; full control over distribution and balance. Weaknesses: takes 1-2 hours of human effort to do well; synthetic examples can subtly miss the long tail of real inputs.
Use when: examples are medical, financial, legal, or otherwise regulated; or when the workflow processes records across many tenants and you cannot risk cross-tenant leakage.
Pattern three: structural anonymization
Keep the structure (input shape, output shape) but remove all content. The few-shot becomes:
Input: ticket from a frustrated user about a billing issue, 80-120 words, mentions a specific dollar amount and a deadline.
Output:{ "category": "refund", "urgency": "high", "amount_mentioned": true }
The model sees the abstract pattern, not the data. This is the strictest anonymization. It works for classifier tasks where the pattern is what matters; it works less well for tasks where the model needs to learn vocabulary or tone from the examples.
Strengths: ironclad privacy; the example contains no customer data at all. Weaknesses: weaker pattern induction; works well for binary or low-cardinality classification, less well for extraction or generation.
Use when: maximum privacy is required; the task is structurally simple; or you are building a template that other teams (with different data) will reuse.
The anonymization checklist
Before adding a real-record few-shot example to a system prompt, ask:
- Does this example contain a name, email, phone, address, or account ID? Replace or remove.
- Does this example contain a company name plus a job title plus a specific event? Generalize or remove — the three together identify a person.
- Does this example contain dollar amounts, medical details, or legal-case details that are themselves sensitive? Replace with magnitude placeholders or generalize.
- Could a knowledgeable insider read this example and identify the underlying customer? If yes, anonymize further.
- Is the example processed by a model provider whose terms allow training on your data? If yes, switch to a tier that does not (Anthropic API, OpenAI Enterprise tier, Google Gemini paid tiers all opt out by default — but verify).
- Is the example logged in your observability tooling? Audit which examples appear in logs and ensure your retention and access policies cover them.
Choosing Examples That Teach
Beyond anonymization, the choice of which examples to include is where most of the accuracy gain comes from. Three principles.
Principle one: cover the class distribution
If your classifier has four labels and 60% of historical tickets are question, do not include three question examples and one refund example. The model over-anchors on the majority class and biases toward it. Instead, include one example per class plus one tricky-edge-case example.
For our four-label account-health classifier (green, yellow, red, insufficient_context), the 5-shot example set is: one green, one yellow, one red, one insufficient_context, and one edge case that the team historically mis-classified. The mis-classified case is where the gain comes from — it teaches the model the boundary explicitly.
Principle two: pick the boundary cases, not the obvious ones
If your examples are obvious (the textbook refund case, the textbook bug case), you teach the model the easy patterns it already knows. If your examples are boundary cases (the ticket that mentions a refund but is actually a question, the ticket that mentions a bug but is actually a feature request), you teach the model the hard distinctions. The accuracy gains come from the hard examples.
How to find boundary cases: ask the ops team for 20 historical examples they remember being unsure about. The model is unsure where they were unsure.
Principle three: keep the examples short
A 400-token example takes up 5x the cache space of an 80-token example. For most classifiers, 60-150 tokens per example is plenty. Trim the input to the signal-carrying parts; the model does not need three paragraphs of pleasantries before the actual complaint.
For extraction tasks, examples need to be longer because the input has to contain the extractable signal. 200-400 tokens per example is typical. Past that, you are paying for tokens without proportional pattern-teaching gain.
Few-Shot in the Three Workflow Tools
n8n
Place few-shot examples at the bottom of the System Message field, separated by clear markers. We use:
Examples:
---
Input: ...
Output: ...
---
Input: ...
Output: ...
The triple-dash separator is reliable; the model recognizes it as a delimiter. Avoid using markdown headers (## Example 1) because some models will mirror them in the output. Avoid using the word "input" or "output" in your downstream parsing because they appear in the system prompt and confuse log filters.
Make.com
Make's Anthropic Claude module lets you build the messages array as separate items. You can put few-shot examples either inline in the system message (recommended for caching) or as separate user/assistant message pairs in the conversation history (recommended for very large or structured examples). For caching to apply, inline in system message wins because the cache_control marker is on the system block.
Zapier
Use the direct Anthropic app (not the wrapped "AI by Zapier" action) so you have access to the System Prompt field. Place few-shot examples in the system prompt, same pattern as n8n. The wrapped action does not give you reliable system-prompt access and the workaround is fragile.
The Five Failure Modes
Failure one: PII in examples
The example contains "Sarah Chen at Acme Logistics, refund $4,200 for Q1 outage." You shipped customer data into every prompt. Fix: anonymize via Pattern One, Two, or Three above. Run a one-time scan of all your existing few-shot examples and replace direct identifiers. Add a pre-commit hook (if you store prompts in Git) that flags candidate PII strings.
Failure two: examples don't match production distribution
You picked five examples that looked good but were all from one customer segment, one product line, one urgency level. Production traffic is more varied. Model under-performs on the missing distribution. Fix: sample examples proportionally from your historical distribution, plus an extra boundary case.
Failure three: too many examples
You added 12 examples because "more is better." Input cost doubled. Marginal accuracy gain was 1%. Cache size grew, eviction risk increased, latency rose. Fix: prune to 5. Keep the most representative + the boundary case.
Failure four: examples that contradict each other
Two examples have similar inputs but different outputs. Model gets confused, picks neither, or invents a hybrid. We saw this in a sales-email tone-rewriter where two on-brand examples actually had subtly different brand voices because they were written by different copywriters at different times. Fix: ensure examples are internally consistent. If your training data has inconsistency, fix the data first, then the prompt.
Failure five: examples never refreshed
Your few-shot examples were chosen 8 months ago. The product has shifted, the customer base has shifted, the support categories have shifted. Examples no longer reflect current reality. Model accuracy drifts. Fix: re-validate examples quarterly. Run them through the current classifier and confirm the model would label them correctly. Refresh any that no longer represent the current distribution.
A Real Workflow from Anonymization to Shipping
One concrete walkthrough from a recent project. A travel-ops team had a classifier that triaged customer emails into seven categories: refund, change, complaint, question, compliment, fraud, other. Volume: ~800 emails/day.
Step 1: Build the holdout set. Pulled 200 historical emails the team had manually triaged in the previous month. Froze the labels. No further look at the holdout until measurement time.
Step 2: Baseline zero-shot. Wrote a clean system prompt with role + contract + forbidden behaviors + fallback. Ran the 200 holdout. Accuracy: 81%. Failure pattern: model confused change and refund when emails mentioned both; defaulted to question when input was a complaint with no explicit refund request.
Step 3: Pick few-shot candidates. Pulled 40 emails from before the holdout time window. Anonymized using Pattern One (replace names, emails, dates, dollar amounts; keep linguistic structure). Picked 7 candidates: one per category plus one boundary case (a complaint that mentions a refund but is really a question about policy).
Step 4: Pre-publication anonymization audit. Two team members independently read the 7 examples. They flagged two: one mentioned a specific destination + travel date that could identify the customer; another mentioned a specific airline + booking number that could identify the customer. Replaced both with generalized placeholders. Final set: 7 examples, fully anonymized.
Step 5: Iterate on count. Tested 3-shot (skipping fraud and compliment and the boundary case), 5-shot (all categories, no boundary), and 7-shot (full set). Measured on holdout. Results: 86%, 91%, 93%. Picked 5-shot for the cost/quality trade.
Step 6: Production shadow run. Ran the 5-shot prompt in shadow mode (decisions logged but not acted on) for 3 days of real traffic. Compared shadow decisions to ops team's live decisions. Agreement rate: 91% — matched the holdout. No drift, no surprise failures.
Step 7: Cutover and monitor. Switched the workflow from rule-based routing to model-based routing. Added a daily eval-set re-run on a 50-example rolling holdout. Alert if accuracy drops below 88%.
Three months later: workflow has processed ~70,000 emails, weekly accuracy has stayed between 89% and 93%, and the ops team has reviewed roughly 5% of decisions (sampling, not exhaustive). One example needed refresh after a new product launched (a new compliment sub-category emerged); the team added a new few-shot and accuracy recovered within two days.
When Not to Use Few-Shot
Few-shot is not always the right tool. Three cases to skip it:
- The task is trivially well-known. Sentiment classification on standard categories (positive/neutral/negative) does not need examples — modern models handle it zero-shot at 95%+. Adding examples adds tokens without accuracy gain.
- The output schema is highly constrained. A binary classifier with two values that you post-validate via a strict allowed-list won't benefit much from examples — the schema does most of the work.
- The examples would be repetitive. If your five candidate examples are basically the same case repeated, you are wasting tokens. Either find more varied examples or skip few-shot.
The decision: zero-shot first, measure on holdout, add few-shot only if accuracy is below your target.
The Build Routine
Whenever you're adding few-shot to a workflow, run these in order:
- Pull 50-200 historical records. Label them as the holdout set. Freeze.
- Pull 30-50 separate historical records for few-shot candidates. Different time window from holdout if possible.
- Anonymize all candidates. Pattern One for low-sensitivity; Pattern Two or Three for regulated data. Two-person audit before promoting any example to production.
- Run zero-shot baseline on the holdout. Record accuracy and the error pattern.
- Pick 5 candidates that cover the class distribution + one boundary case.
- Test 3-shot, 5-shot, 7-shot. Measure on holdout. Pick the count where marginal gain stops being worth the token cost.
- Shadow-run for 3-7 days. Compare to live ops decisions. Confirm holdout accuracy generalizes.
- Cut over with a rolling eval. 50-example holdout re-run daily. Alert below threshold.
- Refresh quarterly. Re-validate examples against current distribution. Replace any that no longer represent reality.
Key Takeaways
- Few-shot prompting typically moves classifier accuracy from 70-85% to 90-95% with 3-5 examples. Past 5 examples, returns diminish fast and input cost rises.
- Examples teach three things simultaneously: the input-output mapping, the output format, and the tone or verbosity of any prose fields. Showing beats telling for all three.
- Measure accuracy on a 50-200 record holdout set that is frozen before prompt iteration. Without it, every prompt change is vibes. Below 30 records, the noise band exceeds typical accuracy gains.
- Real-record few-shot ships customer data into every prompt and into prompt-cache, observability tooling, and provider logs. Anonymize before promoting examples to production.
- Three anonymization patterns: token replacement (Pattern One) for low-sensitivity data, template-based synthesis (Pattern Two) for regulated data, and structural anonymization (Pattern Three) for maximum privacy.
- Choose examples that cover the class distribution and include at least one boundary case the team historically misclassified. Boundary cases produce the most accuracy gain.
- Keep examples short: 60-150 tokens for classifiers, 200-400 tokens for extractors. Past that, you pay for tokens without proportional pattern-teaching benefit.
- Five failure modes: PII in examples, distribution mismatch, too many examples, contradictory examples, and examples never refreshed.
- Skip few-shot when the task is trivially well-known (e.g., standard sentiment), the schema does the constraint work, or your candidate examples would be repetitive.
- Refresh examples quarterly. Product shifts, customer base shifts, support category shifts — examples that worked 8 months ago may no longer represent current reality.
Skill.re