Few-Shot Learning & Examples
Nadia Oyelaran runs product marketing for a mid-sized electronics distributor. Every month she needs descriptions for new items in the catalogue, and every month she gets something different back from her AI assistant: one description opens with a technical spec, the next opens with a lifestyle scene, another runs long and buries the warranty terms. The instructions she typed were the same each time. What was missing was not a better instruction. It was an example.
The Power of Examples
When you teach a person how to do something, you explain it in words, and then you show them. You say here is what a good job looks like, here is what a poor job looks like, and here is what we are aiming for. By the time you have finished showing examples, they understand far better than they would have from the explanation alone, because the example carries information that description struggles to carry: length, rhythm, level of detail, what to leave out.
Language models respond the same way. When you show them examples of what you want, they adapt their behavior to match. This is few-shot learning: providing a small number of examples, often one to three, that demonstrate the format, tone, style and level of detail you are after. What makes it so useful is that it costs nothing beyond the prompt itself. There is no training, no configuration, no waiting. You include the examples in the request and the model adapts immediately, the way someone adapts when you hand them a template and say follow this pattern. The result is a dramatic improvement in consistency across multiple outputs.
Zero-Shot Versus Few-Shot
Zero-shot prompting is asking the model to perform a task with no examples at all. You state what you want and it generates output from its general knowledge. For straightforward one-off tasks this works fine. The problem appears the moment you need several outputs that belong together. Ask for a product description for a wireless headphone model and you will get something usable, but you have no control over its length, its tone, which features it emphasizes or what marketing angle it takes. Ask for a second description and it will not match the first. That is Nadia's monthly problem in one sentence.
The few-shot version changes the request only slightly: here is a product description you wrote earlier that we like, followed by the example, then write a new product description for the new product in the same style and format. Now the model has a reference point rather than a blank page. It matches the style, the length, the detail level and the structure of what you showed it, and because every output in the batch is anchored to the same example, they end up consistent with each other rather than merely individually competent. The golden rule follows directly: if you need outputs that resemble each other in format, style, length and tone, provide examples. If you genuinely do not care about consistency between outputs, zero-shot is fine. For professional work that scales across many items, few-shot is almost always the right call.
Choosing Examples That Teach
Not all examples are equally useful. Some steer the model exactly where you want it; others create confusion or pull it somewhere you did not intend, and the failure is easy to misread as a limitation of the model when it is really a property of what you showed it. The difference lies in how you select and present them, and four habits do most of the work.
Make them representative rather than perfect. Your examples should look like the output you actually want, which is not the same as an idealized specimen. A slightly imperfect example that shows how you handle a trade-off often teaches more than a polished one, because the trade-off is the part the model cannot infer. If you are writing product descriptions, show one that demonstrates how you balance marketing language against factual accuracy, rather than one that has been buffed until no tension remains in it.
Match the complexity of the real task. If you are asking the model to process a complicated input, give it examples of comparable complexity. Simple examples paired with a complex task invite the model to oversimplify, and you will read the result and think the model failed when in fact your example told it to.
Cover the categories you actually have. If the task spans several types of item, include at least one example from each. Descriptions for headphones, microphones and speakers should be demonstrated separately, because that is how the model learns to adapt the template to different contexts rather than forcing every item into the shape of the one case you showed.
Label them unmistakably. Mark each example with something explicit, such as EXAMPLE or the phrase here is how we wrote this previously. Without a clear boundary the model can lose track of what is demonstration and what is the live task, and you will get an output that answers the example instead of the request.
How Many Examples
The instinct is to give more, and the instinct is usually wrong. Each additional example lengthens the prompt, consuming more tokens and more processing time, while the clarity it adds falls off quickly. There is also a subtler cost: every extra example is another pattern the model has to reconcile, and where two of them differ in ways you did not intend, you have handed it a contradiction to resolve rather than a template to follow. The practical shape of the trade-off looks like this.
| Examples | What it buys you | When to use it |
|---|---|---|
| One (one-shot) | Establishes the pattern. A single clear, high-quality example is often enough to lift consistency sharply. | The default. Start here for almost every repeated task. |
| Two (two-shot) | Shows that the pattern holds even when the content changes, which teaches the model what varies and what does not. | When the inputs differ noticeably and you need the format to stay fixed regardless. |
| Three or more | Rarely adds proportional value. Mostly it lengthens the prompt and risks introducing conflicting patterns. | Only when you genuinely have distinct categories that each need demonstrating. |
The underlying point is that examples buy clarity, and clarity saturates. One example of a clear, high-quality output will improve consistency dramatically. A second is worth adding when you need to show variation. Beyond that you are usually paying tokens for confusion rather than precision, and the cure for a still-inconsistent output at that stage is a better single example rather than a third mediocre one.
Enforcing Format Consistency
The highest-value use of few-shot learning is enforcing a consistent output format, and this becomes essential the moment AI output feeds a downstream system rather than a human reader. Suppose you want to extract key information from customer emails into a structure your CRM can ingest. Without an example the model will sometimes include a field, sometimes omit it, and sometimes format the same field two different ways, and every one of those variations is a break in the pipeline. With an example you are not requesting a structure, you are demonstrating it, which is a far stronger instruction.
The pattern in full looks like this. You supply a sample input email, such as "Customer says they have a broken widget and want a refund. They ordered three months ago." You then supply the example output that goes with it, which is the artifact doing the real work here, because it fixes the field names and the value conventions in a single stroke:
- { "issue_type": "defective_product", "requested_resolution": "refund", "order_age_months": 3 }
Then you present the new input, "The software is too complex for my team. Can you simplify it or provide training?", with the instruction to extract key information in the same JSON format as the example above. What comes back follows the demonstrated structure:
- { "issue_type": "usability_concern", "requested_resolution": "training_or_simplification", "specificity_level": "team_level" }
The field names, the naming convention, the use of lowercase underscored values and the general shape are all inherited from the example rather than described in the instruction. That is why this technique holds up across a batch: the structure is anchored to an artifact you control rather than to a sentence the model has to interpret afresh each time.
Combining Few-Shot With Other Techniques
Few-shot learning is at its most powerful layered with other prompting techniques rather than used alone. Combined with chain-of-thought reasoning it produces noticeably better analysis, because the example demonstrates not only what the answer should look like but how the reasoning should proceed: here is an example of how we analyze market opportunities, including the step-by-step reasoning, now analyze this new opportunity in the same format and to the same standard of reasoning. Combined with a system prompt it produces highly consistent specialized roles, where the system prompt sets who the model is and the example sets what its work product looks like. The strongest prompts in most professional libraries are layered in exactly this way.
This is also where the economics turn in your favour. Crafting one genuinely good example takes real effort, but that effort is spent once and reused across dozens or hundreds of tasks. Few-shot learning is, in that sense, a quality multiplier: the care you invest in a single perfect example pays out across every future output that borrows it, which is why the technique scales in a way that repeatedly rewriting instructions does not.
Building Your First Template
The way to learn this is to build one template and use it. Start by choosing a task you already do repeatedly: email responses, meeting notes, customer summaries, data extractions, content generation. Repetition is what makes the investment worthwhile, so pick the thing that lands on your desk weekly rather than the interesting thing that happens once. Then produce one high-quality output by hand, demonstrating exactly the style, format, length and detail you want. This artifact is the whole asset; everything else is packaging.
Structure the prompt around it in this shape: "You are [ROLE]. Here is an example of how we [DO THE TASK]. [INSERT EXAMPLE]. Now, do the same for: [NEW TASK]". Then test it on inputs it has not seen and watch how much steadier the output becomes. Save the template somewhere you will find it again, because the value compounds only if you reuse it. Few-shot learning is one of the most practical and immediately deployable prompting techniques available, and the reason to start small is that a single example on a single recurring task will show you the effect within a week.
Anti-Patterns
- Showing an example of what to avoid. The model treats a demonstrated pattern as the pattern to follow, so a cautionary example is quite likely to be reproduced rather than avoided. Show what you want, and put what you do not want into the instruction text instead.
- Using an unrepresentative example. An example far simpler or far more complex than the live task sets the wrong target, and the output that follows will look like a model failure when it is actually faithful execution of a bad demonstration.
- Being so specific the model copies rather than adapts. A good example demonstrates a pattern; an over-specified one demonstrates a formula, and you will get your own example back with the nouns swapped.
- Letting examples contradict each other. Two examples with different formats or different levels of detail leave the model to guess which one governs, and the inconsistency you were trying to eliminate reappears in a new form.
- Leaving examples unlabelled. Without an explicit marker the model can mistake demonstration for instruction and answer the example rather than the task.
- Adding examples to fix a weak example. When one demonstration is not working, the fix is a better single example, not a third and fourth that lengthen the prompt and multiply the patterns on offer.
- Letting a template outlive the task it was built for. A saved example keeps enforcing the format it captured, so when the requirement changes and the template does not, few-shot will hold your output firmly in the wrong shape.
Practice Prompts
- Take the last output you were dissatisfied with and, instead of rewriting the instruction, write the example you should have supplied. Compare the two versions of the request.
- Run the same task zero-shot several times and few-shot several times, then lay the outputs side by side and mark where the zero-shot batch diverges.
- Pick a task with several distinct input categories and build a one-shot template for each category rather than one example for all of them. Note which categories genuinely needed their own.
- Convert one of your existing extraction tasks to a demonstrated output structure and check whether the field names hold steady across a batch of real inputs.
- Write an example that deliberately includes a trade-off you always have to make, and see whether the model reproduces the judgment or only the format.
- Audit your saved templates and find the one whose example no longer reflects current requirements. That is the template quietly producing outdated work.
Reflection
Nadia's problem was never that her instructions were unclear. They were perfectly clear, and they still produced a different result every month, because a written instruction can describe intent but cannot easily convey length, rhythm, emphasis or the thousand small conventions that make a set of documents feel like they came from the same organization. An example carries all of that without naming any of it. Think about the recurring outputs your own team produces. If someone asked you to explain, in words alone, exactly what makes a good one, how long would that explanation take, and how much of what you know would survive the writing down? Wherever the answer is uncomfortable, you have found a task that wants an example rather than a longer brief.
Glossary
- Few-shot learning: Including a small number of worked examples in a prompt so the model adapts its output to match the demonstrated format, tone, style and level of detail.
- Zero-shot prompting: Asking a model to perform a task with no examples supplied, relying entirely on its general knowledge and your instruction.
- One-shot prompting: Supplying exactly one example, which is usually enough to establish a pattern and is the sensible default for a repeated task.
- Two-shot prompting: Supplying two examples to demonstrate that the format holds even when the content changes, separating what varies from what must not.
- Prompt template: A reusable prompt structure with placeholders for role, example and new task, saved so the effort of crafting a good example is spent once.
- Format enforcement: Using a demonstrated output structure, rather than a described one, to keep fields and conventions stable across a batch of outputs.
- Token cost: The prompt length penalty paid for each additional example, which is the reason example count is a trade-off rather than a free improvement.
Related Lessons
- Anatomy of an Effective Prompt covers the instruction side of the prompt, which examples complement rather than replace.
- Chain-of-Thought & Reasoning Prompts is the technique this lesson layers with when you need the reasoning demonstrated as well as the format.
- System Prompts & Persona Design handles the role-setting half of the combination described here.
- Structured Output Engineering goes further into producing machine-readable output for downstream systems.
- Prompt Libraries & Version Control addresses the storage and maintenance problem that appears once your templates become assets worth keeping.
- Iterative Refinement covers what to do when a single example is not yet producing the output you want.
Closing
Few-shot learning earns its place in a professional toolkit because it is cheap, immediate and disproportionately effective at the one thing plain instructions handle badly: making many outputs resemble each other. It requires no training, no infrastructure and no specialist knowledge, only the discipline to produce one good example and the judgment to know when a second is worth its length. Nadia did not need a more sophisticated tool or a longer brief. She needed one description she was proud of, labelled clearly and placed in front of the request, and a template she could reach for the following month. Identify one task you repeat, build that example, and reuse it. That is the whole technique, and it is available today.
Key Takeaways
- Examples carry what instructions cannot. Format, length, tone and level of detail transfer far more reliably through demonstration than through description, and few-shot learning is simply demonstration inside a prompt.
- Use few-shot whenever outputs must resemble each other. Zero-shot is adequate for a one-off; consistency across a batch is the specific problem examples solve.
- Select examples for representativeness, matched complexity, category coverage and clear labelling. A slightly imperfect example that shows a real trade-off teaches more than a polished one.
- One example is the default and two is the common upgrade. Beyond that, additional examples mostly add length and risk conflicting patterns rather than clarity.
- Demonstrated structure beats described structure for machine-readable output. Showing the exact field names and conventions keeps a pipeline stable in a way that a written specification does not.
- Layer few-shot with reasoning and role techniques. An example that shows both the reasoning and the finished format, inside a system prompt that sets the role, is the shape most strong professional prompts take.
- The effort is spent once and recovered many times. One carefully built example becomes a template that raises the quality of every future output that borrows it, provided you keep it current.
Frequently Asked Questions
Does few-shot learning change the model in any way? No. Nothing is trained and nothing persists between requests. The examples live inside the prompt and influence only the output of that request, which is precisely why the technique is available immediately to anyone with access to a model and costs nothing beyond the extra prompt length.
My output is still inconsistent even with an example. What now? Look at the example before you look at the model. Most often it is not representative of the real task, or it is simpler than the inputs you are actually sending, or it is not clearly separated from the live request. Fixing the single example resolves this far more often than adding a second one does.
Can I reuse the same example across different tasks? Only where the tasks genuinely share a format. An example teaches structure and style along with everything incidental about the case it describes, so reusing one across tasks with different requirements imports the wrong conventions. Build one template per recurring task and keep them separate.
How do I know when a second example is worth the extra length? Add one when your inputs vary in ways the first example does not cover, particularly when you have distinct categories of item or a format that must stay fixed while the content changes substantially. If the inputs are homogeneous, the second example is mostly cost.
Skill.re