←
AI for Government
Capable · M30 · lesson 30 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Prompt Engineering Mastery: Chain-of-Thought and Few-Shot
📖
now learning

Prompt Engineering Mastery: Chain-of-Thought and Few-Shot

15 min

Ravi Krishnaswamy manages data analytics for the Pennsylvania Department of Revenue's audit division. His team of six analysts produces roughly 200 written audit summaries per year, each one synthesizing tax return data, compliance history, and examiner notes into a structured document that a supervisor uses to decide whether to pursue formal enforcement. In March, his division adopted an AI writing assistant to help draft those summaries. The first week's outputs were impressive. By week three, Ravi noticed a pattern: the summaries were well written but shallow. The AI would describe what the data showed but not why it mattered, would list discrepancies without flagging which ones crossed the threshold for referral, and would occasionally omit qualifications that changed the interpretation entirely. The tool was not failing. Ravi was failing to prompt it. When he learned two specific techniques, chain-of-thought prompting and few-shot prompting, the quality of the AI-assisted summaries improved enough that his supervisor asked what had changed. This lesson explains both in concrete terms.

Why Prompting Technique Matters for Government Work

A large language model (LLM), the type of AI behind most writing and analysis assistants, does not automatically produce good government outputs. It produces outputs that are plausible continuations of the prompt you give it. If the prompt is vague, the output is generic. If the prompt does not specify the reasoning standard required, the model will produce whatever reasoning style felt most common in its training data. For government work, where accuracy matters, where legal and regulatory context shapes every analysis, and where a mistake in an audit summary or a procurement memo has real consequences, plausible is not good enough.

Government increasingly uses LLMs for drafting, analysis, research, and decision support, so the gap between a careless prompt and a careful one compounds across a lot of work. Two techniques reliably improve output quality for analytical and drafting tasks. Chain-of-thought prompting makes the model set out its reasoning before producing a conclusion. Few-shot prompting teaches the model what a good output looks like by providing examples. Neither requires technical expertise. Both require deliberate attention to how you write the prompt, which is a skill that develops through iteration rather than arriving fully formed.

Chain-of-Thought Prompting

Chain-of-thought prompting instructs the model to work through a problem step by step before producing its final answer. The reason it helps is that a model asked only for a conclusion will often produce a plausible-sounding conclusion with no visible basis, and you have no way to tell a well-founded answer from a confident guess. Asking for the steps changes what you receive: a set of intermediate claims you can check individually, against the source material, without having to reconstruct what the model might have been doing.

Be precise about what that buys you. Showing the steps makes the reasoning visible and checkable. It does not guarantee that the stated steps are the ones that actually produced the answer, and a chain of reasoning can be fluent and wrong at any link. The benefit is diagnostic rather than protective: when the conclusion is bad, you can usually see which step went wrong and fix the prompt at that point. Treat the visible chain as evidence you now have to read, not as a warranty that the reading can be skipped.

The simplest possible illustration

Start with a benefits question. A basic prompt with no chain-of-thought asks: "Is this benefits application eligible for SNAP benefits? Application: [details]." The response comes back as a single word, "No," and you have learned nothing about why, which means you cannot tell whether the model applied the right rule, the wrong rule, or no rule at all.

Now the same question structured as a chain of thought: "Is this benefits application eligible for SNAP benefits? Let me work through this step by step: What is the household income? What is the household size? Is income below the eligibility threshold for this household size? Are there other disqualifying factors? Please work through these steps."

The response becomes something you can audit. With illustrative figures, it might read: household income $2,400 per month; household size 3; income threshold for a 3-person household $2,250 per month, so income exceeds the threshold by $150; no other disqualifying factors identified; conclusion, not eligible due to income. Those numbers are stand-ins for a worked demonstration, not current program limits, and the actual thresholds must come from your program's own eligibility tables. What the structure gives you is the ability to check each line: is that the household size on the application, and is that the threshold in the table? An error is now locatable instead of invisible.

The same technique on a real workload

Here is the version from Ravi's work. Without chain-of-thought, his prompt was "Summarize this audit finding: [data]," and the result was a well-formatted paragraph describing the data with no analysis of whether the discrepancy was significant enough to warrant referral. With chain-of-thought, the prompt became: "Review this audit data. First, identify each discrepancy between the reported figures and the supporting documentation. Second, for each discrepancy, assess whether it exceeds the referral threshold of $15,000 or 10% of reported income. Third, note any qualifications or explanations in the file that affect interpretation. Fourth, based on that analysis, draft a summary that a supervisor can use to make a referral decision."

The second prompt produces a materially different output, one where the model has been directed to perform specific analytical steps rather than to describe information. Ravi's team found that chain-of-thought prompts reduced the frequency of outputs that missed the referral-threshold analysis from about 35% of cases to under 5%. The remaining failures still had to be caught by a human, which is the point: the technique moved the error rate, it did not remove the review.

Structured chain-of-thought for complex analysis

For heavier analysis, specify the structure of the reasoning explicitly rather than leaving the model to invent one. A reusable template looks like this: "Analyze [document]. Please provide your analysis in this structure: SUMMARY (2 sentences). KEY POINTS (bulleted): Point 1, Point 2, Point 3. EVIDENCE (for each key point): Point 1 is supported by [specific quote or reference], Point 2 is supported by [specific quote or reference]. LIMITATIONS (potential problems with analysis): [Limitation 1], [Limitation 2]. CONFIDENCE LEVEL: [High/Medium/Low] because [explanation]."

Two of those sections do most of the work. The EVIDENCE section forces each claim to be tied to something in the document, which makes unsupported assertions visible as gaps rather than hiding them inside fluent prose. The LIMITATIONS section makes the analysis admit its own weak points, which is the part analysts most often have to add by hand. Keep the field names and the bracketed placeholders as written; the specificity is what produces the structure.

The CONFIDENCE LEVEL field deserves a warning. A self-reported confidence rating is another sentence the model generated, produced by the same process as everything above it, and a stated "High" is not a measurement of anything. It is still worth requesting, for one narrow reason: a "Low" or "Medium" is a useful flag for triage, and an explanation after "because" sometimes exposes an assumption you can check. Never treat "High" as a reason to review less carefully. If anything, treat a confident rating on a complicated question as a prompt to look harder.

When to use chain-of-thought

Chain-of-thought works best for tasks with multiple steps, tasks where the reasoning matters as much as the conclusion, and tasks where errors in reasoning are hard to detect inside a confident-sounding output. Government examples: drafting policy analysis memos, evaluating contractor performance, assessing permit applications against eligibility criteria, and reviewing procurement documents for compliance with Federal Acquisition Regulation (FAR) requirements. It is less necessary for purely descriptive tasks, such as summarizing a meeting transcript, where the model is not making judgments and there is no chain to inspect.

Few-Shot Prompting

Few-shot prompting teaches the model what a good output looks like by including examples in the prompt. The model picks up the pattern from the examples and applies it to the new task. "Few-shot" means a small number of examples, typically two to five. The technique is powerful because examples communicate format, tone, depth, and judgment criteria that are extremely difficult to describe in abstract instructions. Telling a model to be thorough accomplishes almost nothing. Showing it two outputs that are thorough in the specific way your agency means accomplishes a great deal.

Zero-shot against few-shot on a classification task

A zero-shot classification prompt gives the model nothing to work from: "Classify this benefit application as APPROVED or DENIED: [application details]." The model has to guess what you want, including what counts as a reason and whether you want one at all.

The few-shot version supplies the pattern. "Classify benefit applications as APPROVED or DENIED. Here are examples: Example 1: Income: $1,500/month. Household size: 2. Assets: $100. Classification: APPROVED. Reason: Income below threshold for 2-person household. Example 2: Income: $3,000/month. Household size: 2. Assets: $5,000. Classification: DENIED. Reason: Income exceeds threshold; assets exceed allowance. Now classify this application: [New application]." The field names, the label values, and the reason line are all part of what is being taught, and the figures are illustrative rather than program limits.

Note what the second example does that the first does not: it demonstrates a denial, and it shows a reason with two grounds rather than one. That is deliberate. Examples that are all approvals teach a model that approval is the expected answer, and examples that always cite a single reason teach it to stop looking after the first one it finds.

The same technique on drafting

Classification is the easy case. Drafting is where few-shot earns its keep. A zero-shot prompt, "Write an executive summary of this program performance report," produces a summary in whatever format the model's training suggested was typical, which may or may not match your agency's conventions. A two-shot prompt says instead: "Write an executive summary of this program performance report. Use the same structure and tone as the following two examples," followed by two real prior summaries, then "Now write a summary for: [new report data]."

If both examples are two-paragraph summaries built the same way, opening with the program goal, then key metric outcomes, then one noted risk, the output consistently arrives in that shape and includes those elements. The model does not have to guess what an executive summary means at your agency. The examples show it, and they show it more precisely than a style guide would.

Choosing your examples carefully

The examples you include teach the model everything they demonstrate, including their flaws. If your example summaries bury the risk assessment in the third paragraph and your agency actually wants risk in the first, the model will bury it too, consistently, in every output, and nobody will be able to explain why. Review examples before using them. The best few-shot examples are strong outputs you would be proud to send to leadership.

Five practices make the difference between examples that teach and examples that confuse. Provide two to five of them. Make them representative of the cases you actually see rather than the cleanest ones in your files. Include both positive and negative examples, so the model learns the boundary rather than one side of it. Show the exact output format you want, since the examples are the format specification. And include the reasoning behind each classification or judgment, because that is what transfers to the new case; a label with no reason teaches the label and nothing else.

When to use few-shot

Few-shot prompting is most valuable when your agency has a strong existing format you want matched, when the task requires structure or tone that is hard to specify abstractly, and when you are doing the same task repeatedly: audit summaries, grant review memos, legislative briefing papers. Build a small library of two or three strong examples for each recurring task type and reuse them. The investment pays off quickly when you are producing 200 summaries a year rather than 20.

Combining the Techniques

The two techniques are not mutually exclusive, and the most effective government prompts usually combine them: a chain-of-thought structure that specifies the reasoning steps, plus a few-shot example that shows what the finished output should look like. One governs how the model thinks about the problem; the other governs what it hands you at the end. Used alone, chain-of-thought produces good analysis in an unpredictable format, and few-shot produces the right format wrapped around thin analysis.

Ravi's current template for audit summaries uses the combined approach in three parts. First, a context section identifying the type of case, the applicable threshold, and any known complicating factors. Second, a chain-of-thought section listing the four analytical steps in order. Third, a format section containing one strong prior example as the structural template. The combined prompt takes two minutes to fill in with case-specific details. The output requires roughly fifteen minutes of review and correction, against the thirty to forty-five minutes of original drafting the task previously took.

The goal is not to remove the analyst from the process. It is to remove the blank-page problem so the analyst's time goes to review, judgment, and accuracy, which are the things the tool cannot do. Ravi's supervisor still signs the referral decision, and the summary still has to be right. What changed is where the analyst's attention is spent, not whether it is required.

Government-Specific Prompting Requirements

Government AI prompts need explicit instructions that private-sector use cases often skip. Six requirements apply to almost all analytical and drafting work in the public sector, and each one is a sentence you can paste into a prompt.

Accuracy over confidence. LLMs are trained to produce fluent, confident text, and in government an overconfident wrong conclusion is more dangerous than an honest statement of uncertainty. Include the instruction directly: "Please be very accurate. Double-check your work. If you are uncertain about any element of this analysis, say so explicitly rather than stating a conclusion you cannot support." This shifts the default. It does not make the model reliable, and a model that says it is certain has told you nothing you can bank.

No invented facts. LLMs sometimes generate plausible-sounding statute numbers, case citations, and regulatory references that are wrong or do not exist. Citing a nonexistent regulation in a memo creates legal and political problems that outlive the memo. Include: "Cite your sources. Only include information you can justify. Do not invent facts. Do not cite any statute, case, regulation, or fact that is not directly from the material I have provided. If you need to reference external law, note the reference and instruct me to verify it." The critical half is the last sentence. A returned citation can be as fabricated as the claim it supports, so asking for sources creates something to check, not something to trust.

Scope constraints. Government work carries jurisdictional and legal limits on what an output may address. A state procurement analysis should not reference federal acquisition rules that do not apply; a county benefits summary should not describe federal benefits the county does not offer. State the scope explicitly, as in "This analysis applies to Pennsylvania state procurement regulations only," which keeps the model from importing frameworks that confuse the reader or create compliance risk.

Demographic and equity awareness. Government analysis often carries obligations to consider effects across groups, and an analysis silent on differential impact may be incomplete in ways that create legal exposure. Add: "Be aware of potential bias. Treat all parties fairly. After completing the primary analysis, note any dimensions of this issue that might affect different income groups, racial or ethnic communities, or geographic areas differently. Note if you are uncertain." The instruction surfaces considerations the first pass often skips; confirming they are the right ones remains yours.

Completeness. Models optimize for a readable answer, which often means a short one. "Provide comprehensive analysis. Do not omit important considerations" pushes against that, and pairs well with a chain-of-thought structure that names the considerations you refuse to have omitted. Where a required element matters, do not rely on a general instruction; make it a numbered step.

Legal and policy compliance. Where a specific authority governs the output, name it: "Ensure output complies with [relevant law or policy]. Flag if uncertain." The bracketed placeholder is the whole point, because a model asked to comply with the law in general will produce a general answer. Naming the instrument gives it something to check against, and the flag gives you a signal about where to look first.

Iteration as a Skill

No single prompt produces a perfect output on the first attempt. Good prompts develop through a loop: write the initial prompt, get the output, assess its quality, identify what is missing or wrong, refine the prompt, and repeat. The discipline that separates strong AI users from weak ones is not writing a brilliant first prompt. It is diagnosing precisely why an output is inadequate and revising to address that specific inadequacy rather than rewriting the whole thing and hoping.

A worked example of the loop, on the simplest possible task. Version 1: "Summarize this policy." The result is vague and generic, because the prompt is. Version 2: "Summarize this policy, focusing on: 1) Who is affected, 2) What changes, 3) Timeline. Use 2-3 sentences for each section." Better, because the structure is now specified. Version 3: "Summarize this policy for government staff without technical background. Focus on: 1) Who is affected, 2) What changes, 3) Timeline. Use plain language, avoid jargon. 2-3 sentences per section." Better again, because the audience and the register are now specified too. Each revision fixed one identified problem.

Use the diagnosis to pick the fix. If the output is too generic, your prompt did not specify the level of detail required or provide a model through examples; add specificity or few-shot examples. If it misses a required element, your chain-of-thought instructions did not include that element as a step; add it to the chain. If it uses the wrong format, your format specification was abstract rather than demonstrated; replace the description with a concrete example. If it invents facts, the accuracy constraints above were missing; add them and regenerate, then verify the output anyway.

Ravi's team keeps a shared log of prompt failures and the specific fixes that resolved them. After six months, that log has become the best training resource they have, more useful than any vendor guide, because it is specific to the analysis the audit division actually produces. It also does a second job nobody planned: it is a governance artifact, a written record of how the team uses AI tools and what standards it holds them to, which is exactly what an oversight reviewer asks for.

Anti-Patterns

  • Assuming the model knows what you want. Vague requests get generically plausible responses, and the gap between what you meant and what you asked for is invisible until the output is wrong. Be explicit with structure, examples, and named requirements.
  • Single-attempt prompting. Writing one prompt, disliking the output, and concluding the tool is not useful. That was the first draft of a specification, not a verdict on the technology. Iterate against a specific diagnosis.
  • Not asking for reasoning. Accepting a bare conclusion on any task where the reasoning matters. Without the chain you cannot locate the error, and the errors that matter are the ones that read well.
  • Treating a returned citation as verification. Asking the model to cite its sources produces citations, which is not the same as producing correct ones. A fabricated reference looks exactly like a real one until someone opens it. Require sourcing, then check the sources.
  • Trusting a stated confidence level. A self-reported "High" is another generated sentence, not a measurement. Use a low rating as a triage flag; never use a high one as permission to review less carefully.
  • Reading the chain of thought as proof. A visible sequence of steps can be fluent and wrong, and it need not be the process that produced the conclusion. Check the steps against the source material rather than checking that steps exist.
  • Reusing weak examples. Few-shot examples teach their flaws with the same fidelity as their strengths, and a bad example propagates silently through every output that uses it. Review the library, date it, and replace examples when the standard changes.
  • Skipping verification because the prompt was good. Better prompting improves the base rate and makes review faster. It does not make the output correct, and for anything informing a regulatory, benefit, or enforcement decision the human check is the control that actually matters.

Practice Prompts

  • Write a chain-of-thought prompt for an eligibility determination. Pick a determination your office actually makes. List the steps a competent reviewer performs, in order, and turn each into a numbered instruction. Then run it and check whether any step you assumed was obvious was skipped, because those are the ones the model will keep skipping.
  • Design a few-shot prompt for policy analysis. Select two or three strong prior analyses, including at least one where the answer was negative or the recommendation was to decline. Strip anything sensitive, assemble them as examples, and test whether new outputs pick up the structure you intended or some other pattern you did not notice was in them.
  • Build the structured template. Take the SUMMARY, KEY POINTS, EVIDENCE, LIMITATIONS, CONFIDENCE LEVEL template and adapt the section names to a document type your team produces weekly. Run it against real documents and see which sections consistently come back thin; those are where your source material or your instructions need work.
  • Develop an iterative strategy for one use case. Take a task you do repeatedly, write version 1, and deliberately run three revisions, recording after each what specifically was wrong and what you changed. The record matters more than the final prompt, because it teaches the diagnosis.
  • Start the failure log. Create a shared document with three columns: what the prompt was, how the output failed, and what fix resolved it. Ask everyone on the team to add an entry whenever a prompt fails, then read the log as a group and look for the failures that appear more than once.

Reflection

Think about the last AI output you accepted with only a quick look. What made you comfortable? If the answer is that it read well, that is worth sitting with, because fluency is the one quality these tools produce reliably regardless of whether the content is right. Would you have caught a wrong referral threshold, a misapplied eligibility rule, or a citation to a regulation that does not exist?

Then look at your own prompts. Do they specify the reasoning steps a competent colleague would follow, or do they ask for a conclusion and hope? Do they carry the accuracy, sourcing, scope, and equity instructions, or do you rely on the model to supply that judgment? And where does your review time go now: to fixing structure that a better prompt would have produced, or to the substantive checking only you can do?

Glossary

  • Large language model (LLM): the type of AI system behind most writing and analysis assistants, which generates text as a plausible continuation of its input.
  • Prompt engineering: the practice of structuring requests to a language model so that the output is usable for a specific task.
  • Zero-shot prompting: asking the model to perform a task with no examples supplied.
  • Few-shot prompting: supplying a small number of worked examples, typically two to five, so the model learns the pattern before attempting the task.
  • Chain-of-thought prompting: instructing the model to set out its reasoning steps before stating a conclusion, so the intermediate claims can be checked.
  • Structured chain-of-thought: a chain-of-thought prompt that also fixes the sections of the response, such as summary, key points, evidence, limitations, and confidence.
  • Prompt chaining: breaking a large task into a sequence of prompts, where each output becomes input to the next step.
  • Iterative refinement: the loop of writing a prompt, assessing the output, diagnosing the specific failure, and revising to address it.
  • Prompt library: a shared, maintained collection of tested prompts and examples for a team's recurring tasks.

Closing

Chain-of-thought and few-shot are not tricks. They are two ways of writing down what you already know about a task: the order a competent person works through it, and what the finished product looks like when it is right. Most of the difficulty in prompting is that this knowledge is tacit, held by experienced staff who have never had to make it explicit. The prompt is where it becomes explicit, which is why writing a good one so often improves how the team understands its own work.

Ravi's division did not get faster because the tool got better. It got faster because a team that had never written down its analytical steps or its house format finally did, in a form the tool could follow and a new analyst could read. The referral decision still belongs to the supervisor. The judgment still belongs to the analyst. What the prompts took over was the part that was never judgment in the first place.

Key Takeaways

  • Chain-of-thought prompting makes the model reason before concluding. Structuring the prompt as explicit steps, "first assess X, then evaluate Y, then draft based on that analysis," produces conclusions that are easier to verify than a prompt asking for a direct answer.
  • A visible chain is evidence, not a guarantee. The stated steps may not be the ones that produced the answer, and fluent reasoning can be wrong at any link. Check the steps against the source rather than checking that steps exist.
  • Few-shot prompting teaches the model what good looks like. Two to five strong examples communicate format, tone, depth, and judgment standards more effectively than abstract instruction. Include negative examples and the reasoning behind each judgment.
  • Examples teach their flaws too. Whatever your samples demonstrate gets reproduced consistently and invisibly, so review the library before you standardize on it.
  • Combine both techniques for complex government tasks. Chain-of-thought governs how the problem is worked; few-shot governs what you receive. Alone, each fixes half the problem.
  • Six requirements belong in almost every government analytical prompt. Uncertainty disclosure, no invented facts with sources you then verify, scope constraints, equity awareness, completeness, and compliance with a named authority.
  • Requested sources and stated confidence are checkable claims, not assurances. A citation can be fabricated as easily as the sentence it supports, and a self-reported "High" is another generated line. Both are useful because they give you something specific to test.
  • Diagnose failure specifically, not generally. "The output was bad" leads nowhere. "The output did not assess the referral threshold" tells you exactly which step to add to the chain.
  • The goal is to move analyst time from drafting to review. Better prompting removes the blank page and the structural scaffolding. It does not remove the human check on anything informing a regulatory, benefit, or enforcement decision.
  • Shared prompt libraries build institutional capability. A team that shares its best prompts and its log of failures improves faster than individuals experimenting alone, and the library doubles as evidence of how the team uses AI and to what standard.

Frequently Asked Questions

How many examples should a few-shot prompt include? Two to five is the usual range, and the practical constraint is that every example consumes room in the prompt that your actual source material also needs. Start with two, add a third if outputs are inconsistent, and prefer variety over volume: an approval and a denial teach more than three approvals. If you are still adding examples and still chasing consistency, the problem is usually that the task is underspecified rather than underexemplified.

Can I use real case files as few-shot examples? Only under your agency's rules for the tool you are using, and the answer differs sharply between an approved internal system and a public assistant. Examples are input, and input to a public tool leaves your control. The safe default is to build examples from published or synthetic material, or to strip identifying detail so thoroughly that the example teaches structure without carrying anyone's information. Check with your privacy officer before assuming a de-identified file is safe to paste.

The output looks right. Do I still need to check it? Yes, and "looks right" is the specific condition under which these tools cause harm. Fluency is what an LLM produces reliably; correctness is what you produce. For anything feeding a regulatory, benefit, or enforcement decision, verify the figures against the source, verify any citation by opening it, and confirm that the required analytical steps were actually performed rather than merely mentioned.

Should I ask the model how confident it is? Ask, but read the answer correctly. A confidence rating is generated text produced by the same process as the analysis, so it is not a measurement of reliability. Its value is one-directional: a low or medium rating, and the explanation after "because," can point you at an assumption worth checking. A high rating tells you nothing and should never shorten your review.

What do I do when the model keeps making the same mistake? Convert the correction into a numbered step in the chain of thought, then into an example. A general instruction such as "be thorough" does not fix a recurring omission; an explicit step, "second, for each discrepancy, assess whether it exceeds the referral threshold," does. Record the failure and the fix in a shared log so the next person does not rediscover it, and revisit the log when the tool is updated.

Do these techniques transfer between different AI tools? The principles do. Being explicit about reasoning steps, supplying examples, constraining scope, and specifying format help across tools, because they address how these systems work rather than how any product works. The details do not always transfer: field names, section labels, and formatting conventions may need retuning, and a prompt that worked well before a tool update can degrade after one. Date your library entries and retest them when anything underneath changes.