How Transformers and LLMs Work
Aisha Rahman had coded survey responses by hand for six years at the U.S. Census Bureau before the request landed on her desk: evaluate whether a large language model could take over the coding of open-ended answers on a household survey. The survey asked respondents to describe, in their own words, the main reason they had moved in the past year. Roughly 40,000 free-text answers per cycle, each needing to be sorted into one of 23 standard categories. A vendor had pitched a model that promised to do it in an afternoon. Aisha's job was not to be impressed and not to be dismissive. It was to understand what the system actually does, well enough to tell her division chief whether it could be trusted with citizens' survey data. To do that, she needed to open the black box. This lesson is the walkthrough she gave herself.
What an LLM Actually Is
A large language model, or LLM, is a piece of software that predicts the next chunk of text given the text it has seen so far. That is the whole job. It does not look anything up, it does not reason like a person, and it does not know what is true. It produces, one small piece at a time, the continuation that its training has made most statistically likely. Almost every system Aisha's agency was being sold, including the survey-coding tool, is built on an architecture called the transformer, introduced in 2017 and now the foundation of nearly every modern generative AI product.
Understanding that one sentence explains nearly everything that follows: why these systems are fluent, why they sometimes invent facts, why the same prompt can produce two different answers, and why a tool that codes survey responses convincingly may still be coding them wrong. It also explains why this matters for governance rather than only for engineers. Government has relied for decades on deterministic software that follows hard rules and produces the same output every time. Transformers do not work that way. They produce probabilistic outputs, they do not follow rules in the sense a case processing system does, and they sometimes generate plausible-sounding answers that are false. Without at least a conceptual grasp of the machinery, you cannot judge whether a model suits your use case, cannot tell genuine capability from confident phrasing, cannot design prompts that hold up, and cannot design oversight that catches anything.
Most of the generative AI agencies are adopting sits on this architecture. The widely used commercial assistants are transformers, and agencies are deploying them for document drafting, policy analysis, constituent response, and knowledge management. The rest of this lesson unpacks the machinery one stage at a time, in the order the model itself works through it.
How an LLM Produces One Answer: An Annotated Walkthrough
Aisha took a single real-style survey answer and traced it through the system. The respondent wrote: "Moved to be closer to my elderly mom after she had a fall." The correct category is Family reasons. Here is each of the five stages the model passes through to reach that label.
Stage 1: Tokenization, where text becomes numbers
The model cannot read letters. Before anything happens, the text is broken into tokens, which are common chunks of characters, and each token is mapped to a number. The obvious alternative, giving every word its own number, does not work: English dictionaries run past 500,000 words, new words are invented constantly, and misspellings and proper nouns would fall outside the vocabulary entirely. So tokenizers work at the subword level instead. "Running" becomes "run" plus "ning." "Unprecedented" becomes "unprece" plus "dented." Common letter sequences get a single token; rare ones are broken into smaller pieces, which is how a model that never saw "cryptocurrency" in training can still process it as "crypt" plus "o" plus "currency." Roughly speaking, 100 words of English become about 130 tokens.
Tokenization is invisible until it bites, and the places it bites are disproportionately governmental. Acronyms split unpredictably: "OMB" might be a single token or three separate letters, and "SNAP" might be one token or four. The same word tokenizes differently across spelling conventions, so "colour" and "color" are not the same sequence and a model may treat them as less related than they are. A respondent who writes "ssi" in lowercase and one who writes "S.S.I." with periods become different token sequences entirely. This is why prompts sometimes produce strange results on acronyms and specialized terms: the model is processing them as token fragments, not as meaningful units. The practical fix is cheap. Use full terms rather than abbreviations, or define the abbreviation inside the prompt. When Aisha tested abbreviations common to her survey, writing "SSI, meaning Supplemental Security Income" measurably improved consistency.
Stage 2: Embeddings, where numbers become meaning
Each token number is converted into a long list of numbers called an embedding, the model's internal representation of that token's meaning, learned during training. Tokens used in similar ways end up with similar embeddings: "mom," "mother," and "parent" sit near each other in this numerical space, while "mortgage" sits far away. This is how the model can tell that "closer to my elderly mom" is about a person and a relationship, not about geography or finance.
Stage 3: Attention, where words get read in context
This is the core innovation of the transformer, and it is worth slowing down on because it explains most of what these systems do well. When you read a sentence, you do not treat all words equally; some matter more for meaning, and some stand in special relationships to others. Consider: "The bank executive announced the merger. She was excited." You know instantly that "she" refers to the executive and not to the bank. Your brain is attending to which words matter for the reference.
Attention is the mechanism that does the same thing numerically. For each token, the model asks which other tokens in the sequence are most important for understanding it, computes a score for every possible relationship, and weights the inputs by those scores. Tokens with high attention weights influence the output more. Processing "was" in that sentence, the mechanism gives high weight to "she" and "executive," which bear on the state being described, and lower weight to "bank" and "merger." In Aisha's survey answer, when the model processes "her," attention links it strongly to "mom," not to "fall"; when it processes "closer," attention links it to "mom" rather than treating it as a real-estate term.
Attention is also what resolves ambiguity. It is how the model decides whether "bank" means a financial institution or the side of a river, based entirely on the surrounding words. Attention runs in stacked layers, a typical model having dozens, and each layer refines the picture, so that by the upper layers the model holds a context-aware representation: this sentence is about relocating for a family member's health. It is also why these models handle long documents far better than the older approach. Recurrent networks tended to forget earlier context and struggled with long-distance relationships. Attention lets the model directly connect a pronoun on line 40 to the name on line 2.
Stage 4: Next-token prediction, where the answer gets built
Picture the architecture as a stack. At the bottom, raw text arrives as tokens and the first attention layer learns which tokens are relevant to which. The second layer takes those context-refined representations and refines the relationships further, now able to attend to abstract patterns rather than mere proximity. This repeats through the stack until the top layer holds a rich representation in which word meanings are resolved by context and relationships are explicit. Sitting above that is a prediction layer that outputs the most likely next token.
Why predict the next token? Because that is the training task. The model was trained on an enormous number of examples of the form "The capital of France is" leading to "Paris," and what it learned is the general rule: given context X, the most likely next token is Y. At inference time you supply a prompt, the model predicts a token, that token is appended to the input, and the process repeats until a stop token is produced or a length limit is hit.
Concretely, the model produces a probability distribution over every possible next token at each step. If the prompt instructs it to output a category, it might assign 71 percent probability to "Family," 12 percent to "Health," 9 percent to "Housing," and spread the remainder across other options. It picks one, appends it, and repeats. The fluent paragraph you read was assembled one token at a time, each one chosen because it was statistically likely given everything before it. This is also why generation is probabilistic rather than deterministic: at each step the model samples from a distribution rather than mechanically emitting a fixed answer, and different samples at different steps produce different outputs from the same prompt.
Stage 5: The context window, or how much it can hold
The model can only consider a fixed amount of text at once. That limit is the context window, measured in tokens. Everything must fit inside it: the instructions, the survey answer, any examples you provided, and the answer being generated. Push past the limit and the earliest text falls out of view. For Aisha this mattered directly. If she sent 500 survey responses in one batch alongside a long coding rulebook, the rulebook could be pushed out of the window and quietly ignored, with no error thrown and no visible sign that the model was now working without its instructions.
Training Versus Inference: Two Very Different Moments
It is easy to conflate two completely separate phases, and keeping them apart resolves a great deal of confusion in vendor conversations.
Training happened once, in the past, at the vendor, and it comes in two stages. Pre-training exposes the model to massive amounts of text, billions of words drawn from the internet, books, and public sources. For each sequence the model tries to predict the next token, and every wrong guess nudges its internal numbers, its parameters, toward a better guess. This happens an enormous number of times, and it is where the model absorbs grammar, facts, reasoning patterns, even code, along with any biases and errors present in that text. Fine-tuning then takes the pre-trained model and trains it further on curated examples of good responses, adjusting parameters to make similar responses more likely. Instruction-following and safety behavior are added here.
Inference is what happens when Aisha actually uses the model. It is not learning from her survey answers in the moment; it is only predicting text using parameters fixed during training. That distinction has a direct privacy consequence, covered below: whether the vendor later reuses submitted text to train future models is a separate contractual question, and the default for many commercial services is not in the government's favor.
What fine-tuning offers a government agency
Fine-tuning matters for government because it means different organizations can specialize the same base model. You could fine-tune on your agency's internal documentation and processes to produce a model specialized for your operations, on your regulatory corpus to produce one specialized in compliance, or on your domain language and terminology to produce one that understands your context. The same base model could be specialized on the Census Bureau's coding manual.
It is more efficient than training from scratch because you reuse the pre-training, but it is not free. It requires curating examples of good input and output pairs, a training run measured in hours to days on suitable hardware, and validation and iteration afterward. It also requires ongoing maintenance, because a fine-tuned model reflects the data it was tuned on and that data ages. The attraction is real accuracy gains on domain-specific work. The cost is quality curated data and a commitment to keep producing it.
RLHF: Where Human Preferences Enter the Model
After fine-tuning, many organizations apply reinforcement learning from human feedback, or RLHF. It runs in four steps. First, the model generates multiple responses to the same prompt. Second, human raters rank those responses from best to worst. Third, a separate reward model is trained on those rankings to predict, given a prompt and a response, whether the response is good. Fourth, the original model is fine-tuned using reinforcement learning, with the reward model's assessment as the signal it is trying to maximize.
The purpose is alignment with human preferences. A model that is merely good at next-token prediction can still produce rambling, inaccurate, or unhelpful text if that is what was statistically likely in its training data. RLHF pushes it toward being helpful, harmless, and honest. For government this cuts both ways, and the second edge is the one to plan for. RLHF is how safety constraints get added, and it is also where new problems enter: if the raters' preferences carry biases, or if the reward signal is misaligned with the values a public agency is actually accountable to, RLHF amplifies the mismatch rather than correcting it.
The governance implication is that model behavior is not purely a product of training data. It is also a product of somebody's preferences, expressed through a ranking exercise you did not observe. Different vendors apply RLHF differently, which is why two models built from similar starting points can respond differently to the same prompt. That difference is not one being right and the other wrong. It is a difference in how each was tuned, and it is a legitimate question to put to a vendor.
Temperature: The Dial That Controls Variation
Because the model produces a probability distribution at each step, there is a setting called temperature that controls how adventurously it samples. At a low temperature near zero, it almost always picks the single most likely token, so outputs are conservative and highly repeatable. At a higher temperature, it sometimes picks lower-probability tokens, producing more varied and creative output.
For survey coding, Aisha wanted near-zero temperature: the same answer should receive the same category every time. For drafting outreach language, a colleague might reasonably want the opposite. Same model, different dial, very different reliability profile. Two cautions go with this. A near-zero temperature makes output far more repeatable; it does not make it identical run to run, and it does nothing whatever about whether the answer is correct. A confidently repeated wrong category is still wrong. And when a vendor cannot tell you what temperature their product uses, that is itself an answer about how much control you have over the system's behavior.
Why These Systems Invent Things
A hallucination is when the model produces fluent, confident text that is simply false. This is not a malfunction. It is the direct consequence of the sentence at the top of this lesson: the model predicts likely text, and the likely-sounding continuation is not always the true one. The model has no internal check for truth. It has a check for plausibility.
Aisha saw this concretely. When a survey answer was vague, "just needed a change," the model still produced a confident category with no hedging, because confident text is what its training made likely. It does not say it is unsure unless that phrasing was itself made likely. Worse, when she asked it to justify a category by quoting Census coding rules, it occasionally produced an official-sounding rule number that did not exist. The model was not lying in any human sense. It was generating text shaped like a citation because the context called for one.
The pattern generalizes badly for government work. Ask a model for a statistic and it will give you one. Suppose it states that the federal poverty level for a family of four is a specific dollar figure, with no hedging and no source. Under time pressure that is hard not to believe. But the model has not consulted any data source. If it encountered that statistic during training it may reproduce it, possibly out of date; if it did not, it will invent something plausible. Publish that number in a government document and you have a wrong poverty level shaping benefit determinations. Ask it what the current OMB guidance on AI procurement says and a model trained before that guidance was issued will happily produce a detailed, authoritative-sounding answer combining requirements from different eras with details that do not exist, and the agency that relies on it discovers the problem after the procurement.
The trap is that fluency reads as competence. A model that miscodes a survey response sounds exactly as confident as one that codes it correctly. There is no tone change, no flicker of doubt. Detecting the error requires a human who already knows the right answer, which is precisely why blind trust at scale is dangerous. Statistics, dates, percentages, direct quotes, rule numbers, and citations are the highest-risk categories, because they are exactly what the model is best at imitating and worst at getting right.
A Plain-Language Capability and Limitation Card
Aisha distilled the walkthrough into a one-page card she could hand to non-technical reviewers. It deliberately avoids jargon and states each point as something the tool can or cannot do.
| What it is genuinely good at | What it cannot be trusted to do |
|---|---|
| Sorting clearly worded text into predefined categories, fast and at scale. | Knowing whether its own answer is correct; it has no truth check. |
| Recognizing paraphrases and synonyms such as "mom," "mother," and "my mother's care." | Citing a real statute, rule number, or statistic reliably; it may invent plausible ones. |
| Producing highly repeatable output when temperature is near zero and prompts are stable. | Handling vague or contradictory input without quietly guessing. |
| Drafting first-pass language a human will review. | Making final, rights-affecting determinations without human review. |
| Following clear, explicit instructions placed inside the context window. | Remembering instructions that were pushed out of the context window. |
Three Government Scenarios
The mechanics translate into concrete deployment patterns, and each pattern carries its own dominant challenge.
Policy analysis and briefing generation, using fine-tuning. A policy analysis office spends 40 percent of its time drafting briefing documents. It holds 500 historical briefings, internal standards for structure covering executive summary, context, analysis, and recommendations, its own terminology and policy acronyms, and preferences about what counts as adequate evidence. Fine-tuning a base model on those 500 documents produces a model that understands the house structure, the vocabulary, and the evidence standard. A staffer supplies background on a topic and the model drafts in the office's style. It is not making policy decisions; it is accelerating drafting, and staffers still review, edit, and sign. The challenge is that fine-tuning requires high-quality examples: if the historical briefings contain errors or biases, the model learns them, so quality control on the training set is not optional.
Constituent service response, using RLHF and prompt engineering. A state legislative office receives hundreds of inquiries daily across licensing, benefits, permits, and complaints, and staff answer most of them by hand. Deploying a model to generate initial drafts, reviewed by staff before sending, moves the tedious part off the humans. A constituent writes: "I applied for my driver's license three months ago and haven't received it. What should I do?" The model drafts a reply noting that processing times vary by state and workload, that this office's range is 2 to 4 months, that at the three-month mark the application is likely still in process, and giving the office phone number and the service request link. A staff member reviews, perhaps adds a link, and sends. The challenge is that RLHF in this setting requires the office's preferences to be explicit. Does it prefer concise or detailed answers? Does it prioritize being maximally helpful or maximally careful? Those preferences have to drive the tuning, or someone else's will.
Technical jargon translation. A benefits agency publishes eligibility documents written in legal and program language that ordinary applicants cannot parse. A model, possibly with light fine-tuning on simplified-language examples, translates dense policy text into plain language. Original: "Applicants whose modified adjusted gross income exceeds 400% of the federal poverty level as adjusted for household composition shall not be eligible for supplementary nutrition assistance." Output: "You probably don't qualify for nutrition help if your income is more than 4 times the poverty limit for your family size. Check our website to see what the poverty limit is for your family." The percentage converts correctly, and the hedge in "probably" is doing real work. The challenge is that simplification changes meaning if it goes too far, so editors verify accuracy and oversight stays tight, particularly at first.
What This Means for Government Survey Data
The mechanics are not academic. They drive concrete legal and policy obligations for an agency like the Census Bureau.
The Privacy Act and Title 13 confidentiality. Census survey responses are protected by Title 13 of the U.S. Code, and personal records held by federal agencies fall under the Privacy Act of 1974. Because inference does not require the model to learn from the data, the question Aisha had to nail down was what the vendor does with submitted text afterward. If the commercial service retains survey responses to improve its future models, that is a disclosure of protected information, and it is not permissible by default. The control is contractual: a written commitment that submitted data is not retained, not used for training, and is processed inside an approved boundary. For federal cloud use, that boundary should be FedRAMP authorized.
The Paperwork Reduction Act. The PRA governs how federal agencies collect information from the public and requires that collection instruments be approved. Introducing a model does not let an agency change what it asks or how it treats responses outside that approved framework. If AI-assisted coding changes the methodology in a way that affects published statistics, that is a methodological change to be documented and defended, not a quiet back-office swap.
Accuracy as a statutory product, not a convenience. The Census Bureau's output is official statistics that drive funding formulas and apportionment. A coding error rate tolerable in a marketing setting is not tolerable here. Because hallucination and miscoding are silent, Aisha's recommendation was measured: use the model to produce a first-pass code plus a confidence-style flag, route every low-confidence and every rights-relevant case to human coders, and audit a random sample of the high-confidence cases against expert hand-coding on a fixed schedule.
That recommendation flowed directly from the walkthrough. Once you understand that the system predicts plausible text rather than retrieves truth, the governance design almost writes itself: ground it where you can, set temperature low, keep instructions inside the window, test on your own inputs before deploying, and put a knowledgeable human between the model and any official output.
Anti-Patterns
- Treating outputs as authoritative without verification. Fluent, coherent text sounds right, especially under deadline. But the model consulted no source; it predicted likely text. A government document publishes a wrong figure, a policy office cites an invented statistic, a compliance check rests on a fabricated legal reading. Never use model output for a factual claim without verifying it against an authoritative source. Build the workflow so the model drafts and a human checks the facts. For legal or compliance questions, consult your legal team rather than the model. Be especially cautious with statistics, dates, percentages, and direct quotes, which are exactly what models fabricate. Tell users in plain words that the output is a draft and all facts must be verified.
- Fine-tuning on biased or unrepresentative data. Fine-tuning uses smaller curated datasets, so a narrow set may not represent the range of real cases, and a set drawn from historical decisions carries whatever biases those decisions had. An agency that fine-tunes on its 500 most recent performance reviews, written over years in which supervisors were systematically harsher toward female employees, gets a model that writes harsher reviews for women. Audit training data for representation before tuning, deliberately include scenarios involving underrepresented groups, test the tuned model on diverse scenarios to see whether behavior shifts across protected groups, train any human raters in the loop to avoid bias, document what you tuned on so auditors can evaluate it, and retrain periodically on new data so old patterns do not crystallize.
- Assuming consistency and deploying without testing on your own inputs. Transformers are trained on broad internet text and handle general language well, but government language is specialized: acronyms, terms of art like non-appropriated funds or obligated authority, and document genres that are underrepresented in general training data. A model deployed to analyze compliance reports may misread "FY2024" as a technical term and produce confused summaries. A drafting tool may handle 90 percent of tasks and fail on the documents involving your specific vocabulary, so staff spend time correcting output, and may not catch every error. Test on representative examples from your own domain before deployment, deliberately include rare acronyms and unusual structures, gather feedback from the staff who will actually use it, define what error rate is acceptable for the use case in advance, and start with low-stakes drafting work before anything consequential.
- Over-trusting prompt engineering as governance. Careful prompts genuinely improve output, and teams that discover this often conclude that a sufficiently sophisticated prompt solves the problem. It does not. You cannot prompt a model out of a fundamental limitation: if the knowledge is not there, instructions will not supply it, and an elaborate prompt will not eliminate a bias baked into the weights. Prompts are also brittle across model versions. A team can spend weeks perfecting a contract-analysis prompt that works beautifully on the test set, then watch it produce worse results on contracts slightly outside that distribution, or stop working entirely after a version upgrade. Use prompts to improve output but never as the whole control. Keep human review, test on diverse cases rather than a curated set, re-test after every model upgrade, combine prompting with fine-tuning where the case justifies it, and document your prompts so they can be audited.
- Reading repeatability as correctness. A near-zero temperature makes the same prompt produce the same answer far more often. That is evidence of stability, not evidence of accuracy, and a stable wrong answer is the easiest kind to institutionalize. Several clean runs demonstrate that the format holds; they do not guarantee the content is right.
- Ignoring the context window. Instructions, rulebooks, and examples pushed past the limit are dropped silently. No error is raised and the output still looks complete, which is why long batch jobs need a check that the instructions were actually in scope.
- Assuming the tool learns from your corrections. Inference does not update the model. Corrections improve nothing unless someone deliberately incorporates them into a fine-tuning or prompt revision cycle and revalidates the result.
Practice Prompts
- Attention in your work. Think about a complex document you read recently. Identify one moment where you had to work out which words were related to which. Write a sentence or two explaining what attention meant in that context, then ask whether a model needs the same kind of attention and why.
- Tokenization impact. Find a specialized acronym or term common in your agency and run it through a public tokenizer tool. How many tokens does it take? How is the model likely to interpret it? Would defining it in the prompt help? Write a brief note on whether tokenization is affecting your results.
- Fine-tuning assessment. Does your agency have domain-specific language, formats, or standards a generic model would not know? List three to five. Would fine-tuning on your documents improve output, and what would it require in data quality, training time, and expertise?
- Hallucination vulnerability. Pick a use case your agency might deploy. Which factual errors would be most damaging? What oversight would catch them? What error rate would make the system unusable? Document this as your baseline for acceptable performance.
- RLHF and values. If your agency aligned a model with its values, what would those values be? Which responses count as good and which as bad? How would you ensure human raters understood and applied those definitions consistently?
- Context window test. Take a real batch job in your workflow, estimate its size in tokens, and check it against the context window of the model you use. Then run the job with a deliberate instruction near the front and confirm the output still obeys it.
Reflection
Take two minutes on one model you have used recently, commercial or agency-provided. What did you ask it to do, and how accurate was the output? What would have happened if you had used that output without verification? Now that you know the system predicts the most likely next token given context rather than retrieving an answer, does that change how much weight you would put on it?
Push one step further. If your answer is that you would verify everything, ask honestly whether your workflow gives anyone the time to do that at volume. Verification that exists only in principle is the same as no verification, and the number of outputs a team can genuinely check per day is a design constraint, not a matter of diligence. That constraint, not the model's capability, usually determines how much of a process can safely be automated.
Glossary
- Attention mechanism. The core innovation of transformers. It learns to weight relationships between tokens by relevance, enabling context-dependent understanding.
- Tokenization. Breaking text into subword units before processing, which lets models handle rare words, new words, and misspellings.
- Token. A subword unit of text. "Running" might be one token or two depending on the tokenizer. Models process text as token sequences.
- Embedding. The numerical representation of a token's meaning, learned during training, in which similar tokens sit near one another.
- Transformer. An architecture built on attention that processes sequences in parallel rather than word by word, enabling understanding of long-distance relationships.
- Pre-training. The first phase of training, in which a model learns patterns from massive amounts of raw text by predicting the next token.
- Fine-tuning. The second phase, in which a pre-trained model is specialized for a task or domain using curated examples.
- RLHF (reinforcement learning from human feedback). A training technique in which human raters rank outputs, a reward model learns to predict quality from those rankings, and the model is tuned to maximize that reward.
- Inference. The moment of use, when the model predicts text using parameters fixed during training. The model is not learning from your input.
- Context window. The fixed amount of text, measured in tokens, that a model can consider at once. Text pushed beyond it is dropped without warning.
- Temperature. A setting controlling how adventurously the model samples from its probability distribution. Near zero yields highly repeatable output; higher values yield more variation.
- Hallucination. Confident-sounding but false output. The model is not lying; it is predicting plausible text that happens to be inaccurate.
- Large language model (LLM). A transformer-based model trained on very large amounts of text to predict the next token.
Related Lessons
- Supervised vs. Unsupervised vs. Reinforcement Learning supplies the three learning paradigms this architecture sits inside.
- Generative AI Deep Dive widens the view from language to image, code, and video generation, and to how diffusion models differ from next-token prediction.
- Multimodal AI: Text, Image, Audio, Video covers systems that reason across modalities at once, and the failure modes that combination introduces.
- Data Quality and AI Performance is where the fine-tuning warning in this lesson becomes a full treatment.
- Systematic AI Output Validation is the practical answer to hallucination: how to check output at volume.
- Human-in-the-Loop: Design and Implementation covers where to place the knowledgeable human this lesson keeps insisting on.
Closing
The single most useful thing to carry out of this lesson is a reframing. Transformers are pattern-matching systems, not reasoning systems. They are extremely good at pattern-matching, which is why they are genuinely useful for synthesis, drafting, question answering, and explanation, all of which are pattern-completion tasks. They are not thinking, and they do not inherently know what is true. They predict the most likely continuation, and when something is statistically likely but false they will produce it with complete confidence.
That is why government needs governance around these systems, and it is not because the models are malicious or stupid. It is because their nature is probabilistic pattern-matching rather than truth-seeking, and no amount of prompting changes that nature. Understanding the architecture also tells you why vendor choice matters: different vendors fine-tune differently, apply RLHF differently, and add different safety measures, so two products can answer the same prompt differently without either being broken.
Aisha did not end up recommending against the tool. She ended up recommending it with a temperature setting, a routing rule, a sampling audit, and a contract clause, all of which she could justify by pointing at a specific stage of the walkthrough. That is what conceptual understanding buys a government analyst: not the ability to build the system, but the ability to say precisely which controls it needs and why.
Key Takeaways
- An LLM predicts likely next text, nothing more. Every behavior, good and bad, follows from this. It does not retrieve facts or verify truth; it continues patterns.
- Attention is what lets the model read words in context. It resolves pronouns and ambiguous words from surrounding text and handles long documents that older recurrent architectures could not.
- Tokenization is invisible until it breaks something. Subword splitting handles rare and new words, but acronyms and abbreviations split unpredictably. Define specialized terms in full inside the prompt.
- Generation is probabilistic by construction. The model samples from a distribution at every step, which is why the same prompt can produce different answers.
- The context window is a hard ceiling. Instructions or rulebooks pushed past the limit are silently ignored, with no error raised.
- Training and inference are different moments. The model learned in the past at the vendor; using it does not teach it, and your corrections change nothing unless someone runs a deliberate tuning cycle. Whether the vendor reuses your data later is a separate contract question, and the default often favors the vendor.
- Fine-tuning specializes a model and inherits your data's flaws. Tuning on agency documents can improve accuracy and fit, and it will faithfully reproduce any bias or error in the examples you supply.
- RLHF aligns behavior with somebody's preferences. It is how safety measures get added, and if rater preferences are biased or misaligned with public-sector values, it amplifies that.
- Temperature controls variation, not correctness. Set it near zero for repeatable, rights-relevant tasks. Repeatability is stability, not accuracy. If a vendor cannot state their temperature, treat that as a finding.
- Hallucination is structural, not a bug. Fluent confidence is not evidence of correctness. Statistics, dates, percentages, quotes, and citations are the highest-risk outputs.
- Prompt engineering is part of governance, not the whole of it. Prompts are surface-level and brittle across model versions. They do not substitute for testing, human review, or documentation.
- Test on your own inputs before deployment. General fluency does not predict performance on agency acronyms, terms of art, and document genres.
- Govern from the mechanics. Protect survey data under the Privacy Act and Title 13 with a no-retention, no-training, FedRAMP-authorized boundary; respect the Paperwork Reduction Act on methodology; and keep a knowledgeable human between the model and any official statistic.
Frequently Asked Questions
Does the model learn from what I type into it? Not during use. Inference runs on parameters fixed at training time. Whether the vendor stores your inputs and uses them to train future models is a separate question answered by your contract, not by the technology, and many commercial defaults are not favorable to government users.
Why does the same prompt give me different answers? Because generation is probabilistic. At each step the model samples from a distribution over possible next tokens rather than emitting a fixed answer. Lowering temperature toward zero makes it pick the most likely token nearly every time, which makes output far more repeatable.
If I write a really good prompt, is that enough governance? No. Prompting improves output but cannot supply knowledge the model lacks or remove a bias in its weights, and prompts are brittle across model versions. Keep human review, test on diverse real cases, re-test after upgrades, and document the prompts so they can be audited.
Why does the model make up rule numbers and statistics? Because it generates text shaped like the thing the context calls for. A citation-shaped gap gets filled with a citation-shaped string. The model has no truth check, only a plausibility check, which is why numbers, dates, quotes, and references are the outputs most in need of verification.
Should we fine-tune a model on our agency's documents? Possibly. It can make output substantially more accurate and better fitted to your formats and terminology. It also requires curated high-quality examples, a training run, validation, and ongoing maintenance, and it will learn any errors or biases in the documents you supply, so audit the training set first.
Our vendor says the model is safe because of RLHF. Is that reassuring? Partly. RLHF is how safety behavior is added, and it is also where rater preferences enter the model. Ask which preferences were encoded, who the raters were, and how the vendor checked for bias in the rankings. Different vendors tune differently, and that is a legitimate procurement question.
What happens if I paste in more text than the model can hold? The earliest text falls out of the context window. Nothing errors, and the output still looks complete, so instructions placed at the top of a long batch can be dropped without any visible sign. Check total size against the window before running volume jobs.
Can we use a commercial service for data protected by statute? Only under a contract that addresses it directly. For protected records you need written commitments that submitted data is not retained, not used for training, and processed within an approved boundary, and for federal cloud use that boundary should be FedRAMP authorized. Absent those terms, submitting protected records to a commercial service can be a disclosure.
Skill.re