←
AI for Managers
Aware · M3 · lesson 3 of 26 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

How Generative AI Works

16 min

Marcus Delacroix manages a 12-person operations team at a regional logistics company. Three months into using an AI assistant, he stormed into a peer's office genuinely irritated. "This thing is broken," he said. "I asked it the founding date of one of our carriers and it gave me a confident answer that turned out to be completely wrong. Then in a long planning chat it forgot a constraint I'd typed an hour earlier. And when I asked the same question twice, I got two different answers." His peer asked one question back: "Do you actually know how the thing works under the hood?" Marcus didn't. By the end of the afternoon, every behavior that had felt like a malfunction made sense to him. None of it was broken. It was all working exactly as designed. That shift in understanding is what this lesson delivers.

What This Lesson Covers

You do not need to understand neural networks, linear algebra, or calculus to use AI well. You do need an accurate mental model of how generative AI produces its output. Generative AI is software that creates new text (and increasingly images, audio, and code) rather than just retrieving stored answers. Once you understand the mechanism, the behaviors that frustrate most managers stop being mysteries. You will know why AI sometimes invents facts, why it loses track of long conversations, why the same prompt gives different answers, and why it has no idea about anything that happened after its training cutoff.

This lesson covers five things: how the model is trained, how it generates text one piece at a time, what a context window is and why it limits long conversations, why responses vary, and why confidently wrong answers (hallucinations) are a structural feature rather than a bug. Throughout, we follow Marcus as each concept resolves one of his complaints.

How the Model Learned: Pattern, Not Knowledge

Models like ChatGPT, Claude, and Gemini were trained by being shown enormous amounts of text: books, articles, websites, code. Training is not like teaching a person. Nobody explained concepts to the model. Instead, the model played one game billions of times: given a stretch of text, predict the next piece. When it guessed wrong, its internal settings were nudged. Repeat that across trillions of words and the model becomes extremely good at predicting what plausibly comes next in any piece of writing.

The critical insight is that the model did not memorize sentences and it does not "know" facts the way a person does. It learned statistical patterns about language, namely which words tend to follow which, which topics cluster together (the word "hospital" sits near "doctor," "patient," "surgery"), how arguments are typically structured. This is why it can write something it never saw before, and also why it has no built-in sense of whether what it writes is true.

For Marcus, this explained the wrong founding date. The model was not looking up a record. It was generating a plausible-sounding date based on patterns in how founding dates appear in text. Plausible is not the same as true. That gap is the root of most AI surprises.

How It Writes: One Token at a Time

When you ask an AI to write an email, it does not compose the whole thing and hand it over. It builds the text one token at a time. A token is roughly a word or a word-piece, and for practical purposes, think "about three-quarters of a word." "The quick brown fox" is about four tokens. One token is roughly four characters of English. Tokens matter because everything the AI does is measured in them, including the limits we will get to next.

Here is the generation loop in plain terms. You give a prompt. The model reads every token of it, then calculates probabilities for what the next token should be. Asked to "Write a brief email requesting a meeting," it might assign "Dear" a 45% chance, "Hi" 20%, "I" 8%, and so on across thousands of candidates. It picks one, appends it, and then repeats the entire calculation with the new word now part of the context. It keeps going until it generates a special stop signal that means "this looks complete."

Because each pick is a probability, not a lookup, the model deliberately does not always take the single most likely word. That small amount of controlled randomness is what makes its writing sound natural instead of robotic, and it is the direct reason Marcus got two different answers to the same question. Same prompt, different probabilistic draws.

The Context Window: Why It "Forgets"

The AI has no memory in the human sense. It does not carry anything from one conversation into the next, and within a single conversation it can only "see" a limited amount of text at once. That visible amount is called the context window, and it is measured in tokens.

In a chat, each new message is processed together with as much of the prior conversation as fits in the window. Early on, the model can reference everything you said. But once the conversation grows past the window, the oldest text falls out of view, and the model literally cannot see it anymore. It is not choosing to forget; the text is simply no longer in front of it.

Context windows have grown large. As of this lesson, some models offer windows around 200,000 tokens (roughly 150,000 words), others around 128,000 tokens; older models had only 2,000 to 8,000. Even a large window has an edge. The practical move: for anything important, restate it rather than assume the model still remembers it from forty messages ago. This is exactly what bit Marcus: the constraint he typed an hour into a sprawling planning chat had scrolled out of the window by the time he referenced it.

Why the Same Prompt Gives Different Answers

The controlled randomness in token selection is tuned by a setting called temperature. At a low temperature (near 0), the model almost always picks the single most likely next token, so output is consistent and predictable, sometimes even repetitive. At a higher temperature (near 1), it more often picks less-obvious tokens, so output is more varied and creative but less repeatable. Most consumer AI tools default to a middle setting that balances natural-sounding writing against reliability.

So when Marcus asked the AI to summarize the same meeting three times and got three slightly different summaries, nothing was wrong. If he needs consistency, he has options: ask the model to pick and tighten the best of several drafts, request a fixed format that constrains how much can vary, or, in tools that expose the control, lower the temperature.

Hallucinations: Confident, Plausible, and Sometimes Wrong

A hallucination is when the AI generates false information while sounding completely sure of itself. Given everything above, this should now feel predictable rather than alarming. The model's entire job is to produce plausible next tokens. When you ask for a fact it does not reliably "have" (a niche company's founding date, a specific regulation in a small jurisdiction, a detail about your own company's internal process), it still produces something plausible-sounding, because that is all it ever does. Plausible and true only line up when the underlying pattern in the training data happened to be correct and well-represented.

It is not lying; lying requires intent. The danger is precisely that the output reads as confident and complete, because that is how human writing usually reads. A model that says "I am not sure" is actually safer than one that fluently invents. Hallucination risk spikes in three zones: anything recent (after the training cutoff), anything specialized and thinly represented online, and anything proprietary to your organization that was never in the training data at all. Marcus's carrier founding date hit two of those, niche and specialized, which is exactly why it came out wrong.

What the Model Did and Did Not Learn

The model genuinely learned language patterns, topic associations, the shape of reasoning and explanation, and how to vary tone from formal to casual. What it did not learn, or learned only unreliably, matters just as much: whether the facts it discusses are actually true, anything after its training cutoff date, your organization's proprietary details, true cause-and-effect (it learned correlation, not causation), and your personal preferences unless you state them. It also inherits the biases of its training data, which skews toward English-language, recent, and widely published sources. That is not the model "deciding" to be biased; it is the model reflecting the text it was shown.

Worked Example: Marcus Vets a Vendor Brief

A week after his crash course, Marcus had to brief his director on a potential new carrier partner. He asked the AI to draft a one-page summary covering the carrier's founding, service regions, and the relevant transport-safety regulations. The draft came back polished and confident. This time Marcus ran each claim through a simple five-question checklist before trusting any of it:

  1. Is this a factual claim or creative text? The founding date and the regulation citations are factual. The "why this partnership fits our network" paragraph is interpretive. Factual claims need verification; the interpretive framing he can judge himself.
  2. Is it recent? The regulation reference cited a "2024 amendment." His model's training cutoff predated part of 2024, so anything that recent is suspect.
  3. Is it specialized? Regional transport-safety rules are niche. High hallucination risk.
  4. Is it proprietary? The draft confidently described "your company's standard onboarding timeline." His company's process was never in the training data, so that detail was invented and he deleted it.
  5. Can I verify it? He checked the founding date against the carrier's own filings (the AI's date was off by two years) and pulled the actual regulation text from the regulator's site (the "2024 amendment" did not exist).

The numbers tell the story. Of the eight factual claims in the draft, three were wrong. The draft still saved him roughly 30 minutes of structuring and writing, but had he forwarded it unverified, he would have walked into his director's office with a 37% error rate on facts. The lesson is not "do not use AI for this." It is "use AI to draft and structure, then verify every factual claim in a high-stakes or specialized document." That single habit converts a risky tool into a reliable accelerator.

Putting the Model to Work

Match the model size to the task. Bigger models know more and reason better but cost more and run slower. Smaller, faster models are often perfectly good (sometimes better grounded) for narrow, well-defined jobs. You do not always need the largest model in the menu.

Do not overstuff the context window. Dumping an entire codebase, the requirements doc, the project status, and the style guide into one chat leaves less room for the model to focus on your actual question, and quality drops. Give it what the current task needs.

Restate important constraints. The model will not carry a rule from a previous conversation, and within a long one it may lose early instructions over the window edge. Repeat what matters.

Explain the mechanism to your team. When a team member is frustrated that "the AI keeps making things up," a one-line explanation ("it predicts plausible text statistically; it does not look up truth, so we verify facts") builds the right kind of healthy skepticism faster than any policy memo.

Four Anti-Patterns Worth Naming

Most AI frustration among managers traces back to a small set of beliefs that sound entirely reasonable and are wrong. Marcus held all four at one point or another.

  • "I will run this project through one long chat so the AI stays consistent." It can reference earlier messages while they are still inside the window, but generation remains probabilistic, so even a repeated question can come back phrased or weighted differently, and the earliest messages eventually drop out of view altogether. If you need consistency, do not rely on the conversation to hold it. Write decisions down explicitly and restate them: "Based on our earlier decision to ship in March, what is next?"
  • "I will paste in everything so it has full context." The entire codebase plus the requirements document plus the current status plus the design guidelines fills the window and leaves less room for the model to work on the thing you actually asked. Output quality drops rather than rises. Feed it the specific requirement, the specific material, and the specific question.
  • "It should know our process, I explained it once." The model does not learn from your conversations and carries nothing into the next one, and your internal process was never in the training data to begin with. Proprietary context has to be supplied every time: "Here is how we handle escalations. Given that, what should we do?"
  • "The AI said this is how the regulation works, so that is how it works." It learned how regulations are written about, not what any particular law says. Use it to draft, explain, structure, and analyze, then confirm the facts against an authoritative source before anyone acts on them.

Judgment Checkpoints Before You Trust a Claim

The five questions Marcus ran over his vendor brief are not specific to vendor briefs. They work as a standing checklist for any AI output that makes assertions about the world.

  • Is this factual? Is the model asserting something about objective reality, or producing creative, interpretive, or structural text? Only the first kind needs verifying, which is what keeps the checklist quick.
  • Is this recent? If the claim concerns anything after the training cutoff, the model has no basis for it and will still answer with complete confidence.
  • Is this specialized? Niche domains are thinly represented in the training data, and thin representation is exactly where fluent invention lives.
  • Is this proprietary? Anything specific to your own organization was never in the training data. If you did not supply it in the prompt, the model generated it.
  • Can I verify it? Is there a source you trust that settles the question, and will you actually go and check it?

If the answer to that last question is no, it does not mean you cannot use the output. It means you treat it as a draft or an exploration rather than a conclusion you are willing to put your name on.

Talking About AI Honestly With Your Team

Three habits are worth carrying into how you discuss AI with the people you manage.

Treat AI errors as mechanism, not carelessness. When the model gets something wrong, it did not rush, get lazy, or make a typo. It produced the most plausible continuation it could from the patterns it learned. Framing errors that way sets realistic expectations, and it saves your team from the futile move of telling the model to "try harder" or "be accurate this time" as though effort were the missing ingredient.

Be specific about the limitation. Vague warnings produce either fear or nothing at all. Marcus found that one precise sentence did more than any policy memo: the tool is excellent at mimicking how we write, but it does not understand our situation the way we do, so we verify anything factual. That is the kind of statement people can actually act on, and it produces appropriate skepticism instead of blanket distrust.

Remember that better does not mean reliable. Newer and larger models are genuinely better at most things, and it is tempting to conclude that the verification habit can relax as they improve. It cannot. Better is not the same as trustworthy for every task, particularly in the recent, specialized, and proprietary zones. Match the model to the job and keep the checking discipline regardless of which model you are using.

Practice and Reflection

These five questions turn the mental model into working knowledge. Answer them about your own work rather than in the abstract.

  • Build token intuition. Take a task you want AI to help with and roughly estimate the volume involved, using the rule that a typical page runs a few hundred tokens. Does the whole task comfortably fit inside a modern context window, or are you at risk of pushing the early material out of view partway through?
  • Find your highest hallucination risk. What is one question you would plausibly ask AI where a confident wrong answer would do real damage? Why is that topic risky (recent, specialized, proprietary, or more than one)? What would verification actually look like, and how long would it take?
  • Inventory what the model cannot know. List the knowledge about your industry, your organization, or your team that was never in any training data. How will you supply that context in a prompt when you need it, and what happens if you forget?
  • Plan your context management. For a long-running project, how would you structure your AI conversations so critical constraints do not slide out of the window? Consider where you would keep the decisions of record and how you would restate them.
  • Define your verification standard. For one task you are considering handing to AI, write down what checking the output means in practice and which sources you would trust to settle a disputed fact. If you cannot name a source, that is a signal about how much weight the output can carry.

Key Takeaways

  • The model predicts patterns; it does not know facts. Everything it produces is a statistically plausible continuation of text, which is why plausible-sounding output can still be false.
  • Text is generated one token at a time, probabilistically. A token is roughly three-quarters of a word, and the small controlled randomness in each pick is exactly why the same prompt can yield different answers.
  • The context window is a real, hard limit. The model only sees a fixed amount of recent text; beyond that edge it cannot reference what you said earlier, so restate anything important in long conversations.
  • Variation is by design, controlled by temperature. Lower temperature means more consistent output; higher means more creative. Differing answers to the same prompt are expected behavior, not a fault.
  • Hallucinations are structural and concentrated in three zones. Recent, specialized, and proprietary topics carry the highest risk of confident, fluent, wrong answers.
  • Verification is non-negotiable for facts. Run a quick checklist (factual or creative? recent? specialized? proprietary? verifiable?) and treat anything you cannot verify as a draft, not a conclusion.
  • The model has no memory between conversations. It learns nothing from one chat to the next, and it has no access to your organization's internal information unless you provide it in the prompt.