←
AI for Government
Aware · M1 · lesson 1 of 31 · in progress
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
What AI Is and Is Not
📖
now learning

What AI Is and Is Not

10 min

Maria Delgado runs the benefits call center for a state Department of Human Services. One Monday her director forwarded a vendor email with the subject line "AI that ends your backlog forever." The pitch promised an "artificial intelligence brain" that would "understand every caseworker note and decide eligibility automatically." Maria's gut said something was off, but she could not say exactly what. She had forty caseworkers, a 9,000-application backlog, and a sales rep who used the word "AI" eleven times in two paragraphs. Before she could ask the right questions, she needed to know what that word actually means.

This lesson gives you what Maria needed: a plain, honest picture of what today's AI is, what it is not, and how to tell the difference when someone is selling you something or asking you to trust a result. There are no prerequisites here and no assumptions about your technical background. By the end you should be able to explain these ideas to a colleague in plain English, tell hype from actual capability, and know what you are dealing with when your agency considers a tool.

What AI actually is

Strip away the marketing and "artificial intelligence" is a family of computer techniques that find patterns in data and use those patterns to make predictions or generate content. That is the whole idea. The machine is not thinking. It is doing very fast statistics on examples it has seen before. Technically the term covers any computer system performing tasks we associate with human intelligence, which is so broad it stretches from chess programs of the 1980s to systems that write essays.

That vagueness is exactly the problem in government. When a private company's algorithm shows you an advertisement you might like, the stakes are low. When a government system determines whether someone gets unemployment benefits, denies a permit, or flags a person for fraud investigation, it touches constitutional rights, economic security and civil liberties. So when a colleague says "we are implementing AI," the only useful response is a question: what specifically? Because the word can mean any of these, and they are not interchangeable.

  • Automation that executes pre-defined rules somebody wrote.
  • Machine learning that finds patterns in data.
  • Language models that predict text and hold conversations.
  • Computer vision that analyzes images.
  • Decision systems that make or support decisions.
  • Predictive systems that forecast outcomes.

Each of these works differently, fails differently, and needs different governance. Sorting a pitch into the right box is not pedantry. It is the first move in knowing which rules apply, which failure modes to test for, and what you are actually buying. When the vendor told Maria his product had an "AI brain," what he almost certainly meant was a language model reading caseworker notes plus a scoring model ranking eligibility. Useful, possibly. A brain, no.

The family tree in plain English

You will hear several terms thrown around as if they were the same thing. They are not. Here is how they nest, from the broadest approach down to the newest and most oversold branch.

  • Machine learning (ML) is the broad approach where a system learns patterns from examples instead of following rules a programmer typed in. Show it 50,000 paid and denied claims and it learns which features tend to go with each outcome.
  • Deep learning (DL) is a more sophisticated form of machine learning built on neural networks, inspired by but very different from how human brains work. It powers most of the impressive systems in the news.
  • Natural language processing (NLP) is ML applied to human language: reading text, sorting emails, pulling names and dates out of documents.
  • Computer vision (CV) is ML applied to images: reading a license plate, flagging a crack in a bridge photo, sorting scanned forms.
  • Large language models (LLMs) are a recent kind of NLP system trained on enormous amounts of text. Widely used assistants are built on them. They predict the next most likely word, over and over, to produce fluent writing.
  • Generative AI is the umbrella for systems that produce new content: text, images, audio, code. LLMs are the text branch of it.
  • Agentic AI is the newest term. It means an AI system given the ability to take actions in a sequence: not just draft an email but send it, look something up, then update a record. More autonomy means more risk, which matters enormously in government.

How the learning actually happens

Traditional software follows rules a person wrote down: if the temperature is above 32 degrees Fahrenheit, label it "freezing," otherwise do not. Machine learning inverts that. Nobody writes the rule; the system infers it from examples. That inversion is the whole source of both the power and the danger, because a rule you wrote can be read and argued with, while a pattern the system inferred has to be discovered by testing.

It happens in three steps. First, you feed the system examples, lots of them. For a fraud detection system you might feed it thousands of past transactions, each labeled either "fraud" or "legitimate." Second, the system finds patterns, though not the way a detective would. It is locating statistical correlations, something closer to "transactions at 3 AM with amounts over $10,000 from unknown IP addresses tend to be labeled fraud in our training data." Third, shown a new transaction, it applies what it learned: this looks 71% similar to fraudulent transactions we have seen, so flag it.

Hold onto the critical point. The system is doing pattern matching on historical data. It is not reasoning. It does not understand fraud. It finds correlations, and those correlations are only ever as good as the data it learned from. If your historical labels were wrong, or skewed, or collected under a policy you have since abandoned, the system faithfully reproduces that history and presents it as a prediction about the future.

Deep learning adds layers. The name refers to neural networks with many layers that progressively transform data: a first layer might detect edges in an image, the next simple shapes, the next parts of objects, until deep in the network high-level patterns emerge, such as "this is a face." What makes deep learning powerful also makes it mysterious. Even the engineers who build these systems cannot always explain what each layer is doing. It is powerful, it is often accurate, and it is also a black box in many ways. For government, where a citizen may need to be told why a decision went against them, that opacity is genuinely problematic.

Language, vision and generation

Large language models deserve their own explanation because they are what most people now picture when they hear "AI." Here is what they actually do: they predict the next word. That is not poetic license. Give one the text "The capital of France is" and it calculates probabilities for what comes next. Paris gets 99% probability. It emits Paris, then asks the same question again of the longer string, and again, and again. Chain millions of those predictions together and you get something that holds a conversation, answers questions, writes essays and produces code.

The catch sits inside that description. The model is doing what it was statistically trained to do. It is not consulting a knowledge base. It is not looking anything up. It is using patterns from its training data, and when those patterns lead it wrong it still sounds confident while being completely incorrect. Those errors have a name, hallucinations, and they are not lying in any meaningful sense. The system is extrapolating past the edge of what it learned, with no internal signal that it has crossed that edge.

NLP is broader than LLMs. It covers identifying the main topic of a document, sentiment analysis that asks whether an email is angry or satisfied, translation and conversation. In an agency you might use it to extract key information from unstructured documents, classify incoming requests so they route to the right department, detect sentiment in constituent feedback, or translate documents between languages. Computer vision is the parallel capability for images and video: identifying whether a scanned document is a driver's license or a passport, analyzing satellite imagery for environmental monitoring, detecting equipment failure from machinery images, reviewing video from surveillance systems.

Vision has limits worth naming out loud. It can be fooled by anything it was not trained on, and its accuracy is not evenly distributed across the people it looks at. A face detection system trained on a narrow set of images may fail on people with certain skin tones if the training data was not diverse. That is not a bug someone forgot to fix; it is the training data showing through. Generative AI, meanwhile, creates new content rather than sorting existing content, and that changes the risk profile. A system that labels documents "urgent" or "non-urgent" carries one set of risks. A system that drafts a letter to a citizen carries a different set, around accuracy, bias and appropriateness of tone.

Agentic AI and the autonomy question

Agentic systems are the emerging edge: AI that does not merely respond to a request but takes actions on its own. It sets sub-goals, executes steps and reports back. Asked to streamline a vendor approval process, an agentic system might identify bottlenecks, research practices, draft new procedures and flag them for human review, deciding along the way what to do next. It is taking initiative and operating with some degree of autonomy.

This is where government needs to be especially careful. Autonomy is genuinely useful for efficiency, but government decisions require human judgment, accountability and due process, none of which an autonomous loop supplies. An agentic system that approves permits or denies benefits without human review is a governance nightmare, not because the technology is uniquely bad but because there is no longer a person in the chain who can be asked to explain, and no point at which an affected citizen can intervene.

What AI is not

This is the half that protects citizens and protects your career. Four honest statements, each one the mirror image of a promise you will hear in a sales meeting.

AI does not understand. A model that writes a flawless denial letter has no idea what a denial means to the family receiving it. It is matching patterns of words, not grasping consequences. It will produce a confident, well-written wrong answer as happily as a right one, and nothing in its output distinguishes the two.

AI is not always right, and it does not know when it is wrong. Language models invent facts, citations and policy sections that sound entirely real. A system told to summarize a regulation may cite a subsection that does not exist. It will not warn you. The fluency is the trap, because fluency is the one quality humans use as a proxy for competence.

AI is not neutral. A model learns from historical data, and history contains every bias your agency has ever recorded. If past decisions were unfair to a neighborhood or a group, the model treats that unfairness as the pattern to copy. The math launders bias into something that looks objective, which makes it harder to challenge than the original human decision was.

AI is not a decision-maker accountable to the public. You are. A citizen cannot appeal to an algorithm. When a benefit is wrongly denied, the responsibility sits with the agency and the human who signed off. "The system decided" is not a defense a court, a legislator or a journalist will accept. The most dangerous AI output is never the obvious mistake; it is the confident, fluent, plausible answer that happens to be wrong.

Those four statements are why understanding this technology is part of your job rather than the IT department's. Four questions sit under every AI decision an agency makes, and none can be answered by someone who does not know, at a basic level, what the system is doing: what can we actually ask this system to do, what are its real limitations, what happens when it fails, and who is accountable when something goes wrong. Government agencies operate under public scrutiny and accountability requirements that private companies never face, so those answers eventually get demanded in public.

Where government actually uses this

Abstractions get slippery, so hold three concrete cases in mind. The Social Security Administration processes hundreds of thousands of disability benefit claims, many requiring careful review of medical evidence, work history and current job market conditions. AI could plausibly help by scanning medical records and extracting relevant information, flagging claims very likely to be approved or denied, and organizing similar cases for batch processing. What it should not do is make the approval or denial decision without human review, because disability claims involve discretion, contextual judgment and sometimes life-changing consequences. The law requires meaningful human review.

The EPA monitors air and water quality across thousands of locations, and computer vision fits that work well: analyzing satellite imagery to detect pollution sources, identifying changes in waterways over time, detecting problems in monitoring stations. The system analyzes images and finds patterns, squarely inside what the technology does reliably. The agency still validates findings and decides what action to take, but the monitoring process gets more efficient.

The VA receives applications from veterans seeking healthcare, where NLP can extract relevant information from complex application documents, classify applications by type and urgency, and identify missing information. A language model might draft a follow-up letter requesting more detail. Should it? That letter shapes a veteran's experience with the agency, and if it is confusing or incorrect that matters. So the sensible pattern is a machine-generated draft with a human reviewing and sending, which is the same shape as almost every defensible government use of this technology.

Notice the shape all three cases share. The machine handles volume, extraction and sorting; the person handles discretion, context and the decision that carries consequences. That division is not a temporary compromise forced by immature technology. It is the arrangement that keeps a decision explainable and appealable, which is what a citizen is entitled to regardless of how capable the model becomes.

Matching scrutiny to stakes

You do not need a data science degree to ask good questions. The National Institute of Standards and Technology AI Risk Management Framework, the United States government's own guidance, boils down to one instinct: match the level of scrutiny to the level of impact. Use this sorting test on any AI tool or proposal that crosses your desk, and use it before anyone starts arguing about the technology itself.

Green: low stakes, mostly text you will review

Drafting a first version of a routine letter, summarizing a long document you will read anyway, brainstorming meeting agendas. A human reads everything before it leaves the building, and the cost of an error is a quick edit. Go ahead, with review.

Yellow: real impact, human still decides

Ranking applications by likely complexity, flagging cases for review, suggesting an eligibility outcome a caseworker confirms. Useful, but you must know the error rate, who checks the output, and how a citizen can challenge it. Pilot it. Measure it. Do not scale it yet.

Red: the AI decides something that changes a person's life with no human in the loop

Automatically denying benefits, setting bail, scoring fraud risk that triggers collections. Maria's vendor was selling Red and calling it Green. These uses demand a documented impact assessment, fairness testing, an appeals path, and senior sign-off before a single citizen is affected. Most government AI failures you will read about lived in this zone.

Four questions for any pitch

  1. What does this system actually do? Make them say "it ranks," "it summarizes," "it predicts," not "it has AI."
  2. What data did it learn from, and could that data be biased? If they cannot answer, that is your answer.
  3. How often is it wrong, and what kind of wrong? A vendor who claims "99% accuracy" with no definition is hiding something. Accurate at what, on whose population, and measured against which test set?
  4. Who is accountable when it errs, and how does a citizen appeal? If the answer is "the system," stop the conversation.

Maria printed those four questions and brought them to the vendor demo. The rep could not answer the data question or the appeals question. The "AI brain" turned out to be a keyword search with a dashboard. She did not buy it. Three months later a neighboring state did buy a similar tool, auto-denied hundreds of valid claims, and ended up in the newspaper. Maria's four questions saved her agency from being that story, and they cost her nothing but the nerve to ask them in front of a confident salesperson.

Anti-Patterns to Avoid

Every failure below starts with a reasonable person under time pressure, which is why naming them matters more than warning about them.

  • Assuming AI understands. These systems are conversational and articulate, so it feels like talking to a knowledgeable colleague. Ask a language model to interpret policy and you may get a confident, well-articulated answer that is simply wrong, and now your agency is out of compliance. A FEMA employee who asks a chatbot about disaster relief eligibility can be handed a rule that does not exist in the statute and pass it on in guidance to states. Treat outputs as drafts and starting points, never as authoritative guidance on matters of policy, law or procedure, and verify with subject matter experts and official sources.
  • Using AI on sensitive data without authorization. The tools are convenient and the risk is invisible, so people paste classified or PII-laden documents into consumer tools to get a summary. The text is transmitted to an outside vendor's servers, where the terms may allow data use for model improvement even when the conversation looks deleted. A single paste of a veteran's medical record, with SSN, address and conditions, puts that data outside government control. Only use AI systems your agency has vetted and approved, and know where the data goes and whether it is encrypted in transit and at rest.
  • Trusting AI to be factual without verification. Plausible-sounding text is what these systems are built to produce, and checking takes time nobody has. A housing agency generating press talking points can end up distributing a homelessness statistic that is off by half; an employee asking about recent court decisions can be handed a case that does not exist. Make verification a standard step: check facts against authoritative sources, look up every citation, cross-reference official records.
  • Deploying to high-stakes decisions without governance. Budget pressure, technical enthusiasm and the assumption that testing predicts production combine badly. A fraud detection system with 92% accuracy in testing meets a live population that does not match the test set, starts flagging legitimate benefits as fraud, and thousands of people lose money they are entitled to. Michigan's unemployment system flagged thousands of people as fraud perpetrators on a flawed algorithm; people lost benefits they were owed and lawsuits cost the state millions. Require human review for high-stakes decisions, test on data that matches production conditions, monitor outcomes for disparate impact on protected groups, and build a clear appeals process.
  • Selling Red as Green. Maria's vendor described a system that would decide eligibility as though it were a drafting assistant. When the described function and the claimed risk tier do not match, believe the function.

Practice Prompts

Each of these takes minutes and produces something you can show a colleague.

  • Terminology recognition. Find an article from your field about an AI implementation. Identify which category it is actually describing, then write one sentence stating what the system specifically does.
  • Limitation exercise. Take a task in your agency that people have suggested automating. What would the system need to do? What contextual judgment or human expertise would still be required? Is there room for error, and what would the consequences be?
  • Hype calibration. Find a quote about AI from a vendor pitch or industry article and rewrite it in precise language about what the technology actually does, not what it seems to do.
  • Use case assessment. Identify one repetitive document-processing task in your unit. Would language processing help? What would need to be true for it to be worth the effort and the risk?
  • Governance rehearsal. If your agency deployed a new AI system next month, who should know? Who should approve it? Who should check whether it is working fairly, and what would "working fairly" even mean?

Reflection

Pick an AI system you have heard about at work or in the news, perhaps one your agency mentioned adopting or a contractor pitched, and answer these on paper.

  • What does it actually do? Be specific: not "it helps," but "it predicts X" or "it generates Y" or "it classifies Z."
  • Which category does it fit? Machine learning, a language model, computer vision, something else?
  • What would you need to verify before trusting it with something important?
  • How accurate is it in testing versus real-world conditions, and do you know the difference?
  • What happens when it fails, and who decides whether it stays deployed?

Glossary

  • Artificial intelligence (AI). A broad term for computer systems performing tasks associated with human intelligence, covering machine learning, deep learning, language models, computer vision and autonomous systems.
  • Machine learning (ML). A type of AI where systems learn patterns from data rather than being explicitly programmed with rules. Requires training data and does not consult external sources.
  • Deep learning (DL). A sophisticated form of machine learning using neural networks with many layers. Powers image recognition and language models, but is often opaque in how it reaches a result.
  • Large language model (LLM). A system trained on enormous amounts of text that predicts the next word in a sequence. Can write, summarize, translate and code, but is not factual without verification.
  • Natural language processing (NLP). The field addressing human language understanding and generation, including classification, extraction, summarization, translation and conversation.
  • Computer vision (CV). AI's ability to analyze images and video: identifying objects, detecting patterns, reading text in images.
  • Generative AI. Systems that create new content, whether text, images, audio or video, rather than only classifying or analyzing existing content.
  • Agentic AI. Systems that take sequences of actions with some autonomy, setting sub-goals and executing steps rather than answering one request at a time.
  • Hallucination. When an AI system generates false information while sounding confident. It is not lying; it is extrapolating beyond its training data.

This vocabulary is the floor the rest of the curriculum stands on, and the next several lessons each take one piece of it deeper.

Closing

Today's work was building a common vocabulary, and that is not a small thing. Precision about what a system is lets you ask the questions that actually protect people: what specifically is this doing, what data did it learn from, how accurate is it in testing versus real conditions, what happens when it fails, and who decides whether it deploys at all. Those five questions are unavailable to anyone who only knows the word "AI."

Start using this immediately. When you hear "AI" in a meeting, ask what specifically. Use the vocabulary out loud, and help your agency describe what it is actually doing rather than what it is branding. That precision is where good governance starts, and it is the difference between Maria's outcome and the neighboring state's.

Key Takeaways

  • AI is pattern-matching, not thinking. Today's systems find statistical patterns in data to predict or generate; none of them understand the consequences of what they produce.
  • "AI" alone tells you almost nothing. The word covers automation, machine learning, language models, computer vision, decision systems and predictive systems, each with its own failure modes and governance needs.
  • Learn the family tree. ML is the broad approach; deep learning, NLP, computer vision, LLMs, generative AI and agentic AI are branches with different capabilities and limits.
  • Fluent does not mean correct. Language models predict text, cannot look anything up, and hallucinate facts and citations with total confidence, so a polished output still needs a human check.
  • AI inherits bias from history. Models trained on past decisions copy the unfairness in that data and make it look objective, and vision systems inherit the narrowness of their training images.
  • Autonomy raises the governance bar. Agentic systems act in sequence, so the rules about which decisions require human approval have to be explicit before deployment, not after.
  • Accountability stays human. A citizen cannot appeal to an algorithm; the agency and the person who signed off own every outcome.
  • Match scrutiny to stakes. Sort every use into green, yellow or red, and never deploy to a high-stakes government decision without human oversight, testing on realistic data and an appeals mechanism.
  • Four questions cut through any pitch. What does it do, what data trained it, how often is it wrong, and who is accountable when it errs.

Frequently Asked Questions

Is machine learning the same thing as AI? No. Machine learning is one approach within the larger AI umbrella, the one where a system learns patterns from examples instead of following rules a programmer wrote. The umbrella also covers plain rule-based automation, decision systems and predictive systems. When someone uses "AI" and "machine learning" interchangeably, that is usually a signal they have not looked closely at what the product does.

If a language model sounds confident, why can I not trust it? Because confidence is not a signal it computes. The system predicts likely next words from its training data; it does not consult a knowledge base, look anything up, or track whether it has left the territory it learned. When the pattern leads it wrong it produces the same fluent, assured prose it produces when correct. That is why verification against authoritative sources has to be a standard step rather than an occasional one.

My agency wants to use AI for something citizen-facing. Where do I start? Start by classifying the impact, not the technology. If a person's benefits, rights or safety are touched, expect a documented impact assessment, fairness testing, a human who reviews and owns each decision, an appeals path for anyone affected, and senior sign-off before the first citizen is affected. Then ask the four vendor questions, and be suspicious of any answer that describes a decision system as a drafting assistant.

Is deep learning better than ordinary machine learning? Better at some things and worse at one thing that matters in government. Deep networks handle images, speech and language at a level simpler methods cannot reach, but their internal workings are opaque even to the people who built them. When a citizen is entitled to an explanation of a decision, that opacity is a real cost, and it belongs in the decision about which technique to use rather than being discovered afterward.

What makes agentic AI riskier than a chatbot? A chatbot answers and stops. An agentic system chooses its own next step and acts, which means errors compound before anyone sees them and there is no natural pause where a human could intervene. Government decisions require judgment, accountability and due process, so autonomy has to be bounded explicitly: which actions the system may take alone, which require approval, and how a person affected by the sequence can contest it.