Types of AI Systems
At a Tuesday staff meeting, a city's permitting director announced she wanted "AI" to fix the backlog. Around the table, four people heard four different things. The records clerk pictured the chatbot on the state's website. A planner imagined a tool that writes permit summaries. The IT lead thought of the fraud-flagging model the finance office already used. And a new analyst, fresh off a tech podcast, was picturing an autonomous agent that would email applicants and update the database on its own. They argued for forty minutes and bought nothing, because they were never talking about the same kind of system.
"AI" is not one thing. It is a family of very different tools with very different risks, costs, and rules. If you cannot tell them apart, you cannot scope a project, write a requirement, or judge a vendor's pitch. This lesson gives you a working map of the main types of AI systems and shows where each one already lives inside government.
Why naming the type decides everything else
Some systems do exactly one thing, such as recognizing tax forms. Others write essays, debug code, and answer questions on nearly any topic. Some wait to be asked. Others take initiative and set their own intermediate goals. When somebody says "we're adopting AI," the only useful first question is: what kind? Because the type determines what the system can do, what it cannot do, and what governance it requires.
The landscape divides along three dimensions worth holding in your head. Scope asks how many domains the system can operate in. Capability asks what kinds of task it can accomplish. Autonomy asks whether it responds to requests or takes initiative. Those three questions sort almost any system you will meet, and each answer changes the oversight the system needs.
The stakes of getting this right are practical. A narrow classifier that sorts routine documents is relatively low risk when a human reviews its output. An agentic system that restructures government processes on its own authority is a governance problem before it is a technology problem. Naming the type is the first step toward appropriate oversight, and skipping it is how agencies end up buying a tool that solves a problem they do not have.
The big divide: narrow, general, and the hypothetical
Narrow AI, sometimes called weak AI, does one specific task. It sorts emails, reads license plates, predicts which water main might fail, or transcribes a hearing. It excels inside its domain and is utterly useless outside it. It has no idea what it is doing and cannot do anything else. Every AI system in government today, without exception, is narrow AI. The plate reader cannot suddenly decide to write your budget, and a medical AI system asked to classify document types will simply fail.
Narrow AI is genuinely good for government, and for an unglamorous reason: it is easier to understand, test, and control. A system trained to do one thing has one purpose, and you can measure whether it serves that purpose well. A general-purpose chatbot can seem general because it handles many language tasks, but it remains a text-prediction system trained on internet data. Ask it about quantum mechanics and it will generate plausible-sounding text that may be completely wrong.
General AI, also called strong AI or artificial general intelligence, would understand concepts across domains the way humans do, learn from experience, and transfer knowledge from one domain to another. It does not exist. Researchers are working on it. Some think it is decades away and some think it is impossible. When someone says AI will replace all human workers, they are imagining general AI: a system that can learn any human job. That is not where we are. We are in the narrow AI era, and each system is a specialist.
Super-intelligence, a hypothetical system smarter than humans at all tasks, is science fiction territory right now. It is worth naming only because you will hear it in policy discussions, usually in the form "we need to prepare for super-intelligence." Long-term thinking is fair enough. It does not affect today's decisions about the systems your agency is actually deploying, and treating it as if it did is a reliable way to spend a meeting badly.
The practical implication is simple. Do not worry about general or super-intelligence. Evaluate the specific capabilities and limitations of the narrow systems in front of you. And when a vendor's slide hints that their tool "understands" or "thinks," translate that in your head to "this is a narrow tool with good marketing." If anyone tells you their government AI is generally intelligent, you have learned something important, just not about the AI.
Classification systems: sorting and scoring
The oldest and most common government AI. A classification system takes an input and puts it in a bucket or gives it a score. Is this email spam or not? Is this transaction normal or suspicious? Which of nine departments should this 311 complaint go to? How likely is this bridge to need inspection this year? Is this case high risk or low risk?
Classifiers are relatively mature and they work. They suit government for three specific reasons. They are comparatively interpretable, so you can often understand why the system reached a given output. They fit government workflows, because human staff make the actual decision while the system supplies input. And they are falsifiable: you can test whether the classifier is accurate and prove it wrong when it is.
Their core risk is bias. Because they learn from historical data, they can quietly carry forward the patterns in that data, including unfair ones. A classification system that scores benefit applications can absorb and repeat past discrimination if nobody checks. Interpretability helps here but does not cure it: a system whose reasoning you can read can still be reasoning from a biased history.
Where it already lives in government: fraud detection in benefits and tax, routing of citizen service requests, predictive maintenance for infrastructure, risk scoring across many programs, document classification in records management, and flagging high-priority cases for faster processing.
Generative AI: producing new content
This is what most people now mean by "AI." A generative system, built on a large language model, produces new text, images, audio, video, or code in response to a prompt. It can draft a memo, summarize a 90-page report, translate a notice into Spanish, or answer a citizen's question in a chatbot.
Generative systems are newer and trickier for government for four reasons. They are less interpretable, since you cannot easily explain why a piece of writing came out the way it did. They hallucinate, producing confident, fluent, completely false statements. They create new liabilities: who is responsible when a generated government letter is inaccurate? And they require strong human oversight, because the output needs review before it goes anywhere public.
The line that follows is worth writing down. A system that generates a first draft of a briefing paper, which a human then edits and approves, is good practice. A system that autonomously generates and sends government communications without human review is not. It does not know facts; it predicts likely words. That makes it excellent for drafting and dangerous for any task where being wrong has consequences and nobody checks.
Where it already lives in government: public-facing chatbots, drafting routine correspondence, summarizing public comments, translating documents, and helping staff search internal policy.
Reactive and agentic: how much autonomy the system has
A reactive system responds to inputs. You give it a request, it processes it, it returns a result. It does not set its own goals or take initiative. Almost all AI systems today are reactive: the model you ask a question answers your question, the classifier you submit a document to classifies that document. It cannot decide on its own to go reorganize the filing system. That is why reactive systems sit comfortably in government: the human still controls the agenda, decides what to ask for, and carries the responsibility.
An agentic system has autonomy. Given a goal, it decides how to pursue it. It may break the goal into sub-goals, build plans, and execute steps without asking permission at each one. This is the newest and least mature family, and the one the analyst in the meeting was picturing. The promise is real. Imagine a system that takes an incoming permit application, checks it against zoning rules, requests a missing document by email, and schedules a review.
The risk is real too, and larger than the others, because when a system acts on its own its mistakes act as well. An agent that misreads a rule does not merely give a wrong answer. It sends a wrong letter, denies a wrong application, or moves real money. Consider the version a vendor will eventually pitch you: "your goal is to reduce permit processing time, here is your budget and authority, execute." The system analyzes processes, identifies bottlenecks, proposes procedures, implements changes, and reports back. It is taking initiative, making decisions, and executing plans.
For government that raises questions current policy mostly does not answer. Who approves the sub-goals? What authority does the system hold? When must it escalate to a human? What happens, and who is accountable, when it makes a decision that harms someone? Most government AI policy today assumes reactive systems. Agentic systems will require new oversight mechanisms, and until those exist the honest answer about how much a system should decide alone is: usually nothing that touches a citizen.
Where it is starting to appear in government: mostly pilots and internal tools, such as automating multi-step back-office workflows with cautious, narrow scopes.
Rule-based and learning-based: how the system works inside
A rule-based system executes human-defined rules. If temperature is below 32 degrees Fahrenheit, classify as freezing. If income is below X, approve for benefits. If risk score is above Y, flag for review. Rule-based systems are interpretable, because you can read the rules, and they apply the same logic every time. Note carefully what that consistency does and does not buy you: it makes the system predictable, not correct. A wrong rule is wrong every single time, at full scale, with perfect consistency. They are also brittle, working only where someone anticipated the situation, and failing when a case arrives that no rule covers.
A learning-based system, which is machine learning, learns patterns from data. You do not write rules; you provide examples and the system finds patterns. These are more flexible and can handle situations you did not anticipate. They are also less interpretable, to the point where even the designers often cannot explain a specific decision. And they depend entirely on good training data: biased examples produce a biased system.
Most modern AI systems are learning-based because they are more powerful, which creates the explainability problem government keeps running into. The better implementations combine approaches. Rule-based logic for decisions that are well understood and stable, such as whether an application clearly meets a stated eligibility threshold. Learning-based components for nuanced prediction. Human judgment for edge cases and high-stakes decisions, which is where both approaches run out of road.
A usable artifact: the AI type decision map
Bring it together in one reference. When a project lands on your desk, identify the type first; everything else follows from it.
| Type | What it does | Signature risk | Oversight needed | Government example |
|---|---|---|---|---|
| Classification | Sorts or scores inputs into buckets | Bias from historical data | Bias testing; human review of high-stakes calls | Fraud flags, 311 routing |
| Generative | Creates new text, images, or code | Confident false output (hallucination) | A human checks every consequential output | Chatbots, drafting, summaries |
| Agentic | Takes multi-step actions toward a goal | Mistakes become actions, not just answers | Strong limits; human approval before real actions | Early back-office workflow pilots |
| Rule-based | Applies human-written rules exactly | Brittle; a wrong rule is wrong consistently | Review the rules themselves; a path for cases no rule covers | Threshold checks in eligibility screening |
| General | Would match humans at any task | Does not exist; a marketing red flag | None; treat the claim with skepticism | None |
Five systems, classified
Tax form classification. Optical character recognition combined with a document classifier. Given an image of a tax document, it outputs Form 1040, Form 1099, Form W2, Amended Return, and so on. It is narrow, because it classifies tax documents and nothing else. It is reactive, responding to a submission. It is learning-based, trained on examples. The result is faster mail sorting and initial triage, with human staff still reviewing and processing. It speeds the work; it does not replace judgment.
Welfare fraud detection. A system trained on historical claims, both approved and later found fraudulent, predicts whether a new claim is likely fraudulent or likely legitimate. Narrow, reactive, learning-based. It flags cases for human investigation, and a caseworker does the actual investigation and determination, so the system prioritizes work rather than deciding it. The governance requirements attached to it are not optional: regular bias testing, an audit that the system is not disproportionately flagging claims from particular demographics, and an appeals process for people whose claims were flagged.
Environmental monitoring. Satellite imagery combined with computer vision, identifying coastal erosion, illegal dumping, and deforestation. Narrow, since it looks for specific phenomena. Reactive, processing imagery as it arrives. Learning-based, trained on labeled imagery. The result is alerts for human scientists to investigate, at a scale that would be impossible manually.
Benefits chatbot. A language model trained on government benefits policies and frequently asked questions. Citizens ask questions and the system answers. This one is relatively broad, covering benefits across many programs. It is reactive, generative, and learning-based. The governance challenge is that it may generate incorrect information, so the controls are human review of generated answers before deployment, regular audits of the common questions and their answers, and a clear disclaimer telling citizens to verify with official resources.
Procurement automation. A hypothetical system that, given a procurement need, researches vendors, drafts solicitations, manages the bidding process, and recommends winners, with the goal of reducing procurement time and cost. It is broad, potentially agentic, and hybrid, using rules for defined processes and learning for recommendations. The governance questions pile up fast: who approves major decisions, what the escalation process is, how fairness is ensured, and how vendors will respond to being evaluated by a system. The current status is that this is mostly ruled out for core decisions without heavy human involvement. It is plausible for administrative acceleration such as scheduling meetings with potential vendors, while humans make the real decisions.
Back at the permitting meeting
This map would have ended the argument in five minutes. The honest answer to the backlog was a mix: a generative tool to draft permit summaries with staff review, and a classification tool to route applications to the right reviewer. The autonomous agent was a year too early, and the governance to run one safely did not exist in that building. Naming the types turned a forty-minute argument into a scoped, fundable plan, and it also produced the requirements document, because each type came with its own oversight obligations already attached.
Anti-patterns
Deploying agentic capabilities without governance. A system that takes initiative makes consequential decisions without a clear human approval process. It happens because it seems efficient: the system can identify problems and fix them, so why wait for review? What goes wrong is invisible until it is not. The system adjusts staffing or resource allocation on its own analysis, makes a decision that harms a specific group by reducing service in a particular neighborhood, and nobody notices until after the fact. Then nobody is accountable, because "the system decided."
Define clear boundaries before deployment: what the system may decide autonomously, which in government is usually nothing; what requires human approval, which is most decisions; and what requires external approval, which is anything affecting citizens. Build escalation paths and monitor what the system actually does.
Assuming a classifier is objective. Using a learned classifier as though it were neutral, without checking for bias or unfairness. It happens because a mathematical model looks like mathematics. What goes wrong is that a classifier trained on biased historical data perpetuates that bias while wearing mathematical authority, and staff start trusting the algorithm over their own judgment. A hiring classifier trained on past hiring decisions that reflected discrimination will learn to replicate the discrimination and systematically disadvantage the same groups. Audit classifiers for fairness, test performance across demographic groups, compare the system's decisions against human expert decisions, and keep humans in the loop for high-stakes choices.
Publishing generative output as an official statement. Releasing AI-generated letters, guidance, or official statements without human review. It happens because it is fast and the output looks professional. What goes wrong is that official-sounding but inaccurate guidance reaches citizens who act on it, and legal problems follow. In one case an agency automated correspondence using a language model, which generated a response about benefits eligibility that sounded authoritative and was factually wrong, and citizens relied on it. Require human review for all external communications, especially anything that looks official or could shape a citizen's decisions. Use generative systems for drafting, not for publishing.
Practice prompts
- System classification. Take an AI system your agency uses or is considering and classify it on every dimension: narrow or broad, generative or classification, reactive or agentic, learning-based or rule-based. Write down what each answer implies for governance.
- Governance drafting. For each type of system, draft the governance requirements. Who reviews the output? When? Which decisions require human approval before anything happens?
- Bias test design. If your agency deployed a classification system, how would you test it for bias? What would fair performance look like in numbers, and how would you know if the system was working unfairly?
- Agentic boundary setting. Imagine an agentic system aimed at one of your agency's processes. What should it be allowed to decide independently, what requires human approval, and what requires public notice?
- Type evolution. Systems that start narrow face pressure to expand. Pick one in your agency, describe how that pressure would arise, and name the safeguard you would want in place before the scope grew.
Reflection
Think of a routine decision your agency makes: approving something, denying something, prioritizing something. Could an AI system help, and if so, what type would actually be appropriate? A classification system flagging cases for human review? A generative system drafting correspondence? Something else entirely, or nothing at all? Write the answer as a sentence naming the type, not as a sentence naming "AI."
Then write the second half, which is the half people skip. What safeguards would you want in place before that system touched a real case? Who reviews, on what schedule, with what authority to switch it off? If you can answer both halves, you have written the beginning of a requirements document. If you can only answer the first, you have written a wish.
Glossary
- Narrow AI. A specialized system that performs well in one domain and is useless outside it. All AI systems today are narrow.
- General AI. A hypothetical system that understands and learns across domains the way humans do. It does not exist.
- Super-intelligence. A hypothetical system smarter than humans at all tasks. Discussed in policy circles; not a factor in today's deployment decisions.
- Classification system. A system that assigns inputs to categories, such as approve or deny, urgent or routine.
- Generative system. A system that creates new content: text, images, code, audio, or video.
- Reactive system. A system that responds to inputs without taking initiative of its own.
- Agentic system. A system with autonomy that sets sub-goals, makes decisions, and takes action without explicit human direction at each step.
- Rule-based system. A system that executes human-defined rules. Interpretable and predictable, but brittle.
- Learning-based system. A system trained on examples to find patterns. Flexible but less interpretable.
- Hallucination. Fluent, confident output from a generative system that is factually false.
Related lessons
- What AI Is and Is Not sets the boundary between what these systems are and what the marketing claims, which this lesson subdivides.
- How AI Actually Works explains the shared mechanism underneath every type on this map, including why learning-based systems resist explanation.
- What AI Does Well and Where It Fails takes the next step: matching a specific task to the capability it needs.
- AI in Government Today shows these types running in real agencies, with the deployments that worked and the ones that harmed people.
- The Human in the Loop develops the oversight column of the map into a working practice rather than a policy sentence.
Closing
You can now classify any AI system you encounter, and the classification tells you which questions to ask. For a narrow classifier: is it accurate in its domain, and is it biased? For a generative system: what human review process exists, and has it been tested for accuracy? For an agentic system: what are the approval boundaries, and what happens when it fails? Different types are different jobs of oversight, and asking a generative system's questions about a classifier will leave you comfortable and unprotected.
Every system fits into these categories, and each category brings its own strengths and its own governance challenge. That is the whole value of the map. It does not tell you whether to buy something. It tells you what you are buying, which is the fact that everything else in the procurement depends on and that the four people at that Tuesday meeting did not have.
Key Takeaways
- "AI" is a family, not a thing. Scope, capability, and autonomy are the three dimensions that sort it, and the risks, costs, and rules differ sharply across the family.
- All government AI today is narrow AI. It does one task with no understanding. General AI that matches humans broadly does not exist, and any claim that it does is a marketing red flag.
- Classification sorts and scores, and its risk is bias. It is interpretable, falsifiable, and fits government workflows, but because it learns from history it can repeat past unfairness unless someone tests for it.
- Generative AI drafts and summarizes, and its risk is confident falsehood. Use it for drafting, not publishing, and put a human review step in front of anything external.
- Agentic AI acts, which makes its mistakes act too. It is the least mature family, most policy assumes reactive systems, and in government the list of decisions it should make alone is usually empty.
- Rule-based means predictable, not correct. Consistent logic applies a wrong rule just as consistently as a right one, and it fails outright on situations nobody anticipated.
- Identify the type first; everything else follows. The right oversight, the right requirements, and the right vendor questions all flow from knowing which kind of system you are buying.
Frequently Asked Questions
A vendor says their system is "AI-powered." What do I ask next?
Ask which type, using the three dimensions. Is it narrow or broad in scope? Does it classify inputs or generate new content? Is it reactive or does it take actions on its own? Is it rule-based, learning-based, or a hybrid? Every answer changes the oversight you need to specify in the contract, and a vendor who cannot answer these plainly has told you something useful about the product.
Is a classification system safe because it is interpretable?
Interpretability helps and it is not a safety guarantee. You may be able to read why the system scored a case the way it did, and the reasoning can still be inherited from biased historical data. A classifier trained on past hiring decisions that reflected discrimination will replicate that discrimination legibly. The controls that matter are fairness testing across demographic groups, comparison against human expert judgment, and human review of high-stakes calls.
Should my agency be using agentic AI yet?
The current practice in government is narrow, cautious pilots on internal back-office workflows, not anything touching a citizen. The blocker is not capability, it is governance: most existing policy assumes reactive systems, so the approval boundaries, escalation paths, and accountability rules an agent needs mostly have not been written. Before an agent acts, someone has to answer who is accountable when it acts wrongly, and that answer cannot be the system.
Are rule-based systems old-fashioned?
They are older and still the right tool for a real class of problems. Where a decision is well understood and stable, such as checking whether an application meets a stated threshold, a rule you can read and audit beats a model nobody can explain. The strong implementations are hybrids: rules where the logic is settled, learning where prediction is genuinely useful, and human judgment for edge cases and anything high-stakes.
Do I really need to worry about general AI or super-intelligence?
Not for any decision you will make this year. They belong in long-term policy conversations, and they are useful mostly as a diagnostic: a vendor invoking them to sell you a permitting tool is telling you how they expect the conversation to go. Spend your attention on the narrow systems your agency is actually deploying, where the risks are documented and the controls are known.
Skill.re