The AI Timeline: Expert Systems to LLMs
Marcus Bell runs the records division at a mid-sized county clerk's office. In 1998, his predecessor bought a $40,000 "expert system" that was supposed to flag fraudulent deed filings. It worked for about a year, then a new state form broke it, and nobody could fix it because the rules lived in a consultant's head. In 2026, a vendor walked into Marcus's office promising an AI that "understands" every document it reads. Marcus had one question the vendor did not expect: "Is this the same thing that broke last time, or something genuinely different?" Knowing the answer is the difference between a smart purchase and a repeat of 1998.
You do not need to write code to make good decisions about AI. But you do need a mental map of how this technology got here, because the word "AI" has covered several very different things over seventy years. Each generation failed in its own way and got replaced for specific reasons. History also tells you which approaches worked and which were dead ends, which is the fastest way to know what hype to ignore. When a vendor or a headline says "AI," your first job is to ask: which kind, and what does it actually do?
This matters more in government than almost anywhere else, because agencies make five to ten year commitments to technology. If you adopt something that is about to become obsolete, that is wasted public money. If you refuse something that has genuinely matured, you fall behind on service delivery you are accountable for. Understanding the arc of AI development is how you invest wisely in either direction.
Four eras hiding behind one word
The story starts in the summer of 1956 at Dartmouth College, where a small group of researchers, among them Marvin Minsky and John McCarthy, coined the term "artificial intelligence" and predicted human-level machines within a generation. They were wrong by about seventy years and counting, though they were right that it was possible in principle. The field they started then moved through distinct phases. Here is the map, told through what each era could and could not do for someone in Marcus's chair.
Era one: rule-based expert systems (1970s to 1990s)
Marcus's broken 1998 system was a classic "expert system." A human expert sat with a programmer and wrote down hundreds of IF-THEN rules: if the grantor name is blank, then flag for review. Medical versions worked the same way: if patient has fever and cough, suspect flu. The computer simply followed the rules. Governments used early AI of this kind for military targeting systems, medical diagnosis assistants, and administrative automation. It was the first serious attempt to encode human expertise in a machine, and inside a narrow, stable domain it genuinely worked.
It was also brittle. Change the form, and the rules no longer matched. Nobody could update them without the original expert. The systems could not learn from new examples; they only knew what someone had typed in by hand. A medical system built on 1980s knowledge was wrong by 1990 as new conditions emerged, because knowledge moves and hand-coded rules do not. That brittleness is exactly why Marcus's system died when the state changed a form.
The first AI winter (1970s to 1980s)
Expectations had been set far too high. Expert systems were expensive to build and expensive to maintain, and when they failed to deliver, funding dried up. Entire research areas became unfashionable, and AI spent years widely regarded as a failed technology. This is worth holding onto as you listen to current claims. AI has boomed and busted before, and the pattern of inflated expectation followed by disappointment followed by realistic assessment has repeated more than once. Recognising it does not mean dismissing the current wave. It means evaluating maturity separately from enthusiasm.
Era two: classic machine learning (1980s to 2010s)
The next leap was teaching computers to find patterns in data rather than following hand-written rules. Instead of telling the machine the rules for spam, you showed it ten thousand emails labeled "spam" or "not spam," and it learned the pattern itself. This is "machine learning." Early on, progress was real but slow, because computing power was limited and training datasets were small by today's standards. The internet changed that: it created massive datasets at the same time as more computing power became available, and machine learning systems started winning at spam filtering, recommendations, and image recognition.
Governments adopted these systems widely for fraud detection, pattern analysis, and decision support, including predicting which tax returns to audit or which bridges needed inspection first. The constraint was the shape of the problem. Machine learning needed clean, labeled data and a narrow, well-defined question. It could rank audit risk; it could not read a contract and summarize it.
Era three: deep learning (2010s)
Around 2012, "deep learning," meaning machine learning built on neural networks with many layers, suddenly got very good at images and speech. The ingredients were the networks themselves, massive computing power in the form of GPUs, and massive datasets. Image recognition abruptly worked. Computers played Go at superhuman levels. Machine translation improved dramatically. For government this powered passport photo matching, call-center transcription, and a wave of document processing and predictive analytics. It still needed huge labeled datasets and did one task at a time. A face-matching model could not also answer questions about policy.
Era four: large language models (2017 to now)
In 2017 a research design called the "transformer" changed the economics. It let models learn from enormous amounts of ordinary text without hand-labeling. The result is the "large language model," or LLM, the technology behind tools like ChatGPT, Claude, and Microsoft Copilot. An LLM is trained to predict the next word across a huge slice of human writing, and out of that simple goal it gains a striking ability to draft, summarize, translate, answer questions, and write code. ChatGPT launched in late 2022 and made this visible to everyone at once.
That was an inflection point, but be precise about why. LLMs are not fundamentally a different kind of thing from earlier machine learning; they are still finding patterns in training data. What changed is scale and breadth of capability. Suddenly AI was not a specialist tool bought for one narrow job. It could do broad tasks, and agencies began exploring use cases immediately. This is what the 2026 vendor is selling Marcus, and it is genuinely different from the 1998 system: it was not hand-coded, and it can handle forms it has never seen.
Since then, roughly 2024 onward, LLMs are everywhere and nearly every technology vendor has added "AI" to its product line, usually meaning some form of LLM integration. Government is exploring cautiously: chatbots for citizen service, document drafting, research assistance, analysis. The hype cycle is in full swing. Some claims are credible. Many are not, and telling them apart is the job.
Which parts are actually new
The single most useful thing the progression gives you is the ability to tell which pieces of a "new" technology are genuinely new. Almost every product now sold to agencies carries an AI label, and in most cases that label means a language model has been attached to a feature that already existed. That is not dishonest and it is often exactly what you want. But it changes the question you should be asking, from "is this AI" to "which part of this is the AI, and what is it doing that the old version could not?"
Run that test against the eras. If the answer is that a fixed set of rules now runs faster, nothing has changed except speed, and the maintenance problem of era one is still yours. If the answer is that patterns are learned from your own historical data, you are in era two or three, and your exposure is the quality and the bias of that history. If the answer is that a general-purpose language model is reading and writing text, you are in era four, and your exposure is fluent wrongness plus the difficulty of explaining any single output.
Each of those three answers implies a different control. Rules need an owner and an update path. Learned patterns need testing against the population they will be applied to. A language model needs human verification of anything that leaves the building. Sorting a product into the right era before you write requirements is what stops you buying a control that does not match the failure mode you actually face.
What genuinely changed, and what did not
Here is the honest answer Marcus needs. The new system is more capable and far more flexible than the one that broke. It will not shatter the first time the state tweaks a form. That is real progress, and it is the reason this era is worth taking seriously rather than dismissing as another cycle of noise.
But three old problems did not go away, and one new problem appeared.
- It still depends on data. Every era, including this one, is only as good as what it learned from. An LLM trained mostly on general web text may not know your county's specific deed rules.
- It can still be confidently wrong. The 1998 system failed loudly, in that it stopped working. The 2026 system can fail quietly, producing a fluent, professional-sounding answer that is simply false. This is called a "hallucination," and it is the single most important risk for government use.
- It still needs a human accountable for the output. No era of AI ever removed the need for a person to own the decision. A deed gets rejected on the clerk's authority, not the machine's.
- The new problem: you often cannot see why it decided something. Hand-written rules were auditable line by line. LLM reasoning is far harder to inspect, which matters enormously when a citizen has a right to know why government said no.
One more thing did not change, and it is the least intuitive. Context decay is permanent work. A system that was brilliant in 1985 could be useless in 2025 because the world around it moved, and the same is true of a model deployed today. Updates, retraining, and revalidation are not a project phase you finish. They are an operating cost you carry for as long as the system runs, and they belong in the budget from the first year.
A timeline-aware vendor screening tool
Use this short checklist the next time someone pitches "AI" to your office. It turns the history above into a set of questions that reveal what you are actually buying, and each row maps to a failure that some earlier era of AI actually suffered. Write the vendor's answers in the right column and keep it in the procurement file, because a written answer is a commitment in a way a demo never is.
| Question to ask the vendor | What a weak answer sounds like | What a strong answer sounds like |
|---|---|---|
| Which kind of AI is this: fixed rules, pattern-learning, or a language model? | "It's just AI, very advanced." | "It's a large language model fine-tuned on permitting documents." |
| What data was it trained on, and does it know our specific forms and rules? | "It knows everything." | "General training plus your last 5,000 filings; here's the data sheet." |
| How often is it confidently wrong, and how do we catch that? | "It doesn't make mistakes." | "Here is our measured error rate on documents like yours, and here is the human-review step." |
| Can we get a record of why it made each decision? | "The model is proprietary." | "Each decision logs the source documents it used." |
| When a form or rule changes, who updates it and how fast? | "Call us, we'll schedule it." | "You update a settings file; no consultant required." |
Notice the last row. The reason Marcus's 1998 system died was the answer in that bottom-left cell. If the new vendor gives the same answer, he is buying the same risk in a shinier box. Notice the third row too: a vendor who will not put a number on how often the system is wrong is either not measuring it or not willing to tell you, and both answers should change what you are willing to sign.
What might come next, and what probably will not
Experts disagree about the next decade, and anyone who tells you otherwise is selling something. Informed speculation points in a few directions. Agentic systems, meaning AI that takes initiative, sets its own sub-goals, and acts rather than only answering, would raise governance questions that current oversight structures were not built for. Multimodal systems that handle text, images, video, and audio together would be more powerful and correspondingly harder to govern.
Two other directions are less dramatic and more likely to touch your budget. Specialized models trained on single domains such as law, medicine, or finance may prove more accurate than one general model while being less flexible, which suits the narrow, well-defined problems agencies actually have. And efficiency improvements matter quietly, because today's models are computationally expensive to run, and anything that delivers similar work for less compute changes what a small office can afford without changing what it can do.
The list of things that will probably not arrive soon is just as useful in a procurement conversation. General AI, meaning a system as capable and flexible as a person across all domains, is not imminent, and experts disagree about whether it takes decades, centuries, or never happens. Super-intelligence, a system beyond all humans at all tasks, is further away still. Neither belongs in a business case you have to defend.
Two more items on that list deserve saying plainly. AI without bias is not coming: no system removes bias entirely, it only changes which biases are present, which is why testing for disparate outcomes is permanent rather than a launch task. And fully autonomous government, with AI deciding without human oversight, is a governance choice rather than a technical inevitability. Nobody will ever be forced into it by the technology, and any vendor who implies otherwise is describing a decision your agency gets to make.
Five lessons the history actually teaches
Strip away the dates and the arc leaves a handful of durable lessons that apply directly to how you evaluate and buy AI today.
- Hype cycles are real. AI has boomed and busted before. The current enthusiasm will almost certainly produce both real successes and expensive failures, and being hot is not the same as being ready for high-stakes use.
- Slow and steady beats magic. The most successful AI systems are not magical. They are careful applications of well-understood techniques to well-defined problems, which is also what makes them defensible when questioned.
- Context changes the value. A system brilliant in one decade can be useless in the next once the surrounding facts move. Updating and retraining are perpetual work, not a one-time project.
- Governance matters more than capability. Well-governed government AI succeeds and poorly governed ambitious AI fails dramatically, and the difference has repeatedly turned out to be organisational rather than technical.
- Specialization beats generalization. A system built for document classification outperforms a general system told to "classify documents." Depth beats breadth, and narrow scope is easier to test, explain, and defend.
Why this moment matters for government
Three forces have lined up at once. The technology crossed a usefulness threshold around 2023, when LLMs became good enough to draft real work. The cost dropped sharply, so a small county office can now use tools that once needed a national lab. And federal guidance arrived: the White House issued government-wide direction in 2024 requiring agencies to inventory their AI uses and protect the public from harmful ones. That means AI is no longer an optional experiment for many agencies. It is something you will be expected to use responsibly and account for.
For Marcus, the takeaway is not "buy" or "don't buy." It is that he can now have the right conversation. He knows the 2026 tool is a real advance over 1998, that its biggest danger is quiet wrongness rather than loud failure, that whatever he buys will need maintenance for as long as it runs, and that a human in his office still owns every deed decision. That clarity is worth more than any single product.
Anti-Patterns
Each of these is a way the history gets ignored, and each has a specific correction from the eras above.
- Assuming hype equals readiness. A technology being fashionable says nothing about whether it is mature enough for a decision that affects someone's benefits or permits. Evaluate maturity independently of enthusiasm, using evidence about performance on work like yours rather than the market's general excitement.
- Failing to plan for obsolescence. Government commits to systems for five to ten years, and an AI investment sized for that horizon may be outpaced well before it ends. Build flexibility, exit options, and data portability into long-term plans instead of assuming today's architecture survives.
- Forgetting past failures. "We learned our lesson from expert systems" gets said sincerely and forgotten immediately, every time a new approach promises to solve everything. Study your own agency's history with previous systems before declaring the current one different.
- Treating a language model like a rules engine. Expecting deterministic, identical output every time from a system that works by prediction leads teams to build controls that do not fit the failure mode. The risk is a plausible wrong answer, not a crash, and the control for that is human review of substance.
- Buying general when you need specific. A broad tool bought because it can do everything usually does your particular job worse than something scoped to it, and it is far harder to test and explain when a citizen asks why.
Practice Prompts
These exercises turn the timeline into something you can use in your own office this month.
- Hype recognition. Pull up current marketing from any AI vendor selling to government. Mark each claim as credible, unverifiable, or hype, and write the question you would ask to move each unverifiable claim into one of the other two columns.
- Timeline extrapolation. Based on the arc in this lesson, write down where you think AI will be in 2030. Note which of your predictions rest on evidence and which rest on the current mood, then keep the note.
- Governance comparison. Describe how expert systems were governed in the era they were deployed and how LLMs are governed in your agency now. What changed, what did not, and which gap worries you most?
- Failure study. Pick one AI approach that failed, such as expert systems or early neural networks, and write a paragraph on why it failed. Then name the specific practice that would stop you repeating it.
- Department history. Find out whether your agency adopted an earlier AI or automation system. What happened to it, who maintained it, and how does that story change the questions you would ask a vendor today?
- Run the vendor screen. Take the five-question table above to your next demo and record the answers in writing. Any question the vendor cannot answer on the spot is your follow-up list.
Reflection
Take a few minutes with this one. Think about a technology decision your office made years ago that you now regard as a mistake. Was the mistake the technology itself, or was it the assumption that the world around it would hold still? Most failures in this history were the second kind. Now apply that to whatever AI your agency is considering. If the model, the vendor, or the surrounding rules changed without warning, who in your organisation would notice, and what would they be able to do about it? If the honest answer is nobody and nothing, that is the finding, and it is more useful than any opinion about the technology. Then ask yourself where you sit personally: do you lean toward dismissing this wave because you watched one fail, or toward accepting it because the demonstrations are impressive? Knowing your own lean is what lets you correct for it.
Glossary
- Expert system: An AI system that encodes human expert knowledge as explicit logical rules, popular in the 1980s and 1990s and dependent on someone maintaining the rules by hand.
- First AI winter: The period in the 1970s and 1980s when AI fell out of favour and funding collapsed after early systems failed to meet inflated expectations.
- Machine learning: Learning patterns from data rather than hand-coding logic, which emerged as the alternative to expert systems and removed the dependence on a single human expert.
- Deep learning: Machine learning using neural networks with many layers, which enabled major capability improvements in images, speech, and translation starting in the 2010s.
- Transformer: The research design introduced in 2017 that let models learn from large amounts of ordinary text without hand-labeling, making today's language models practical.
- Large language model (LLM): A system trained on enormous amounts of text to predict and generate text, the currently dominant approach and the technology behind most tools now marketed as AI.
- Hallucination: A confident, fluent output that is simply false, which is the characteristic failure mode of language models and the main reason government output needs human verification.
- Inflection point: A moment when the rate of change accelerates. The arrival of widely usable LLMs in 2022 and 2023 was an inflection point in both capability and adoption.
- Hype cycle: The pattern of inflated expectations followed by disappointment followed by realistic assessment. AI has moved through this cycle more than once.
Related Lessons
This lesson gives you the chronology; three others give you the substance underneath it. What AI Is and Is Not draws the boundary between what the current generation can genuinely do and what it is only claimed to do, and How AI Actually Works explains the prediction mechanism that makes hallucination a structural feature rather than a bug to be patched. Types of AI Systems maps the categories a vendor may be selling under one label, which is the distinction the first row of the screening table depends on. For where the technology stands in agencies right now, AI in Government Today covers current adoption, and Government AI Policy Landscape takes up the rules that arrived alongside it.
Closing
Marcus did not become a technologist over the course of an afternoon, and he did not need to. What he gained was the ability to place a sales pitch on a seventy-year map and see which promises had been made before and how they turned out. That is the practical value of this history: it converts a conversation you cannot evaluate into one you can. Your next step is small and specific. Find out what automated or AI-driven systems your office already runs, when they were last updated, and who would fix them if a form changed tomorrow. The answer to that question tells you more about your agency's readiness than any vendor demonstration ever will.
Key Takeaways
- "AI" names several different technologies. Rule-based expert systems, classic machine learning, deep learning, and today's large language models each work differently and fail differently, so always ask which one is in front of you.
- AI has boomed and busted before. The first AI winter followed inflated expectations and expensive, brittle systems, which is why current enthusiasm should be evaluated separately from current capability.
- Old systems were brittle; new ones are flexible but fuzzy. Hand-coded rules broke when the world changed; language models adapt but can produce fluent, confident, false answers.
- The transformer in 2017 was the real inflection point. It let models learn from raw text without hand-labeling, which is why today's tools can draft, summarize, and answer across tasks rather than doing one job.
- LLMs are a change of scale, not of kind. They are still pattern-matching from training data, which is exactly why they inherit the old dependence on what they learned from.
- Hallucination is the defining government risk. A polished wrong answer is more dangerous than an obvious crash, because nobody questions it.
- Auditability got harder, not easier. When a citizen has the right to know why government decided something, demand decision logs and traceable sources before you sign.
- Maintenance is permanent. Context moves, models drift, and retraining is an operating cost rather than a project phase, which matters when the commitment runs five to ten years.
- A human still owns every decision. No era of AI removed accountability; it stays with the public servant who signs off.
- Screen vendors with history-aware questions. Which kind of AI, what data, how often wrong, can you audit it, and who maintains it when rules change.
Frequently Asked Questions
If LLMs are just pattern matching like older machine learning, why did they feel like such a sudden change? Because scale changed what the pattern matching could do. Earlier systems needed hand-labeled data and a narrow question, so each capability had to be built separately. Training on vast quantities of ordinary text removed the labeling bottleneck and produced one system that handles many tasks. The underlying idea is continuous with what came before; the breadth of usable output is not.
Should I expect another AI winter? Nobody can tell you that, and treat confident predictions in either direction with suspicion. What history supports is narrower: enthusiasm has previously outrun delivery, funding has previously collapsed when it did, and the current wave will likely contain both genuine successes and expensive failures. The practical response is to buy things that still make sense if the market cools, which usually means narrow scope, portable data, and short commitments.
Does the history mean expert systems were a mistake? No. They worked well in narrow domains with stable knowledge, which is a real category of problem that still exists. The mistake was expecting them to hold up when the knowledge moved and nobody was funded to maintain the rules. That failure mode is not extinct; it just wears different clothing now.
How do I tell a genuine AI product from a relabeled one? Ask the first question in the screening table and insist on a specific answer. Many products now marketed as AI are conventional software with a language model attached to one feature, which may be entirely appropriate for your need. The problem is not relabeling; it is paying for capability you are not receiving, and a vendor who describes the architecture plainly is telling you what you are buying.
Is agentic AI something my agency should be planning for now? Planning, yes; deploying, that depends entirely on what it would touch. Systems that take initiative and act rather than answer raise oversight questions that most current governance structures were not designed for, so the useful work today is deciding in advance what an automated system in your agency would never be permitted to do without a human authorising it. That decision is a governance choice, and it does not become inevitable because the technology arrives.
Where does this leave the 1998 system in Marcus's story? As the most useful thing in the room. It gave him a concrete failure to test the new promise against, which is why his question to the vendor was better than most. If your agency has its own version of that story, it is an asset rather than an embarrassment, and it belongs in the room when the next pitch happens.
Skill.re