Legacy System Modernization for AI
Daniel Okafor became CIO of a state unemployment insurance agency the year after the pandemic exposed everything wrong with its technology. The core benefits system ran on a mainframe programmed in COBOL in 1987. When the agency's new chief data officer proposed an AI model to detect fraudulent claims, the demo was dazzling. Then the engineers tried to connect it. The mainframe had no API. Claimant data lived in fixed-width files refreshed once a night by a batch job nobody fully understood. The fraud model needed data in seconds; the system could deliver it in hours. Six months and $1.4 million later, the pilot was shelved, not because the AI was wrong, but because the plumbing it depended on did not exist.
This is the quiet truth of government AI. Most failed AI projects do not fail at the model. They fail at the seam where a modern model meets a legacy system that was never designed to feed it. Daniel's agency is not unusual, and neither is its mainframe. As an agency technology leader, your job is to find and fix those seams before they sink your investment, and to sequence the unglamorous plumbing work ahead of the exciting model work even when the political incentives run the other way.
Why Legacy Systems Are the Real AI Constraint
Government runs on old code. Federal agencies alone operate thousands of systems a decade or more old, many written in languages few people still learn. These systems work, which is exactly the problem; they are reliable enough to keep, but rigid enough to block everything new. The same pattern repeats at state and county level, where a single mainframe often carries eligibility, payment and case management for an entire benefits program, and where the appetite for touching it is understandably close to zero.
AI is uniquely demanding of the things legacy systems do worst. A model needs clean, well-labeled data; legacy systems store data in cryptic codes and undocumented fields. A model needs data fast and on demand; legacy systems move data in slow nightly batches. A model needs to be retrained and updated; legacy systems were built to never change. The gap is not a detail. It is the project. You do not have an AI problem so much as a data-access problem wearing an AI costume.
The federal picture is documented. In a 2019 report, Agencies Need to Develop Modernization Plans for Critical Legacy Systems, the Government Accountability Office identified ten federal systems it judged highest risk, with ages the report gave as ranging from eight to fifty-one years. The examples named there are the ones agency leaders still cite: the Internal Revenue Service Individual Master File, Treasury's Direct Express hardware and software, the Defense Department's Strategic Automated Command and Control System that ran on eight-inch floppy disks until 2019, Social Security's Title II systems, the Office of Personnel Management retirement systems, an immigration enforcement case module at Homeland Security, air traffic systems at the Federal Aviation Administration, and federal personnel pay and time systems.
These systems share a family resemblance. They are written in languages such as Assembly Language Code, COBOL, MUMPS, Jovial and PL/I, for which the available workforce keeps shrinking. Their data formats were optimized for batch processing rather than for streaming interfaces. Their identity models predate the federation protocols that modern platforms assume. And they sit inside a regulatory footprint, from the Federal Information Security Modernization Act to OMB Circular A-130 to program specific statutes, that makes every change slow on purpose. None of that is an argument against modernizing. It is an argument for planning the modernization as carefully as you would plan the AI.
What AI Demands That Legacy Systems Cannot Supply
It helps to name precisely what the model needs, because that list is what your modernization has to deliver. AI systems want low-latency read and write access, while mainframe transactions run on a schedule and inference runs per-request. They want abundant structured and unstructured data with clear provenance, aligned to the agency's evidence and data strategy obligations, while mainframe schemas were designed for storage efficiency rather than downstream analysis. Bridging those two worlds is expensive, and the bridge itself becomes a governed component under the NIST AI Risk Management Framework's manage function.
Two more demands are easier to overlook and harder to retrofit. AI platforms expect identity federation through modern assertion and token protocols such as SAML 2.0, OAuth 2.0 or OpenID Connect, not the mainframe access control products, RACF or ACF2 among them, that guard the systems of record. And they expect policy as code: inventory obligations, monitoring expectations and continuous monitoring practices are difficult to satisfy when the governed system is a batch process with no introspection hooks. If a memo requires you to show how a system is monitored, and the system cannot report on itself, you will be writing that evidence by hand forever.
Assessment Before Architecture
Before choosing a modernization path, you must honestly assess what you have. Most agencies skip this and pay for it later. Score each system that an AI initiative would touch on five dimensions, each from 1 (worst) to 5 (best).
- Data accessibility. Can other systems read this system's data through a documented interface, or only through batch files and manual exports?
- Data quality. Is the data clean, consistent, and labeled, or full of blank fields, duplicate records, and codes nobody can decode?
- Documentation. Does anyone understand how this system works, or does the knowledge live in one retiring employee's head?
- Change tolerance. Can you modify this system safely, or does every change risk breaking benefits payments?
- Security posture. Does it meet current standards, including federal cloud authorization where relevant, or is it a compliance exception waiting to be found?
A system scoring 1 or 2 on accessibility and quality is not a candidate for AI; it is a candidate for modernization first. Daniel's mainframe scored a 1 on accessibility and a 2 on documentation. Had anyone scored it before the pilot, they would have known the fraud model was premature. Be careful with the opposite reading, though. High scores across the board mean the system will not obstruct an AI initiative; they are not a finding that any particular model built on it will be accurate, lawful or fair. The score measures the plumbing, not the water.
Modernization Pathways and How to Choose
You do not have to replace a legacy system to make it AI-ready. Practitioners group the options differently, and you will hear at least two vocabularies in the same meeting: an architecture crowd that talks about encapsulating, re-platforming, re-architecting and replacing, and a federal modernization crowd that talks about encapsulate, extend, replace in place and sunset. They overlap more than they conflict. What matters is matching the option to the system in front of you.
Encapsulate: wrap it
Leave the legacy system running and build a modern interface, an API layer, in front of it. The mainframe still does the work; the API translates modern requests into something it understands. This is the fastest path to AI consumption, it preserves the system of record, and it fits the incremental delivery that federal IT management law favors. For Daniel's agency, wrapping the mainframe with an API that could serve claimant data on demand would have unblocked the fraud model for a fraction of a rewrite. The cost is honest: encapsulation does not reduce technical debt, and a clean facade over a messy interior can hide the mess from exactly the people who need to see it.
Extend: build alongside
Add a modern service next to the legacy system. New capability runs on the modern side while existing workloads stay where they are. Social Security's incremental modernization has used this shape, putting modern online access in front of core processing that still runs on COBOL. Extend is attractive when the new AI feature is genuinely new work rather than a replacement for something the old system already does, and when you can define a clean boundary between the two.
Re platform and re-architect: move it, then redesign it
Re-platforming moves the system to modern infrastructure, usually a cloud environment, with minimal code change, then improves from there. Use it when the application logic is sound but the hardware is at end of life. Re-architecting redesigns the system around modern patterns with data services and interfaces built in. It is expensive and slow, and it produces a genuinely AI-ready foundation. Reserve it for systems central to your mission that will run for the next twenty years, and phase it. A single big-bang cutover is how most large government rewrites fail.
Replace in place: the strangler fig
Gradually replace modules behind a stable facade, a pattern Martin Fowler named after a vine that grows over a tree. Callers keep talking to the same front door while what sits behind it changes piece by piece. It is slow, and it is the only one of these options that permanently retires debt rather than deferring it. Long running electronic health record and benefits modernizations have attempted this shape across multiple administrations, which tells you both that it is the right pattern and that it outlives the leaders who start it.
Sunset: retire it
Turn the old system off. This is realistic only when a successor fully serves the mission, and the risk that decides the outcome is data migration: decades of records have to move without loss or corruption, with retention obligations intact. Retirement modernization at the Office of Personnel Management has sunsetted components in phases rather than all at once, which is the pattern to copy.
Sequencing the Portfolio
No agency picks one pathway. You need a portfolio and an order: encapsulate the critical-path system that is blocking work now, extend with modern services for the new AI capability, pursue strangler fig replacement on the multiyear horizon, and sunset only when the successor has proven itself in production. The architect's real job is sequencing, and the inputs to that sequence are not architectural. They are the security categorization of the system, the readiness of the platform that will receive the workload, the availability of people in the relevant mission-critical occupations, the stability of appropriations across your multiyear capital plan, and the agency's own AI use case inventory.
That last input is the one technology leaders forget. If the modernization roadmap and the AI use case pipeline are planned by different people on different cycles, you will finish a beautifully governed API two years after the program office abandoned the use case that justified it. The way to prevent that is unglamorous: one roadmap, reviewed by both groups, with each modernization item traced to the use cases it unblocks.
Migrating a Live System Without Going Dark
The safest way to modernize a live government system is not to replace it all at once. It is to build the new system alongside the old and migrate one function at a time, retiring pieces of the legacy system as the new one takes over.
For Daniel, that meant leaving the mainframe paying benefits untouched while building a modern data service that mirrored claimant records in near real-time. The fraud model read from the modern service; the mainframe kept paying claims. Over eighteen months, more functions moved to the new service, and the mainframe's footprint shrank. Nothing ever went dark during a cutover, because there was no single cutover to fail. For government, where downtime means people miss rent, that risk profile is not optional.
State the benefit precisely, though, because incremental migration is often oversold. Running two systems in parallel removes the single big-bang failure and adds a new component you did not have before: the synchronization layer that keeps them consistent. That layer can drift, lag or silently drop records, and when it does, the symptom appears in the new system while the cause sits in the old one. Parallel running protects payments only if you test the mirror against the system of record on a schedule and treat a reconciliation break as an incident rather than a nuisance.
Data Liberation and Records Integrity
The first technical move in almost every AI-enablement program is data liberation: getting a governed copy of the data off the system of record and into a place where it can be queried, joined and served. In practice that means an offload from the mainframe into a governed data lake, a schema discovery exercise to work out what the undocumented fields actually contain, and transformation from legacy formats into modern columnar formats that analytical tools can read. None of that is glamorous, and all of it is prerequisite.
Everything in this paragraph is legally sensitive and none of it is optional. Moving data out of a system of record does not move it out of its obligations. Records must remain retained and disposable according to the applicable National Archives records schedules, including the general records schedules, and a copy in a data lake is a federal record like any other. If the source system is covered by a system of records notice under the Privacy Act, the new store and the new uses have to sit inside that notice or the notice has to be updated first. Tax return information carries its own confidentiality regime under the Internal Revenue Code, and health data carries its own. Data integrity obligations survive the copy as well: if the lake and the system of record disagree, the system of record is authoritative and you need a documented way to prove which one a given decision used. Do not let anyone tell you that a data lake is a staging area exempt from any of this.
Zero Trust as the Connective Tissue
Once legacy data is reachable, the question becomes how a model is allowed to reach it. Zero trust architecture, described in NIST Special Publication 800-207, is the connective pattern most federal programs are building toward, and CISA's zero trust maturity model gives a progression view consistent with the OMB zero trust strategy. Applied here it means something concrete: when a model needs to read from a legacy system, it does so through a policy enforcement point, with authorization decided per-request rather than per connection, with audit logging that satisfies your control baseline under NIST Special Publication 800-53, and with data transformed from legacy formats into modern schemas at the boundary rather than inside the model.
Be careful how you sell this internally. A zero trust design makes AI access to legacy data governable, and it produces much of the evidence an auditor will want. It does not by itself make you compliant with anything. Compliance is a set of documented decisions, tested controls and monitored outcomes; an architecture diagram is an input to that, not a substitute for it. The agencies that get burned are the ones that told an oversight body they had adopted zero trust and could not then show which policy decisions were actually enforced for which model.
Managing Technical Debt During AI Adoption
Technical debt is the accumulated cost of shortcuts: the undocumented field, the batch job nobody understands, the integration held together with a nightly script. AI does not tolerate debt; it amplifies it. A model trained on dirty legacy data will produce dirty decisions at scale, and now they carry the false authority of the algorithm having said so.
The discipline is to pay down the debt the AI will touch, and only that, before you deploy. You do not need to clean every record in a thirty-year-old database. You need to clean the specific data fields your model depends on, document them, and monitor them so they stay clean. Set that expectation honestly with your board: monitoring surfaces the problems you configured it to look for, on the fields you chose to watch, which is why the choice of what to watch matters more than the dashboard. Tie the work to the federal AI risk management framework's emphasis on data validity. A model is only as trustworthy as the data feeding it, and in government a bad data field becomes a wrongful denial.
Funding, Governance and the People Who Decide
Modernization needs money and political air cover, and both come through named vehicles. The Technology Modernization Fund, established by the Modernizing Government Technology Act of 2017, has funded modernization work across agencies including agriculture service delivery, background investigation systems, tax processing components and immigration services. Its awards are repaid from the savings the project generates and require a documented business case, which is a discipline worth adopting even if you never apply. Capital planning obligations under OMB Circular A-11 shape what you can propose and when, and cost estimates should follow the GAO Cost Estimating and Assessment Guide, published as GAO-20-195G, because that is the standard an auditor will measure your estimate against.
Governance is a short list of officials with real authority. Federal IT management law gives the agency chief information officer authority over IT spending and puts agency performance on a scorecard that deputy secretaries pay attention to. The chief information security officer owns the security requirements under the Federal Information Security Modernization Act and the zero trust direction. When AI is in scope, OMB M-24-10 adds chief AI officer accountability on top. Procurement runs through governmentwide acquisition contracts and agency schedules; name the vehicle in your plan, but confirm with your contracting officer which vehicles are currently open, because they consolidate and get replaced more often than architecture documents get updated.
The team shape matters as much as the funding. A modernization program needs a program manager fluent in technology business management, enterprise architects who know both the legacy and the target stack, security architects fluent in the relevant NIST control and zero trust publications, data engineers who can write extraction and transformation from Assembly Language Code or COBOL into modern analytical formats, and product managers who can write requirements as agile increments rather than waterfall blocks. Put the chief AI officer in the room from day one alongside the CIO, the chief data officer, the CISO and the human capital officer. Otherwise the program delivers a clean interface nobody uses, because the AI use case that justified it moved on while the contract was being awarded.
What the Case Files Teach
Federal modernization has a long enough record to learn from, and the lessons are about governance more than technology. Tax processing modernization has spent years unlocking master file data with partial success, and later public-facing services depended on that adjacent work. Disability claim modernization at Social Security has been incremental, with COBOL still running behind modern front ends. Electronic health record modernization at Veterans Affairs has run across administrations, and it gates clinical AI: you cannot deploy AI-assisted care coordination on a record system that is mid migration. An immigration case system moved from paper to digital and made triage automation possible at all.
The failures teach faster. An agriculture modernization program was cancelled outright and is the textbook case of modernization without governance. A major law enforcement case management rebuild was delivered late and over budget, and the governance reforms that followed shaped later department programs. A housing agency replaced single-family systems built in the 1980s and only then had data that supported modern analytics. Read these as a set and the pattern is consistent. Pattern choice, funding stability and governance decide outcomes; the technology decision is rarely the one that sinks the program.
The Modernization Decision Worksheet
For each legacy system in scope for an AI initiative, complete one row before committing budget.
- System name and mission-criticality. What it does and what breaks if it stops.
- Assessment scores. The five scores above; flag any system scoring under 3 on accessibility or quality.
- AI need. What the AI initiative actually requires from this system: real-time data, historical data, write access, or a combination.
- Recommended pathway. Encapsulate, extend, re-platform, re-architect, replace in place or sunset, with the reason.
- Records and privacy position. The applicable retention schedule, whether a system of records notice covers the intended use, and any statutory confidentiality regime that follows the data.
- Sequencing. Whether this work must finish before the AI deploys or can proceed in parallel, and which use cases it unblocks.
- Debt to retire. The specific data fields, documentation, and integrations to fix, and who owns each.
- Cost and timeline. A realistic estimate with a contingency, built to the federal cost estimating guide, because legacy work always reveals surprises.
Run this worksheet and the decision often inverts. The exciting AI project waits a quarter while the unglamorous data-access work goes first. That sequencing is not a delay. It is the difference between Daniel's shelved pilot and a fraud model that actually ships.
Anti-Patterns to Avoid
Buying the model before you can feed it. An agency signs for an AI capability, then discovers the data it needs arrives nightly in fixed-width files. The money is committed, the clock is running, and the only remaining options are bad ones. Score the systems the initiative touches before the acquisition, not after.
Treating a good assessment score as a safety finding. A system that scores well on accessibility, quality and documentation will not obstruct your model. It says nothing about whether the model is accurate or whether its decisions are lawful and fair. Teams that conflate the two skip the impact assessment because the infrastructure review came back green.
Calling parallel running a guarantee of continuity. Building the new alongside the old removes the big-bang failure and introduces the synchronization layer as a new place to fail. If nobody reconciles the mirror against the system of record on a schedule, the first sign of drift will be a wrong benefit decision, not an alert.
Presenting an architecture as compliance. Zero trust patterns, a governed data lake and a documented API are inputs to an accountable program. They are not evidence that policy was enforced for a given model on a given day. Auditors ask for the enforcement record, and a diagram cannot produce one.
Modernizing on a schedule nobody else shares. A modernization roadmap planned separately from the AI use case pipeline delivers interfaces for use cases that no longer exist. One roadmap, two owners, every item traced to the work it unblocks.
Practice Prompts
1. Score your blocking system. Pick the legacy system most likely to block your next AI initiative. Score it on all five assessment dimensions, write one sentence of evidence for each score, and name the person who would have to agree with your scores.
2. Choose and defend a pathway. For that same system, choose encapsulate, extend, re-platform, re-architect, replace in place or sunset. Write the case for your choice and the strongest case against it, including what the choice does to technical debt over the life of the system.
3. Trace the records obligations. For the data your AI initiative needs, list the applicable retention schedule, whether an existing system of records notice covers the intended use, and any statutory confidentiality regime that follows the data into a new store. Identify who in your agency signs off on each.
4. Design the reconciliation. Assume you will run new and legacy in parallel. Specify what you will compare, how often, what counts as a break, who is paged when a break occurs, and what happens to decisions made from the copy while the break is open.
5. Build the sequencing brief. Draft a one-page brief for your deputy that says which modernization work must finish before which AI use case can deploy, what it costs, and what happens if the sequence is inverted. Write it so a non-technical executive can act on it.
Reflection
Think about the last AI proposal your agency reviewed. How far into the approval process did anyone ask whether the underlying systems could actually deliver the data the model needed, and who asked it? If the answer is that nobody asked until an engineer tried to build it, that is not an engineering failure. It is a governance gap in your intake process, and it is cheap to close. Consider also which of your legacy systems you are quietly hoping to leave to your successor, and what it would cost the agency if that decision were made by a failure rather than by you.
Glossary
- Encapsulation. Placing a modern interface in front of a legacy system so other systems can read and write through a controlled boundary without changing the underlying code.
- Strangler fig pattern. Incremental replacement of a system module by module behind a stable facade, so callers see one unchanging front door while the implementation is replaced piece by piece.
- Technical debt. The accumulated future cost of past shortcuts, such as undocumented fields, unowned batch jobs and integrations that only one person understands.
- Data liberation. Extracting a governed copy of data from a system of record into a store where it can be queried, joined and served, with its records and privacy obligations intact.
- Zero trust architecture. A security model, described in NIST Special Publication 800-207, in which no request is trusted by virtue of its network location and every access decision is authorized individually at a policy enforcement point.
- System of record. The authoritative source for a given data element. When a copy and the system of record disagree, the system of record governs.
- Technology Modernization Fund. A central federal fund established by the Modernizing Government Technology Act of 2017 that finances modernization projects against a documented business case, with awards repaid from resulting savings.
- Schema discovery. The investigative work of determining what undocumented legacy fields actually contain, usually by combining code reading, data profiling and interviews with long-serving staff.
Related Lessons
Modernization sits inside a wider architecture curriculum. Designing Agency-Wide AI Platforms covers the target state that modernization is building toward, and Data Mesh and Data Fabric for Government goes deeper on how liberated data is organized once it leaves the mainframe. Data Governance for AI covers the quality and stewardship practices that keep the fields your model depends on trustworthy. AI Interoperability Across Agencies extends the interface problem across organizational boundaries, and Sovereign AI: Data Residency and National Security addresses where the data is allowed to live. For the money side, AI Infrastructure Cost Optimization pairs directly with the funding and estimating discipline described here.
Closing
Legacy modernization is the unglamorous precondition for durable agency AI. The instinct is to start with the model, because the model is what gets funded and what demonstrates well. Start instead with the data, the identity and the interface. Match the pathway to the system, fund it through a documented business case, govern it with the officials who actually hold authority, secure it with per-request authorization, and carry the records and privacy obligations through every copy you make. The agency that modernizes thoughtfully this decade will operate AI effectively. The agency that bolts AI onto unmodernized systems will spend the decade in pilots that demo beautifully and never ship.
Key Takeaways
- Most AI projects fail at the seam, not the model. The gap between a modern model and a legacy system that cannot feed it is where government AI investments die.
- Assess before you architect. Score each system on data-accessibility, quality, documentation, change tolerance, and security; systems scoring low on access and quality need modernization before AI, not AI.
- Match the pathway to the system. Encapsulate for fast access, extend for genuinely new capability, re-platform for aging hardware, re-architect or replace in place for mission-critical systems, sunset only when a successor has proven itself.
- Wrap before you rebuild, and admit what wrapping does not fix. An API layer in front of a stable legacy system often unblocks AI for a fraction of a rewrite, but it defers technical debt rather than retiring it.
- Migrate gradually, and reconcile relentlessly. Parallel running removes the single cutover failure and adds a synchronization layer that must be tested against the system of record on a schedule.
- Records and privacy obligations survive every copy. Retention schedules, system of records notices and statutory confidentiality regimes follow the data into the data lake; the copy is not a staging area exempt from any of them.
- Pay down only the debt the AI touches. Clean and document the specific data fields your model depends on, and remember that monitoring only surfaces problems on the fields you chose to watch.
- Sequencing is the real decision. The unglamorous data-access work usually has to go first, planned on one roadmap shared with the AI use case pipeline, and accepting that is what separates a shipped system from a shelved pilot.
Frequently Asked Questions
Can we just put an AI layer on top of the mainframe and skip modernization? Sometimes, and that is exactly what encapsulation is. If the legacy system is stable, the data you need is retrievable, and the only obstacle is that nothing outside can reach it, an interface layer may be all you need. What encapsulation cannot fix is data quality, undocumented fields or a system that cannot respond fast enough for the use case. Score the system honestly first; the assessment tells you whether wrapping is sufficient or merely postponing.
How do we justify modernization spending when the AI project is what leadership wants to fund? Tie every modernization item to the use cases it unblocks and present them together. A request to modernize a claims interface is abstract; a request to modernize a claims interface so that a named fraud detection capability can deploy in this fiscal year is a decision an executive can make. Build the estimate to the federal cost estimating guide so the number survives scrutiny, and document the business case, which is required for central modernization funding anyway.
Is moving legacy data into a data lake a records or privacy problem? Yes, and treat it as one from the first design meeting. The copy is subject to the same retention and disposition schedules as the source. If the source is covered by a system of records notice, the new store and the new purpose must fall inside that notice or the notice must be updated before the data moves. Data with its own statutory confidentiality regime, such as tax or health information, carries that regime with it. Bring records management and privacy counsel in before the pipeline is built, not after.
Which modernization pathway is best for AI specifically? There is no single answer, which is why the assessment comes first. Encapsulation is usually fastest to an AI outcome, extend is best when the AI capability is new work rather than a replacement, and replace in place is the only option that permanently retires the debt. Most agencies end up running all of them at once across different systems, which is fine as long as the sequence is deliberate and written down.
How do we keep a multiyear modernization alive across leadership changes? Anchor it in things that outlast individuals: a documented business case, a funding vehicle with repayment expectations, a roadmap traced to mission outcomes rather than to technology, and a governance body that includes the CIO, the CISO, the chief data officer and the chief AI officer. Programs that depend on one champion end when that champion leaves. Programs with a written sequence and a shared roadmap survive the transition, because the next leader inherits a plan rather than a rumor.
Skill.re