←
AI for Government
Strategic · M44 · lesson 44 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Technical Debt Management in AI
📖
now learning

Technical Debt Management in AI

15 min

Tomas Reyes is the chief AI officer at a state revenue department. Three years ago, under pressure to modernize, his teams stood up eleven different AI systems fast: a chatbot here, a document classifier there, a fraud scorer, a forecasting model. Each launched on time. Each was a small win. Then the bills came due. Two models had been quietly degrading for a year. One ran on a data pipeline only a departed contractor understood. The chatbot used a vendor model that was being retired in ninety days. And nobody could produce a single inventory of what was running, on what infrastructure, trained on what data.

Tomas was not facing eleven projects. He was facing eleven piles of technical debt: shortcuts that felt free at launch and now charged compounding interest. Technical debt is the hidden cost of choices that made something faster to build but harder to maintain, and the metaphor is a loan. You borrow speed today and repay it later with interest, in the form of fragile systems, rework, and risk. Ordinary software carries technical debt. AI systems carry all of that plus several kinds unique to AI that most leaders never see coming. This lesson, built around Tomas's eleven systems, gives you a way to find AI technical debt, measure it, pay it down, and design so that it stops accumulating.

Why AI Debt Behaves Differently

In ordinary software, debt is code that works but is not maintainable. You took shortcuts, and you pay interest in bugs, slowness, and difficulty adding features. The code itself is static: unless somebody changes it, a system that works today generally works tomorrow, and testing catches most problems before production. That stability is what makes conventional technical debt tolerable. It gets worse slowly, and it gets worse only when you touch it.

AI systems break all three of those assumptions. Models change as the data around them changes, so a system nobody has touched can degrade on its own. Data quality degrades over time whether or not anyone is looking. Testing is harder because performance metrics are probabilistic rather than binary, so there is no green build that means the model is correct. And one broken data pipeline can silently corrupt months of output before anyone notices, because the system keeps producing confident answers the entire time it is wrong.

Consider the shape the source describes. An agency has a model that has been in production for two years. It started at 95 percent accuracy. It is now at 87. Nobody knows exactly why. The data scientist who built it has left, the training code is undocumented, and the model draws on data from three systems in formats nobody fully understands anymore. Nothing in that story involves anyone making a mistake after launch. That is what makes AI debt distinctive: the interest accrues without anybody borrowing again.

Government portfolios have a particular vulnerability to this pattern, and it is a staffing pattern rather than a technical one. Much AI delivery runs through contractors and term-limited technical staff, so the person who understands a pipeline is frequently on a contract that ends. Tomas's forecasting model illustrates the whole mechanism: the system did not fail when the data changed, it failed when a person left, because the documentation that would have made the departure survivable was the shortcut taken to hit the launch date. Knowledge that lives only in a person is debt with a resignation date attached.

The Debt That Only AI Carries

Naming the forms of debt is the first step, because each behaves differently and each needs a different fix. Five forms recur across government AI portfolios, and the vocabulary matters mainly because it lets you tell a budget office precisely which one you have and what it will cost to service. A request for engineering time to address dependency debt on a ninety-day vendor clock is a fundamentally different conversation from a request to document a pipeline, and collapsing both into a plea to work on technical debt is how neither gets funded.

  • Data debt. The pipelines feeding the model are undocumented, fragile, or dependent on one person. Tomas's forecasting model effectively died the day a contractor left, because the data flow lived in that person's head.
  • Model drift debt. A model degrades silently as the world changes, and nobody is monitoring. Two of Tomas's models had been getting worse for a year. The debt is the gap between knowing it worked at launch and having no answer to whether it still works.
  • Dependency debt. Your system relies on an external model or service you do not control. When the vendor retires the model, as Tomas's chatbot vendor was doing in ninety days, your system breaks on someone else's schedule.
  • Documentation debt. No record of what data trained the model, what it must not be used for, or how to retrain it. Without that, every fix begins with archaeology.
  • Reproducibility debt. You cannot rebuild the model from scratch. If it breaks, you cannot recreate the version that worked, because the data, the settings, and the code were never captured together.

These compound rather than add. A model with drift debt and no monitoring, running on a pipeline with data debt, documented nowhere, is not three problems. It is one system that will fail in a way nobody can diagnose, at a moment nobody chose. In ordinary software, technical debt is bad code. In AI, it is also a model quietly getting worse, a pipeline only one person understands, and a vendor about to withdraw a service, none of which announce themselves until they come due together.

Where the Debt Actually Sits

The five forms above describe how debt behaves. A second and complementary view describes where it lives, and that view is more useful the moment you start assigning remediation work to actual teams. Debt accumulates in the data, in the model, and in the infrastructure, and in most agencies those three layers are owned by three different groups with three different reporting lines and three different backlogs. A remediation plan that does not say which layer it is addressing will be read by each group as somebody else's responsibility, which is a reliable way to have nothing happen.

Data. The model's output is only as good as the data behind it. Data debt shows up as stale training data that no longer matches the world, missing features that would materially improve performance, quality problems including nulls, duplicates, and internal inconsistencies, systematically biased data that favors one group over another, and undocumented datasets nobody can trace to a source. The consequences are degraded performance, unfair outcomes, and outputs that make no sense with no obvious explanation for why.

Model. The model accumulates its own debt: no documentation of how it works or why it is structured as it is, no systematic tests of its performance or behavior, hard-coded magic numbers for thresholds and weights, no versioning so nobody is certain which version is actually in production, and undocumented dependencies on preprocessing, feature engineering, and other models upstream. The result is a system that is hard to debug, hard to retrain, and hard to improve, which in practice means it does not get improved.

Infrastructure. This layer is the one most often left out of AI debt conversations, and it is where outages come from. The symptoms are no monitoring, so nobody knows whether the model is performing in production; fragile serving infrastructure that is slow or crashes; no documentation of how to deploy a new model version; manual retraining, deployment, and rollback; and tight coupling, where a change in one system breaks another. The consequences are deployment times measured in weeks rather than days, outages when anything breaks, and an inability to scale.

Measuring It: From Vague Worry to a Ranked List

You cannot manage what you do not measure. Tomas's first move was an inventory, a single register of every AI system, scored for debt. This mirrors what the 2024 government-wide AI guidance already asks of agencies, which is to maintain an inventory of AI use cases. Tomas turned that compliance obligation into a management tool by adding a debt score to each entry, which cost him nothing he was not already required to do.

Score each system across the five debt types from 0, meaning none, to 3, meaning severe, then add a column for impact, meaning how much harm results if it fails. A high-debt, high-impact system is your top priority. A high-debt, low-impact system might simply be retired. The scoring is deliberately coarse because precision here is false comfort: the purpose is to rank, and a rough ranking everyone can produce beats a precise one nobody has time to compute.

SystemDataDriftDependencyDocsReproducibilityImpact if it failsPriority
Fraud scorer23122High1 (fix now)
Vendor chatbot11312Medium2 (ninety-day clock)
Forecasting model32033Medium3
Doc classifier11121Low4 (monitor)

This table did more for Tomas than any vendor pitch. It converted a bad feeling about the department's AI into a ranked, defensible list he could take to the budget office. The fraud scorer, high drift and high impact, went to the top. The chatbot's dependency debt got a hard ninety-day deadline set by someone else. The classifier, low everywhere, went on a watch list. Note that the register above shows the four named systems rather than all eleven; the point of the exercise is that every system in the portfolio gets a row, including the ones nobody remembers owning.

Three questions turn the register from a technical artifact into a budget document. Which systems carry the most debt, which is what the scores answer. How is that debt manifesting right now, whether as degraded performance, difficulty modifying the system, or outright crashes, which is what makes it concrete to a reader who does not write code. And what is the debt actually costing, measured in staff time consumed working around it, reliability lost, and performance the agency is not getting. The third question is the one that funds the work, and it is the one most registers leave blank.

The Metrics Underneath the Scores

A debt score is a summary judgment, and behind it should sit metrics somebody actually tracks. Three families matter, on three different cadences, and the cadence is as much of the design as the metric itself: a data quality figure checked once a quarter tells you only that a pipeline broke sometime in the preceding quarter, which is not information you can act on. Set each family's frequency to match how fast the thing it measures can move, and to match how long you could tolerate not knowing.

For data quality, track each dataset for completeness, meaning the proportion of fields populated; accuracy, meaning the proportion of values that are correct; consistency, meaning the proportion that are logically consistent with each other; timeliness, meaning how fresh the data is; and uniqueness, meaning the proportion of records that are not duplicates. The source sets targets of above 95 percent for completeness, above 98 percent for accuracy, above 95 percent for consistency, above 99 percent for uniqueness, and an update within 48 hours for timeliness. Treat those as the source's starting points rather than as universal thresholds, set your own against your own baseline, and track them daily so that a drop triggers an investigation rather than a discovery.

For model performance, track accuracy against the baseline established during development, precision, meaning the proportion of positive predictions that are correct, and recall, meaning the proportion of actual positives the model finds. The source suggests holding precision and recall within five percentage points of their development baselines. Track fairness as well, meaning whether performance varies across demographic groups. The source's fairness target arrived corrupted and no threshold can be reconstructed from it, so none is offered here. Define the fairness measure and its acceptable range with your legal and equity stakeholders, because that threshold is a policy decision rather than a technical one.

Tracking metrics is not the same as acting on them, and the gap between the two is where most measurement programs die. A metric earns its place only if somebody is named against it, a movement past the threshold triggers a defined action rather than an email, and the resulting investigation has somewhere to go. The most common failure is a dashboard that is technically complete and organizationally inert: every number is present, nobody owns any of them, and the drop that mattered sat visible on a screen for months.

For code quality, track documentation coverage across functions and modules, complexity, meaning whether functions are simple or convoluted, and when the code was last reviewed. The source's targets are above 90 percent documentation coverage, average complexity below 10, and a review within the last six months. Track these quarterly rather than daily, since code quality moves slowly, and when it drops, allocate time to fix it rather than noting it.

Remediation: Paying It Down Without Stopping Everything

You cannot fix all debt at once and you should not try. The disciplined sequence is measurement, then prioritization by impact on users, then allocation of real engineering time, then execution, then verification that paying the debt down actually improved the system. That last step is the one most often skipped, and skipping it means you never learn whether the work was worth doing.

Tomas used a tiered strategy any agency can copy. The fraud scorer, high debt and high impact, was fixed immediately with drift monitoring, a documented retraining process, and a named owner, because it was the system that could hurt the public most. The chatbot's retiring vendor model was replaced before the ninety-day cliff, with the switch tested in shadow mode first. The forecasting model's undocumented pipeline was rebuilt with full documentation, and where a system was barely used it was simply shut down. The classifier needed only a quarterly check, which freed resources for the real problems.

The source works a fourth example in enough detail to be worth following. A model has drifted from 95 percent to 87 percent accuracy. Measurement shows 87 percent on recent test data and 95 percent on the original training data, which supports the conclusion that the model has not changed and the world has. Prioritization ranks retraining on recent data as high, since it could restore accuracy; improved monitoring as high, since it would have caught the drift earlier; and documentation as medium, since it helps during maintenance. Allocation is two weeks to retrain and validate, one week for monitoring, and one week to document.

Verification is where the example earns its place. The retrained model reaches 94 percent on recent data, against the 95 percent the original achieved on its own training data, and drift detection is now active. The honest conclusion is the one the source draws: the debt is partially paid, not paid. The gap that remains is small and it is real, and reporting it as a full recovery would have been the easier and worse choice. Note also how the allocated effort divides: of the two weeks for retraining, one for monitoring, and one for documentation, exactly half goes to monitoring and documentation, which fix nothing today and are the only reason the next drift gets caught early.

The last point about remediation is the one leaders most often get wrong. Refactoring a pipeline, migrating to a new platform, or replacing a vendor model does not eliminate debt. It relocates it. A rebuilt pipeline is a new pipeline that will need documentation, monitoring, and an owner of its own, and a migration trades a dependency you understand for one you do not yet. That is frequently the right trade, and it should be made with open eyes and a plan for the debt the new arrangement will start accruing on day one.

One more sequencing point is worth stating, because it is where remediation programs most often stall. Fixing the system is the visible work, and confirming that the fix held is the work that gets cut when the schedule tightens. Verification means measuring the same metrics you measured before, on the same cadence, and reporting what actually changed, including when the answer is that the debt was only partly paid. An agency that never verifies cannot tell a successful remediation from a busy one, and it will make the same allocation decision next year with no better information than it had this year.

Prevention: Designing So Debt Stops Piling Up

Paying down debt while new debt accumulates faster is a losing game. The durable fix is architectural, which means making the cheap and fast choice also the maintainable one. Tomas established four rules for every new AI system, enforced at the design-review gate before any project is funded, because a rule enforced at funding is a rule and a rule enforced at launch is a negotiation.

  • No system ships without monitoring. If you cannot tell whether it is still working, it does not launch. Monitoring does not prevent drift; nothing does. It converts a silent failure into a visible one, which is the entire difference between a fix and an incident.
  • Every model has a documented data sheet. What it was trained on, what it must not be used for, and how to retrain it, written before launch rather than after a crisis.
  • Reproducibility by default. Data version, code, and settings captured together so that any model can be rebuilt. This is the AI equivalent of keeping the receipts.
  • Name an owner and a sunset date. Every system has a human accountable for it and a scheduled review at which it is renewed, refactored, or retired.

Underneath those gates sit the ordinary engineering practices that prevent most debt from forming: assigning data ownership so that somebody is responsible for data quality, requiring model documentation before production, testing at all three levels with unit tests for code, integration tests for pipelines, and performance tests for models, versioning data as well as models and code so that you always know what is in production, monitoring data quality and model performance continuously, and reviewing models and code before they reach production. None of this is novel, and the reason it is worth restating is that AI projects routinely skip the parts that feel like software engineering.

Allocating time is the rule that makes the others real. The source's guidance is to dedicate at least 20 percent of engineering capacity to debt paydown and to enforce it, because a team that is always busy building new features will never find the time otherwise. These practices also align with the NIST AI Risk Management Framework, which is voluntary rather than binding, and whose govern and manage functions are precisely about sustaining systems over time rather than only launching them. Tomas did not invent new bureaucracy. He turned an existing framework into four gate questions and a budget line.

The deeper shift he made was cultural. His teams used to celebrate launches. Now they also celebrate retirements and clean documentation, because he made total cost of ownership rather than time to launch the measure of success. The eleven hurried systems taught him an expensive lesson: in AI, the fast win and the lasting win are rarely the same choice, and the interest on the difference is paid by the public.

Anti-Patterns

No debt accounting at all

You build a model, it works, and you move on. Nobody ever measures its debt, and years later the system is undocumented, unmonitored, and fragile, with no one who understands it still on staff. This is the default outcome rather than an unusual failure, because nothing in the ordinary course of work forces the question. Assign an owner to every model by name, put every model on the register, and run regular health checks on a fixed cadence so that the question gets asked whether or not anyone is worried.

Paying down debt without prioritizing

A team decides to address all debt equally and spends weeks improving code quality on a marginal model while a critical one drifts unnoticed. Equal treatment feels fair and is the least effective possible allocation of scarce engineering time. Prioritize by impact: the systems affecting the most people, or the most consequential decisions, come first, and low-impact systems with high debt are candidates for retirement rather than repair.

Allocating no time to debt at all

The team is always busy building new features, no capacity is ever set aside for debt, and the portfolio gradually becomes unmaintainable. The source's answer is a standing allocation of at least 20 percent of engineering time for debt paydown, enforced rather than aspirational. Whatever number your organization can defend, the allocation has to be explicit and protected, because unprotected time for maintenance is time that gets reassigned to whatever is on fire.

Keeping debt invisible to leadership

Engineering knows about the debt. Leadership does not, so leadership presses for features rather than quality, and the system degrades exactly as the incentives dictate. This is a communication failure rather than a technical one. Make debt visible with tracked metrics, report to leadership quarterly, and translate the debt into the terms leadership actually decides on, which are risk to the public, staff time consumed, and the cost of the outage that has not happened yet.

Selling a refactor or migration as debt elimination

Rebuilding a pipeline or moving to a new platform is often the right decision, and it is frequently oversold. A migration exchanges debt you understand for debt you have not yet discovered, and a rebuilt system begins accruing documentation, monitoring, and ownership debt from its first day in production. Budget for the new system's maintenance in the same paper that proposes the migration, and describe the work as relocating debt to a place where it is cheaper to service rather than as removing it.

Treating monitoring as a guarantee

A monitored model is not a safe model. Monitoring watches the metrics you chose to watch, at the frequency you chose, against thresholds you set, and it is silent about every failure mode outside that set. Tomas's rule that nothing ships without monitoring is right, and it buys visibility rather than safety. Review what your monitoring does not cover as deliberately as you review its alerts, and treat a quiet dashboard as an absence of evidence rather than as evidence of health.

Practice Prompts

Each of these produces an artifact rather than an opinion, and each is deliberately sized to be done with the staff you have rather than the team you would like. Do them in order, since the first supplies the input to all the others, and expect the first one alone to change how the rest of the exercise looks, because almost every agency that builds a register discovers systems it did not know it owned.

  1. Build the debt register. List every AI system your agency runs, including models embedded in vendor software, and score each across the five debt types from 0 to 3, then add an impact rating. Note which systems you had to ask around to discover, because that fact is itself a finding.
  2. Establish a measurement baseline for one critical model. Record its current performance metrics, its data quality metrics, and its code quality, note where you cannot measure something and why, and set a monthly tracking cadence.
  3. Write the paydown plan for your worst system. State the priority in terms of impact on residents, the specific work required across retraining, refactoring, documentation, and monitoring, the time each piece takes, and whether that time is worth the benefit. Include how you will verify afterwards that the system actually improved.
  4. Cost the debt in the terms leadership decides on. For the two or three worst systems, estimate the staff time consumed by working around them, the reliability and performance the agency is losing, and what the failure would cost the public. This is the version of the register that gets funded.
  5. Design your design-review gate. Write the questions a new AI system must answer before funding, covering monitoring, the data sheet, reproducibility, and a named owner with a sunset date, and identify who has the authority to refuse funding when the answers are missing.

Reflection

What is the oldest model your agency has in production, and is it still performing the way it did at launch? If you cannot answer the second half from a dashboard, you have drift debt regardless of how the model is actually behaving, because the debt is the absence of the answer rather than the presence of the problem. Then ask the harder question: if you retired one of your current models tomorrow, would anyone notice? A system nobody would miss is not a system to maintain forever out of inertia.

Consider two more. How much of your engineering time currently goes to paying down technical debt, honestly measured rather than estimated, and how does that compare to what you would defend in front of an oversight body? And what would actually happen if a critical data pipeline failed for a week, not in terms of the alert that fires, but in terms of the decisions your agency would keep making on stale or missing data while nobody noticed?

Glossary

  • Technical debt. Shortcuts and compromises made during development that create future costs, repaid with interest in the form of rework, fragility, and risk.
  • Model drift. Degradation of a model's performance in production because the data or the world it operates in has changed, even though the model itself has not.
  • Data debt. Undocumented, fragile, stale, biased, or single-person-dependent data pipelines and datasets feeding a model.
  • Dependency debt. Reliance on an external model or service you do not control, where the provider's decisions set your timeline.
  • Reproducibility. The ability to rebuild a model from scratch because the data version, code, and settings were captured together.
  • Data quality. The completeness, accuracy, consistency, timeliness, and uniqueness of a dataset, each measurable and each trackable over time.
  • Cyclomatic complexity. A measure of code complexity based on the number of independent paths through a piece of code.
  • Test coverage. The proportion of code paths exercised by automated tests.
  • Data sheet. A document recording what a model was trained on, what it must not be used for, and how to retrain it, written before launch.

Closing

Technical debt in AI is invisible until it becomes a crisis. A model drifts slowly, nobody notices, and then it is producing bad recommendations to the public and everyone is scrambling. The solution is not to avoid all debt, which is impossible and would mean shipping nothing. The solution is to measure it, understand what it costs, and pay it down before it becomes critical, which requires only that somebody be accountable for asking on a schedule.

Tomas's eleven systems were each a reasonable decision at the time, made under real pressure to modernize, and together they produced a portfolio nobody could account for. What changed his situation was not new technology or new staff. It was a register, a ranking, a protected share of engineering capacity, and a gate at the point where projects get funded. Those four things are available to any agency, they cost attention rather than money, and they are the difference between an AI portfolio that gets better with age and one that quietly gets worse while everyone celebrates the launches.

Key Takeaways

  • AI carries debt ordinary software does not. Data, drift, dependency, documentation, and reproducibility debt each fail in their own way and compound when combined, and unlike ordinary code, an AI system nobody has touched can degrade on its own.
  • Debt sits in three layers with three different owners. Data, model, and infrastructure debt need different fixes, and infrastructure debt, meaning missing monitoring, fragile serving, manual processes, and tight coupling, is the layer most often left out of the conversation.
  • Drift debt is the silent one. A model degrades as the world changes, and without monitoring you find out only after the damage. The debt is the missing answer, not the missing performance.
  • Turn your AI inventory into a debt register. Score every system across the five debt types plus impact and you convert a vague worry into a ranked, defensible priority list, using an inventory federal guidance already asks agencies to keep.
  • Measure with real metrics on real cadences. Data quality daily, model performance against its development baseline, code quality quarterly. Set thresholds against your own baseline rather than inheriting someone else's.
  • Remediate in tiers and verify afterwards. Fix high-debt high-impact systems now, replace severe dependencies on a deadline, refactor or retire the rest, monitor the healthy ones, and confirm the work actually improved the system rather than assuming it.
  • Retiring a system is a legitimate fix. A low-value, high-debt system is often cheaper to shut down than to maintain forever.
  • A refactor relocates debt, it does not eliminate it. Budget the new system's monitoring, documentation, and ownership in the same paper that proposes the migration.
  • Prevent debt at the funding gate. Require monitoring, a data sheet, reproducibility, and a named owner with a sunset date before any AI system is funded, and protect a standing share of engineering capacity for paydown.
  • Measure success by total cost of ownership. Celebrating only launches breeds debt. Celebrate clean documentation and clean retirements too.

Frequently Asked Questions

Where do we start if we have no idea how much debt we have?

With the register, which requires only attention. List every AI system, including models embedded inside vendor software, and score each from 0 to 3 across data, drift, dependency, documentation, and reproducibility debt, then add an impact rating. The scoring is deliberately coarse. Its purpose is to produce a ranking you can defend to a budget office, and a rough ranking that exists beats a precise one nobody has time to build.

Should we use the metric targets in this lesson?

Use them as starting points and set your own against your own baseline. The source offers targets for data quality of above 95 percent completeness, above 98 percent accuracy, above 95 percent consistency, above 99 percent uniqueness, and updates within 48 hours, along with precision and recall held within five percentage points of their development baselines and code quality targets of above 90 percent documentation coverage, average complexity below 10, and review within six months. What matters more than any specific figure is that a baseline exists and that somebody investigates when a metric moves away from it.

What fairness threshold should we set?

That is a policy decision rather than a technical one, and this lesson deliberately does not supply a number. The source's fairness target arrived corrupted and no defensible threshold can be reconstructed from it, so inventing one here would be worse than leaving the gap visible. Track whether model performance varies by demographic group, then set the measure and the acceptable range with your legal, equity, and program stakeholders, and document who made the decision.

How much engineering time should go to debt paydown?

The source's guidance is at least 20 percent of capacity, enforced rather than aspirational. The exact figure matters less than two properties: the allocation must be explicit, so it appears in a plan rather than in an intention, and it must be protected, because unprotected maintenance time is the first thing reassigned when something breaks. If your organization cannot defend 20 percent, defend a smaller number you will actually hold, and report what the gap is costing.

Does retiring a system count as fixing it?

Yes, and it is frequently the cheapest fix available. A system with high debt and low impact costs real staff time to maintain and produces little in return, and the decision to keep it usually reflects inertia rather than an assessment. Every system should carry a named owner and a scheduled review at which it is explicitly renewed, refactored, or retired, which turns retirement from an admission of failure into a normal outcome of a review.