←
AI for Government
Proficient · M21 · lesson 21 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Continuous Improvement for AI Systems
📖
now learning

Continuous Improvement for AI Systems

15 min

Priya Nadar, a program director at a municipal transit authority, launched an AI tool that predicted which bus routes would run late so dispatchers could pre-position spare drivers. At launch it was 88 percent accurate, and everyone celebrated. Eighteen months later nobody had touched it, and dispatchers had quietly stopped trusting it. When Priya finally pulled the numbers, accuracy had slid to 71 percent. A new bus-rapid-transit corridor, a fare change that shifted ridership, and weather patterns the model had never seen had all moved the ground underneath it. The system had not broken. It had stopped improving while the world kept changing, which in government amounts to failing slowly.

This lesson is about the discipline that prevents Priya's slide. A deployed model is not a finished product like a paved road. It is closer to a garden: alive, drifting, and degrading the moment you stop tending it. Deployment is not the end of the work, it is the point at which you finally have real data about how the system behaves, how people actually use it, and where it could be better. Your job as a strategist is to build the tending into the operation so that improvement happens on purpose rather than as an emergency.

Why AI Systems Quietly Get Worse

Traditional software is stable. The payroll system you deployed in January behaves identically in December. AI is the opposite, because its accuracy depends on the world matching the data it learned from, and the world refuses to hold still. That is drift, and it comes in two flavours worth naming separately, because they call for different responses.

  • Data drift: the inputs change. Priya's model started seeing a new corridor and new ridership patterns it had never trained on.
  • Concept drift: the relationship changes. What counted as a late route shifted when the fare change altered who travelled and when.

Drift is not a bug. It is the normal condition of a deployed model, and treating it as an incident rather than an expectation is what leaves agencies surprised by it. The mistake Priya made was treating launch as the finish line. Launch is the starting line. Almost everything that makes the system worth the public's money happens in the months afterwards, and almost none of it happens automatically.

The mindset that works sits between two failures. One is deploy and forget, which is what happened here. The other is perfect before deploying, which sounds responsible and produces systems that never ship, or that ship years late against requirements that have themselves drifted. The usable position is to deploy the best version you can defend, then improve it continuously against real-world data and feedback.

Why This Is Harder at Government Scale

Government serves everyone, which means your systems have to work at scale, across diverse populations, in contexts that differ enormously from one another. An improvement approach that suits a fifty-person organisation does not transfer to a service accountable to millions of residents. The volume of feedback is larger and noisier, the segments that matter are more numerous, and a change that improves the aggregate can quietly worsen conditions for a group that is small in the data and significant in law.

Two consequences follow for the improvement loop. First, segment analysis is not optional at this scale, because the populations most likely to be poorly served are usually the ones least represented in the training data and least visible in the headline metric. Second, the change process needs more formality than a small team would tolerate: a documented proposal, a review, a rollback path and a record of what was authorised. That is not bureaucracy for its own sake. It is what lets you answer, long afterwards and under different leadership, why the system behaves differently than it did at approval, and who agreed to the change.

The Continuous Improvement Loop

Continuous improvement is a loop you run on a schedule, not a project you complete. Four stages, turning indefinitely: measure, learn, change, verify. Each one is a discipline most agencies skip, and skipping any single stage tends to break the others. Measuring without learning produces dashboards nobody acts on. Changing without verifying produces superstition.

Measure: watch the right signals

You cannot improve what you do not track, and measurement here has two layers. The first is system performance: accuracy, drift, and error patterns over time. The second, which agencies forget, is real-world outcome. Priya's model was 88 percent accurate at predicting delays, but the question that mattered was whether pre-positioning drivers actually reduced passenger wait times. A model can be accurate and useless if it predicts the wrong thing well.

Production data is the point. It shows how the system actually behaves rather than how you expected it to, which is why the measurement plan should be built before launch and instrumented at launch. Retrofitting observability onto a system already in service is possible and always more expensive, and in the interval you are running a public function without knowing whether it is working.

Learn: find where it fails and why

Aggregate accuracy hides the failures that matter. Break performance down by segment: which routes, which times of day, which conditions, which populations. Priya eventually found her model was still 90 percent accurate on the older routes and around 40 percent on the new corridor. Which aggregate that produces depends on how much of the volume runs where, and the single figure of 71 percent told her nothing actionable. The breakdown told her exactly what to fix, and told her that most of the system was fine.

Incidents are the other learning source, and they are underused because they arrive labelled as problems rather than as data. A wrong output that reached a resident, a complaint, an override by a caseworker: each is a specific, dated example of the model failing in a way your metrics may not capture. Reviewing them systematically is usually the fastest route to knowing what to change.

Change: make a deliberate adjustment

Improvement is a change you choose: retraining on recent data, adjusting a threshold, adding an input feature, or fixing an upstream data feed. The key word is deliberate. Each change should be stated as a hypothesis, such as expecting that retraining on recent months recovers corridor accuracy, rather than as a general hope that a newer model will be better. A change with no stated expectation cannot be verified, because nothing was predicted.

Verify: show that the change helped

This is the stage that separates real improvement from superstition, and it is where testing earns its place. It is also where most improvement programs quietly stop, because by the time the change ships everyone has moved on to the next problem. Build the verification into the change, with its own date and owner, or it will not happen.

A/B Testing, Done Carefully in Government

A/B testing means running two versions side by side and comparing results: the current model against the proposed improvement, on comparable real cases, to see which performs better before you commit. In the private sector this is routine. In government it carries an ethical weight the private sector can ignore, and you have to hold both ideas at once.

The benefit is real. Testing stops you from rolling out a fix that quietly makes things worse, and it catches unexpected negative effects during the test rather than after full deployment, at least for the effects you thought to measure. Priya's retrained model could be tested on a sample of routes before going authority-wide. The caution is equally real: you cannot run an experiment that gives some residents a worse service in order to learn something. Testing two ways of pre-positioning spare buses is fine, because both groups still get service. Testing a benefits-denial model where one group receives a less accurate eligibility decision is not, because you would be experimenting on people's livelihoods.

The governing rule is that you may experiment on the model's internal performance, subject to your normal privacy and governance review, but you may not run an experiment whose losing arm denies a person a right, a benefit or a safety protection they would otherwise receive. Where you are unsure, test against historical data instead: does the new version beat the current one on last quarter's real cases? That is a good default and it has a limit worth stating plainly. Historical testing tells you how a version would have handled conditions that have already happened. It cannot tell you how it will handle the drift that is currently underway, which is precisely the problem you were trying to solve. Treat a strong backtest as a reason to proceed carefully, not as a result.

Choosing an Iteration Cadence

How often should you iterate: weekly, monthly, quarterly? There is no universal answer, and the trade-off is straightforward. Frequent iteration enables fast learning and keeps the model close to current conditions. It also increases operational complexity, because every change needs validation, documentation, an approval path and a rollback plan, and in government it may need a governance review as well.

Pick a cadence your organisation can actually sustain, then hold it. A quarterly cycle that happens every quarter is worth considerably more than a monthly cycle that happens twice and lapses. The cadence should also be tiered: a routine review on the calendar, plus a defined trigger that pulls the next cycle forward when a drift alert fires. Priya's rebuilt program used a weekly accuracy check, a monthly segment review and quarterly retraining, with the retrain moving earlier if the drift alert fired first.

Release management is the other half of cadence and it is usually the part that is missing. Every change needs a version number, a record of what changed and why, a person who approved it, and a retained previous version that can be put back. Without that, an improvement that turns out to be a regression cannot be undone cleanly, and nobody can reconstruct which model made a given decision. In a government setting that reconstruction is not a convenience, it is frequently what an audit or an appeal requires.

Build Feedback Loops From the People Who See Errors First

The richest improvement signal is free and usually ignored: the staff and residents who interact with the system every day. Priya's dispatchers knew the corridor predictions were wrong long before her metrics did. They had no channel to report it, so they worked around the tool instead, which is the quietest and most common form of failure in deployed government AI. A working improvement program turns those people into sensors.

Feedback comes from at least four sources and each sees something different. Users of the system notice individual wrong outputs. Operators and administrators notice patterns and workload effects. Affected citizens notice outcomes, especially bad ones. Advocates and community organisations notice distributional problems that no individual complaint would reveal. Collecting from only the first source, which is the default, gives you a view of the system that systematically misses its worst failures.

  • Staff loop: a one-click way for a frontline worker to flag that a prediction was wrong, feeding a queue somebody reviews weekly. This is usually your earliest drift warning, and it costs almost nothing to build.
  • Resident loop: a clear, accessible route for the public to question or challenge an AI-influenced outcome, with corrections fed back into the learning stage. This is both an improvement source and, for rights-affecting systems, an obligation.

The appeals channel deserves particular attention because it is a feedback loop most agencies already operate without using it. When an applicant appeals an eligibility decision and a human review determines the correct outcome, the comparison between the original decision and the reviewed one is a labelled example of where the model was wrong. One benefits agency built exactly this cycle: decision made, applicant appeals, human review establishes the correct answer, the comparison identifies the error, the model is improved. Appeal volume subsequently fell, which is consistent with a more accurate model and, as with any before-and-after comparison, also consistent with other things changing at the same time. Worth building; worth describing carefully when you report it.

The Office of Management and Budget's memorandum M-24-10, Advancing Governance, Innovation, and Risk Management for Agencies' Use of Artificial Intelligence, issued in 2024, requires covered agencies to monitor rights- and safety-impacting AI for performance and to provide affected people with ways to seek redress. The NIST AI Risk Management Framework, which is voluntary guidance, puts ongoing monitoring and measurement at the centre of its Manage function. Feedback loops are how you satisfy both, not as paperwork but as the mechanism that keeps the system honest.

An Experimentation Framework, So Changes Get Reviewed Once

Improvement proposals arrive constantly and reviewing each one from scratch is how agencies end up either blocking everything or waving everything through. A standing framework fixes this. One government required every proposal to change a system to state four things: the expected impact, meaning what will improve and by how much; the experiment design, meaning how the change will be tested; the success criteria, meaning what evidence would count as working; and a risk assessment, meaning what could go wrong. Proposals were approved only where an experiment could be run at acceptable risk.

The value is not the paperwork, it is that the argument happens before the change rather than afterwards. It also produces the record you will want later, when someone asks why the model behaves differently than it did at authorisation. Note the honest limit: a framework like this catches the negative effects your success criteria and risk assessment thought to look for. Effects nobody anticipated still reach production, which is why monitoring after rollout matters as much as testing before it.

The Culture That Makes This Possible

Continuous improvement is a cultural problem before it is a technical one. It needs psychological safety, so that a staff member can say the tool is producing bad results without it being heard as a criticism of the people who bought it. It needs systematic learning, so that findings are recorded rather than remembered. It needs genuine experimentation, which means accepting that some changes will not work. And it needs adaptation, meaning that findings actually change what the organisation does.

The payoff is cumulative rather than dramatic. One government employment service matching job seekers to openings tracked which matches led to placements, which did not, how long placements took and whether job seekers were satisfied. Analysis showed that including remote roles increased placements by 12 percent, and that insight was built into the model. Where categories showed lower success rates, investigation revealed why, whether a skills gap or a geographic mismatch, and led to targeted interventions. Placement rates improved 20 percent over the year across all of that work combined. No single change was transformative. The accumulation was.

The AI Continuous Improvement Plan

Stand up one of these per production AI system. It assigns owners and cadences so that improvement is a routine rather than a rescue. The worked column shows the transit tool, run the way it should have been.

ElementWhat to defineWorked example (delay prediction)
Performance metricThe model number you trackDelay-prediction accuracy, weekly
Outcome metricThe real-world result that mattersAverage passenger wait time on covered routes
Drift watchSignal and alert thresholdAccuracy alert if it drops more than 5 points against the 90-day baseline
Segment reviewBreakdowns you checkBy route, time of day and weather; monthly
Incident reviewHow wrong outputs are studiedFlagged predictions and overrides reviewed as labelled failure cases
Staff feedback loopHow frontline flags errorsOne-click "prediction wrong" button, reviewed weekly
Resident feedback loopHow the public is heardService-quality form; complaints tagged and routed to the data team
Change cadenceHow often you retrain or adjustQuarterly retrain, or sooner on a drift alert
Change proposal standardWhat a change must stateExpected impact, test design, success criteria, risk assessment
Verification methodHow you show a change helpedTest the new model on last quarter's real cases, then watch it closely in production
Ethics gateWhat you may not test liveNo live test that worsens service for any rider group
Rollback pathHow you undo a bad changePrevious model version retained and redeployable
Owner and review forumWho runs it, where it is reportedData lead owns it; reported in the monthly operations review

Make Improvement a Calendar Item, Not a Crisis

The reason Priya's system rotted was structural rather than technical. Nobody owned its ongoing performance, and nothing on anyone's calendar forced a look. Continuous improvement survives only when it is a named person's responsibility, reported in a recurring forum, with a budget line for the compute and staff time that retraining costs. A model with no improvement owner will drift, because the world does not stop changing out of courtesy to your launch metrics.

When Priya rebuilt the program she set a weekly accuracy check, a monthly segment review, quarterly retraining and a one-click dispatcher feedback button. Within two months corridor accuracy was above 85 percent. The retrain now happened on a schedule rather than after the decay had been running unnoticed, and that is the most plausible explanation for the recovery, though it is one authority's experience rather than a controlled comparison. The model was the same. The difference was that somebody was finally tending the garden.

Anti-Patterns

  • Treating deployment as completion. The launch is celebrated, the team disperses, and nothing on any calendar forces a look at the system again. Drift does the rest. Assign the improvement owner in the same document that authorises the deployment.
  • Waiting for perfect before deploying. The mirror image failure. A system held back until it is flawless never ships, or ships against requirements that have themselves moved. Deploy a version you can defend, with monitoring, and improve it.
  • Trusting the aggregate. A single headline accuracy figure can hide a segment where the model has failed almost completely. Always break performance down by route, time, condition and population before concluding the system is healthy.
  • Measuring accuracy instead of outcome. A model can predict its target accurately and still deliver nothing, because the target was not the thing that mattered. Track the operational result alongside the model metric.
  • Changing without verifying. A retrain ships, everyone assumes it helped, and nobody checks. Improvement becomes folklore. State the expected effect before the change and schedule the check that tests it.
  • Treating a strong backtest as proof. Historical testing shows how a version would have handled conditions that already occurred, which is exactly not the drift you are worried about. Use it to decide whether to proceed, then monitor closely after rollout.
  • Experimenting on people's entitlements. Running a live A/B test whose losing arm gives residents worse decisions is not a methodology question, it is an ethics failure. Test internal performance freely and never a person's right, benefit or protection.
  • Ignoring the workaround. When staff quietly stop using a tool, that is the strongest possible signal about its performance and it appears in no dashboard. Build the channel that captures it before you need it.
  • Improvement without accountability. A plan exists, nobody owns it, no KPI tracks it, and nothing happens. The strategy stays aspirational. Assign responsibility, define the measures, and review them in a standing forum.
  • Copying another agency's cadence wholesale. A monthly cycle that works for a large data team will collapse in a two-person shop. Use another jurisdiction's design as a guide, then set the cadence you can actually sustain.

Practice Prompts

  • Assess the current state. For one production AI system, find out when its performance was last measured, by whom, and against what baseline. If nobody can answer, you have found the gap.
  • Separate the two metrics. Write down the model metric and the real-world outcome metric for that system. Are you tracking both, and would improvement in the first necessarily show up in the second?
  • Run a segment breakdown. Split the system's performance by the segments that matter for your service and identify the worst one. Decide whether that segment is a fixable data problem or a scope problem.
  • Design a feedback loop. Specify how a frontline worker would report a wrong output without leaving their workflow, who reviews the queue, how often, and what happens to what they find.
  • Draft a change proposal. Take one improvement you want to make and write the four elements: expected impact, test design, success criteria, and risk assessment. Note which one was hardest, because that is usually the weak point of the proposal.
  • Set the cadence. Decide the review, retrain and reporting cadence you can genuinely sustain with your current staffing, and name the trigger that pulls the next cycle forward.
  • Cost the improvement. Estimate the annual compute and staff time your improvement plan requires and check whether that money exists in a budget line. If it does not, the plan is a wish.

Reflection

Pick the production AI system your agency depends on most. Who owns its ongoing performance by name, and where is that ownership written down? When was it last retrained, and was that a scheduled event or a response to a complaint? If its accuracy had fallen by a third since launch, as Priya's did, what signal would have told you, and how long would it have taken? Then consider the softer question, which is often the real one: if a frontline member of staff believed the tool was producing bad results, is there a route by which that belief would reach you, and would raising it feel safe? Most improvement programs fail at that step rather than at the technical one.

Glossary

  • Continuous improvement. The iterative enhancement of a deployed system based on data and feedback, run as a standing routine rather than as a project with an end date.
  • Drift. The gradual divergence between the conditions a model was trained on and the conditions it now operates in. The normal state of a deployed model, not an incident.
  • Data drift. A change in the distribution of inputs reaching the model, such as a new population, a new location or a new document type.
  • Concept drift. A change in the underlying relationship the model learned, so that the same inputs should now produce a different answer.
  • Performance metric. The measure of how well the model does its stated task, such as accuracy on a defined test set.
  • Outcome metric. The measure of the real-world result the system exists to produce. A system can improve on the first while doing nothing for the second.
  • Segment analysis. Breaking performance down by route, time, condition or population to find failures that the aggregate conceals.
  • A/B testing. Running two versions in parallel on comparable cases to compare their performance before committing to one.
  • Backtesting. Evaluating a candidate version against historical cases with known outcomes. Safe, useful, and blind to the drift currently underway.
  • Feedback loop. A defined channel through which errors observed by staff or residents return to the team that can act on them.
  • Model versioning. Retaining previous model versions so that a change which turns out worse can be rolled back quickly.
  • Learning organisation. An organisation that systematically learns from its own experience and changes what it does as a result.
  • Data quality. The accuracy, completeness, consistency and reliability of the data feeding a system. An upstream data problem is a common cause of what looks like model degradation.

The measurement layer this lesson depends on is built in Continuous Monitoring Fundamentals, AI Metrics and KPIs for Government and Measuring AI Impact. Moving from Pilot to Production covers the transition that has to happen before an improvement program is meaningful, and Data Infrastructure for Enterprise AI and Data Quality and AI Performance address the upstream feeds that cause much of what gets blamed on the model. When degradation becomes an availability problem rather than a quality one, Continuity of Operations with AI takes over. Bias Detection and Mitigation at Scale is the segment analysis applied to fairness, Cultural Transformation for AI and Building an AI Quality Culture cover the psychological safety this work needs, and Establishing an AI Governance Board is where change proposals are reviewed.

Closing

Continuous improvement is unglamorous and it is the difference between a system that keeps earning its funding and one that quietly stops. The mechanics are not complicated: measure the model and the outcome, break the numbers down far enough to see where they fail, make deliberate changes with stated expectations, verify that those changes helped, and keep a channel open to the people who notice errors first.

What makes it hard is that none of it is urgent on any given day. Nothing breaks. No alarm sounds. The accuracy slides a point at a time while everyone is busy with the next launch, and the cost only becomes visible when someone finally pulls the numbers, as Priya did after eighteen months. Put the review on the calendar, put a name against it, and fund it, because those three steps are what turn improvement from an intention into an operation.

Key Takeaways

  • Launch is the starting line. Deployed AI drifts as the world changes, and treating launch as completion produces slow, invisible failure rather than a clean one.
  • Run the loop on a schedule. Measure, learn, change, verify, indefinitely. Improvement is a routine somebody owns, not a project with a completion date.
  • Track outcomes as well as accuracy. A model can be accurate and useless. Measure the operational result the system was funded to deliver.
  • Break performance down by segment. Aggregate accuracy hides the specific routes, times or groups where the model has failed, and the aggregate itself depends on the mix.
  • Every change is a hypothesis. State the expected effect before you make the change, and schedule the verification, or improvement becomes folklore.
  • Backtesting decides whether to proceed, not whether it worked. Historical cases cannot exercise the drift currently underway, so watch the new version closely in production.
  • A/B test the model, never the public's rights. Experiment freely on internal performance under normal governance, and never run a test whose losing arm denies someone a benefit or protection.
  • Turn frontline staff into drift sensors. The people using the system see errors before the metrics do, and a quiet workaround is the strongest signal you will ever get for free.
  • Use the appeals channel you already have. Every reversed decision is a labelled example of where the model was wrong, and most agencies collect them already without feeding them back.
  • Name an owner and fund the cadence. Continuous improvement needs a responsible person, a recurring forum, and a budget line for retraining, or it will not survive the first busy quarter.

Frequently Asked Questions

How often should we retrain?

Often enough that drift is corrected before it affects service, and rarely enough that each change can be validated, documented and rolled back if needed. Set a routine cadence you can sustain and pair it with a drift trigger that pulls the next cycle forward. A quarterly cycle that actually happens beats a monthly cycle that lapses after the second month.

We have no data science team. Can we still run this?

Most of the loop does not require one. Tracking a performance metric and an outcome metric, breaking results down by segment, collecting frontline reports and reviewing them on a cadence are organisational disciplines. The retraining step may need a vendor or a partner, in which case the cadence, the rollback path and the verification standard belong in the contract rather than in an informal understanding.

Is A/B testing on residents ever acceptable?

Where both arms receive service and neither is disadvantaged in a right, benefit or protection, comparative testing is reasonable and should still go through your normal privacy and governance review. Where the losing arm receives a worse decision on something that matters to a person's life, it is not, regardless of how much you would learn. When in doubt, backtest, and treat the result as a reason to proceed carefully.

Our staff have stopped using the tool. Is that an improvement signal or an adoption problem?

Assume it is both and investigate the performance question first. Quiet abandonment is what happens when people who see wrong outputs have no channel to report them, and it usually means the model is failing on a segment your metrics do not isolate. Ask which cases they stopped trusting it for, because that answer is a segment definition.

How do we justify the ongoing cost of improvement?

By treating it as part of the cost of running the system rather than as an enhancement. A model with no improvement budget degrades until it is either abandoned or replaced at full cost, and both outcomes are more expensive than the maintenance. Put the compute and staff time in the operating budget line at authorisation, when it is much easier to defend than it will be later.

What should we report to leadership, and how often?

The two metrics, the worst segment, what changed since the last review, and whether any change was verified. A standing slot in a monthly operations review is usually enough. The reporting matters as much as the analysis, because a metric with no forum where it is discussed is a metric nobody acts on.