←
AI for Government
Proficient · M35 · lesson 35 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Moving from Pilot to Production
📖
now learning

Moving from Pilot to Production

15 min

Aisha Bello ran a pilot that everyone loved. As a program director at a state revenue department, she had tested an AI tool that flagged likely errors in business tax filings before they became audits. Over four months on a sample of 2,000 returns, it caught errors a third faster than staff and recovered an extra $380,000. The leadership team applauded. The press release wrote itself. Then she tried to scale it to the full 1.2 million annual filings and everything broke. The model that ran fine on a laptop choked at volume. Staff who had been hand-picked for the pilot were replaced by an entire division that had never been trained. Filers started calling to contest flags, and there was no appeal process. The error rate that looked tiny on 2,000 returns became thousands of wrong flags at scale. Six months later, the celebrated pilot was paused. The pilot had succeeded. The transition to production had failed.

This is the most common graveyard in government AI: the successful pilot that never becomes a real service. Pilots and production are not the same activity at different sizes. They are different disciplines, with different data, different people, different consequences and different legal obligations. This lesson is the playbook for crossing the gap, the place where most government AI quietly dies, and it is deliberately organised around the things that actually break rather than the things that look good in a steering committee slide.

Why a Successful Pilot Can Mislead You

A pilot is designed to answer one question: can this work? It answers it under ideal conditions, with small volumes, motivated staff, clean sample data, and forgiving stakes. Production asks a completely different question: can this work reliably, at full scale, with ordinary staff, messy real data, real consequences, and no special attention? A pilot that proves the first tells you almost nothing about the second. Pilot success is evidence gathered under pilot conditions. It is not a guarantee of behaviour under production load, and it is not a prediction about a population the pilot never saw.

Aisha hit all four classic traps. Performance that holds at small volume collapses at scale. Staff who were enthusiastic volunteers are replaced by an entire workforce that needs training and may resist. Error rates that are negligible on a sample become large absolute harms across a population. And the governance that you could improvise for a pilot, meaning oversight, appeals and monitoring, has to become real, durable infrastructure. The pilot proved the idea. Production is a different build, and treating it as a bigger version of the same build is how the pilot becomes the high-water mark.

The two settings differ on almost every axis that matters, and it helps to see the contrast laid out rather than described. Pilots are controlled experiments. Production is the opposite of a controlled experiment in every dimension at once, which is why weaknesses that were invisible for months surface in the first week after cutover.

DimensionPilotProduction
DataCurated and preparedMessy, incomplete, sometimes adversarial
TeamSmall and hand-pickedWhole operating workforce
ScopeNarrowUniversal
DowntimeTolerableDamages public trust
FailureInstructiveTriggers congressional oversight and inspector general inquiries

The published experience is blunt about how often the gap swallows programs. Analyses from the General Services Administration AI Centers of Excellence and from the Office of Management and Budget document that up to 80 percent of federal AI pilots never reach sustained production. The common causes are not exotic. Infrastructure was insufficient for the load, staffing was misaligned with operations, governance gaps surfaced late during authorization review, and change management failures left frontline employees unable or unwilling to trust the system.

Accuracy behaves the same way. A pilot that classified 10,000 Department of Veterans Affairs disability claims per month at 94 percent accuracy may drop to 78 percent on the full annual workload of 1.3 million claims, because edge cases, data drift and integration errors surface only at full scale. Notice what that pilot volume represents: 10,000 claims a month is roughly 120,000 a year, under a tenth of the annual workload. The pilot was never asked the question production asks.

What the Record Already Shows

You do not have to speculate about how scaling failures unfold, because several have played out in public. The Internal Revenue Service deployment with ID.me verified a few thousand taxpayers in pilot and collapsed when tens of millions of citizens tried to authenticate, forcing Treasury to reverse course in 2022 after civil rights objections from the Electronic Privacy Information Center and from members of Congress. The system had scaled without adequate accessibility testing, pushing taxpayers who could not pass facial recognition into long phone queues.

The Michigan Integrated Data Automated System performed acceptably in testing but issued more than 40,000 false fraud determinations after scaling, requiring a court-ordered multi-million dollar settlement. The Dutch childcare benefits scandal showed an apparently successful algorithmic risk model at national scale wrongly accusing 26,000 families of fraud, and it ultimately brought down the Rutte government. In each case the model was not the point of failure. The scaling process was: no load discipline, no meaningful oversight of automated adjudication, no redress path that worked at population scale.

Two further cases sharpen the lesson in different directions. The Houston Independent School District teacher evaluation model scaled without adequate transparency and was struck down in Houston Federation of Teachers v HISD for due process violations. The COMPAS recidivism tool scaled across jurisdictions without public sector validation and was found by ProPublica to exhibit racial disparities in false positive rates. Scaling multiplies whatever you already have, including the parts you never examined, and it multiplies them in front of an audience that did not consent to the experiment.

There is a counterexample worth studying as a model rather than a warning. The Allegheny County Department of Human Services child welfare screening tool published its validation study, engaged community advisors, ran in shadow mode before production, and carved out the most sensitive decisions explicitly. That is what mature scaling looks like: the same tool, deployed with the evidence and the escape hatches built before anyone needed them.

The Four Dimensions of Production Readiness

Before scaling anything, assess readiness across four dimensions. A weakness in any one is enough to sink the transition, and in practice agencies assess the first dimension carefully and the other three casually.

Technical readiness

Can the system handle full production volume, with acceptable speed, reliability, and the data infrastructure to feed it continuously? Aisha's model worked on a laptop and a sample. It had never been tested against 1.2 million live filings or integrated with the production data pipeline. Load test at full scale before you commit, not after, and test the pipeline as well as the model, because integration failures are at least as common as model failures and are far harder to diagnose once the whole division is depending on the output.

Operational readiness

Are the people and processes ready? This is the dimension that fails most often and gets attention least. The full workforce that will use the system daily needs training, new procedures and support. Aisha trained a handful of pilot staff; she needed to prepare an entire division and redesign their workflow around the tool. Operational readiness also means a staffing model that survives contact with reality, including site reliability and machine learning operations coverage on a rotation rather than one analyst who happens to care.

Governance readiness

Are the safeguards real and durable? At production scale, with real consequences for real people, you need monitoring, an appeal process, bias testing and clear human accountability. Aisha had none of these because a pilot did not seem to require them. At scale they are not optional; they are what stands between you and a wrongful-flag scandal. This is exactly where federal guidance to agencies requires impact assessments and a human alternative for systems that affect people's rights before deployment.

Risk readiness

What happens when it goes wrong, and it will? You need to know your error rate at production volume, who bears the harm, how people are made whole, and how you would roll back the system without disrupting the service. A plan for failure is a precondition for going live, not an admission of doubt. Risk readiness is also where you write down, in advance, the conditions under which you would stop, because a threshold invented during an incident is a negotiation rather than a control.

Infrastructure and Authorization Readiness

Before a government AI system moves to production it needs a complete authorization boundary, and this is usually the longest pole in the tent. For cloud-hosted systems that means a FedRAMP moderate or high authorization for the underlying platform and for any AI interfaces it calls, a FISMA system security plan, continuous monitoring, and an authorization to operate signed by the agency authorizing official. None of this is paperwork you can run in parallel with go-live. The authorizing official cannot responsibly sign without the underlying controls actually in place.

The engineering checklist underneath the authorization is specific. Compute is load tested at two to three times peak expected volume. Storage accommodates the historical records required by the Federal Records Act and the applicable National Archives retention schedules. Networking is benchmarked for tail latency rather than mean latency, because the citizen who waits is at the tail. Redundancy includes active-active or hot standby failover across at least two FedRAMP regions. Disaster recovery includes tested restore procedures with recovery time and recovery point objectives set by mission criticality.

Observability deserves particular attention in government, because your logs are also your accountability record. Production observability means structured logs, per-decision audit records sufficient to answer Freedom of Information Act and Privacy Act requests, real-time metrics, alerting runbooks, and dashboards visible to program leadership rather than buried in an engineering tool. Security means encryption in transit and at rest, role-based access control tied to agency identity providers, continuous vulnerability scanning, and alignment with the binding operational directives issued by the Cybersecurity and Infrastructure Security Agency.

One integration reality deserves naming. Many production systems in agencies such as the Internal Revenue Service, the Social Security Administration and the Department of Veterans Affairs must interoperate with legacy mainframes that were never designed for continuous model inference. Vendor management matters here too, since your AI capability inherits the authorization posture of the providers behind it. Treat the FedRAMP authorization of a provider as a continuing control you monitor, not a certificate you file.

The Compliance Layer You Cannot Skip

Scaling in the federal environment is not only an engineering exercise. Office of Management and Budget guidance issued in 2024, memorandum M-24-10, requires agencies to designate AI as rights-impacting or safety-impacting where it meets those definitions, to complete impact assessments, to maintain AI use case inventories, and to provide public notice before production use. The inventory obligation traces to Executive Order 14110, the 2023 executive order on artificial intelligence, which is no longer in force; the inventory practice it established persists in agency processes and in the memorandum. Confirm current requirements and deadlines with your agency's chief AI officer and counsel rather than relying on any course summary.

Two further frameworks shape what the auditors will ask. The NIST AI Risk Management Framework 1.0 is voluntary and non-binding, but its MANAGE function is the clearest available statement of what documented controls for drift monitoring, human oversight and decommissioning should contain, and agencies increasingly write their internal standards against it. The Government Accountability Office AI Accountability Framework sets expectations for performance monitoring, data quality tracking and transparency obligations for any AI that supports agency decisions, and it is the lens a GAO engagement will use on your program.

Read those obligations as a sequence rather than a pile. Designate first, because the designation determines which safeguards attach. Assess impact before deployment, because the assessment is a gate rather than a report. Publish notice before production use, because notice issued after the fact is what turned the ID.me deployment into a reversal. Then monitor continuously, because every one of these frameworks treats a one-time sign-off as insufficient.

The Scaling Playbook: Phased, Never Big-Bang

The cardinal rule: do not flip a switch from pilot to full production. Scale in deliberate stages, expanding only when each stage proves stable. Had Aisha followed this, the story ends differently.

  1. Harden the pilot. Re-engineer the system for production: real infrastructure, production data pipelines, monitoring, and load testing at full volume. The laptop model is rebuilt for 1.2 million returns.
  2. Limited production. Run live on a small but real slice, say 5 percent of filings, with real consequences and real staff. This surfaces the problems the sterile pilot hid, while the blast radius is still small.
  3. Build the governance. Stand up the appeal process, monitoring dashboards, and bias testing before expanding further. Filers who are flagged now have a real path to contest it.
  4. Staged expansion. Grow to 25 percent, then 50 percent, then full coverage, training the workforce in waves and pausing at each step to confirm performance, fairness and operations hold.
  5. Full production with continuous monitoring. Once at full scale, monitoring never stops. You watch for the slow accuracy decline that creeps in as filing patterns change over time.

Each gate has a go or no-go decision. The willingness to stop at a gate is what makes the playbook work; a staged rollout you never pause is just a slow big-bang. Note what those percentages are and are not. They are thresholds the agency sets for itself and records in advance, so that the decision to proceed is made against a written standard rather than against the mood of the room. They are not legal tests, and they do not separate lawful deployment from unlawful deployment. That line is drawn by the safeguards and statutes that apply to your use, not by a rollout schedule.

Rollout Patterns and What Each One Buys You

Gradual rollout is the primary mitigation for scaling risk, and the patterns are not interchangeable. Choosing the wrong one wastes the caution you paid for. Each pattern answers a different question, and mature programs use several in sequence rather than picking a favourite.

PatternWhat it doesWhat it protects against
Shadow modeRuns the new system beside the legacy process, logging outputs without using them for decisionsDeploying a system whose real-world accuracy you have never compared to current practice
Canary deploymentRoutes a small percentage of real traffic to the new systemFailures that only appear with live data and live users
Geographic rolloutStarts with one state, region or office before wider expansionPopulation differences and local process variation
Temporal rolloutIncreases the share of traffic on a schedule with go or no-go gatesLoad and operational capacity problems
Parallel operationMaintains the prior system long enough to fall back to itBeing unable to retreat when something goes wrong

Shadow mode is the underused one. It costs compute and patience but nothing else, and it is the only pattern that produces a direct comparison between the model and the process you are proposing to replace, on the same cases, before anyone is affected. If your program cannot articulate why it skipped shadow mode, it probably skipped it because the comparison was inconvenient.

Monitoring, and the Arithmetic of Error at Scale

Scaling introduces risks that were invisible at pilot, and the clearest way to see this is to do the multiplication. A one percent error rate on 10,000 cases is 100 affected citizens, which an appeals team can absorb. The same one percent on 10 million cases is 100,000 affected citizens, which is a crisis, a hearing and a settlement. The error rate did not change. The population did. For rights-impacting AI, that arithmetic is the whole argument for building redress capacity before expanding rather than after.

Monitoring during rollout must therefore cover more than uptime. Track model performance, fairness across demographic groups in line with the protections against algorithmic discrimination described in the Blueprint for an AI Bill of Rights, incident counts, user satisfaction, and downstream outcomes for the people the system touches. Monitoring detects only what you chose to watch, so the choice of metric is itself a control decision, and the group you did not disaggregate is the group whose harm you will learn about from a reporter.

A pause-and-fix posture is essential and has to be established before launch, when it costs nothing politically. Agencies that promise leadership they will never pause a rollout are setting up the next Michigan MIDAS. Write down which observations trigger a pause, who can call one, what happens to in-flight cases during it, and how quickly the prior process can carry the full load again. Rehearse the rollback at least once while the stakes are still small.

Change Management: The Human Half of Scaling

The technology is the easier half. The harder half is the people who must change how they work. Aisha's pilot staff chose to participate; her production division was told to. People who fear an AI tool will judge their work, replace their jobs or expose their mistakes will quietly resist or undermine it, and a tool the workforce resists will fail regardless of how good it is. Scaling changes jobs, and it changes them in ways that touch workflows, performance metrics and sometimes bargaining unit protections.

Effective change management starts before deployment and has a specific playbook. Run an impact assessment that maps every role affected. Hold collaborative design sessions with line staff. Consult the American Federation of Government Employees, the National Treasury Employees Union or the relevant union where collective bargaining agreements apply. Differentiate training by audience, since operators, supervisors and executives need different things. Pilot the training itself with a small cohort and iterate. Roll out in waves so support can scale with demand, and fund a dedicated help desk for the first ninety days.

Two points get missed most often. First, training must emphasise the human oversight expectation in federal guidance so that staff understand not only how to use the output but when and how to override it. Second, employees need explicit psychological safety to flag problems. Resistance is natural, and it is frequently the earliest signal of a real defect in the system rather than a communications problem to be managed. The Department of Labor and the Equal Employment Opportunity Commission have both issued guidance that AI workplace tools must not produce disparate impact on protected classes, which makes staff concerns a compliance input as well as an adoption one.

The goal is for the workforce to experience the tool as something that helps them rather than something done to them. Be honest about what changes and what does not, especially about jobs. Frame the tool as augmenting judgment rather than replacing it, and make that true by keeping humans accountable for final decisions. Support generously in the first weeks, when frustration is highest and abandonment is most tempting.

The Pilot-to-Production Readiness Gate

Run this gate before approving any scale-up. Score each item green, yellow or red. Any red blocks the next stage. The gate is an agency standard you set and record in advance, not a legal test, and passing it does not discharge any obligation that applies to your system.

  • Load tested at full scale. Performance, speed and reliability are proven at production volume, not extrapolated from the pilot.
  • Production data pipeline live. The system is fed by real, continuous, monitored data, not a clean sample.
  • Authorization boundary complete. Platform and interfaces are authorized, the system security plan exists, and the authorizing official has signed.
  • Workforce trained and bought in. The full set of users, not just pilot volunteers, are trained, supported and involved in the workflow design, with union consultation where agreements apply.
  • Appeal process operating. People affected by the system have a real, accessible path to a human review.
  • Public notice issued. Affected people know the system is in use before it decides anything about them.
  • Monitoring and drift detection. Dashboards track accuracy, fairness by group and decline over time, with someone accountable for watching them.
  • Error impact quantified. The production-scale error rate is known, who it harms is identified, and there is a way to make people whole.
  • Rollback plan rehearsed. You can revert to the prior process without disrupting service, and you have practised it.
  • Impact assessment complete. A pre-deployment assessment exists per federal guidance, with a documented human alternative for rights-affecting decisions.
  • Staged, gated rollout defined. Expansion proceeds in slices with explicit go or no-go gates, not a single cutover.

Anti-Patterns

Each of these has a documented case behind it, which is why they are worth naming rather than assuming your program is immune.

  • Treating pilot success as proof the system will work at scale. A pilot answers whether the idea can work under pilot conditions. It is evidence, not a guarantee, and it says nothing about production load, messy data or a population the pilot never sampled. Aisha's four months on 2,000 returns were real evidence and still did not predict 1.2 million.
  • Scaling without load testing. Extrapolating from pilot throughput is a guess dressed as a projection. Test at a multiple of peak volume that you set in advance.
  • Scaling without union consultation. Where collective bargaining agreements apply, this is not a courtesy, and discovering it late converts a rollout into a dispute.
  • Scaling without a rollback plan. A system you cannot turn off is a system that will keep making the same error while you debate.
  • Scaling without required public notice. Notice issued after deployment is the pattern that forced a reversal at the Internal Revenue Service.
  • Treating authorization as a box check. FedRAMP and FISMA posture is a continuous control. A certificate on file describes the day it was issued.
  • Promising leadership you will never pause. The promise sounds like confidence and functions as a commitment to keep going through evidence of harm.
  • Monitoring only aggregate accuracy. Monitoring surfaces what you chose to watch. An aggregate number can hold steady while one group's error rate doubles.

Practice Prompts

Work these against a system you actually own. The value is in the specifics, and the discomfort of the first prompt is the point.

  1. Do the error arithmetic. Take your pilot error rate and multiply it by your production population. Write the resulting number of affected people on one line. Then answer who absorbs those cases, how long redress takes, and what your appeals capacity actually is.
  2. Write your authorization critical path. List every authorization artefact your system needs, who signs each, and the current status. Identify which one is furthest from done and what it depends on.
  3. Choose your rollout pattern and justify it. Pick from shadow mode, canary, geographic, temporal and parallel operation. Say what each one would buy you for this system, and if you are skipping shadow mode, write down why.
  4. Define your pause triggers. Name the three observations that would stop your rollout, the person who can call the pause, and what happens to in-flight cases when they do.
  5. Map the workforce impact. List every role whose work changes, whether a collective bargaining agreement applies, what training each audience needs, and who staffs support in the first ninety days.
  6. Rehearse the rollback. Describe the steps to return to the prior process today, how long it takes, and what breaks. If nobody can answer, that is your gate result.

Reflection

Think about the last system your agency scaled, whether or not AI was involved. What was the first thing that broke, and was it a technology problem or an operations problem? Most people answer operations, and yet most scaling plans allocate most of their pages to technology.

Then ask the harder question about a system you are scaling now. If the evidence turned against it at the second gate, could you actually stop? Who would have to agree, what would it cost you politically, and have you told anyone in advance that stopping is a permitted outcome? A gate nobody believes you will use is not a gate. If your honest answer is that stopping would be impossible, the time to change that is before the rollout starts, not during the incident.

Glossary

  • Shadow mode: running a new system alongside the existing process, logging its outputs without using them for decisions, to compare performance directly before deployment.
  • Canary deployment: routing a small share of real traffic to a new system so that failures surface at limited scale.
  • Parallel operation: keeping the prior system running long enough that you can fall back to it if the new one fails.
  • Authorization to operate: the decision, signed by an agency authorizing official, that a system may run in production given its documented security posture.
  • FedRAMP: the federal program that standardises security authorization for cloud services, at levels including moderate and high.
  • FISMA: the Federal Information Security Modernization Act, which requires agencies to secure their information systems to a documented standard.
  • Tail latency: response time at the slow end of the distribution rather than the average, which is what the worst-served users experience.
  • Drift: the gradual decline in model performance as the real-world data or population changes away from what the model was trained on.
  • Recovery time and recovery point objectives: how quickly a system must be restored after failure, and how much recent data the agency can afford to lose.
  • Rights-impacting AI: under federal guidance, AI whose output is a principal basis for decisions affecting people's rights, benefits, services or opportunities, which triggers additional required safeguards.

Closing

Aisha's story is not a story about a bad model. Her model was good, and the pilot result was real. The failure was that nobody translated a promising experiment into an operating service with the infrastructure, authorization, workforce, oversight and redress that an operating service requires. That translation is a distinct project with its own budget, its own timeline and its own leadership, and treating it as a deployment step is the single most reliable way to lose a good idea.

The programs that make it across the gap tend to look slower from the outside and are usually faster in the end, because they do not spend six months recovering from a public reversal. They designate early, assess before deploying, shadow before switching, expand in slices, watch the groups they might harm, and keep the ability to stop. None of that is exotic. It is simply the difference between proving something can work and running it for everyone, every day, in public.

Key Takeaways

  • The pilot-to-production gap is the graveyard. Analyses from the General Services Administration AI Centers of Excellence and OMB document that up to 80 percent of federal AI pilots never reach sustained production.
  • Pilot success is evidence, not a guarantee. It was gathered under pilot conditions and predicts nothing about production load, adversarial data or a population the pilot never sampled.
  • Assess four kinds of readiness. Technical, operational, governance and risk; a weakness in any one is enough to sink the rollout.
  • Authorization is the long pole. FedRAMP posture, a FISMA system security plan, continuous monitoring and a signed authorization to operate cannot be run in parallel with go-live.
  • Do the error arithmetic before you expand. One percent of 10,000 cases is 100 people; one percent of 10 million is 100,000. The rate stays the same and the harm does not.
  • Scale in gates, never in one jump. Shadow, canary, geographic, temporal and parallel patterns each buy a different protection, and thresholds are standards you set in advance, not legal tests.
  • Change management is half the job. Map affected roles, consult unions where agreements apply, train by audience, staff a help desk for the first ninety days, and treat resistance as a signal.
  • Have a plan for failure before you go live. Know your error impact, your pause triggers and your rollback path, and rehearse the rollback while the stakes are still small.

Frequently Asked Questions

How long should the transition from pilot to production take? There is no standard duration, and any number you are given without reference to your authorization status is guesswork. The honest answer is that the timeline is set by your longest dependency, which in federal environments is usually the authorization boundary rather than the engineering. Map that critical path first, then build the rollout schedule around it.

Can we start the staged rollout while the impact assessment is still in progress? No. Under federal guidance the pre-deployment impact assessment is a gate rather than a report, which means it precedes production use of the system on real people. Running a limited slice with real consequences is production use. If timing is genuinely impossible, that is a conversation with your chief AI officer about the formal waiver and extension processes, not a decision a program manager makes quietly.

Our pilot showed 94 percent accuracy. Why would that change? Because pilot data is curated and production data is not, because edge cases appear in proportion to volume, and because the population using the service at scale differs from the sample you tested. The documented Veterans Affairs example in this lesson shows a pilot at 94 percent on 10,000 claims a month dropping to 78 percent on an annual workload of 1.3 million. Expect a drop, measure it deliberately in shadow mode, and decide in advance what level would stop you.

What is the difference between a waiver and an extension? An extension is deferred compliance: you expect to meet the practice and you are committing to milestones and a date. A waiver is a determination that a specific practice will not be met, with justification and alternative mitigations. Neither one removes the underlying civil rights, privacy or records obligations that apply to your system, and neither is a decision a project team makes on its own.

How do we handle union concerns without stalling the rollout? Start earlier than feels necessary. Where a collective bargaining agreement applies, consultation is an obligation rather than an option, and discovering that at cutover costs far more time than engaging at design. Treat the concerns raised as operational intelligence: staff who work the process daily usually identify the failure modes your test plan missed.

Who should own the production system after launch? A named program owner with budget authority, supported by a defined operations model that includes site reliability and machine learning operations coverage and an on-call rotation. Ownership by a project team that disbands at go-live is the most common way monitoring quietly stops within a year.