←
AI for Government
Proficient · M41 · lesson 41 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Scaling Capstone: From Your Pilot to Enterprise
📖
now learning

Scaling Capstone: From Your Pilot to Enterprise

15 min

Raymond Cho is a program director at a large city's department of transportation. His team ran a six-month AI pilot that read 311 service requests and routed them to the right crew automatically. It was a hit: routing time dropped from two days to under an hour, and the pilot covered one district. Then the deputy commissioner said the words that sink most government AI: "Great, roll it out citywide by the end of the quarter." Raymond knew the trap. A pilot that works in one district with a hand-picked team and a vendor on call is not a system that survives across twelve districts, four union jurisdictions, a procurement office and a public records request. Scaling is not the same thing, only bigger. It is a different discipline.

This is your capstone. You have learned to run a pilot. Now you build the scaling plan that turns a promising experiment into durable enterprise capability, or, just as valuably, the plan that tells you honestly when not to scale. The spine of this lesson is Raymond's 311 routing tool and the real-world gauntlet it has to pass to go citywide, and the output is a document you could actually hand to a budget office, a union representative and an auditor without rewriting it three times.

Why Most Government Pilots Die at the Scaling Step

Pilots are designed to succeed. They run in friendly conditions: eager staff, clean data, a vendor watching closely, and a narrow scope. Those same conditions are exactly what disappears at scale. Government also scales differently from the private sector, because government serves everyone. Management approaches that work for a fifty-person organisation do not survive contact with a population in the millions, across diverse communities and wildly different local contexts, and the difference is one of kind rather than degree.

The honest reasons government AI pilots fail to scale are predictable, and naming them early is what makes a scaling plan credible instead of aspirational.

  • The pilot ran on heroics. One enthusiastic analyst fixed problems by hand every morning. Twelve districts cannot run on one person's mornings.
  • The data was cleaner than reality. The pilot district happened to have well-tagged requests. Other districts free-text everything.
  • The money was a grant, not a budget line. Pilots are often funded by one-time money; scaling needs a recurring appropriation and a sustainment plan.
  • Nobody designed for oversight. A pilot can dodge procurement, accessibility and audit questions. An enterprise system cannot.

The fix is to stop asking whether the pilot worked and start asking what it would take to run this for five years, for everyone, under public scrutiny. That question reorganises everything, because it pulls in operations, money, governance and people at the same time instead of leaving them to be discovered in sequence by whoever is unlucky enough to be on duty.

The Five Dimensions of a Real Scaling Plan

A credible scaling plan addresses five dimensions. Raymond's plan has a section for each, and a go or no-go answer for each. If any one is a clear no, the project does not scale yet.

Operations: who runs it at 3 a.m.?

At pilot scale, the vendor and one analyst handled everything. At enterprise scale you need defined operations: a help desk path, a named owner per district, monitoring for when routing accuracy drifts, and a documented procedure for when the model is wrong. Raymond's plan replaces "the analyst checks it each morning" with an automated daily accuracy report and an on-call rotation. The test of an operations section is whether it still works when the person who built the pilot leaves, because at some point that person always does.

Data and technical readiness: does it work outside the friendly district?

Before going wide, test the model on the messiest district, not the cleanest. Raymond's team ran the routing tool against six months of unedited requests from the district everyone considered worst. Accuracy fell from 94% to 71%. That is not a failure. It is the single most valuable number in the whole plan, because it tells him exactly what to fix before scaling rather than after a public embarrassment. Data consistency across districts is usually the deepest technical problem in a government rollout, and it is rarely visible until someone deliberately goes looking for it.

Money: pilot cost versus true cost of ownership

Pilots hide the real bill. Total cost of ownership at scale includes licensing across all districts, staff time to monitor, retraining the model as request patterns change, and the cost of the human review step. Budget for hidden costs as well as direct ones: integration work, help desk load in the first months, and the staff time that oversight actually consumes. A scaling plan that shows only licence fees is not a business case, it is a quote.

Governance and compliance: surviving public scrutiny

Enterprise scale invites oversight the pilot avoided. This is where federal frameworks become practical tools rather than paperwork. The NIST AI Risk Management Framework, which is voluntary and non-binding, gives Raymond a structure to document risks and controls that auditors already recognise. Federal AI guidance issued in 2024 requires agencies to inventory AI uses and assess impact on the public, and a citywide routing tool that shapes where crews are sent is exactly the sort of system whose rights and safety impact needs to be assessed rather than assumed. Anything citizen-facing must meet Section 508 accessibility requirements. Answer these now, on your terms, rather than under a council inquiry later.

People and change: will the crews actually use it?

A pilot team volunteers. An enterprise rollout meets people who did not ask for change, including unionised crews whose work assignments the tool now influences. Raymond's plan includes early engagement with union representatives, training for dispatchers across every district, and a clear message that the tool routes requests but humans still own assignments. Adoption, not technology, is where most scaled systems quietly fail, and the training that worked for six volunteers rarely survives being delivered to six hundred people who were told to attend.

What a Complete Scaling Plan Contains

The five dimensions tell you what to think about. The plan document itself has a standard shape, and using it saves you from discovering at the review that you have three pages on architecture and nothing on contingency. A complete scaling plan covers risk assessment, infrastructure requirements, change management, resource allocation, metrics and monitoring, contingency plans, timeline and success criteria.

Written out as deliverables, that becomes an executive summary that states why scaling matters and gives the high-level plan; a detailed scaling plan covering current state, future state, the phases with their timelines, go or no-go criteria for each phase, and the risks and mitigations attached to each; a business case with projected costs, benefits and payback period; an organisational changes section covering staffing, training and change management; a technical architecture section covering infrastructure, data, integration with existing systems and scalability analysis; a governance and monitoring section defining how you will monitor during scaling, which indicators you will watch and who decides; and a risk register listing your top risks with likelihood, impact and mitigation.

Two of those sections carry most of the weight in a government review and are usually the thinnest. Contingency planning asks what you do when a risk materialises, not merely how likely it is. Go or no-go criteria per phase force you to write down, in advance, the condition that would stop you. Those criteria are thresholds your agency sets for itself and records before the evidence arrives. They are not legal tests, and clearing them does not make a deployment lawful or discharge any obligation that applies to your system; they exist so that the decision to continue is made against a written standard rather than against the mood in the room.

The Business Case, Done Honestly

A business case has to demonstrate that scaling is justified: projected costs, projected benefits, and a risk-adjusted view of the return. Risk adjustment is the part teams skip. A benefit that depends on adoption reaching every district should be discounted by the chance that adoption stalls in three of them, and a plan that presents the optimistic case as the expected case will be read that way by exactly one audience, which is the one that approves it and then holds you to it.

Raymond's five-year cost picture, which is illustrative rather than a benchmark for your agency, breaks down like this.

  • Software licensing across twelve districts: about $140,000 per year.
  • Monitoring and operations staff at half a full-time equivalent: about $70,000 per year.
  • Periodic retraining and tuning: about $60,000 per year.
  • Human review step: about $10,000 per year.
  • One-time integration and rollout: about $300,000.
  • Comparison point: an estimated $4.2 million in routing labour over the same five years.
  • The pilot itself cost $80,000, which is the number everyone remembers and the least useful one in the table.

Do the arithmetic yourself rather than quoting a headline. The recurring lines add to $280,000 per year, which over five years is $1,400,000, and the one-time integration and rollout sits on top of that. An earlier version of this lesson reported a five-year total of $1.4 million, which matches the recurring lines alone and omits the one-time cost, so the stated total does not follow from the stated inputs and has been removed. The line items are what matter: they are what a budget office will interrogate, and a total that nobody can reconstruct from them is the fastest way to lose the room.

The point survives the correction. A pilot funded from leftover grant money becomes a seven-figure five-year commitment, and even a strong comparison against the labour it displaces does not conjure a recurring appropriation. Scaling is a budget conversation before it is a technology conversation, and the plan that names the sustainment line is the one that gets funded twice.

Resource requirements deserve their own answer rather than being folded into the cost figure. Be specific and realistic about what you need in four categories: money, people, infrastructure and training. People is the one that gets estimated worst, because monitoring, retraining, help desk coverage and the human review step are all staff time that has to come from somewhere, and half a full-time equivalent on a spreadsheet is a real person whose other duties do not disappear. Infrastructure needs a scalability analysis rather than an assurance, meaning a stated view of what happens to capacity and response time at full volume. Training needs a delivery plan for every district and every audience, not a single session recorded once.

Stakeholders, Risk and Contingency

Stakeholder analysis is not a slide with logos on it. Identify each stakeholder, what they actually want, the objection they are most likely to raise, and your strategy for engaging them. For Raymond that means the deputy commissioner who wants speed, the budget office that wants the recurring line justified, the union representatives whose members' assignments the tool influences, the district supervisors who will absorb the disruption, the procurement office, and the residents who file the requests and never see any of this. Each of those has a different question and none of them are satisfied by the same briefing.

Risk work has three parts that teams routinely collapse into one. Assessment asks what could go wrong and which risks carry the highest impact. Mitigation asks what you will do to reduce likelihood or severity in advance. Contingency asks what you will do when it happens anyway. The third is the one that turns a risk register from a document into a control, because it forces a named owner and a trigger rather than an adjective. For the 311 tool, the obvious candidates are the system becoming a bottleneck at citywide volume, data quality varying across districts, and adoption resistance from dispatchers and crews.

Define success metrics before you scale rather than after. Deciding what good looks like once the numbers are in is how a rollout gets declared a success on the strength of whichever indicator happened to move. Write down the indicators you will watch during and after scaling, who reviews them and how often, and what level would trigger a pause. Then keep watching the ones that could embarrass you, not just the ones that could flatter you.

A Worked Example From Another Domain

Scaling plans read better with a second example that is not your own. Consider a national employment services capstone. The pilot serves 10,000 job seekers a month in one region with a 65% successful placement rate. Full scale means 1 million job seekers a month across the country, with a target placement rate of 75%. That is a hundredfold increase in volume, which checks out against the stated figures and is the number that drives everything else in the plan.

The key challenges follow directly from the scale change: infrastructure upgrades to carry a hundred times the load, data consistency across regions that record things differently, training employment officers in many locations at once, and integration with different existing systems in each region. The phasing runs in three stages: months one to six expand to five regions and 100,000 seekers a month; months seven to twelve expand to fifteen regions and 500,000 a month; months thirteen to eighteen reach all regions and 1 million a month. Note that the plan does not assume regions are the same size, and neither should yours.

The go or no-go criterion is the sharpest part of the example: if phase one does not achieve a 70% placement rate, the program does not proceed to phase two. That threshold sits between the pilot's 65% and the target of 75%, which is what makes it a real gate rather than a formality. It was written down before the evidence arrived, it is a standard the agency set for itself, and it names a specific number that a specific person would have to overrule in public.

One caution about the business case attached to that example. Its investment figure, its annual economic value and the resulting return ratio arrive without units and cannot be reconstructed, so they are not reproduced here, and its stated benefit of improved outcomes for 100,000 job seekers a year does not follow from its own inputs: a ten point improvement applied to a million seekers a month is an order of magnitude larger. The inputs are sound and are given above. Build the benefit figure from them yourself, state the units, and show the arithmetic, because a reviewer who catches one unsupported number stops trusting all of them.

The Scaling Decision Scorecard

Use this scorecard to convert a gut feeling into a defensible decision. Score each dimension red, yellow or green. Any red means stop and fix before scaling. Mostly yellow means a limited expansion rather than a citywide one. All green means go. Like every threshold in this lesson, the scorecard is a standard you adopt and record in advance, not a legal test.

DimensionThe hard questionRed (stop)Green (go)
OperationsWho runs it without heroics?Depends on one personNamed owners, monitoring, on-call
Data readinessDoes it hold up in the worst district?Untested on messy dataTested; accuracy known and acceptable
MoneyIs the five-year cost funded?One-time grant onlyRecurring budget line approved
GovernanceCan it survive an audit today?No risk assessment, no inventory entryDocumented risks and controls, impact assessment done
PeopleWill the affected staff use it?No engagement, no trainingStakeholders engaged, training ready

Scale in Stages, Not in One Leap

Even with green across the board, never go from one district to twelve overnight. Raymond's plan expands in waves: fix the messy-data problem, then add three districts, watch for a month, then add four more, then the rest. Each wave has a checkpoint where the project can pause or stop. This staged approach means a problem affects three districts for a month rather than the whole city for a year. It also generates the evidence, real accuracy and cost figures at scale, that justifies the next wave to the deputy commissioner and the budget office.

Each wave should also produce the artefacts the next reviewer will want: updated accuracy by district, incident counts, help desk volume, and the operational cost actually incurred rather than the one forecast. A staged rollout that produces no new evidence is just a slower single cutover, and it spends the caution without buying anything with it.

Presenting and Defending the Plan

A capstone is not finished when the document is written. It is finished when it survives being presented and improved. Put the plan in front of peers and people who will be affected by it, and treat the objections as findings rather than as obstacles. The three questions that expose a weak plan fastest are what happens when the champion leaves, where the recurring money comes from, and which single observation would make you stop.

Incorporating feedback is part of the deliverable, not an optional courtesy. A plan revised after review is stronger evidence of judgment than a plan defended intact, and in government the people who raise the sharpest objection early are usually the people who will have to operate the thing. Their objection is your cheapest test.

The deepest lesson of the capstone is that a confident "not yet" is a successful outcome. Raymond's worst district showing 71% accuracy did not kill the project; it focused it. Scaling well means knowing exactly which question would make you stop, and being honest enough to ask it before the public does.

Anti-Patterns

  • Treating pilot success as proof the system will scale. A pilot's result is evidence gathered under pilot conditions: friendly data, volunteers, close vendor support and narrow scope. It says little about production load or about districts and populations the pilot never touched.
  • Treating scaling as a non-critical activity. It looks like a nice-to-have next to delivery pressure, so it gets no resources, and the problems accumulate until they undermine the initiative. Allocate resources to the transition and measure its progress like any other program.
  • Copying a plan from another jurisdiction unchanged. It is tempting to reuse what worked elsewhere, and it fails because your context is different. Use frameworks as guides rather than templates, and adapt them deliberately.
  • A strategy with no accountability. Accountability feels like overhead, so nobody is named and nothing is tracked. The plan then remains aspirational. Assign responsibility, define the indicators, track progress and hold people to it.
  • Testing only on your best data. The accuracy number from the tidiest district is the least informative one you will produce. Test where it will hurt, before the public does it for you.
  • Quoting a total nobody can reconstruct. A five-year figure that does not follow from the line items destroys the credibility of the whole business case. Show the arithmetic and the units.
  • Go or no-go gates you were never willing to use. A gate that has never stopped anything is decoration. Name in advance who can call the pause and what happens to work in flight.
  • Funding sustainment from one-time money. A grant buys a pilot. Only a recurring line keeps a system monitored, retrained and supported in year three.
  • Engaging unions and affected staff after the rollout schedule is fixed. By then the only available response is resistance, and the operational knowledge you needed arrives as an objection instead of as a design input.

Practice Prompts

Work these against a pilot you actually own. The capstone deliverable is the assembled answer, and each prompt maps to a section of it.

  1. Assess the current state. Write the pilot honestly: what worked, what did not, what the conditions were, and which of those conditions will not exist at scale.
  2. Describe the future state. What does full deployment look like in operations, coverage, staffing and citizen experience? Be concrete enough that someone could disagree with you.
  3. Phase it. Define phases with timelines, what happens in each, the go or no-go criterion for each, and the risks and mitigations attached to each phase.
  4. Build the business case. List projected costs including hidden ones, projected benefits, and the payback period. State units on every figure and show how each derived number was computed.
  5. Do the stakeholder map. For each stakeholder, record their interest, their likely objection and your engagement strategy. Include the people who will operate the system and the people it acts upon.
  6. Write the risk register. List your highest risks with likelihood, impact, mitigation, contingency and a named owner. Contingency and owner are the columns that make it real.
  7. Define the metrics before you scale. Name the indicators, the review cadence, the decision-maker, and the level that would trigger a pause. Include at least one metric that could embarrass you.
  8. Test on your worst data. Run the model against the messiest real data you can obtain and record the accuracy drop. That number belongs in the plan whatever it says.
  9. Present it and revise it. Take the plan to peers and to the people it affects, capture the objections, and record what you changed as a result.

Reflection

Look at the pilot closest to you and ask where its heroics are hiding. Somebody is almost certainly doing something manually every morning that nobody has written down, and at scale that person becomes a single point of failure with a calendar. Naming that task is often the fastest route to an honest operations section.

Then ask the question that the whole capstone turns on. What is the one observation about this system that would make you recommend stopping, and have you told anyone what it is? A plan that cannot answer that has no gates, only checkpoints that get waved through, and the difference only becomes visible when the evidence turns bad and everyone discovers that pausing was never really an option.

Glossary

  • Total cost of ownership: the full five-year cost of running a system, including licensing, monitoring staff, retraining, human review, integration and support, rather than the pilot price.
  • Sustainment funding: a recurring budget line that keeps a system operating after the one-time money that built it runs out.
  • Go or no-go criterion: a condition written down in advance that must be met before a phase proceeds, set by the agency for itself.
  • Risk register: the list of a program's significant risks with likelihood, impact, mitigation, contingency and a named owner for each.
  • Contingency plan: what the program will do when a risk materialises, as distinct from the mitigation intended to stop it happening.
  • Risk management: identifying, assessing and mitigating potential problems as a continuing discipline rather than a one-time exercise.
  • Data governance: the rules and processes for managing organisational data responsibly, including ownership, access and quality.
  • Data quality: the accuracy, completeness, consistency and reliability of the data a system depends on.
  • Data pipeline: the system that collects, transforms and moves data from source to destination on a continuing basis.
  • Portfolio management: managing multiple initiatives together so that scarce resources go where they produce the most value.
  • Shared assets: models, data or systems used and maintained collaboratively by more than one department or agency.
  • Continuous improvement: iterative enhancement of a system based on data and feedback after it is in operation.

Closing

The capstone is not really a document exercise. It is the habit of asking, before anyone is committed, what this system will cost every year, who runs it when the champion moves on, which population it has never been tested against, what would make you stop, and who has to agree that stopping is allowed. Programs that can answer those five questions tend to scale slowly and successfully. Programs that cannot tend to scale quickly and then appear in a hearing.

Raymond's plan did not promise a citywide rollout by the end of the quarter. It promised a fixed data problem, then three districts, then a checkpoint, with the accuracy in the worst district on the front page rather than buried. That is a less exciting document than the deputy commissioner asked for and a considerably more persuasive one, because every number in it can be traced to something the team actually measured. Build yours the same way, and be willing to write "not yet" where the evidence says so.

Key Takeaways

  • Scaling is a different discipline from piloting. A pilot proves the idea can work under pilot conditions; scaling proves you can run it for everyone, for years, under public scrutiny, without heroics.
  • Government scale is different in kind. Approaches that work for a fifty-person organisation do not transfer to a population in the millions across diverse communities and contexts.
  • Address all five dimensions. Operations, data readiness, money, governance and people; any single red means stop and fix first.
  • Test on your worst data, not your best. The accuracy drop in the messiest district is the most valuable number in the plan, and finding it early is what makes the plan credible.
  • Build the business case from line items with units. Show projected costs including hidden ones, projected benefits and the payback period, and make sure every derived figure follows from its inputs.
  • Budget for sustainment, not just deployment. Licensing, monitoring, retraining and human review turn a pilot funded from one-time money into a seven-figure five-year commitment that needs a recurring line.
  • Write go or no-go criteria in advance. They are thresholds your agency sets for itself, not legal tests, and their value comes entirely from your willingness to use them.
  • Separate mitigation from contingency. One reduces the chance a risk occurs; the other says what you do when it occurs anyway, with a named owner.
  • Roll out in waves that produce evidence. Staged expansion limits the blast radius and generates the real accuracy and cost figures that justify the next wave.
  • A confident "not yet" is success. The best scaling plans name the exact condition that would make you stop, and ask it honestly before the public does.

Frequently Asked Questions

How long should a scaling plan be? Long enough that each of the standard sections is genuinely answered and short enough that a deputy commissioner will read it. The structure matters more than the page count: an executive summary, a phased plan with go or no-go criteria, a business case, organisational changes, technical architecture, governance and monitoring, and a risk register. If any of those sections is missing, the plan has a hole regardless of its length.

What if leadership will not accept a phased rollout? Bring evidence rather than caution. The strongest argument is a measured accuracy drop on your messiest real data, because it converts an abstract argument about pace into a specific defect with a fix and a timeline. The second strongest is the blast radius comparison: a problem affecting three districts for a month against the same problem affecting the whole city for a year.

Our pilot was funded by a grant. How do we handle the money question? Treat the recurring cost as the headline rather than the footnote. Build the annual figure from line items, name the year the one-time money runs out, and put the sustainment ask in the same document as the expansion ask. A scaling plan that quietly assumes next year's funding is a plan to stop the system in eighteen months.

Do we need an impact assessment for a routing tool that does not decide anything about a person? Ask it as a question about consequence rather than about the tool. Federal guidance issued in 2024 requires agencies to inventory AI uses and assess impact on the public, and a system that shapes where crews are sent affects which neighbourhoods get service and when. Take the designation question to your governance body with the analysis written out rather than deciding it inside the project team.

How do we engage unions without stalling the schedule? Start before the schedule is fixed, because engagement after the dates are set can only produce resistance. Where a collective bargaining agreement applies, consultation is an obligation rather than a courtesy. In practice the crews and dispatchers also know the failure modes your test plan missed, so the earliest conversation is usually the most valuable one you will have.

What makes a capstone plan good rather than merely complete? Traceability and honesty. Every number can be tied to something the team measured, every derived figure shows its arithmetic, every phase has a criterion that could actually stop it, and the plan says clearly what it does not yet know. A plan that survives peer review with its weaknesses visible is worth more than one that hides them and fails in production.