Public-Private Innovation at Scale
When Daniel Osei became chief innovation officer for a state with twelve million residents, he inherited a graveyard of pilots. Eleven AI proofs of concept with eleven different vendors, each impressive in a demo and none of them in production. A chatbot that answered four hundred questions but could not connect to the real benefits system. A document-processing tool the vendor priced at $40,000 for the pilot and $3.2 million to scale. The state had spent two years and roughly $6 million proving that AI could work, and had nothing that served a citizen at scale. Daniel called it innovation theater, and he set out to end it.
Public-private innovation at scale is the discipline of turning the genuine capability of the private sector into durable public value without surrendering control, accountability or the public interest along the way. The hard part is not finding a clever vendor. It is designing a partnership that survives a budget cycle, an election, a vendor acquisition and an inspector general review, and still delivers. This lesson covers the failure patterns that kill pilots on their way to production, how to match a partnership model to a problem, and the governance scaffold you put in front of a procurement board.
One boundary is worth setting at the start. The structural anatomy of a government partnership, the ownership terms, the audit rights, the data sovereignty questions and the conflict-of-interest machinery, is worked in detail in Public-Private Partnerships for Government AI, and this lesson does not repeat it. What follows is the portfolio problem that sits on top: why individually reasonable engagements fail to add up to anything in production, and what changes when the unit of decision is a program of work rather than a single contract.
Why Pilots Die and Partnerships Scale
Daniel's eleven dead pilots were not technical failures. Every one of them worked in the demo, which is exactly what made the pattern hard to see from inside. They were design failures, and they shared a structure. Four failure modes account for most of what goes wrong between a working proof of concept and a service the public can rely on.
- The integration cliff. A pilot runs in a sandbox with fake data. Scaling means connecting to live systems of record, meeting security requirements and handling real volume. That work is often ten times the pilot effort and rarely budgeted.
- The pricing trap. Vendors price pilots to win and production to profit. The $40,000 pilot that becomes a $3.2 million deployment is not a surprise to the vendor; it is the plan. Without pricing the full lifecycle up front, leaders get committed before they see the real bill.
- The lock-in risk. A partnership in which the vendor owns the model, the data and the integration leaves the government unable to switch, unable to renegotiate, and often unable to explain what it bought.
- The accountability gap. When the system makes a flawed decision about a citizen, the vendor's algorithm did it is not a defensible answer to a legislator or a court. The government owns the outcome regardless of who built the model.
A demo proves the technology can work. A partnership proves the public can rely on it. Leaders who confuse the two fund the first and never reach the second, and they do it repeatedly, because each new pilot looks like evidence that the last one was simply the wrong vendor. Eleven proofs of concept are not eleven experiments. They are one experiment run eleven times with the same design flaw.
The fourth failure mode is the one that reaches you personally, so it is worth stating plainly. The government owns the outcome of a citizen-facing decision whoever built the model, whoever hosts it and whoever tuned it. A partner can be liable to you under a contract, and that is a separate matter from the agency being answerable to the person whose benefit was denied, to the oversight body that asks how the decision was made, and to the court that will not accept a proprietary system as an explanation. Any partnership structure that leaves you unable to answer those questions has misplaced a responsibility you cannot actually delegate.
What a Partnership Has to Survive
Designing for the signing day is easy. The test is whether the arrangement is still working years later, and there are four predictable stressors between here and there. A budget cycle can cut the operating line while leaving the service running and the public expecting it. An election can change the priorities that justified the program, its executive sponsor, or both at once. A vendor acquisition can replace the people who understood your deployment and the commercial logic that made your account worth serving. An inspector general review can ask for evidence about decisions made two years ago by a system nobody documented.
Each of those has a design answer, and each answer costs something at the outset. Operating cost stated up front survives the budget cycle better than a cost that materializes as a surprise. A program defined by the citizen outcome it delivers survives an election better than one defined by the technology it uses. Assignment, continuity and exit terms are what an acquisition tests. Documentation, logging and audit access are what a review tests. None of this is expensive to specify at the beginning and all of it is close to impossible to add once the relationship is under strain.
The Cliff and the Bill
The integration cliff deserves its own treatment because it is where the money actually goes. A sandbox pilot is a demonstration of model capability. Production is an exercise in systems of record, identity, error handling, throughput, monitoring, records retention and the security authorization that has to be in place before the system touches live data. None of that is what the pilot proved, and none of it is what the pilot budget covered. The claim that the work often runs to ten times the pilot effort is a warning about proportion rather than a planning figure. Budget it as a distinct phase with its own estimate.
Daniel's chatbot shows the cliff in miniature. It answered four hundred questions and could not connect to the real benefits system, which means it could tell a resident what the rules were and could not tell them anything about their own case, let alone act on it. That gap is not a shortfall in the model. It is the entire distance between an information service and a transaction, and it is bridged by integration work, identity handling and authority to write to a system of record. A capability that can only answer is a capability that leaves the citizen exactly where they started.
Read Daniel's numbers carefully, because they answer different questions. The $40,000 and the $3.2 million are two quotes for one capability at two stages, and the second was a price for scaling rather than money the state spent. The roughly $6 million is the portfolio total across two years and eleven pilots, and the source does not break it down, so do not derive a cost per pilot from it and present that as a finding. What the figures support is the qualitative point: a pilot price tells you almost nothing about a production price, and a portfolio can consume real money while producing nothing a citizen can use.
Matching the Model to the Problem
Not every problem calls for the same structure. Match the model to the stakes, to the maturity of the technology, and to how much control you need to retain. The choice is made early, often implicitly, and it constrains everything that follows.
- Procure a finished service. Buy a mature, commercially available capability through standard acquisition. Best when the need is common and the technology is proven. Lowest risk, least customization.
- Co-development partnership. The government and a partner jointly build something for a specific public need, with shared milestones and the government retaining rights to the model and the data. Best for problems no off-the-shelf product solves. Higher effort, far less lock-in.
- Challenge or prize model. Pose a public problem with a defined reward and let many private and academic teams compete. Best for early-stage problems where you want options before you commit. Spreads risk across competitors.
- Shared infrastructure or consortium. Several agencies or jurisdictions pool resources to build common capability once and reuse it. Best when many bodies face the same need. Highest coordination cost, strongest economies of scale.
Daniel's mistake had been running every problem as an isolated co-development pilot, which is the most expensive structure available and the one that maximizes the number of distinct integrations you must eventually pay for. Some of his problems needed only a procured service. One was a strong candidate for a multi-county consortium. Matching model to problem cut his portfolio from eleven scattered bets to four focused programs, which is a consolidation rather than a cancellation: the same needs, carried by structures that could reach production.
The model choice is worth making explicitly because it decides things you will not revisit. It sets who does the integration work and therefore where most of the cost lands. It sets what you own at the end and what you can take to another provider. It sets how many separate relationships your team has to manage, which is a real constraint in an office that is smaller than its portfolio. An agency that never names the model has usually assembled one by default, and the default is whatever the first enthusiastic vendor proposed.
Where you procure a mature service, security clearance status is part of the assessment. A capability already assessed against a government cloud security baseline, such as FedRAMP, starts from a stronger position than one that has never been examined. Be precise about what that tells you. An authorization describes a provider's service as assessed against a security baseline, at a point in time and within a defined boundary. It does not establish that the model is accurate for your decisions, lawful for your data, or accountable for your outcomes, and it does not replace your own authorization decision. Confirm scope and current status with your security officer rather than reading a listing as a clearance for your use case.
The Consortium Move
The consortium model is the one most often skipped, and it is the one that most directly addresses scale. Its logic is straightforward. Where many bodies face the same need, building the capability once and reusing it converts a set of individually unaffordable projects into one that several budgets can carry together. It also concentrates the integration work, which is the expensive part, into a place where it can be done properly once instead of badly many times.
The cost is coordination, and it is real. A shared capability needs an owner, a funding mechanism that survives the departure of any single participant, a way of settling requirements when participants disagree, and an answer to what happens when one member wants to leave. Those questions are harder than the technology and they are the reason consortia stall. Settle them before the build rather than during it, because a shared service without a governance answer degrades into whichever participant is willing to keep paying for it, which is a dependency the others did not choose and cannot control.
A Governance Scaffold for Partnerships That Scale
Whatever model you choose, the governance belongs in the agreement rather than improvised afterwards. Use the scaffold below as the spine of any AI partnership above trivial scale. It is what Daniel now requires before a deal reaches the procurement board, and its function is to force a set of questions while the government still has the leverage to insist on the answers.
- Full-lifecycle cost, in writing. Require the partner to price the pilot, the integration, production, and five years of operation and support up front, so that the small pilot with the large deployment attached is visible before you commit.
- Data and model rights. Specify who owns the trained model, the training data and the outputs. Insist on the right to export your data and to obtain enough documentation to switch providers. Treat lock-in as a contract term rather than as an accident.
- Accountability clauses. Write in that the government retains decision authority over citizen-facing outcomes, that the partner must support audits and explain model behavior, and that performance and fairness standards are contractual obligations with consequences attached.
- Security and compliance gates. Require the system to meet your security categorization and the relevant standards before it touches live data, not after. Make production access contingent on passing those gates.
- Exit and continuity plan. Define what happens if the partner is acquired, fails or is terminated. Citizens cannot lose a service because a company changed hands.
- Public-value metrics. Set success criteria in terms of citizen outcomes, such as faster processing, fewer errors and equitable results across groups, rather than vanity measures such as queries answered. Review them with the partner on a fixed cadence.
Understand what a clause is and is not. Every item on that list is a remedy, a right or a trigger, and none of them prevents the failure it addresses. An exit clause does not keep a vendor solvent; it determines what you hold when the vendor is not. An audit right does not make a model explainable; it obliges someone to try. The scaffold is worth building because rights you did not reserve are rights you will not have in the conversation where they matter, not because a signed agreement makes any of these outcomes impossible.
The corollary is that the scaffold needs someone to operate it. Name the person who reviews the lifecycle cost against actuals, who exercises the audit right on a schedule rather than only after an incident, and who tests the exit plan against a partner that no longer exists. A governance scaffold with no owner after award is a set of unexercised rights, and unexercised rights have a way of turning out to be unusable in exactly the circumstances they were written for.
Measuring Public Value Rather Than Activity
The deepest fix Daniel made was changing what counted as success. His predecessor had measured activity: pilots launched, vendors engaged, demos held. Daniel measured public value instead, which meant how many citizens were served at production scale, how much faster a real transaction completed, whether outcomes were equitable across communities, and the total cost per citizen served. Under the old measures, eleven pilots looked like a thriving innovation program. Under the new ones they looked like what they were, which was money spent to serve nobody at scale.
Activity metrics are not chosen because anyone is dishonest. They are chosen because they are available early, they move quickly, and they can be reported before any outcome exists. That makes them genuinely useful for tracking whether a program is moving, and useless for deciding whether it is working. Keep them as operational telemetry, and keep them off the slide that answers whether the public got value. Pick measures that a skeptical legislator and a served citizen would both recognize as real, and publish them at a cadence that does not let a bad quarter disappear.
Define the public-value measures precisely enough that they cannot drift. Cost per citizen served depends entirely on who counts as served, and a figure computed over completed transactions is not comparable with one computed over everyone eligible or everyone who made an attempt. The same trap catches transaction speed, which improves whenever the difficult cases are routed somewhere the clock does not follow. Write the definition down with the numerator and the denominator stated, keep it stable across reporting periods, and never compare two figures whose denominators were not the same.
Anti-Patterns to Avoid
- Treating a demo as evidence of feasibility at scale. A proof of concept on sample data establishes that a model can produce plausible output in favourable conditions. It says nothing about live systems of record, real volume, error handling or security authorization, which is where the effort and the money actually sit.
- Accepting a pilot price as an indication of the production price. The two are set by different logic, and a partner with a low entry price and no published lifecycle cost has an incentive you have not priced. Ask for pilot, integration, production and multi-year operating costs before the pilot, when you can still walk away.
- Running every problem as a bespoke build. Co-development for a need that already has a mature commercial answer spends public money reproducing something you could configure, and it multiplies the integrations you will eventually maintain. Choose the structure from the problem, not from the habit.
- Letting the portfolio grow by addition. Each new pilot is defensible on its own and the collection is not. When nothing is ever consolidated or stopped, a program accumulates commitments that no budget can carry to production, and the choice about what actually ships gets made by attrition.
- Reading a security authorization as a clearance for your use. An assessment against a baseline covers a defined service boundary at a point in time. It is not a statement about the model's accuracy for your decisions, the lawfulness of your data flows, or who answers when the system gets a citizen's case wrong.
- Reporting activity as public value. Pilots launched, vendors engaged and demos held describe what the innovation office did. Citizens served at production scale, transaction time, equity of outcomes and cost per citizen describe what the public received, and only the second set can close a program down.
Practice Prompts
- Inventory the graveyard. List every AI pilot your organization has run in recent years and mark each one as in production, abandoned, or still running as a pilot. For everything not in production, write the single sentence that explains why, and see how many of your sentences are the same sentence.
- Price the cliff. Take one live pilot and estimate, with your technical and security staff rather than with the vendor, what production would require: systems integration, security authorization, monitoring, records handling and support. Compare that estimate with the pilot cost and put both in front of whoever approved the pilot.
- Reclassify your portfolio. Assign each active engagement to one of the four models. Where a co-development effort could have been a procured service, or where several units are separately building the same thing, propose the consolidation and name what would have to be true for it to happen.
- Test one exit. Choose your most consequential AI partnership and write what the agency would hold if the partner ceased trading next month: which data, in what format, with what documentation, and how long the service would be degraded. Turn each gap into proposed contract language.
- Rebuild the scorecard. Replace the activity measures in your current innovation reporting with citizen-outcome measures, then present both versions side by side to your leadership. Note which programs look different under the second version, because those are the ones your governance has not been seeing.
Reflection
How many of your organization's AI engagements have reached production, and how many of the rest failed for the same structural reason rather than for a technical one? If your largest partner disappeared next quarter, what would your agency still hold and what would citizens lose? Which of your current measures would keep improving even if no citizen was ever served at scale? And when a partnership was last renewed, who asked whether that structure was still the right one, or was the renewal simply the path of least resistance?
Glossary
- Innovation theater. A portfolio of visible AI activity, pilots, demonstrations and announcements, that produces no service the public can actually use at scale.
- Integration cliff. The step change in effort between a sandbox pilot and a production system connected to live systems of record, at real volume and under security requirements.
- Full-lifecycle cost. The priced total of pilot, integration, production deployment and multi-year operation and support, established before commitment rather than discovered afterwards.
- Vendor lock-in. The condition in which a partner's ownership of the model, the data or the integration removes the government's practical ability to switch, renegotiate or independently understand the system.
- Consortium model. A structure in which several agencies or jurisdictions pool resources to build a shared capability once and reuse it, trading coordination cost for economies of scale.
- Public-value metric. A success measure expressed in citizen outcomes, such as time to complete a transaction, error rates, equity across communities, or cost per citizen served.
Related Lessons
- Public-Private Partnerships for Government AI covers the partnership structures, IP non-negotiables and conflict-of-interest management that this lesson deliberately does not repeat.
- Moving from Pilot to Production works the integration cliff at the level of a single system.
- Vendor Lock-In Prevention goes deeper into data portability, exit planning and switching rights.
- OMB M-24-18 and AI Procurement Governance and AI Contract Negotiation address the acquisition instruments that carry these terms.
- Shared Services and Infrastructure Models expands the consortium approach and its funding mechanics.
- Managing AI Vendor Performance covers operating the relationship once the scaffold is in place.
- AI Metrics and KPIs for Government supplies the measurement discipline behind public-value reporting.
Closing
Daniel's problem was never a shortage of willing partners or of capable technology. He had eleven of each. What he lacked was a way of deciding which structure fitted which problem, what the whole thing would cost before he was committed, and what he would count as success. Those three decisions are what separate a program that reaches production from one that generates a steady supply of impressive demonstrations and an eventual audit finding.
If you take one thing from this lesson into your next engagement, take the sequence. Decide the model from the problem before you meet the partner. Price the lifecycle, not the pilot. Reserve the rights you will need in the conversation you hope never to have. Name someone to operate the scaffold after award. And measure the thing a citizen would recognize as a result, because a portfolio measured on its own activity will keep telling you it is succeeding right up to the moment somebody asks who has actually been served.
Key Takeaways
- Pilots die from design, not technology. The integration cliff, pilot pricing, lock-in and the accountability gap kill scaling far more often than any technical limit does.
- A demo is not a partnership. Proving the technology works is the easy part. Proving the public can rely on it across budget cycles, elections and audits is the actual work.
- Match the model to the problem. Procure mature services, co-develop for genuinely novel needs, run challenges when you want options, and form consortia when many bodies share one need.
- Price the full lifecycle up front. Pilot, integration, production and multi-year operating costs in writing, so the small pilot with a large deployment attached is visible before commitment.
- Make lock-in a contract term. Ownership of model, data and outputs, plus export and switching rights, settled before signature rather than negotiated during a dispute.
- Accountability stays with the government. The public owns citizen-facing outcomes whoever built the model, so decision authority, audit support and fairness standards belong in the agreement.
- A clause is a remedy, not a control. The scaffold determines your position when something fails and needs a named owner to exercise it, since unexercised rights tend to be unusable when they are finally needed.
- Measure public value, not activity. Citizens served at production scale, real transaction speed, equity across communities and cost per citizen are the measures that can close a failing program.
Frequently Asked Questions
Our pilot succeeded. Why would we not simply scale it with the same vendor?
You might, but decide it rather than drift into it. A pilot establishes that a model can work on a bounded problem; it does not establish that the partner can integrate with your systems of record, meet your security requirements, operate at volume, or price production reasonably. Ask for the full lifecycle cost and the exit terms before you extend, because the moment the pilot becomes the production path is the moment your negotiating leverage disappears.
Is a security authorization enough to let a service handle our data?
No. An assessment against a government cloud security baseline tells you a provider's service was examined against that baseline within a defined boundary and at a point in time. It says nothing about whether the model is accurate for your decisions or whether your particular data flows are lawful, and it does not substitute for your own authorization decision. Treat it as a starting position and confirm the scope and current status with your security officer.
How do we justify consolidating a portfolio when every pilot has a sponsor?
By changing the comparison. Individually, each pilot is defensible; collectively, they represent commitments that no budget can carry through integration and production. Show the lifecycle cost of taking all of them to production against the available funding, and show what the same money buys if the duplicated needs are carried once. Consolidation is easier to defend as a route to production than as a cancellation.
What if we cannot get the data and model rights we want?
Then you have learned the price of that structure before signing, which is the point of asking early. Sometimes the right answer is a different model: a procured commodity service where lock-in matters less, or a consortium with enough combined demand to change the terms available. What you should not do is proceed on the assumption that the terms can be improved later. They are hardest to change once your data and workflows are inside the partner's platform.
Who should own the governance scaffold after the contract is signed?
A named person with the authority to act, not a committee that receives a report. Someone has to compare lifecycle cost against actuals, exercise the audit right on a schedule rather than only after an incident, and check the exit plan against a partner that no longer exists. If nobody holds those tasks, the agreement still contains the rights and the agency will discover it cannot use them at the moment it needs them most.
Skill.re