←
AI for Government
Strategic · M15 · lesson 15 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Government AI Venture Creation
📖
now learning

Government AI Venture Creation

15 min

Tomás Warrick had run the City of Phoenix Office of Digital Innovation for four years when the city manager asked him a question he could not answer: "What has this office actually produced?" Tomás could name the projects. An AI triage tool for the water-complaint hotline. A predictive maintenance model for the fleet. A document-review assistant for the city attorney. But when pressed on which of those were still running, who owned them, and whether any had expanded, the answers were uncomfortable. The triage tool was live, but nobody had reviewed its accuracy in eight months. The maintenance model had been paused when its developer left. The document assistant had been used by two attorneys and then forgotten.

Phoenix had spent $1.8 million on AI experimentation in three years and had accumulated a graveyard of prototypes. The problem was not bad projects; several of them had worked. The problem was that nothing existed around them to keep them alive, and a city that funds experiments without funding what happens next will keep producing demonstrations rather than services. The problem was no ecosystem.

The Graveyard Problem

Government AI graveyards, meaning collections of pilots that launched and stalled, share a common structure. A problem gets identified. A solution gets funded. A vendor or an internal team builds something. The something works well enough to demo. Leadership celebrates. The team moves to the next problem. Six months later, the something is still technically running, but the person who understood it has left, the data feed it relied on has changed, and the accuracy has drifted without anyone noticing. Citizens are still affected by its outputs, but nobody is checking.

That last sentence is the part that matters most and the part that gets least attention. A stalled internal tool is a waste of money. A stalled tool that still makes decisions about permits, complaints, or maintenance priorities is an unmonitored system acting on the public, and its continued operation is a choice the organization is making by omission. The graveyard is not only inefficient. It is a live risk surface that nobody has been assigned to watch.

The decay is worth separating into its parts, because each part has a different remedy. Knowledge leaves when the person who built the tool moves on, which is a documentation and handoff problem. Inputs change when an upstream data feed is altered by a team that does not know the tool exists, which is a dependency problem. Performance drifts as conditions move away from the ones the tool was built for, which is a monitoring problem. And nobody notices any of it, which is an ownership problem. Fixing only the last of the four is the common half-measure, and it is also the one that makes the other three visible.

This is not a technology failure. It is a governance failure, and specifically a failure to build the infrastructure that turns individual innovations into durable services. That infrastructure is what an innovation ecosystem is. The core analogy is a seed and a greenhouse. A single AI project is a seed. An ecosystem is the greenhouse: it does not guarantee that every seed grows, but it provides consistent conditions, funding, attention, accountability, and physical infrastructure, so that the ones worth growing can survive long enough to produce.

Building the Greenhouse

Tomás rebuilt Phoenix's approach around five structural elements. Each one addresses a specific failure mode he had watched play out, which is the only defensible reason to add structure to a government process. Read them as a set rather than a menu; the scaling gate without named ownership produces decisions nobody executes, and named ownership without a scaling gate produces owners of systems that should have been stopped.

A dedicated pipeline team

Phoenix's original model assigned innovation work to people who had other primary jobs. The result was that innovation happened when everything else was caught up, which is almost never. The new model created a three-person pipeline team with a single job: move ideas from intake to evaluation to pilot to scaling decision. Three full-time equivalents cost approximately $380,000 per year in loaded salary. Phoenix was already spending more than that annually on stalled projects that nobody was managing, which is the comparison to put in front of a budget committee rather than the cost in isolation.

Pre-positioned procurement capacity

The most expensive part of government AI innovation is time, not money. A procurement that takes 14 months kills the momentum that makes experimentation possible, because the staff who identified the problem have moved on and the conditions that made it urgent have changed. Phoenix negotiated two master service agreements, meaning pre-competed contract vehicles, with AI service vendors before identifying specific projects. When an idea cleared evaluation, the team could issue a task order in two to three weeks rather than starting a new solicitation. This does not bypass procurement law. It front-loads the competitive process so individual experiments move quickly inside an already-vetted structure.

A problem registry

Twice a year, the pipeline team runs structured interviews with frontline staff in each department. Not executive presentations, but working sessions with the people who process permits, handle complaints, and manage caseloads. The outputs go into a shared problem registry: a database of documented service problems with estimated scale, meaning how many people are affected, estimated cost in staff hours, error rates, and citizen burden, and a feasibility rating that asks whether the agency actually holds the data a solution would require. When an idea is proposed it goes into the registry, and ideas already there with strong feasibility ratings move to the front of the evaluation queue.

An explicit scaling gate

Every pilot runs for 90 days and produces a scorecard against pre-specified success criteria. A three-person panel, the innovation director, the department program lead, and the budget officer, reviews the scorecard and makes one of three decisions: scale, extend for another 90 days with a defined test, or stop. The stop decision is documented as a learning outcome, with a summary added to the problem registry so that the next team approaching the same problem starts with more information. The fleet maintenance model that had been "paused" was formally stopped, its lessons documented, and its data assets transferred to the team that later built a replacement that worked.

Named system ownership

Every AI tool that scales to production gets a named system owner, a department employee rather than a vendor contact, whose performance objectives include quarterly reporting on the system's accuracy, usage, and any complaints or errors. The owner attends the annual review where the innovation panel decides whether to continue, upgrade, or retire each production system, and ownership cannot transfer without a 30-day transition period and a written handoff document. Naming an owner is necessary but not sufficient: an owner with no allocated time, no authority to pull a system, and no reporting obligation in their objectives is a name on a spreadsheet rather than accountability.

Where to Start When You Cannot Build All Five

Few organizations get to stand up five structural elements at once, and the order in which they arrive changes how much each one is worth. The cheapest and most immediately useful is the ownership register, because assigning a named owner to every production system requires no new budget, no procurement action, and no software. It also produces, as a side effect, the inventory that every subsequent argument depends on. An organization that cannot say what it is running cannot make a case for anything else.

The scaling gate comes next, because it is the element that stops the graveyard growing while everything else is being built. It costs three people a meeting per pilot and a written scorecard, and it converts the most expensive habit in government innovation, letting a pilot end without a decision, into a documented outcome. The registry follows naturally once stops are being documented, since the lessons have to go somewhere, and a registry that starts life holding what was learned from stopped work is more credible than one that starts as a list of aspirations.

The two elements that need real money, a dedicated pipeline team and pre-competed contract vehicles, are best argued for last, on evidence the first three produce. A budget committee responds differently to a request framed as capacity for future ideas than to one framed as the current annual cost of unmanaged systems that the inventory has just made visible. That was the shape of the Phoenix argument: the pipeline team's loaded cost was set against what the city was already spending on stalled projects nobody was managing.

Venture Models for Smaller Jurisdictions

The innovation models used by large federal agencies do not scale directly to a city or county with 2,000 to 20,000 employees and a $300 million operating budget. Phoenix is a large city, and much of what worked there assumes a dedicated office and a budget line that smaller jurisdictions will never have. Three models travel better, and the useful question when choosing among them is what the jurisdiction is actually short of: money, capacity, or scale.

Innovation funds work at any size. A competitive internal grant programme, funded from a budget set-aside at $15,000 to $75,000 per award, lets department staff propose solutions to their own problems. The process is lighter than a full procurement: a two-page proposal, a scoring rubric, a panel review. Successful applicants get money and a quarter of the pipeline team's time, and failed projects produce a lessons-learned report for the registry. Several Colorado counties have run programmes like this with annual budgets between $100,000 and $300,000, which is the scale at which an internal grant round is worth the administrative effort.

Shared-service platforms suit jurisdictions short of scale rather than money. A city with 300 employees cannot afford to build its own AI infrastructure, but a consortium of fifteen small cities in the same region can negotiate a shared contract for AI-assisted permit review or a shared document-drafting tool, spreading the cost across the group. The National League of Cities has supported cooperative procurement agreements for exactly this purpose. The governance question is the hard part rather than the contract: who decides when the shared tool changes, and who is accountable when it errs in one member city.

University partnerships suit jurisdictions short of technical capacity. A partnership with a state university's public-policy school or computer-science department can supply graduate student teams to prototype under faculty supervision. The city provides the problem definition and the data access; the university provides technical execution. The city must retain ownership of the resulting system, and that has to be specified in the partnership agreement rather than assumed. These arrangements typically produce results in four to six months and cost $20,000 to $60,000, mostly in data-access infrastructure and staff time for oversight.

None of the three removes the need for the greenhouse. A grant-funded prototype with no scaling gate becomes a graveyard entry faster than a procured one, because nobody was ever assigned to decide its fate. Whatever the jurisdiction's size, something has to own the decision about what happens after the prototype works.

The three models also combine more usefully than they compete. A university partnership can produce the prototype, an innovation fund can pay for the work that makes it deployable, and a shared-service consortium can carry the running cost across jurisdictions too small to bear it alone. What breaks that chain is nobody holding the handoffs: a prototype delivered to a city with no one funded to operationalize it is a graveyard entry with an academic paper attached.

Measuring Ecosystem Health

An innovation ecosystem that cannot report on its own health will not survive the first budget review that asks what it has produced. Tomás could not answer that question in year four, and no amount of project enthusiasm substitutes for the answer. Phoenix now tracks three categories of metric in a quarterly report to the city manager, chosen so that each one answers a different sceptic: the one who thinks the office is idle, the one who thinks it is producing nothing useful, and the one who suspects it is automating existing inequities.

Pipeline health shows the ecosystem is functioning: ideas evaluated, pilots launched, scaling decisions made including the ratio of stops to scales, and average elapsed time from intake to scaling decision. A healthy pipeline for a city of Phoenix's size processes 15 to 25 ideas per year and produces 3 to 6 production deployments, which is a local benchmark from one large city rather than a national standard. The stop-to-scale ratio is the number worth watching, because an ecosystem that never stops anything is not evaluating, it is approving.

Service impact shows the ecosystem is producing value: processing time changes, error rate changes, cost per transaction, and citizen satisfaction, each tied to the original problem statement for the tool in question rather than to a general claim about efficiency. Phoenix also reports an aggregate, that tools currently in production affect an estimated 34,000 citizen interactions per month, which is useful for conveying scale and useless for conveying benefit. Both kinds of number belong in the report, clearly labelled as what they are.

Equity tracking shows the ecosystem is not automating inequity. For each production tool affecting citizens, Phoenix tracks outcomes disaggregated by neighbourhood, used as a proxy for income and race, and by language. A tool that processes permit applications 40% faster on average but takes 20 additional days for applications submitted in Spanish has not improved the service; it has made a disparity worse while producing a headline improvement. Equity metrics belong in the quarterly report alongside the efficiency metrics, in the same table, read by the same audience.

The framing that keeps all three honest is simple enough to say in a budget hearing. The innovation ecosystem's job is not to produce AI. Its job is to produce better public services. AI is one method, and the method is not the outcome.

Cadence matters as much as content. A quarterly report to the city manager means the numbers arrive before they are demanded, on a rhythm that lets a trend become visible while it can still be acted on, and it removes the incentive to assemble favourable figures in response to a specific question. It also means the equity breakdown is published in quarters when nothing is wrong, which is what makes it credible in the quarter when something is. A report produced only at budget time is a defence, and everyone in the room reads it as one.

Anti-Patterns

  • Counting launches instead of survivals. An office that reports pilots started is reporting activity. The number that answers the city manager's question is how many tools are in production, owned, and still monitored, which is a much smaller and much more useful figure.
  • Treating a named owner as sufficient. Ownership only functions when it comes with allocated time, a reporting obligation written into performance objectives, and the authority to pull a system. Without those, the name in the register changes nothing about what actually gets watched.
  • Pausing instead of stopping. A paused project is an undocumented stop that keeps its budget line and loses its lessons. Make the stop formal, write down what was learned, and move the data assets somewhere the next team can find them.
  • Starting the procurement after the idea. If every experiment needs its own solicitation, the pace of experimentation is set by the acquisition calendar. Pre-competed vehicles move that cost to a point where it is paid once rather than every time.
  • Running the registry as a wish list. A problem registry without scale, cost, and feasibility ratings is a list of complaints. The ratings are what turn it into a queue that can be defended when somebody's favourite idea is not at the front.
  • Reporting the aggregate and omitting the breakdown. An average improvement can conceal a group for whom the service got worse. Publishing the aggregate without the disaggregation is not a reporting shortcut; it is the mechanism by which a disparity survives review.

Practice Prompts

  • Inventory your graveyard. List every AI or automation tool your organization has built or bought in the last three years. For each, record whether it is running, who owns it, when its accuracy was last reviewed, and whether it still affects the public.
  • Cost the alternative. Estimate what your organization currently spends annually on unmanaged or stalled projects, then compare that to the loaded cost of a small dedicated pipeline team. Bring both figures to the same conversation.
  • Draft the scorecard. For a pilot you are planning, write the success criteria before it starts, name the three people who will read the scorecard, and specify what evidence would justify each of scale, extend, and stop.
  • Write one system owner's objectives. Take a production tool and draft the ownership paragraph you would add to a named employee's performance objectives, including the reporting cadence and the authority to suspend the system.
  • Pick your constraint. Decide whether your jurisdiction is most short of money, technical capacity, or scale, then design the corresponding model: an internal innovation fund, a university partnership, or a shared-service consortium with neighbouring agencies.

Reflection

  • If your city manager or agency head asked what your innovation function has produced, what would you say, and how much of the answer would be projects rather than services?
  • Which of your production systems is currently affecting the public without anyone having reviewed its accuracy this year?
  • When your organization last stopped an AI project, was the stop documented as a learning outcome, or did the project simply go quiet?
  • Who would have to agree before a shared tool with neighbouring jurisdictions could be changed, and has anyone written that down?
  • Of the metrics your innovation function reports, which one would reveal a disparity if one existed, and would anybody read it?

Glossary

  • Innovation ecosystem. The standing infrastructure of funding, attention, accountability, and process that turns individual innovations into durable services, as distinct from the projects themselves.
  • Problem registry. A shared database of documented service problems recorded with estimated scale, estimated cost, and a feasibility rating, used to queue ideas and to preserve lessons from stopped work.
  • Scaling gate. A pre-specified review at the end of a pilot at which a panel decides to scale, extend with a defined test, or stop, against criteria written before the pilot began.
  • Named system owner. A department employee accountable for a production system's accuracy, usage, and complaints, with reporting obligations in their performance objectives and a documented handoff on transfer.
  • Master service agreement. A pre-competed contract vehicle negotiated before specific projects are identified, allowing task orders to be issued quickly within an already-vetted competitive structure.
  • Stop-to-scale ratio. The proportion of pilot decisions that end a project against those that expand it, used as an indicator of whether evaluation is genuine.

Closing

The question Tomás could not answer was fair, and it had a structural answer rather than a defensive one. Phoenix was not short of ideas, funding, or capable staff. It was short of the unglamorous machinery that decides what happens to an idea after it works: who evaluates it, who stops it, who owns it afterwards, and who notices when it drifts. Building that machinery is slower than building tools and considerably less interesting to talk about, and it is the difference between an office with a portfolio and an office with a graveyard.

The test to apply to your own function is the one the city manager applied. Not what have you started, but what is running, who owns it, and what has it done for the public this quarter. An ecosystem that can answer those three questions in a routine report will survive budget scrutiny. One that cannot will keep producing prototypes until somebody stops paying for them.

Key Takeaways

  • Prototype graveyards are a governance failure. Tools that launch and stall are not a technology problem. They are the result of missing infrastructure: no named owners, no scaling gates, and no accountability for what happens after the demo.
  • A stalled tool that still runs is a live risk. Systems that keep affecting the public while nobody monitors them are not dormant. They are unsupervised decisions, and the organization is making them by omission.
  • Pre-positioned procurement buys the time that makes experimentation feasible. Master service agreements let task orders move in weeks rather than months. This does not bypass competitive procurement; it front-loads it so individual pilots move inside a legal framework.
  • A problem registry operationalizes institutional knowledge. Documenting problems with scale, cost, and feasibility ratings, linked to lessons from prior attempts, means each team starts with more information than the last instead of rediscovering the same problem every few years.
  • Explicit scaling gates prevent both zombie projects and premature cuts. A 90-day evaluation with a documented stop, scale, or extend decision treats every outcome as informative, and it requires success criteria written before the pilot rather than after.
  • Named ownership is necessary and not sufficient. Every production system needs one named government employee, not a vendor, responsible for accuracy, reporting, and complaints, with the time and authority to act on what they find.
  • Small jurisdictions have real options. Internal innovation funds at $100,000 to $300,000 per year, cooperative procurement with neighbouring agencies, and university partnerships all work for organizations that cannot fund a dedicated innovation office.
  • Equity metrics belong in every production report. Aggregate improvements that mask demographic disparities are not improvements. Track outcomes disaggregated by neighbourhood, income proxy, and language from the first pilot forward.

Frequently Asked Questions

Where do you start if there is no innovation function at all?

Start with the inventory and the ownership register, because both are cheap and both produce the argument for everything else. Listing what your organization has already built, who owns each item, and when its accuracy was last checked usually surfaces enough unmanaged production systems to justify a pipeline team without any further business case. Building a registry and a scaling gate before there is a queue of ideas tends to produce process nobody uses.

Does pre-competing contracts really stay within procurement rules?

Front-loading is not avoidance. The competitive process still happens; it happens once, before specific projects exist, so that individual task orders can be issued inside a structure that has already been competed. Phoenix negotiated two master service agreements this way and moved from cleared idea to task order in two to three weeks. Confirm with your own acquisition office which vehicles are available to you, because the answer varies by jurisdiction.

What makes a scaling gate more than a rubber stamp?

Two things: success criteria written before the pilot starts, and a stop decision that carries no penalty for the team that proposed the idea. A panel that has never stopped anything is approving rather than evaluating, which is why the stop-to-scale ratio belongs in the quarterly report. Documenting the stop as a learning outcome, with the lessons filed in the registry, is what makes the decision worth making.

How small is too small for any of this?

A city with 300 employees cannot build its own AI infrastructure, but it is not out of options. It can join a consortium negotiating a shared contract, run a small internal fund, or partner with a state university for prototyping at $20,000 to $60,000 and four to six months per engagement. What it cannot skip is deciding who owns anything that reaches production, which costs nothing and is the element most often missing.

What single metric would you report if you could only report one?

The count of tools currently in production with a named owner and a review in the last quarter. It is unglamorous and it is the only number that distinguishes a portfolio from a graveyard. Everything else, including citizen interaction volumes and processing time improvements, means something only once you know the systems producing those numbers are being watched by someone with a name.