←
AI for Government
Proficient · M6 · lesson 6 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Pilot Program Design
📖
now learning

AI Pilot Program Design

15 min

Marcus Reilly, a program director at a mid-sized state Department of Motor Vehicles, signed a $2.4 million contract for an AI chatbot to handle license renewal questions. The vendor demo was flawless. Eighteen months later the chatbot was answering 12 percent of inquiries correctly, call center volume had not dropped, and a legislative auditor was asking Marcus to explain how the money was spent. The painful part was not the failure. It was that nobody had ever defined what success looked like before the contract was signed. There was no pilot, no exit ramp, and no way to know the thing was broken until it was fully deployed.

That is the problem a well-designed pilot solves. A pilot is a small, time-boxed, low-stakes test of an AI system against criteria you agree on before you start. It is your chance to validate vendor claims on your data, your infrastructure, and your success metrics rather than theirs. Done well it tells you the difference between a vendor who can demo and a vendor who can deliver inside your actual environment. Done badly it wastes months and delays the decision it was supposed to inform, which is why the design of the pilot deserves as much care as the evaluation of the system.

What a Pilot Is Actually For

Pilots serve three purposes. First, they validate technical claims: does accuracy match the vendor's promises on your data rather than their test set? Second, they surface implementation challenges: does the system integrate with your infrastructure, and how long does integration actually take once real systems and real approvals are involved? Third, they enable a go or no-go decision grounded in evidence from your own operational environment instead of a demonstration staged in someone else's.

Be precise about the limit of that, because it is the assumption most likely to fail. A pilot surfaces problems. It does not prevent a bad deployment on its own. Preventing one requires someone obliged to read the result and empowered to act on it, which is an organizational property rather than a property of the test. Agencies have run rigorous pilots, documented clear failures, and deployed anyway because the procurement clock was running. The pilot did its job; nobody was accountable for the finding. Design the decision alongside the test, and name the person who owns it.

Government raises the stakes on both sides. Deploy the wrong system and you damage public trust in a way a commercial buyer does not have to absorb; fail to deploy a good one and you have missed a real chance to improve service to people who have no alternative provider. Pilots reduce both risks. Government procurement is also transparent, so you may have to defend the pilot's design and its results to an inspector general, to GAO, or to Congress. A pilot designed rigorously is defensible. A pilot designed hastily, or designed to produce a particular answer, is not.

Why Government AI Pilots Quietly Fail

Pilots fail in predictable ways. They run forever with no decision point and become pilot purgatory, a system in permanent test mode that nobody dares cancel or scale. They use clean demonstration data instead of your messy real records. They measure whether the technology works instead of whether the mission outcome improves. And they almost never define a no-go, meaning the conditions under which you walk away, which turns walking away into a personal failure rather than a planned outcome that was priced in from the start.

Federal acquisition offers mechanisms for running a small paid test before committing to scale, including phased task orders and, where the agency holds it, Other Transaction Authority. Office of Management and Budget guidance on government AI use, issued in 2024 as memorandum M-24-10, expects agencies to test rights-impacting and safety-impacting AI before deployment. A pilot is how that expectation gets satisfied in practice rather than on paper, and how the resulting record becomes something you can hand to an auditor.

Setting Objectives and Scope

Start by writing down what you are trying to learn, because a pilot that tries to learn everything learns nothing on schedule. Typical objectives fall into six categories. Technical validation asks whether accuracy matches the claim and whether latency is acceptable. Integration asks whether the system can connect to your data infrastructure at all. Fairness asks whether the system performs equally across demographic groups. Operational asks how the system behaves in practice and what training staff will need. Economic asks what the real cost is, not just the license fee. Organizational asks whether staff will accept the system and what change management is required.

Scope then makes those objectives affordable. Set a duration long enough to test real usage patterns rather than a first impression. Choose a data volume that is a representative sample of your actual workload. Decide whether the pilot runs at a single location or several, knowing that multiple sites buy you generalizability at the cost of complexity. Name which teams participate and how much of their time you are actually asking for. And budget honestly for vendor support during the pilot, any third-party evaluation, and staff time, which is the line most often left out and most often the largest.

Success Criteria, Written Before You Start

Define what success means before the pilot begins. This is the single most effective step for preventing the argument that otherwise arrives at the end, when the vendor says the pilot succeeded and the agency says it did not and neither side can point to a document that settles it. Split the criteria into hard and soft, and be explicit about which is which.

Hard success criteria are objective and measurable, and failing them fails the pilot. Accuracy must reach a stated percentage measured on your data rather than the vendor's test set. Fairness requires that accuracy variance across demographic groups stays within a stated bound. Latency requires that 95 percent of decisions complete within a stated number of seconds. Integration requires the system to process your data without errors. Compliance requires demonstrated conformance with your security and privacy requirements. And cost requires that actual pilot spend does not exceed the budget estimate by more than 10 percent.

Soft criteria are qualitative. They influence the decision without triggering it automatically: whether staff feel confident using the system, whether training a person takes under 10 hours, whether the system handles edge cases gracefully rather than failing loudly, whether vendor support is responsive, and whether you can actually see what the system is doing from the outside. Soft criteria are where a system that passes every number still turns out to be unusable, so record them rather than treating them as atmosphere.

For every metric, decide in advance how it will be measured. Who measures it: your team, the vendor, or an independent evaluator? How often: daily, weekly, or once at completion? What sample size, given that the result has to be statistically meaningful rather than anecdotal? And against what baseline: the manual process, the previous system, or nothing at all? A metric without a named measurer and a stated baseline is a number you will end up arguing about.

The Five-Gate Pilot: One DMV Example Carried Through

Redesign Marcus's chatbot the right way and it runs through five gates, each ending in a written go or no-go decision. The gates are the decision spine; the phases in the next section are the schedule that sits inside them.

Gate 1: Problem and baseline

Before any technology, write the problem in mission terms and measure today's performance. Marcus's real problem was that 40 percent of his call center's volume came from three repetitive questions about renewal documents. His baseline: average wait time 14 minutes, 22 percent of callers abandoning the call, cost per handled call $6.80. If you cannot state a baseline, you cannot prove improvement, and you have no business buying anything yet.

Gate 2: Success criteria, written and signed

Define numeric thresholds and record who agrees to them. For the chatbot: correctly resolve at least 70 percent of those three question types, cut average wait by 30 percent, and hold accuracy above 95 percent on document-requirement answers. Equally important, define the no-go: if accuracy on rights-affecting answers drops below 90 percent, the pilot stops. Sign these with your finance officer and your legal counsel now, while everyone is calm and nobody has a sunk cost to defend.

Gate 3: Scoped pilot on real data

Run small and run dirty. Marcus pilots at one regional office, on real customer questions with real edge cases, for 90 days, with a hard budget of $85,000. Real data is the whole point. The vendor's demo never saw the customer typing "my license got wet in the lake, is it still good."

Gate 4: Measure against the criteria, not the vibe

At day 90 you compare results to the Gate 2 thresholds, with no moving of the goalposts. The chatbot resolves 58 percent, missing the 70 percent target, while accuracy reaches 96 percent and wait time drops 35 percent. Two thresholds met, one missed, and the no-go floor untouched. That is a partial result, which is exactly what gates are for.

Gate 5: Go, no-go, or pivot

Three honest options. Go means scale as designed. No-go means stop, document the lesson, and reclaim the budget. Pivot means the result is promising but incomplete, so you adjust scope and run a second short pilot. Marcus pivots: he narrows the bot to the two question types it handled well and adds a clean human handoff for the third. That is an $85,000 lesson rather than a $2.4 million one, and the difference is entirely in the design, not in the technology.

The Four Phases Inside the Pilot

Within Gate 3 the pilot itself has a structure, and running it in phases is what lets you find integration problems before they touch a member of the public. Phase 1, roughly weeks 1 and 2, is preparation: finalize success criteria and metrics, stand up a test environment representative of production, prepare test data representative of your actual workload, brief staff and recruit participants, and establish governance so it is clear who makes daily calls and how issues escalate.

Phase 2, roughly weeks 3 and 4, is controlled testing. The system operates on historical data and makes no live decisions, which lets you observe its behavior with nothing at stake. You measure accuracy, fairness, and reliability, discover integration issues in a safe environment, and refine configuration on the basis of early findings. Phase 3, roughly weeks 5 through 8, is parallel operation: the system runs alongside the existing process, producing recommendations while humans continue to make the actual decisions. This captures real usage patterns and edge cases, and lets you compare the system's recommendations against human decisions and measure the override rate.

Parallel operation is the most informative phase, and also the one most often misread. Comparing recommendations to human decisions surfaces errors and edge cases far better than testing in isolation, but a high agreement rate tells you the system agrees with your staff, not that either is right. Where the historical decisions carry a disparity, a system that reproduces it will look like a success. Treat agreement as a signal to investigate disagreements, and measure fairness against an external standard rather than against the pattern you already have.

Phase 4, roughly weeks 9 through 12, is evaluation and decision: complete the analysis of all metrics, make the structured go, no-go, or conditional-go call, document the findings and the rationale, and plan next steps. Those four phases sum to 12 weeks, which is why the working rule for pilot length is a minimum of 8 to 12 weeks. Anything shorter misses seasonal variation, long-tail problems, and the organizational friction that only appears once the novelty wears off. Marcus's 90-day pilot sits at the top of that range, which is appropriate for a system touching the public directly.

The Go, No-Go, Conditional-Go Framework

A go decision means all hard criteria are met, no major issues emerged, and the soft criteria are generally positive: accuracy meets or exceeds target on your data, fairness metrics are acceptable with no demographic group significantly disadvantaged, integration succeeds, cost lands within budget, and staff confidence is high with training proving effective. The next step is planning full deployment and, critically, establishing the monitoring metrics that will tell you whether the pilot result holds once the conditions change.

A no-go means hard criteria are not met and fixing them would require substantial work: accuracy remains significantly below target even after optimization, fairness problems cannot be resolved because the system inherently discriminates, integration is impossible without major infrastructure rework, cost substantially exceeds budget, or staff cannot be trained to use the system effectively. The next step is evaluating alternatives, which may mean a different vendor or a different approach entirely, including the approach of not using AI for this problem.

A conditional-go, which the five-gate framing calls a pivot, means the hard criteria are mostly met but specific issues must be fixed before full deployment. Accuracy may be acceptable but below target with the vendor committing to improvement on a specific timeline. Integration may work for most of the data with a workaround covering the rest. Fairness may be acceptable for the primary use case but concerning for a secondary one, in which case you limit deployment to the primary use initially. Cost may sit slightly above budget with savings elsewhere offsetting it. The next step is defining exactly which conditions must be met before the next phase and establishing monitoring to verify they stay met.

Notice the tension between a hard integration criterion of processing all data without errors and a conditional-go that accepts a workaround for a small remainder. That tension is real, and resolving it is your job before the pilot rather than after. Set hard criteria at levels you would actually enforce. A hard criterion you intend to waive is a soft criterion wearing a costume, and the waiver will arrive at exactly the moment when the pressure to deploy is highest.

Weighting the Decision Beyond Accuracy

A go/no-go decision is credible only if it is settled before you see the results and only if it covers more than accuracy. Weight it across the five things that actually sink government AI: mission impact, cost realism, equity and rights, security and data handling, and operational fit. A model that is 94 percent accurate but fails Section 508 accessibility requirements, the rules requiring government digital services to work for people with disabilities, is a no-go regardless of how good the accuracy looks. The same is true of a system that cannot meet your data handling obligations, however well it performs on the metric the vendor chose to feature.

Mission impact asks whether the outcome you named at Gate 1 actually moved, not whether the technology functioned. Cost realism asks what the system costs in practice rather than what the license costs, which means counting integration effort, staff time, the monitoring you will have to run afterward, and the share of the pilot budget consumed by problems the vendor said would not occur. This is the dimension where a pilot most often changes the answer, because a system that works and costs three times the projection is a different proposition from the one that was approved.

Equity and rights covers the fairness variance across demographic groups you set as a hard criterion, plus language access and accessibility, and it belongs in the decision rather than in a separate report filed afterward. Security and data handling asks whether the system demonstrated compliance with your requirements under real conditions, not whether the vendor asserted it would. Operational fit is the quietest of the five and the most common cause of a technically successful system going unused: whether the staff who have to work with it can be trained in a reasonable time, whether it handles edge cases without producing work, and whether you can see what it is doing well enough to supervise it.

Score each dimension separately and record the scores. A single blended verdict lets a strong accuracy number carry a weak equity result across the line, which is exactly the trade the decision framework exists to prevent. If one dimension fails, say which one failed and what would have to change, because that sentence is what makes the pilot defensible to an auditor a year later and what tells the next team where to start.

Your One-Page Pilot Design Canvas

Fill this out before you spend a dollar. If any box is blank, you are not ready to pilot. Adapt the rows to your mission.

ElementWhat to writeMarcus's example
Mission problemThe outcome you want to change, in plain termsCut wait time and cost for 3 repetitive renewal questions
Today's baselineCurrent numbers, measured not guessed14 min wait, 22% abandon, $6.80 per call
Pilot objectivesThe 3 or 4 things this pilot must answerAccuracy on real questions, integration, staff acceptance
Go thresholdNumeric bar that triggers scaling70% auto-resolve, 95% accuracy, 30% faster
No-go thresholdThe line that stops the pilot, no debateRights-affecting accuracy below 90%
Measurement planWho measures, how often, sample size, baselineInternal team, weekly, all pilot-office calls
Scope and durationOne site, real data, fixed time box1 regional office, 90 days
Pilot budget capHard ceiling, separate from scale budget$85,000
Data planReal records, privacy and security handlingLive questions, no PII stored by vendor
Equity and access check508 accessibility, language access, fairnessTested with screen readers and Spanish queries
Decision ownerWho signs go/no-go, with finance and legalProgram director plus CFO plus counsel
Exit rampHow you stop cleanly if no-goRevert to current IVR, no lock-in clause

Managing Risk Inside the Pilot

Five risks recur, and each has a mitigation that has to be in place before the risk materializes. Scheduling risk is the pilot running long and delaying the deployment decision; mitigate with tight project management, weekly status reviews, a clear critical path, and early identification of slippage. Scope creep is the pilot expanding beyond its original bounds, adding time and cost without adding learning; mitigate with a written scope document and a change control process in which any expansion requires approval.

Data quality risk is pilot data that is not representative of actual operations, so that nothing you learn transfers; mitigate by comparing the pilot data distribution to the production distribution and validating representativeness before you begin rather than explaining it afterward. Vendor support risk is a vendor who does not support the pilot adequately, leaving problems unresolved; mitigate with a service level commitment covering pilot support, including consequences if support proves insufficient. Staff burnout risk is real and underrated: parallel operation means people run the old process and the new system at once. Reduce other responsibilities for participants, keep the pilot to its time box, and recognize the extra effort explicitly.

Contract the Pilot So You Can Actually Walk Away

The exit ramp lives in the contract, not in good intentions. Three provisions carry the weight. A phased structure makes the pilot a separate, fully funded phase, with scaling as a priced option you may decline. Data and model portability means you keep your data and any fine-tuning in standard formats, so declining does not strand you. Performance acceptance tied to your Gate 2 criteria means payment for the scale phase depends on hitting the thresholds you wrote down in advance. Marcus's original contract had none of these, which is precisely why he was trapped. With them, a no-go is simply one of the deal terms rather than a crisis.

Anti-Patterns

  • Pilot theater. The pilot exists to justify a decision already made, and does not actually inform it. Procurement timeline pressure creates the incentive: the schedule assumes deployment, so the pilot becomes a box. The failure shows up when the pilot finds major issues and the agency deploys anyway, then discovers them again at scale. Make the go/no-go contingent on results in writing, and do not let anyone decide before the data exists.
  • Treating pre-deployment testing as a clearance. A completed pilot is evidence about a system's behavior under the conditions you tested. It is not permission, and it is not a guarantee that the system is safe to launch. Testing surfaces the problem; stopping a launch requires a named person who is obliged to read the result and empowered to act on it. Where nobody holds that obligation, a rigorous pilot and no pilot at all produce the same deployment.
  • Pilots too short to surface anything. Impatience and the desire to deploy quickly compress the pilot until it cannot see a full cycle of work. The problems then appear during full deployment, when remediation is expensive. Hold to a minimum of 8 to 12 weeks, long enough to observe real work patterns, seasonal variation, and the long tail.
  • Piloting on clean data. Vendors offer their best data for pilots, and it is always tempting to accept because it makes setup faster. The system then works perfectly in the pilot and fails in production on real records. Insist on your actual data, or a realistic simulation of it, and treat a vendor's suggestion to use curated data as information about what the system cannot handle.
  • Vague success criteria. Precise criteria are hard to write, so it is easier to be vague. Then the pilot ends, the vendor calls it a success, the agency does not, and there is no document that resolves it. Write specific, measurable criteria before the pilot starts and record them where both parties can see them.
  • No decision framework, so no decision. The pilot ends and the organization cannot decide, so it defers, and the learning decays while the system sits in permanent test mode. This is pilot purgatory, and it costs more than a no-go because it consumes staff attention indefinitely. Define the decision framework upfront and decide against the pre-agreed criteria rather than through post-hoc debate.
  • Reading agreement with human decisions as accuracy. In parallel operation, a system that matches your staff's decisions 95 percent of the time has demonstrated agreement, not correctness. If the existing process carries a disparity, matching it will read as a pass. Measure fairness against an external standard, and treat the disagreements as the most informative part of the dataset rather than as noise.
  • Waiving a hard criterion under deployment pressure. A criterion you will not enforce was never hard. If you would deploy at 88 percent against a 92 percent target, the target was 88 percent and you should have said so. Decide before the results arrive which criteria are genuinely disqualifying, and put the waiver authority somewhere other than with the person whose schedule the delay would hurt.

Practice Prompts

  1. Design a complete pilot. For an AI system you are considering, define the three or four most critical objectives to validate, the success criteria that would prove the system works, the pilot duration, the risks you would monitor, and your go/no-go decision framework. Then check whether any of the criteria are ones you would waive under pressure.
  2. Metric definitions for a benefits system. For an AI benefits determination system, define the accuracy metric including how and on what data it is measured, the fairness metric including which demographic groups and what variance is acceptable, the integration metric as a share of applications processed successfully, the cost metric as actual against budget, and the user acceptance metric including how it is measured and what threshold counts as success.
  3. Make the call. An AI hiring system pilot has completed. Accuracy came in at 88 percent against a 92 percent target, 4 points below. Fairness shows women at 85 percent and men at 90 percent, a 5 point disparity, where historical hiring showed 6 points. Integration works for 98 percent of applications. Cost landed 12 percent above estimate. Staff report confidence using the system. Is this go, no-go, or conditional-go? Check each result against the hard criteria listed earlier before you answer, including the cost threshold. If conditional-go, write the conditions, and say what the fairness result does and does not establish.
  4. Risk register. Identify five or six risks that could occur in your pilot. For each, rate probability and impact as low, medium, or high, state a mitigation strategy, and say specifically how you would monitor for it while the pilot is running rather than discovering it in the retrospective.
  5. Stakeholder plan. Plan stakeholder involvement: which teams need to participate, what time commitment you are actually asking of them, how you will communicate progress, how you will gather feedback, and how you will address concerns that emerge mid-pilot without expanding scope.

Reflection

Take an AI acquisition you are planning. What would the critical success criteria be, stated as numbers rather than adjectives? How would you design the pilot to test them, and which phase would catch the failure you most fear? What specifically would make you say no-go, and would you actually say it if the schedule had already been briefed to leadership? Which risks would you monitor while the pilot runs? How would you ensure the pilot data is representative of production rather than a curated slice? And who, by name, is obliged to read the result and empowered to stop the deployment if it comes back bad?

Glossary

  • Go/no-go decision. The point decision on whether to proceed to full deployment, based on pilot results measured against criteria set in advance.
  • Conditional-go. A decision to proceed subject to specific, documented conditions being met and verified before the next phase. Called a pivot in the five-gate framing.
  • Parallel operation. Running the new system alongside the existing process, comparing outputs while humans continue to make the actual decisions.
  • Hard success criteria. Objective, measurable criteria stated as numeric thresholds; failing them fails the pilot.
  • Soft success criteria. Qualitative criteria that inform the decision without triggering it automatically, such as whether staff accept the system.
  • Representativeness. The extent to which the pilot's data and environment match actual production conditions.
  • Baseline. The measured performance of the current process, without which improvement cannot be demonstrated.
  • Pilot purgatory. A pilot with no decision point that persists indefinitely in test mode, consuming staff attention without ever being scaled or cancelled.
  • Exit ramp. The contractual and operational means of stopping cleanly after a no-go, including phased funding, data portability, and a defined fallback to the current process.

The pilot sits inside a wider acquisition. Federal Acquisition of AI: FAR/DFARS covers the framework the phased structure depends on, Writing AI Requirements in RFPs and SOWs is where your success criteria should first appear, and AI Contract Negotiation covers the phasing, portability, and acceptance clauses that make a no-go survivable. AI Vendor Evaluation Methodology and Evaluating AI Vendor Claims address the claims the pilot is testing. AI Metrics and KPIs for Government helps define measurable criteria, Algorithmic Impact Assessments covers the rights-impacting analysis that runs alongside the pilot, and Moving from Pilot to Production takes over the moment the decision is go.

Closing

A well-designed pilot answers one question: should we deploy this system? It answers it by creating a safe environment to test real-world performance, surfacing problems while they are still cheap, and producing evidence against criteria that were agreed before anyone had a stake in the answer. Define objectives and metrics upfront, use representative data, run long enough to see a full cycle of work, manage the risks actively, and decide with the framework you wrote in advance.

The cost of a thorough pilot is trivial next to the cost of deploying a system that does not work. Marcus learned that in the most expensive available way: eighteen months, $2.4 million, and a legislative auditor. The version of his project that had a pilot cost $85,000 and ended with a narrower system that actually worked. What separated the two was not better technology or a better vendor. It was writing down what success meant while everyone was still calm, and giving someone the authority to act on the answer.

Key Takeaways

  • Define success before you spend. Numeric go and no-go thresholds, split into hard and soft, signed by finance and legal before any technology touches your environment. Vague criteria guarantee an unresolvable argument at the end.
  • Validate the vendor's claims on your data. A claim can be true for the vendor's test conditions and false for yours. Pilot on real, messy records including edge cases, and treat a suggestion to use curated data as a finding.
  • Pilot small, dirty, and time-boxed. One site, a fixed end date, and a hard budget cap separate from the scale budget turn a multi-million dollar risk into a lesson you can afford.
  • Hold to 8 to 12 weeks minimum, structured in phases. Preparation, controlled testing on historical data, parallel operation, then evaluation. Shorter pilots miss seasonal variation, long-tail problems, and organizational friction.
  • Parallel operation is the most informative phase and the easiest to misread. Agreement with human decisions is agreement, not accuracy, and reproducing an existing disparity will look like a pass. Investigate the disagreements.
  • Testing surfaces problems; people prevent deployments. A pilot is not a clearance. Name the decision owner, make the deployment contingent on the result in writing, and put waiver authority away from the person whose schedule a delay would damage.
  • Measure mission outcomes, not technology impressions. Resolution rate, wait time, and cost per case matter more than whether the demo impressed the room, and none of them mean anything without a measured baseline.
  • Include equity, accessibility, and security in the decision weighting. Section 508 accessibility, language access, fair treatment across communities, and data handling are pass/fail alongside accuracy, not considerations to be traded against it.
  • Put the exit ramp in the contract and end every pilot. Phased funding, data and model portability, and acceptance tied to your criteria are what make no-go possible. Every pilot ends in go, no-go, or one tightly scoped pivot, never a permanent test nobody can cancel.

Frequently Asked Questions

How long should a pilot run? Plan for a minimum of 8 to 12 weeks, structured as roughly two weeks of preparation, two of controlled testing on historical data, four of parallel operation, and four of evaluation and decision. Marcus's 90-day pilot sits at the top of that range, which suits a system facing the public directly. The reason for the floor is not thoroughness for its own sake: shorter pilots systematically miss seasonal variation, long-tail cases, and the friction that appears only after staff stop being on their best behavior.

The vendor wants to run the pilot on their sample data because our records need cleaning. Is that reasonable? No, and the request is itself informative. Clean curated data is exactly the condition under which the demo already succeeded, so the pilot would add nothing. If your records genuinely need work before the system can use them, that is a finding about the total cost and timeline of deployment, and you want it now rather than after signature. Pilot on your actual data or a realistic simulation of it.

Our pilot results are mixed. How do we avoid arguing about them? You avoid it by having written the criteria down beforehand and by classifying each as hard or soft at that time. Mixed results are normal and are what a conditional-go exists for: name the specific conditions to be met, the timeline, and the monitoring that verifies they stay met. What you cannot do is renegotiate a threshold after seeing the number it failed, because that converts the whole exercise into an opinion.

What if the pilot passes but we still have doubts? Take the doubts seriously and look at the soft criteria, which are usually where the discomfort lives: staff confidence, support quality, error handling, and whether you can observe what the system is doing. A pass on every hard number alongside consistent unease from the people who used it is a reason for a conditional-go with monitoring, not a reason to override the framework. Write down what specifically you would watch for after deployment.

Can we skip the pilot if the vendor already has federal customers? Another agency's experience tells you the system can work somewhere. It does not tell you it works on your data, with your integration, for your population, under your compliance obligations. Those are the four things a pilot tests and a reference call cannot. If time is genuinely short, narrow the objectives rather than skipping the test, and keep the phased contract structure so declining to scale remains an option.

Who should own the go/no-go decision? Someone senior enough to stop the project and not the person whose delivery date the delay would damage, with finance and legal signing the criteria alongside them. In Marcus's redesign the program director signs with the CFO and counsel. The point of the joint signature is not ceremony: it means no single person can quietly relax a threshold, and the decision survives the departure of whoever championed the system.