Agile and Iterative AI Development
Marcus Trujillo had been an IT project lead at the state Department of Motor Vehicles (DMV) for nine years, and he had buried two large IT projects in that time. Both had died the same way: eighteen months of requirements documents, a fixed-price contract, a system delivered to spec, and a spec that turned out to be wrong. So when the DMV director asked him to lead an AI pilot to verify the authenticity of uploaded identity documents, automatically flagging likely forgeries before a clerk ever saw them, Marcus said yes, and then said something that made the director pause: "I am not going to write you a plan that tells you exactly what we will deliver in month twelve. I cannot, and anyone who says they can is guessing." This lesson is about why Marcus was right, and the charter he wrote instead.
Why AI Projects Are Not Like Traditional IT
Traditional IT projects in government usually run on a waterfall model: gather requirements, design, build, test, deploy, each phase finishing before the next begins. Waterfall works when the problem is deterministic, meaning the same input always produces the same, knowable output. A payroll system that calculates a paycheck is deterministic. You can specify it completely in advance because the rules are fixed.
AI is probabilistic, not deterministic. Marcus's document-verification model does not "know" a license is forged. It produces a probability, a confidence score, that a document is genuine or fake. Three things follow from that, and each one breaks waterfall.
- You cannot fully specify performance in advance. Nobody can promise 95 percent accuracy in a requirements document, because the achievable accuracy depends on the data, and you do not know what the data can support until you build something and test it.
- The system is data-dependent. An AI model is shaped by the examples it learned from. If the DMV's historical fraud cases skew toward one document type, the model will be strong there and weak elsewhere, and you only discover that by measuring.
- Performance drifts. Forgery techniques evolve. A model that performs well at launch degrades as the real world changes underneath it. A deterministic system does not quietly get worse; an AI system does.
Set the two worldviews side by side and the mismatch is stark. Waterfall assumes you can specify requirements completely upfront, that implementation follows the design once the design is settled, that testing validates against the original design, and that problems discovered late are cheap to fix. For AI, requirements discover themselves as you learn what is achievable, systems rarely follow a predictable design trajectory, testing often reveals that the fundamental approach needs to change, and late-stage problems are catastrophically expensive. Every one of the four assumptions inverts. That is not a reason to manage AI work loosely. It is a reason to manage it differently.
Two Agencies, One Problem, Two Outcomes
Marcus had watched a neighboring state specify a benefits-eligibility AI to hit 92 percent accuracy, build it to spec, and then discover at acceptance testing that their data could only support 78 percent. The shape of that failure is worth tracing precisely. The agency spent six months gathering requirements, naming exact accuracy targets, the fairness metrics that mattered, and how the system should work. They designed it, contracted for development, and built for twelve months. Testing revealed the gap. At that point they were eighteen months in with a redesign ahead of them, more than two million dollars spent, and roughly two years behind. The spec was not wrong about what they wanted. It was wrong to assume anyone could promise it before measuring.
A different agency faced the same business problem and ran it the other way. In four weeks they built a very simple eligibility screening model, a baseline that did nothing more than flag obviously-eligible cases. It reached 40 percent accuracy, which sounds like a failure until you notice it already solved the core problem: it was faster than manual review. They deployed it to two pilot offices. Over the next eight weeks they added features based on pilot feedback, brought accuracy to 68 percent, and refined their fairness metrics against real outcomes. Within four months they had a deployed system serving six offices at 82 percent accuracy, with a fairness analysis showing no significant disparities.
Neither team was smarter than the other. The difference was where each one put its learning. Waterfall is assumption-driven, so problems surface late. Iteration is learning-driven, so problems surface early and get fixed cheaply. Waterfall in government typically means planning for eighteen months, spending twelve gathering requirements and designing, twelve building, six testing, then missing the timeline and going over budget anyway. When the work is as uncertain as AI work, that pattern does not merely disappoint. It fails catastrophically, and it fails at the most expensive possible moment.
Why Agile and Iterative Fits AI
Agile development treats learning as a deliverable. Instead of one long arc toward a fixed spec, the work proceeds in short cycles called sprints, typically one to two weeks, each producing a working model and a measurement. You build the simplest thing that tests your biggest unknown, measure it, and decide what to do next based on evidence rather than assumption.
For AI this is not a stylistic preference. It is the only honest way to manage a probabilistic, data-dependent system. The question "can our data support a useful forgery detector at all?" should be answered in week three for a few thousand dollars, not in month eighteen for the price of the whole contract.
The loop itself is not complicated. Start with clarity on the business problem and the success metrics, because requirements still matter and iteration is not an excuse to skip them. Build a minimum viable version quickly, in weeks rather than months. Get it in front of users or run it against real data. Measure performance against the success criteria. Learn what is working and what is not. Iterate and improve on that learning. Maintain regular stakeholder visibility throughout, so that nobody has to wait for a milestone to find out where things stand.
The benefits are concrete rather than cultural. Time to value shortens, with a first deployment in about three months rather than eighteen. Problems surface earlier, while they are still cheap to fix. Stakeholder alignment improves because the learning happens in the open. Risk drops because you are not betting the entire program on one big design. And cost estimates get more realistic, because you are estimating against what you have proved feasible rather than against what you hoped. In a resource-constrained agency that last point is decisive. Iteration lets you prioritise and make deliberate tradeoffs instead of discovering halfway through that the approach will not work.
MVP, pilot, and production are three different things
Marcus kept his team disciplined about three terms people tend to blur:
- Minimum viable product (MVP): the simplest version that tests the core assumption and produces real metrics. For Marcus, the MVP was a model that flagged only the most obvious forgeries, built in two weeks. It was not impressive. It answered the only question that mattered first: is there signal in this data at all?
- Pilot: a bounded real-world deployment to a small set of users or cases, with a defined evaluation window. Marcus's pilot ran in two field offices for 90 days, with clerks reviewing every AI flag.
- Production: the system operating at scale as the standard process. You do not get here until the pilot has produced evidence against pre-registered thresholds.
What Minimum Viable Actually Means
An MVP for AI is not a perfect system scaled down. It is the simplest version that tests your core assumptions and provides value to users. The distinction matters because the wrong reading of "minimum viable" produces two opposite errors: a team that builds everything, and a team that builds something unusable and calls the shortfall a virtue.
Take fraud detection as the worked case. The wrong approach is to build a system with every feature you eventually want, network analysis, temporal patterns, transaction behaviour modelling, test it thoroughly, and deploy. That takes eight months and might not work at all. The right approach starts with the simplest thing that could possibly help: rules-based flagging of obvious fraud patterns, which does not require machine learning at all. Deploy that baseline in two weeks. Then measure. Is it catching obvious cases? Is it reducing manual review time? If yes, start adding machine-learning features incrementally. If no, either the data is fundamentally unsuitable or the basic approach is wrong, and you have learned that early and cheaply instead of eight months in.
A well-built MVP answers five questions that no requirements document can answer on its own. Does the basic approach solve the business problem? Can we get access to the data we need? Do users actually want what we are building? What is the achievable accuracy with our data? Are there fairness concerns we did not anticipate? Notice that three of the five are about the world rather than about the software, which is exactly why they cannot be settled in advance.
Five properties separate a real MVP from a demo. It should be deployable, meaning it runs against real data or real users under whatever authorization applies, not a prototype that only works in a notebook. It should be measurable, so you see genuine performance metrics rather than impressions. It should be minimal, the simplest version that tests the core assumptions. It should be valuable, so users and stakeholders see something worth having immediately. And it should be iteratable, with a clear path to improvement in the next sprint.
Sprint Structure for an AI Pilot
Marcus structured the pilot as a sequence of two-week sprints, each ending in a demo and a measurement, not a document. The shape looked like this:
- Sprints 1 to 2, baseline. Explore the data, build the simplest model, get a first accuracy number on held-out cases. Deliverable: a working script and an honest baseline.
- Sprints 3 to 4, improve and check fairness. Add features, raise accuracy, and run a disaggregated analysis to check whether error rates differ across document types and demographic groups. Deliverable: improved model plus a fairness readout.
- Sprints 5 to 6, build the human workflow. Design the clerk review screen, the override path, and the monitoring dashboard. The model never decides alone. Deliverable: the human-in-the-loop interface.
- Sprints 7 onward, pilot in the field. Deploy to two offices, measure against pre-registered thresholds, gather clerk feedback, decide go or no-go. Deliverable: pilot evidence.
Filled in with goals, deliverables and a demo line, the same cadence looks like the table below. The demo line is the discipline that keeps a sprint honest, because it forces the team to say out loud what a stakeholder will actually see at the end of two weeks.
| Weeks | Goals | Deliverables | Demo |
|---|---|---|---|
| 1 to 2 | Data exploration, simple baseline model, initial validation | Working script producing predictions on sample data; initial accuracy metrics | "Here is our data, here is what a simple model produces" |
| 3 to 4 | Add features, improve accuracy, begin fairness analysis | Improved model with a 5 to 10 percent accuracy gain; fairness analysis on protected classes | "Accuracy improved, here is our demographic performance analysis" |
| 5 to 6 | Build serving infrastructure, design human review workflow, plan monitoring | Model-serving interface, human review screen, monitoring dashboard design | "Here is what the system will look like in production and how reviewers will use it" |
| 7 to 8 | Pilot with a small user group, gather feedback, refine on real usage | System deployed to one pilot office; initial user feedback; performance metrics | "The system is live with pilot users, here is what we are learning" |
| 9 to 12 | Address pilot findings, optimise performance, prepare broader rollout | Improvements from feedback; expanded pilot; final validation | "Pilot results and user satisfaction, and whether we are ready for broader rollout" |
Five principles hold the cadence together regardless of the calendar. Each sprint produces working software or a working model, not just documents. Sprints have clear, measurable goals. Measurement happens at the end of every sprint, covering accuracy, fairness and user feedback. Learning from each sprint informs the next sprint's plan. And stakeholders see progress on a fixed rhythm rather than on request.
Human-in-the-loop during iteration, not just at the end
A mistake Marcus refused to make was treating human oversight as something bolted on at the finish line. From the first field sprint, every AI flag went to a clerk who could accept or override it, and every override was logged. Those overrides were not failures. They were the richest data the team had, showing exactly where the model was wrong and why. Human-in-the-loop was both a safeguard and a learning engine.
Measuring Progress Across Four Dimensions
You cannot manage what you do not measure, and in iterative AI development a single accuracy number is not enough to steer by. Four families of measurement matter, and a team that tracks only the first will optimise the wrong thing for months without noticing.
Accuracy metrics answer whether the model is getting better and how fast. A running readout might show sprint 1 at 65 percent, sprint 2 at 71 percent, a gain of six points, and sprint 3 at 74 percent, a gain of three, against a target of 85 percent. Resist the temptation to draw a straight line through those points and announce an arrival date. The gains in that sequence are decelerating, six points then three, which is the normal shape of model improvement and the reason naive projections overshoot. Report the trajectory and the deceleration, and let the next sprint's measurement update the estimate rather than confirming a promise.
Fairness metrics ask how accuracy varies across groups and whether iteration is introducing or reducing disparity. A sprint 2 readout showing 71 percent overall, 68 percent for one group and 74 percent for another carries a six-point disparity, which is a signal to investigate before proceeding rather than a number to note and move past. Aggregate accuracy can rise while a gap widens underneath it, so both have to be read together.
Business metrics connect the model to the mission. How much reviewer time is the system saving? How has cost per unit of work changed? What has happened to the experience of the people the process serves? A sprint 3 impact readout might show manual review time falling from 40 minutes per case to 15 minutes, and cost per decision falling from $12 to $4, with user satisfaction improving on iterative feedback. These are the numbers an executive can act on.
Engagement metrics tell you whether anyone is actually using it. How many cases is the system processing? What is the override rate, meaning how often do humans disagree with the AI? Are users confident? A sprint 4 readout of 150 cases per week, an 8 percent human override rate and 95 percent completion of user training suggests a system people trust and know how to operate. Update all four families every sprint, track them visibly, and use them to drive sprint planning rather than to decorate a status report.
Governance Gates That Coexist With Agile
Here is the objection Marcus heard constantly: agile means moving fast, government means oversight, and the two cannot mix. He found that false. Agile and governance coexist when you build the gates into the sprint cadence rather than treating them as a separate, slower track.
The NIST AI Risk Management Framework helps here. Its MEASURE function asks you to assess performance, fairness, and robustness on a regular cadence, which maps cleanly onto end-of-sprint measurement. Agile does not weaken measurement; it schedules it every two weeks instead of once. For a rights-affecting system, OMB Memorandum M-24-10 requires an AI impact assessment and meaningful human oversight before a rights-impacting AI is deployed. Document verification that can block someone's license is plausibly rights-impacting, so Marcus treated the impact assessment as a hard gate before the field pilot, not paperwork to backfill afterward.
The reframe that unlocked it: governance gates are not the opposite of kill-switches. They are the kill-switches. A gate that says "we do not proceed to field deployment until the impact assessment is signed and fairness thresholds are met" is exactly an agile stop-or-go decision, written in advance so it cannot be argued away under deadline pressure.
Running Agile Inside Government Constraints
Marcus could not pretend the DMV was a startup. Three real constraints shaped how agile he could actually be.
- Procurement. He could not spin up a new contract every sprint. He scoped the entire pilot to a vendor already on an existing statewide contract vehicle, so iteration happened inside a signed agreement rather than waiting on a new one.
- Security authorization. The system needed an Authority to Operate (ATO) before touching real resident data. Marcus ran early sprints on a small set of de-identified historical documents so the team could learn while the ATO process ran in parallel, then moved to live data only once authorization cleared.
- Change control. Production systems sit under change-control boards. Marcus kept the pilot environment separate from production so the team could iterate freely, and routed only vetted, stable releases through change control. Iteration speed and change discipline lived in different environments.
Managing Stakeholder Expectations
One of the hardest parts of iterative AI work has nothing to do with models. Stakeholders arrive expecting waterfall-style progress, a plan with dates and a line that only goes up, and the expectation has to be reset explicitly rather than absorbed quietly. Marcus had the conversation once, at the start, in plain terms, and referred back to it every time the pressure rose.
The substance of that conversation is worth having word for word. We are taking an iterative approach. We will release a very simple version in four weeks; it will not be perfect, but it will work and it will teach us what is achievable. We will improve it every sprint based on what we learn. You will see progress every two weeks through demos and metrics. We may discover that the original timeline was optimistic or pessimistic, and we will know that after sprint 2. We will make tradeoffs in the open: we can improve accuracy further, but it means delaying launch, so what is your priority? And success means achieving the agreed success criteria, not following the original plan.
When the first results disappoint
It is common for sprint 1 to look bad. An MVP might show 55 percent accuracy against a target of 85. That is normal, and more importantly it is information. The framing that keeps a project alive at that moment is honest rather than defensive: this is expected, the MVP exists to test whether the basic approach works and what the data can support, and a result in that range shows there is signal to improve on. Say what you learned, say what you are doing differently next sprint, show the path you believe leads to the target, and make sure everyone understands this was the learning goal of the MVP rather than a missed deadline. Do not promise the arrival sprint you cannot yet evidence.
Demo discipline
Demos every two weeks are what make expectation management possible, but only if they are run with discipline. An effective demo states the sprint goal, what was accomplished, the metric and how it moved, what the team learned, the plan for the next sprint, and then takes questions. An ineffective demo is a long technical deep-dive nobody follows, a tour of code or model internals, a report on effort with no results, or a vague claim that the system is working better. The second kind erodes exactly the confidence the cadence exists to build.
Knowing When to Iterate and When to Ship
At some point the question becomes: do we stop iterating and deploy for real? The criteria for shipping are a checklist, and every item is a gate rather than a preference. Success criteria are met, or an agreed acceptable tradeoff has been reached. Fairness analysis shows no unacceptable disparities. The human override rate is reasonable, typically in the 5 to 15 percent range. User acceptance testing is positive. Monitoring and alerting are in place. Rollback procedures are documented. Stakeholders have signed off. And there is a plan for monitoring and updating after launch.
The criteria for continuing to iterate are their mirror image. Success criteria are not yet met but look achievable with more work. Fairness concerns exist and are addressable. User feedback says the system is not solving the problem. The technical architecture needs to change before you can proceed. Or you are learning something critical that changes the design. None of these is a failure; each is a reason the next sprint is worth more than a launch.
The ship decision is not a single moment either. You might ship an MVP to two pilot offices, iterate for two sprints on what the pilot tells you, and only then decide whether to expand or redesign. Treating the launch as one irreversible door is what pushes teams to argue about readiness instead of measuring it.
The Artifact: An AI Iteration Charter With Kill-Switch Gates
Marcus's central deliverable was a one-page charter that any executive, auditor, or successor could read. It named the unknown the pilot existed to resolve, the sprint cadence, the pre-registered success thresholds, and the explicit gates where the project would stop if it failed to clear a bar. Pre-registering the thresholds mattered enormously: writing down "success means X" before you see the results prevents the all-too-human habit of redefining success after the fact to declare victory.
| Charter element | For the DMV document-verification pilot |
|---|---|
| Core unknown | Can our document data support a forgery detector accurate and fair enough to reduce clerk workload without harming applicants? |
| Sprint cadence | Two-week sprints; demo and measurement at each sprint close; biweekly stakeholder visibility. |
| Pre-registered success thresholds | At least 85 percent agreement with expert clerk judgment on a held-out set; false-positive rate (genuine document flagged as fake) at or below 5 percent; no error-rate gap greater than 4 percentage points across document types or demographic groups. |
| Human oversight | Every flag reviewed by a clerk with authority to override; overrides logged; no autonomous rejection of any document. |
| Gate 1 (proceed to field pilot) | Baseline shows usable signal; M-24-10 impact assessment signed; ATO granted. If not met, stop or rescope. |
| Gate 2 (proceed to production) | Pilot meets all three pre-registered thresholds over the 90-day window. If false-positive rate exceeds 5 percent or any fairness gap exceeds 4 points, do not promote; iterate or retire. |
| Kill-switch | If at any sprint the false-positive rate on genuine documents exceeds 10 percent, pilot pauses immediately pending review. Authority to pull the switch: project lead, no committee required. |
| Drift plan | Quarterly re-measurement post-deployment; retrain trigger if accuracy falls below 85 percent. |
That charter did two jobs at once. It let Marcus's team move fast and learn, and it gave the director, the privacy office, and a future auditor a single artifact showing that speed never came at the cost of a documented stopping point. When the legislature's oversight staff later asked how an agile AI project could be accountable, Marcus handed them the charter. The gates were the answer.
Anti-Patterns to Avoid
- Waterfall in agile clothing. The team claims to be agile but operates waterfall: six weeks of planning before any code, nothing deployed until it is "done", and one big-bang delivery at the end wearing sprint vocabulary. Ship something to real users or real data by week four, and make shipping part of every sprint cycle rather than a phase at the end.
- Iteration without measurement. The team iterates constantly but never measures whether anything improved. Six months in there have been many changes and no evidence that accuracy or fairness moved. Define the metrics upfront, measure them every sprint, and let the measurements set iteration priorities.
- MVP as an excuse to ship poor quality. "It is only an MVP, so we do not need monitoring, fairness analysis or documentation." Minimum viable means viable: it works, it has a fairness analysis, it has monitoring. It does not mean half-baked, and in a government context the shortcuts most often skipped are the ones that protect the public.
- Ignoring pilot feedback. The team deploys to a pilot group, collects feedback, and discards it because iteration was already planned. Complaints get dismissed with "they will adjust once we train them." If multiple users report a problem, the problem is real; if users hate something, change it before broader rollout.
- Premature optimisation. Six sprints spent moving accuracy from 85 percent to 87 percent when the real constraint is that nobody is using the system. Measure what matters. If the problem is adoption, improve usability, not the model.
- Selling a trajectory as a schedule. Two good sprints produce a straight-line projection and the projection becomes a commitment in someone's briefing. Model gains typically decelerate, so a projection is an estimate to be updated at the next measurement, not a date to be defended.
- Treating a sandbox result as an authorization. Iterating quickly on de-identified data is exactly right, and it says nothing about whether the system may touch live resident data. Sandbox learning and the authorization to operate are separate things, and only one of them is a decision the project team can make.
Practice Prompts
- Take a government AI project you know. Design six two-week sprints that move it from problem statement to deployed MVP. For each sprint, write the goal, the deliverable, what gets measured, and what decision the sprint informs.
- For a real or hypothetical government AI system, define the MVP. What is the simplest version that tests the core assumptions? What would success look like for the MVP specifically? What would you measure, and what result would prompt you to pivot rather than persist?
- Design the measurement dashboard you would show stakeholders each sprint. Which accuracy, fairness, business and engagement metrics would appear, and how would you present them so people stay informed and confident without being misled?
- Write the expectation-setting conversation you would have with your own executive sponsor, in their language, before sprint 1 begins. Then write the version you would give at sprint 1 review if the result came in far below target.
- Draft a one-page iteration charter for a project you are close to: core unknown, cadence, pre-registered thresholds, gates, kill-switch, and who holds the authority to pull it.
Reflection
- Think of a project where the initial planning turned out to be wrong. Why did the plan fail? What did you learn only after work started, and how would an iterative approach have changed the outcome?
- In your agency, which governance step is treated as paperwork to backfill rather than as a gate? What would change if it were written into the sprint calendar as a stop-or-go decision?
- If your current project failed to clear its own success criteria, who would have the authority to stop it, and would that person find out in time?
- Which of your constraints, procurement, authorization or change control, actually limits your iteration speed, and which one has simply never been tested?
Glossary
- Sprint. A time-boxed period, typically one to two weeks, during which a team completes a defined set of work and produces a measurable increment of the product.
- MVP (minimum viable product). The simplest version of a system that solves the core business problem and tests fundamental assumptions, providing value to users while maximising learning per unit of effort.
- Iterative development. An approach where the system is built incrementally in multiple cycles, each cycle incorporating learning from the previous ones and producing working software.
- Pilot deployment. A limited deployment to a small group of users or offices to test the system in real conditions and gather feedback before broader rollout.
- Sprint metrics. Measurable indicators, covering accuracy, fairness, business impact and user adoption, tracked at the end of each sprint to assess progress and inform the next sprint's plan.
- Waterfall. A sequential development model in which requirements, design, build, test and deployment each finish before the next begins.
- Pre-registered threshold. A success criterion written down before results are seen, so that success cannot be redefined after the fact.
Related Lessons
- How AI Projects Differ from Traditional IT sets out the underlying differences this lesson turns into a working method.
- Requirements Gathering for AI covers the part iteration does not replace: getting clarity on the problem and the success metrics before sprint 1.
- Working with AI Vendors and Contractors deals with the procurement constraint Marcus worked around, and how to scope a contract that permits iteration.
- Moving from Pilot to Production picks up where the field pilot ends and the promotion decision begins.
- Testing and Validating AI Systems is the discipline behind the end-of-sprint measurement that makes iteration meaningful.
- Systematic AI Output Validation is the next step for teams who now need to prove that an improvement is genuinely an improvement.
- Communicating AI Projects to Leadership extends the expectation-setting conversation into the executive briefing.
Closing
The shift from waterfall to agile thinking is one of the most important mindset changes for government AI teams. Waterfall assumes planning and prediction are possible; agile assumes they are not and structures the work accordingly. In practice that means being comfortable shipping something imperfect because you are learning, measuring everything because metrics drive decisions, staying flexible about scope and timeline because you are discovering what is achievable, involving users and stakeholders continuously because their feedback is data, and treating failure as information rather than catastrophe. It is harder than waterfall in one respect, because you cannot hide behind a plan, and easier in another, because you are not betting the farm on one design. Marcus's charter is what makes both true at once: it lets the team move, and it names in advance the moment they stop.
Key Takeaways
- AI is probabilistic and data-dependent, so waterfall hides the most important fact until the most expensive moment. You cannot promise accuracy in a requirements document; you discover it by building and measuring.
- Treat learning as the deliverable. Answer your biggest unknown in week three with a throwaway MVP, not in month eighteen with the whole budget.
- Keep MVP, pilot, and production distinct. An MVP tests an assumption, a pilot is a bounded real deployment with an evaluation window, and production only follows evidence against pre-registered thresholds.
- Measure across four dimensions, not one. Accuracy, fairness, business impact and engagement each catch a failure the others miss, and a team that tracks only accuracy will optimise the wrong thing for months.
- Pre-register success thresholds before you see results. Writing down what success means in advance stops anyone from redefining it after the fact to claim a win.
- Governance gates are kill-switches, not the opposite of agile. Map NIST AI RMF MEASURE onto your sprint cadence, and treat the M-24-10 impact assessment as a hard gate before any rights-affecting deployment.
- Put the human in the loop during iteration, not just at the end. Logged overrides are both a safeguard and the richest signal you have about where the model is wrong.
- Reset stakeholder expectations explicitly and demo on a fixed rhythm. Say at the start that success means meeting the agreed criteria rather than following the original plan, then show goal, result, metric, learning and next step every two weeks.
- Run agile inside government constraints by separating environments and using existing contract vehicles. Iterate on de-identified data in a sandbox while ATO and change control run in parallel; promote only vetted releases.
- The one-page charter is the artifact that proves accountability. Core unknown, cadence, thresholds, gates, and a named kill-switch authority turn a fast-moving pilot into something an auditor can trust.
Frequently Asked Questions
Does agile mean we skip requirements? No. Iteration replaces the assumption that requirements can be fully specified in advance, not the requirements themselves. You still start with clarity on the business problem and the success metrics. What changes is that performance targets are treated as hypotheses to be tested against your data rather than as promises to be written into a contract.
Our oversight bodies require documented plans. Can we still iterate? Yes, and the charter is how. Documenting the cadence, the pre-registered thresholds, the gates and the kill-switch authority gives reviewers something more auditable than a schedule, because it states in advance what would cause the project to stop. Marcus's experience was that oversight staff found the gates more reassuring than a Gantt chart, not less.
What if the MVP result is terrible? That is a normal sprint 1 outcome and it is information rather than failure. A weak first number tells you whether there is any signal in the data to build on. Say what you learned, say what changes next sprint, and show the path you believe leads to the target. What you should not do is announce the sprint in which you will hit the target, because model gains typically decelerate and a straight-line projection will overstate.
How do we iterate when we cannot touch real data until the ATO clears? Split the environments. Early sprints can run on a small set of de-identified historical records so the team learns while authorization proceeds in parallel, and live data only enters once authorization is granted. The sandbox result is learning, not permission; the authorization decision belongs to the people who own it.
When do we stop iterating and ship? When the success criteria are met or an agreed tradeoff is documented, fairness analysis shows no unacceptable disparities, the override rate is reasonable, user acceptance testing is positive, monitoring and rollback are in place, stakeholders have signed off, and there is a post-launch monitoring plan. If any of those is missing, the next sprint is worth more than the launch.
Is an override rate a good thing or a bad thing? Both, read carefully. Overrides are your richest signal about where the model is wrong, so a rate of zero usually means reviewers have stopped reading rather than that the model is perfect. A rate in the typical 5 to 15 percent range suggests reviewers are engaged and mostly agreeing. What matters more than the number is that every override is logged and fed back into the next iteration.
Skill.re