←
AI for Government
Capable · M21 · lesson 21 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
How AI Projects Differ from Traditional IT
📖
now learning

How AI Projects Differ from Traditional IT

15 min

Priya Raman had shipped two dozen IT systems for a county's health and human services department. New case-management portal? She knew the drill: gather requirements, sign off on a spec, build to the spec, test against the spec, deploy, done. So when the department asked her to lead a project using AI to flag benefit applications that needed extra review, she ran the same playbook. Eight weeks in, the vendor demoed a model that flagged applications. It worked. Then it stopped working, not because the code broke, but because a new state form changed the data slightly and the model's accuracy quietly fell from 91 percent to 68 percent. Nothing had "failed." No error appeared in any log. Priya stared at the dashboard and realised her trusted playbook had a hole in it the size of a freeway.

This lesson is for project leads, analysts and supervisors who manage technology work in government. You already know how to run a traditional IT project. The goal here is narrow and practical: to show you exactly where AI projects break the rules you rely on, so that you can adjust your plan, your contract and your expectations before the hole swallows your timeline. It is also, increasingly, a compliance question rather than a philosophical one, and we will get to why that matters for the way you write your next statement of work.

Software follows rules, AI learns patterns

Traditional software is a set of rules a programmer wrote. If income is below a threshold, mark eligible. The behaviour is deterministic: the same input produces the same output today, tomorrow and next year. A correctly built payroll system pays the same net amount for the same inputs every pay period. A query against a stable table returns identical rows. Behaviour like this can be unit-tested with fixed assertions, regression-tested against a golden dataset, and accepted against a binary pass or fail criterion. That is the world every federal IT acceptance process was designed for.

An AI model is different in kind. Instead of following written rules it learns patterns from examples, thousands of past applications and their outcomes, and then makes a prediction about new cases. That prediction is probabilistic: a best guess with a confidence level, not a guaranteed answer, and its correctness is defined in aggregate rather than case by case. A benefits-eligibility classifier operating at 92 percent accuracy will be wrong eight times in a hundred, and which eight is not predictable in advance. Priya's model did not break. The world it learned from shifted, and the model silently fell behind.

Generative systems add a second source of variation. A large language model summarising Freedom of Information Act requests may vary its wording, its emphasis or its factual grounding between calls even with an identical prompt, because of sampling temperature, context window reshuffling, or the provider's own testing of model variants. The practical consequence is that "we ran it once and it was right" is not a test result. Traditional software does what you told it. AI does what the data taught it, which means that when the data changes the system changes even though nobody touched the code.

Four things that follow for acceptance and operations

First, acceptance criteria must be stated as distributions rather than points. Not "the model shall classify correctly" but something in the shape of the source's worked example: the model shall maintain at least 90 percent precision and 85 percent recall on the validation set, with no demographic subgroup dropping more than five percentage points below the aggregate. Second, monitoring must continue after go-live, because performance can degrade while the code sits untouched. The MEASURE function of the NIST AI Risk Management Framework addresses this explicitly, and the framework is voluntary guidance rather than binding law, which makes citing its specific subcategories in your contract the thing that gives it teeth.

Third, the Authority to Operate has to be extended with conditions that trigger re-evaluation, rather than resting on a fixed reauthorisation clock. OMB Memorandum M-24-10 requires this for rights-impacting AI. Fourth, incident response must be designed for statistical failures and not just outages. The traditional operations question is "is it up?" The AI operations question is "is the error distribution still within tolerance?" A system that is up, responsive and confidently wrong will pass every availability check you own.

Four rules of traditional IT that AI breaks

1. Lock the requirements, then build

In traditional IT you can write a precise specification up front. With AI you often cannot, because nobody yet knows whether the model can reach the accuracy you need on your actual data. Priya could not promise "flags 95 percent of risky applications" before testing, any more than a doctor can promise a diagnosis before the examination. AI projects start with a hypothesis and a question, not a fixed spec. Plan for a proof-of-concept phase whose job is to discover what is achievable, before you commit scope and budget to a number you cannot yet defend.

2. Test once, and if it passes it works

Traditional software that passes its tests keeps working until someone edits it. AI degrades on its own. As real-world data drifts away from training data, accuracy falls, and there is no error message. The system keeps producing confident answers that are increasingly wrong. AI projects therefore need ongoing monitoring rather than one-time acceptance testing. Be precise about what that buys you: monitoring catches degradation in the measures you instrumented, on the population you sampled, where you have ground truth to compare against. It does not catch a failure mode nobody thought to watch.

3. The code is the product

In AI, the data is the product. A mediocre model on well-governed data will usually outperform a sophisticated model on poor data, and Priya's eight weeks of vendor work were undone by a data problem, a changed form, not a code problem. This reorders the project. Securing, cleaning, labeling and governing data becomes the main event rather than a setup step you rush through to reach the "real" work. The source's planning figure is to budget 40 to 60 percent of total effort for data work, and that is a floor to argue up from, not a ceiling.

4. Right or wrong is binary

Traditional software is correct or it has a bug. AI is rarely simply right or wrong; it is right most of the time, with a known error rate. The hard question is not "is it perfect?" but "what error rate is acceptable, and what happens to a person when it is wrong?" An AI that wrongly flags an eligible family for extra review imposes a real delay on a household that needed the money this month. Define, in advance and in writing, the acceptable error rate and the human safety net for the errors that will certainly happen.

Data dependency and data debt

Traditional IT depends on code; AI depends on data. That sounds glib until you consider what it does to debugging. A code bug is usually reproducible, localisable and fixable by a developer with repository access. A data bug shows up as degraded output for a subset of cases, and tracing it back to a mislabeled training example, a shift in upstream collection, or a population that was never represented at all requires tooling most agencies do not have: data lineage, data quality monitoring, drift detection and model explainability. Priya had none of the four when her accuracy fell.

The consequences are well documented outside government too. The Epic Sepsis Model, deployed to hundreds of hospitals, was found in 2021 by University of Michigan researchers, published by Wong and colleagues in JAMA Internal Medicine, to perform far below vendor claims, with a measured sensitivity of only 33 percent against a vendor-reported 76 percent; treat that pairing carefully, because a vendor headline figure and a measured sensitivity are not always the same statistic. IBM Watson for Oncology, deployed at MD Anderson Cancer Center from 2013, was shelved in 2017 after training on hypothetical rather than real patient cases produced recommendations described internally as unsafe and incorrect. Optum's chronic care algorithm, analysed by Obermeyer and colleagues in Science in 2019, encoded racial bias by using healthcare spending as a proxy for medical need, systematically understating the need of Black patients who had historically received less spending.

In each of those cases the failure was in the data, not the code. Government agencies carry the same exposure with an extra complication: their data was typically collected for a different purpose, is governed by Privacy Act System of Records Notices, and falls under the E-Government Act Privacy Impact Assessment regime. Using pre-existing administrative records to train a model often requires a fresh privacy impact assessment, updated notices, and sometimes new Paperwork Reduction Act clearance. Data debt, the accumulated cost of cleaning, documenting and governing data that was collected with no AI use in mind, commonly dwarfs the cost of model development, and traditional IT project budgets almost never carry a line for it.

Iterative, hypothesis-driven development

Traditional IT projects follow a system development lifecycle with reasonably predictable phases: initiation, requirements, design, construction, integration, operations, disposition. Even the agile variants now dominant in federal digital work still assume each sprint produces demonstrable, accepted functionality. AI projects instead run a research loop closer to applied science: hypothesis, dataset, baseline, experiment, evaluate, iterate. The source's estimate is that ninety percent of experiments fail or underperform the baseline, and it offers no citation for that figure, so treat it as a practitioner rule of thumb rather than a measured rate. A team may spend six weeks demonstrating that an approach does not work, which is a valuable negative result in science and very hard to justify under classical earned-value management.

Jennifer Pahlka, former U.S. Deputy Chief Technology Officer and author of Recoding America, argues that the mismatch between probabilistic discovery and waterfall governance is the single largest obstacle to useful AI in the federal government. Three concrete adjustments follow. Contracting vehicles must allow exploratory phases where the deliverable is validated learning rather than built functionality. Program reviews must be structured around go and no-go gates at experiment boundaries rather than Gantt-chart milestones. And budget rules must tolerate an unknown number of experiments before production investment is committed.

This is not theoretical. The Defense Innovation Unit and, historically, the General Services Administration's digital service teams piloted modular contracting and agile blanket purchase agreement vehicles built for exactly this pattern. Confirm which of those vehicles and support organisations currently exist and are available to you before you plan a dependency on one. Most civilian agency contracting officers still default to fixed-price, fixed-scope structures, which are hostile to discovery work, and the default will not change because you wish it would.

Ambiguous requirements and measurable outcomes

Traditional IT requirements are written in terms of features and data flows: the system shall allow the user to upload a PDF; the system shall process 10,000 transactions per hour. Both are verifiable by direct observation. AI requirements are frequently aspirational: the system shall help claims adjusters identify fraud. What does help mean? At what precision and recall tradeoff? Compared with which baseline? Translating aspiration into a measurable, contestable specification requires a discovery phase that traditional IT methodologies assume is already complete before contract award.

OMB M-24-10 implicitly acknowledges the problem by requiring agencies to document the intended use, expected benefits, risks and mitigations for each AI use case before deployment, which forces the specification work to the front. It does not remove the need for that specification to evolve as experiments reveal what is actually possible. Good program managers therefore plan two phases explicitly: a discovery or pilot phase whose acceptance criterion is a well-specified production requirement, and a production phase whose acceptance criterion is the operationalisation of that requirement. Collapsing them into one contract produces either an unsuccessful product or a painful modification, and usually both.

Validation, testing and red teaming

Traditional IT validation rests on unit tests, integration tests, system tests, user acceptance tests and operational readiness reviews, all well understood through the Capability Maturity Model Integration, the Systems Engineering Body of Knowledge and, for defence software acquisition, DoD Instruction 5000.87. AI validation demands techniques that are foreign to most government quality assurance teams: statistical performance evaluation across stratified subgroups; bias and fairness testing against protected classes under authorities such as the Civil Rights Act and the Equal Credit Opportunity Act; robustness testing against adversarial inputs; red-teaming for prompt injection and jailbreaks; data-leakage testing for membership inference; and continuous monitoring for concept drift.

Executive Order 14110, signed in 2023, directed NIST at Section 4.1 to publish guidelines for AI red-teaming, and that direction produced the NIST AI 600-1 Generative AI Profile in 2024. The order set the broader federal direction of its period and should be read historically rather than as a live instruction; confirm the current status of any executive order before you cite it in a procurement document. The Department of Defense maintains its own Responsible AI Strategy and Implementation Pathway. The procurement consequence is blunt: your request for proposals must specify which validation techniques are required, how the evidence will be delivered, and who bears the cost of red-teaming. Historically this has been a recurring contract dispute, with vendors arguing that extensive red-teaming was out of scope and agencies arguing that it is implicit in a duty of safe operation.

Governance, workforce and operating model

Traditional IT governance is organised around chief information officers, information system owners, system security officers and privacy officers, with responsibilities codified in FISMA and the Privacy Act. AI adds roles that did not previously exist: the Chief AI Officer, which M-24-10 requires at each CFO Act agency; an AI governance board; an AI model owner; a data steward; an algorithmic accountability officer; and a human oversight officer for high-risk systems. Agencies including GSA, DHS, VA, HHS, CMS, FDA, NSA, DOD, State and the IRS have stood up governance bodies of varying maturity, and CISA has published AI security guidance drawing on the MITRE ATLAS adversarial machine learning taxonomy.

The workforce implications are just as large. Federal IT workforces are structured around the 2210 information technology specialist series, while AI work needs competencies in statistics, machine learning, data engineering, MLOps, behavioural science and human-computer interaction, which sit across series such as 1515 operations research, 1530 statistics and 0110 economist. OPM has begun updating qualification standards and the pace lags demand. Culturally, AI teams work better with researcher-like autonomy and tolerance for failed experiments, while traditional IT shops reward predictability; forcing either model onto the other is a common cause of attrition among scarce specialists. One further operational difference is easy to miss: AI systems need live feedback loops from end users to stay aligned with reality, which creates a dependency on product management skills that are weakly represented in many agencies.

Procurement and contracting implications

Under FAR Part 15 negotiated procurements and Part 16 contract types, the default for defined-scope work is fixed-price, which minimises agency risk but presumes the scope is knowable. AI discovery work rarely is. Time-and-materials or cost-reimbursement arrangements fit the exploratory phase better, transitioning to firm-fixed-price for production operations once requirements have stabilised. Other Transaction Authority, cited by the source as 10 U.S.C. 4022 for the Department of Defense and 42 U.S.C. 7256b for the Department of Energy, allows more flexibility and has been used extensively by the Defense Innovation Unit, DARPA and the Advanced Research Projects Agency for Health. Verify the current citation with your contracting officer rather than quoting it from a slide.

Contracting officers and their technical representatives have to understand AI-specific acceptance criteria or they risk two symmetric failures: accepting a system that does not work, or rejecting a system that works well but was measured against the wrong benchmark. Intellectual property clauses need particular attention. Who owns fine-tuned model weights trained on government data? The rights-in-technical-data frameworks at DFARS 252.227 and FAR 52.227 predate modern AI and are being reinterpreted case by case. Data use and reuse clauses must state whether vendor models may be trained on agency data, whether agency data may leave the authorisation boundary, and whether the vendor's base model was itself trained on data your agency would consider sensitive.

What goes wrong when the IT playbook is applied anyway

The failure mode repeats across sectors and decades, and it is always the same: traditional acceptance thinking applied to a probabilistic system, with the fairness, drift and stratified-performance analyses skipped because the old playbook had no place for them. The cases below are the ones the source treats as canonical. Read them as a checklist of the analyses that were not run rather than as a gallery of other people's mistakes.

CaseWhat happenedThe assumption that caused it
Michigan MIDAS, 2013 to 2015More than 40,000 unemployment claimants flagged as fraudulent, with a false positive rate later measured at roughly 93 percent; more than $20 million paid in settlements, with litigation continuing as of 2023Acceptance tests treated automated determinations as deterministic, and no human-in-the-loop review caught the statistical error pattern for about two years
IRS ID.me, 2021 to 2022Facial-recognition identity proofing procured on a compressed timeline; millions of taxpayers lost account access and Treasury reversed course toward Login.gov after civil society objections and pressure including from Senator Ron WydenTreated as a classic vendor integration, without a privacy impact assessment update adequate to the scale of deployment
Dutch toeslagenaffaire, 2005 to 2021An opaque risk-scoring algorithm flagged immigrant families for fraud investigation at disproportionate rates; tens of thousands had childcare benefits clawed back unjustly, and the scandal brought down the Rutte III cabinet in January 2021No contestability, no disparate-impact analysis, and an algorithm treated as an administrative tool rather than a decision affecting rights
COMPAS, documented by ProPublica in 2016A recidivism scoring tool used across state court systems shown to have disparate error rates by race, with Black defendants more likely to be falsely labeled high risk; the ensuing academic debate clarified how fairness metrics trade offAggregate accuracy accepted as sufficient evidence, with no stratified error analysis by group
Epic Sepsis Model, 2021Widely deployed across hospitals and found by Wong and colleagues in JAMA Internal Medicine to perform far below vendor-reported accuracyVendor marketing claims accepted without independent validation on local data
IBM Watson for Oncology, 2013 to 2017Trained on hypothetical rather than real cases at MD Anderson; recommendations described internally as unsafe and incorrect; project terminatedTraining data provenance never validated against the population the system would serve
Optum chronic-care algorithm, Obermeyer and colleagues, 2019Encoded racial bias by using cost as a proxy for medical need, in an algorithm class affecting more than 200 million AmericansA convenient proxy variable accepted as equivalent to the thing it stood in for

Applying the NIST AI Risk Management Framework

The NIST AI Risk Management Framework 1.0, published in 2023, offers four functions that replace the traditional IT risk register as the organising structure. GOVERN establishes policies, accountability and culture, and is where Chief AI Officer authorities and governance boards live. MAP contextualises risk to the specific use case, stakeholder community and lifecycle stage, and is the right place to introduce historical failures such as MIDAS and Optum as analogue cases rather than as war stories. MEASURE quantifies the trustworthiness characteristics: validity and reliability, safety, security and resilience, accountability and transparency, explainability and interpretability, privacy, and managed bias. MANAGE prioritises responses, allocates resources and makes the go or no-go calls.

Crosswalks between the framework and the NIST Cybersecurity Framework, the Secure Software Development Framework at SP 800-218, and ISO/IEC 42001:2023 already exist and should be used to avoid duplicating compliance work you have done once. The Generative AI Profile, NIST AI 600-1, adds a set of generative-specific risks that the source numbers at twelve, including chemical, biological, radiological and nuclear information hazards, confabulation, dangerous recommendations and information integrity. Applying the framework is more than a paper exercise: it changes which documents get produced, which committees review them, and which decisions are delegated rather than retained at Chief AI Officer or deputy secretary level.

The planning comparison to run with your vendor

Before kicking off any AI project, walk your team and your vendor through the comparison below. For each row the question is the same: have we planned for the AI column rather than the traditional column? Disagreement in a row is more useful than agreement, because it locates the assumption that would otherwise surface at acceptance.

DimensionTraditional ITAI project
RequirementsFixed spec up frontHypothesis first, then a proof of concept that discovers what accuracy is achievable on your data, then scope
Success criteriaPasses or fails defined testsA target accuracy stated as a distribution and a defined acceptable error rate, agreed by program owners and not only by IT
Data planData is an inputData quality, labeling, rights and governance are the core deliverable; budget 40 to 60 percent of effort here
TestingOne-time acceptance testContinuous monitoring for drift on instrumented measures, with alerts when a measure crosses the agreed floor
Failure handlingLog an error, fix the bugDefine what happens to a person when the model is wrong, and guarantee a human can review and override
Contract termsBuild-to-spec acceptanceAccuracy thresholds, a duty to retrain on drift, data ownership, red-teaming responsibility and audit rights
Maintenance modelPatch as neededScheduled retraining and revalidation as a recurring, funded activity
AccountabilityIT owns the systemA named program official owns the decision the model informs, because the model affects real people

The program manager's checklist

The source sets out ten steps, counted from its own numbered list, as a floor rather than a substitute for agency policy. Before the contract is signed, articulate what will be known at the end of discovery that is not knowable today, and structure the contract so that discovery produces a validated production requirement. Draft acceptance criteria as statistical distributions rather than points. Budget 40 to 60 percent of effort for data collection, cleaning, labeling, documentation, lineage and ongoing governance. Register the use case in the agency AI inventory as M-24-10 requires. Classify it as rights-impacting, safety-impacting or neither, and apply the associated minimum practices.

Then complete or update the privacy impact assessment and the relevant System of Records Notices before production. Specify red-teaming and adversarial testing responsibilities in the statement of work. Plan for Authority to Operate conditions that trigger re-authorisation on drift, not only on a fixed clock. Establish a human-in-the-loop or human-on-the-loop design for every rights-impacting and safety-impacting system. Finally, define sunset criteria: the conditions under which the AI will be turned off rather than upgraded. That last one is the step most often missing, and it is the one that makes every other step enforceable.

Why this matters for government

Government agencies have spent four decades refining frameworks for traditional IT: the Clinger-Cohen Act, FITARA, TechFAR, FISMA, OMB Circular A-130 and FedRAMP. All of them assume that a system, once defined and authorised to operate, behaves predictably: patch it, monitor it, reauthorise on a three-year cycle, done. AI breaks most of those assumptions. A large language model that passes an Authority to Operate on Monday may produce different outputs on Tuesday because of prompt injection, context drift or a silent vendor update, with no change to anything the agency controls.

The compliance regime has begun to catch up, and it is concrete rather than aspirational. OMB Memorandum M-24-10, issued in 2024, is the first federal instrument that treats AI as categorically different, requiring each CFO Act agency to designate a Chief AI Officer, publish an annual use case inventory, and implement minimum practices for rights-impacting and safety-impacting AI before December 1 of each year. The EU AI Act, which entered into force in 2024, creates parallel obligations for high-risk systems and extraterritorial exposure for United States vendors selling into Europe. ISO/IEC 42001:2023 provides a certifiable AI management system standard. GAO published its AI accountability framework in 2021, and inquiries arrive under it.

Priya did not need to become a data scientist. She needed to retire one assumption, that software once built and tested stays built and stays correct, and replace it with a plan treating data quality and continuous monitoring as the work itself. The county's second attempt scoped a proof of concept first, budgeted for monitoring, wrote a retraining duty into the contract, and named a program official who owned the decision the model informed. It shipped slower. It did not silently fail. A program manager who cannot articulate these differences cannot negotiate a realistic contract, cannot staff a team with the right mix, cannot brief an inspector general credibly, and cannot answer a GAO inquiry when it arrives.

Anti-Patterns to Avoid

  • Treating monitoring as a safety net rather than a sampling instrument. Continuous monitoring reports on the measures you instrumented, for the population you sampled, where ground truth is available. Priya's dashboard was green throughout the fall from 91 percent to 68 percent, because nothing on it measured accuracy against fresh labels.
  • Accepting a vendor accuracy claim as a test result. The Epic Sepsis case is the canonical warning. Independent validation on your own data, with your own population, is the only claim you can defend, and a headline vendor figure may not even be the same statistic as the one you measure.
  • Buying discovery work on a fixed-price, fixed-scope contract. It presumes the scope is knowable, which is the exact thing discovery exists to determine, and the result is either a padded price or a modification fight.
  • Reporting a single aggregate accuracy number. COMPAS passed on aggregate. The disparity was in the error rates by group, which nobody was required to report until researchers computed them from the outside.
  • Substituting a convenient proxy for the thing you actually care about. Optum used spending as a proxy for medical need, and the proxy carried the historical inequity of who had received spending. Ask of every feature what it is standing in for and who is under-represented in it.
  • Assuming an Authority to Operate remains meaningful for its full term. Model behaviour can change without any change the agency controls, so re-evaluation conditions matter more than the reauthorisation date.
  • Budgeting data work as a setup task. Data debt is the accumulated cost of governing records collected for another purpose entirely, and it commonly exceeds model development cost while appearing nowhere in the plan.
  • Letting a demonstration substitute for a test. Generative output varies between calls with identical prompts, so a successful demo establishes that a good output is possible, not that it is typical.
  • Leaving the accountable official unnamed. If the answer to "who owns this decision" is "IT," then nobody in the program owns the consequence for the person the decision lands on.

Practice Prompts

These are drafting and stress-testing aids for your own project documents. Use them on redacted material only, and treat every output as a draft you are responsible for verifying.

  • "Here is a draft acceptance criterion for an AI system. Rewrite it as a distribution rather than a point, including a precision target, a recall target, and a subgroup floor. Then list what the rewritten criterion still fails to specify."
  • "Review this statement of work. Identify every place it assumes the scope is knowable at award, and mark which of those belong in a discovery phase instead."
  • "Given this system description, list the monitoring signals that would have to be instrumented to detect a silent accuracy decline, and state for each what ground truth it needs and where that ground truth would come from."
  • "Take this set of model features and, for each one, tell me what real-world quantity it is a proxy for and which population is likely under-represented in it."
  • "Draft a set of sunset criteria for this system: the specific, measurable conditions under which it would be switched off rather than retrained."

Reflection Questions

  • Think of the last IT project you ran. Which of the four broken rules would have hurt you most if that project had been an AI project?
  • For a system you are responsible for now, could you detect a fall from 91 percent to 68 percent accuracy? What specifically would tell you, and how long would it take?
  • Who in your organisation is named as the owner of the decision an AI system informs, as distinct from the owner of the system? If nobody is, who should be?
  • What proportion of your last project's effort actually went into data work? How does that compare with the 40 to 60 percent planning figure, and what explains the gap?
  • What would have to be true for your agency to turn off a deployed AI system? Who could make that call, and has anyone written it down?

Glossary

  • Deterministic. Producing the same output for the same input every time, which is the behaviour traditional acceptance testing is built around.
  • Probabilistic. Producing a best guess with an associated confidence, whose correctness is assessed in aggregate rather than case by case.
  • Model drift. The gradual decline in accuracy as live data diverges from training data, with no error and no code change.
  • Data debt. The accumulated cost of cleaning, documenting and governing data that was collected without AI use in mind.
  • Precision and recall. Two complementary measures: precision is the share of flagged cases that were genuinely positive; recall is the share of genuinely positive cases that were flagged.
  • Stratified performance evaluation. Reporting accuracy separately for defined subgroups, so an aggregate figure cannot conceal a disparity.
  • Red teaming. Structured adversarial probing of a system, covering prompt injection, jailbreaks and data-leakage attempts, before an outsider does it for you.
  • Membership inference. An attack that determines whether a specific record was part of a model's training data, which is a privacy failure even when no output is wrong.
  • Authority to Operate. The formal authorisation permitting a federal system to run, traditionally on a fixed reauthorisation cycle and, for rights-impacting AI, requiring conditions that trigger re-evaluation.
  • Other Transaction Authority. A flexible acquisition instrument outside standard FAR contract types, used where conventional structures do not fit exploratory work.

Closing Thoughts

Nothing in this lesson asks you to abandon what you know about running projects. Stakeholder management, schedule discipline, contract hygiene and clear acceptance criteria matter more in AI work, not less, because the technical uncertainty is higher and the accountability lands in the same place it always did. What changes is where you place the uncertainty. In traditional IT the unknown is how long the build takes. In AI the unknown is whether the thing is achievable on your data at all, and you buy that answer deliberately in a discovery phase rather than discovering it during acceptance.

The failures in this lesson span three decades, four countries and both public and private sectors, and they rhyme. Aggregate accuracy accepted without subgroup analysis. Vendor claims accepted without local validation. Deterministic acceptance testing applied to a statistical system. A proxy accepted as the thing itself. None of those is a technical mistake; each is a governance habit carried across from a world where it worked. Retiring the habit is the whole job, and Priya's second project shows it can be done by the same person with the same skills and a different plan.

Key Takeaways

  • AI learns patterns; software follows rules. Outputs are probabilistic best guesses whose correctness is defined in aggregate, and a classifier at 92 percent accuracy is wrong eight times in a hundred without telling you which eight.
  • You cannot fully lock requirements up front. Buy a discovery phase whose deliverable is a validated production requirement, then contract for production separately.
  • Testing is continuous, and it is not a safety net. Monitoring reports on the measures you instrumented where ground truth exists; it will not catch a failure nobody thought to watch.
  • Data is the product, and data debt is real. Budget 40 to 60 percent of effort for data work, and expect privacy impact assessments, System of Records Notices and possibly Paperwork Reduction Act clearance before you may use existing records.
  • Right or wrong becomes what error rate is acceptable. Define the tolerable rate, the subgroup floor, and the human safety net for the errors that will happen.
  • The failure cases share one cause. MIDAS, ID.me, the toeslagenaffaire, COMPAS, Epic Sepsis, Watson Oncology and Optum all applied deterministic acceptance thinking to a statistical system.
  • Contracts and authorisations must reflect AI's nature. Accuracy thresholds, a duty to retrain on drift, red-teaming responsibility, data ownership, audit rights, and Authority to Operate conditions that trigger on drift rather than only on a clock.
  • Governance and workforce change too. Chief AI Officer, governance board, model owner, data steward and human oversight roles sit alongside the 2210 series that federal IT was built on.
  • A named program official owns the outcome. Accountability rests with whoever owns the decision the model informs, because that is where the harm lands.

Frequently Asked Questions

Does this mean we should never use fixed-price contracts for AI?

No. The distinction is between phases. Discovery, where the achievable accuracy on your data is genuinely unknown, fits time-and-materials or cost-reimbursement structures. Production operations, once requirements have stabilised and you know what the system does, fit firm-fixed-price well. The mistake is applying one structure across both phases, which either prices the unknown into the fixed price or forces a modification when the unknown resolves differently than assumed.

How would we have caught Priya's accuracy fall earlier?

By holding back a stream of freshly labeled cases and scoring the model against them on a schedule, and by monitoring input distributions for change independently of output accuracy. The new state form altered the data before it altered the outcome, so a data-drift signal would have fired before an accuracy signal. Neither is automatic. Both have to be instrumented and funded, and somebody has to be assigned to read the result.

Is the NIST AI Risk Management Framework mandatory?

The framework itself is voluntary guidance rather than binding law. That is exactly why it is useful in contracts: naming specific functions and subcategories converts a general expectation into a checkable obligation between you and your vendor. Separately, OMB memoranda impose requirements on federal agencies through the executive branch policy chain, and those are not voluntary for the agencies they cover. Confirm which instruments currently apply to your agency with your general counsel rather than assuming.

Our AI project is small and internal. Do all these steps apply?

Scale the artifacts, not the questions. A small internal tool still needs a stated acceptable error rate, a named owner of the decision it informs, a way to notice degradation, and sunset criteria. What it may not need is a full quasi-experimental evaluation or a governance board. The trigger for the heavier apparatus is not project size but whether the output materially affects a person's rights, benefits or safety.

Why do so many of the failure cases come from outside government?

Because the private cases were studied and published, and because the mechanism is identical. Epic Sepsis, Watson Oncology and the Optum algorithm all failed on data provenance and validation, exactly as MIDAS and the toeslagenaffaire did. The difference in government is that the affected person usually cannot take their business elsewhere, which is why contestability and appeal paths carry more weight in public-sector design than in commercial deployment.

What is the single highest-value change to make first?

Split discovery from production in the contract. Almost every other problem in this lesson gets easier once you have bought the right to find out what is achievable before committing to what will be delivered. It also creates the natural moment to write the acceptance criteria as distributions, register the use case, complete the privacy work, and decide who owns the decision, because all of those become inputs to the production requirement rather than afterthoughts.