←
AI for Government
Capable · M38 · lesson 38 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
The AI System Lifecycle
📖
now learning

The AI System Lifecycle

15 min

Priya Nair, a program analyst at a state unemployment insurance agency, inherited an AI tool that nobody could explain. It scored new claims for fraud risk. It had been "in production" for fourteen months. When an auditor asked who tested it, who approved it, and how anyone would know if it started failing, the answers were a shrug, a missing email, and a dashboard that had not loaded since the contractor left. The model was quietly rejecting eligible claimants at a higher rate than a year earlier, and no one had noticed because no one owned the question. Priya's job became untangling a system that had been launched but never actually managed.

Most government AI trouble does not come from bad algorithms. It comes from treating AI as a project that ends at launch, instead of a system that lives, drifts, and must eventually be retired. The fix is to think in terms of a lifecycle, a set of stages from first idea to final shutdown, each with its own owner, its own checks, its own evidence, and its own government-specific traps. This lesson walks the whole arc using Priya's fraud-scoring tool as the spine, then maps that arc onto the artifacts an oversight body will actually ask for.

Why a Lifecycle and Not a Project Plan

An AI system in government is not simply software. It is a decision-making or decision-supporting apparatus that operates inside three overlapping structures at once. There is a statutory structure: the authorizing act, the Administrative Procedure Act, the civil rights statutes, the Privacy Act. There is an executive structure: OMB Memorandum M-24-10, agency directives, and the policy direction set by Executive Order 14110 when it was in force. And there is an oversight structure: the Government Accountability Office, Offices of Inspector General, congressional committees, and the OMB Office of Information and Regulatory Affairs.

Each phase of the AI lifecycle is the point at which one or more of those authorities attaches. Missing a phase means missing a duty, and the duty does not disappear because nobody performed it. That is the difference between a lifecycle and a project plan. A project plan ends when the deliverable ships. A lifecycle has no natural end point, only a deliberate one, and every stage in between produces evidence that someone outside your agency is entitled to see. Priya's tool had a project plan. It had never had a lifecycle.

What a Traditional Authorization Does Not Cover

Contrast this with traditional federal software. A conventional system is authorized under the NIST SP 800-53 control catalog and receives a FISMA Authority to Operate. That authorization focuses on confidentiality, integrity, and availability. For a payroll system or a case management system that does not make or influence consequential determinations, that focus is sufficient. Agencies have decades of practice running that process, and the temptation is to assume that an AI system that clears the same gate is therefore cleared.

An AI system adds failure modes the traditional authorization does not fully address, and each one attaches to a different legal duty. A model can be accurate on average and systematically wrong for a subpopulation, which is a Title VI problem. A model can drift silently as the world changes while its inputs still look in-distribution, which strains the relevance-and-accuracy standard implicit in the Privacy Act. A model can be manipulated by adversarial inputs the SP 800-53 controls never contemplated, which is a FISMA problem. A model can be explainable for one population and opaque for another, which strains the Administrative Procedure Act's reasoned-explanation requirement for agency action.

None of these are discovered by the traditional authorization process alone. They require AI-specific phase gates, which is precisely the context OMB Memorandum M-24-10 was drafted against. That memorandum, issued in 2024, imposes minimum practices on rights-impacting and safety-impacting AI because the existing lifecycle under FISMA did not reach these failure modes. Read carefully, its minimum practices are not a separate compliance regime bolted on top. They are lifecycle obligations, each demonstrated at a particular stage.

The Six Stages of an AI System's Life

A useful way to picture any AI system is as six stages, each flowing into the next, with monitoring looping back continuously: development, training, testing, deployment, monitoring, and retirement. Skipping or rushing a stage is what created Priya's mess. Walk each one against a system you actually run, not a hypothetical.

1. Development: deciding what the system is for

This is where you define the problem before touching any data. For the fraud tool, the real question was never "can we build a fraud scorer?" It was "what decision will this score drive, and who is accountable for that decision?" In government, the first question is not "what dataset?" but "under what authority?" Does the agency have statutory authority to use AI for this purpose? Does the use trigger a System of Records Notice under the Privacy Act? Is it in scope for the current AI use case inventory, and is it rights-impacting or safety-impacting? Priya discovered the original team never wrote down whether a high score would delay a payment or deny it, a difference with enormous consequences for a family waiting on rent money.

2. Training: teaching the model from data

Training is where the model learns patterns from historical data, and the central government risk lives here: the model learns whatever bias is baked into the past. The fraud tool was trained on three years of historical investigations. But investigators in those years had disproportionately scrutinized claims from one zip code. The model dutifully learned that pattern and amplified it. Chain of custody for the data matters as much as its content, and so do Privacy Act limits on how data collected for one purpose may be used for another. This stage demands a documented record of what data was used, where it came from, and what is known to be missing.

3. Testing: proving it works before anyone is affected

Testing means checking the system against held-back data, across demographic subgroups, against adversarial inputs, and in scenarios that approximate real conditions, all before it touches a real claimant. The right test for a fraud tool is not just "how accurate overall?" It is "does accuracy hold for every group, and what happens to the people it gets wrong?" A model that is 92% accurate overall but 70% accurate for one community is a civil rights problem, not a success. This is the stage where red-teaming belongs, and where the model card stops being a draft. Priya found no record any of it was ever run.

4. Deployment: the controlled launch

Deployment is the actual go-live, and it is where the authorization to operate must explicitly address AI-specific controls rather than inheriting a generic one. The mistake Priya inherited was a big-bang launch: the tool went straight to scoring every claim on day one. The disciplined approach is a staged rollout. Run the model in shadow mode first, where it scores claims but a human makes every decision and you compare the two. Then release it to a small share of the caseload, watch, and expand. This is also the stage where notice to affected individuals, an opt-out route where one applies, and human review procedures have to exist in writing rather than in intention, and where any citizen-facing element must meet federal accessibility requirements under Section 508.

5. Monitoring: the stage everyone forgets

This is where Priya's system actually failed. Models degrade over time because the world changes, a phenomenon called drift. A pandemic, a new benefit program, or a fraud ring shifting tactics all change the incoming claims, and a model trained on old patterns slowly becomes wrong. Monitoring means watching accuracy, fairness, drift, usage, user feedback, and incidents in live operation, and alerting a named human when any of them slips. The fraud tool's rejection rate had crept up for months. With a working monitor and an owner, that would have triggered a review in week two, not month fourteen.

6. Retirement: knowing when to turn it off

Every AI system eventually outlives its usefulness or its trustworthiness. Retirement means a planned shutdown: notifying affected stakeholders, preserving records under the Federal Records Act, documenting why it was turned off, and reviewing the historical decisions the system made for any outstanding harm still owed a remedy. That last item is the one agencies skip, and it is the one that generates FOIA and Inspector General exposure for years afterward. The opposite of a planned retirement is a system that runs forever because no one is willing to own the decision to stop it, which is exactly the zombie Priya found.

The Federal Seven-Phase Decomposition

The six stages are how a practitioner thinks about the arc. Federal lifecycle documentation usually cuts the same arc into seven phases, and the difference is worth understanding because oversight bodies use the longer list. The seven are: requirements and authority to use; data preparation and model training; evaluation and red-teaming; deployment and Authority to Operate; production monitoring; maintenance and retraining; and retirement and decommissioning. The seventh stage in the practitioner's list is the same; the extra phase is maintenance and retraining, split out from monitoring because it behaves differently.

Splitting it out matters. Retraining is not a maintenance ticket. Models age, the world changes, and you retrain on updated data, validate, and redeploy. But that sequence is a return to phases two through four under change-management discipline, not a patch applied in place. Every retraining produces a retraining trigger record, a new evaluation report, an updated model card, a change-management ticket, and an updated authorization. An agency that retrains quietly has, in effect, deployed a new system without the gates the first one had to clear.

Artifacts, Not Slogans

Every phase produces artifacts. This is the practical test of whether a phase actually happened: not whether someone remembers doing the work, but whether the document exists, is dated, and is signed by someone with authority. When GAO, an Inspector General, or a congressional committee opens an inquiry into a federal AI system, what they request is lifecycle artifacts. Agencies that can produce complete artifacts usually close the inquiry quickly. Agencies that cannot spend the intervening period reconstructing what happened, from the memories of people who have often moved on.

PhaseWhat you are decidingArtifacts that prove it happened
Requirements and authorityWhether the agency may do this at all, and to whom it appliesRequirements document, authority-to-use memo, risk tier classification, stakeholder analysis, initial Privacy Threshold Analysis
Data preparation and trainingWhat the model learns from, and what that data carries with itData governance plan, data quality assessment, bias audit of training data, data-use agreements with source systems, training logs, draft model card
Evaluation and red-teamingWhether it works, for whom, and under attackEvaluation report, demographic disparity analysis, adversarial robustness report, red-team findings, remediation plan, final model card
Deployment and authorizationWhether it may go live, and under what human controlAuthorization to Operate, deployment plan, human-review procedures, notice materials, opt-out procedures, consultation record
Production monitoringWhether it is still working, and who is watchingMonitoring plan, dashboards, quarterly performance report, incident log, records of incident-response activations
Maintenance and retrainingWhether the current model is still the approved modelRetraining trigger record, new evaluation report, updated model card, change-management ticket, updated authorization
Retirement and decommissioningHow it stops, and what survives itDecommissioning plan, stakeholder notice, records retention schedule, historical-decision review report

Notice what this table is not. It is not a list of things to write once and file. Several artifacts are living documents that a later phase invalidates: a model card written at evaluation is stale the moment the model is retrained, and a monitoring plan written before deployment is a hypothesis until dashboards exist to test it. Treat the column as a set of documents with expiry dates rather than a checklist to clear.

Minimum Practices, Mapped to Phases

OMB Memorandum M-24-10 sets minimum practices for two categories of AI. For rights-impacting AI the memorandum lists pre-deployment testing, ongoing monitoring, human review, notice, opt-out, and consultation. For safety-impacting AI it lists pre-deployment testing, ongoing monitoring, real-world testing where feasible, independent evaluation, contingency plans, and human decision-making. Read as a list, these look like six and six separate compliance items. Read against the lifecycle, they are the same seven phases seen from the oversight side.

The mapping is direct. Pre-deployment testing is demonstrated in the evaluation phase. Ongoing monitoring is demonstrated in production monitoring. Human review and notice are established at deployment and sustained through monitoring. Opt-out is established at deployment. Consultation begins at requirements and is reinforced through the lifecycle rather than performed once. Independent evaluation and real-world testing sit alongside evaluation. Contingency plans belong with deployment and are exercised through the incident-response playbook. If you know which phase you are in, you know which minimum practice you are currently being asked to evidence.

Crosswalks to NIST and GAO

Two other frameworks describe the same territory with different vocabulary, and knowing the translation saves a great deal of argument in a review meeting. The NIST AI Risk Management Framework organizes AI governance into four functions: govern, map, measure, and manage. It is voluntary and non-binding, but it is the shared language most federal AI documentation is written in. The lifecycle stages above are simply where those functions get applied: mapping happens at requirements and data, measuring at evaluation, managing at monitoring and maintenance, and governing runs across all of them.

GAO's AI Accountability Framework, published in 2021 as GAO-21-519SP, organizes federal oversight of AI around four principles: governance, data, performance, and monitoring. Those principles also land on phases. Governance covers requirements and deployment authorization; data covers requirements and training; performance covers evaluation and production; monitoring covers production and maintenance. You do not have to memorize any of these frameworks. You have to be able to say, when asked, which phase your system is in and which artifact answers the question.

A Lifecycle Ownership Chart

The root cause of Priya's problem was not technical. It was that no one owned each stage. A simple RACI chart, naming who is Responsible for doing the work, Accountable for the outcome, Consulted, and Informed, prevents this. Adapt the table below for any AI system in your agency and fill in real names, not titles. A title is not a person; when the incumbent leaves, a title-based assignment quietly becomes nobody.

StageKey question to answerWho is Accountable (one name)Evidence to keep on file
DevelopmentWhat decision does this drive, and is it legal?Program ownerProblem statement, legal authority memo
TrainingWhat data, and what bias does it carry?Data leadData source sheet, known-gaps note
TestingDoes it work fairly across all groups?Quality leadTest results by demographic group
DeploymentDid we roll out gradually and accessibly?IT and operations leadRollout plan, Section 508 check
MonitoringWho gets alerted when it drifts?Named monitorLive dashboard, alert thresholds
RetirementWhen and why do we turn it off?Program ownerShutdown plan, archive location

Sector Layers on Top of the Enterprise Lifecycle

The seven phases are the enterprise baseline. Individual agencies and data types layer additional obligations on top, and those layers change what a phase must produce rather than changing the phase itself. Health data brings HIPAA obligations, education data brings FERPA, federal systems bring FISMA, and criminal justice data brings CJIS requirements. A benefits system handling all four data types inherits all four sets of constraints at the data preparation phase, and the data governance plan has to show it.

Agency-specific direction layers the same way. DoD Directive 3000.09 governs autonomy in weapon systems. FDA guidance on software as a medical device shapes clinical AI. The HHS Trustworthy AI Playbook, IRS model governance practices, and the DHS AI Roadmap each set expectations for their own agencies. None of these replace the enterprise lifecycle. They add phase-specific requirements to it, which is why the first question in any new AI effort is which layers apply, asked before the requirements document is drafted rather than after.

How Lifecycles Fail in Practice

The failure modes are predictable enough to name. Agencies that skip requirements produce drifting systems, because nothing was ever specified precisely enough to drift from. Agencies that skip red-teaming produce embarrassing public incidents. Agencies that skip monitoring produce discrimination cases, because the disparity accumulates in silence. And agencies that skip retirement planning produce FOIA and Inspector General exposure for years, because the records needed to explain a past decision were never scheduled.

The public record supports the pattern. The Michigan unemployment MiDAS system was deployed without adequate evaluation or human review and produced tens of thousands of false fraud accusations, followed by settled litigation. Federal facial recognition programs at CBP and TSA were both adjusted after the NIST Face Recognition Vendor Test demonstrated significant demographic disparity at specific algorithm and vendor combinations, which is an evaluation-phase gap surfacing after deployment. The course materials also cite VA pausing several predictive analytics tools after an Inspector General review found lifecycle documentation gaps, and IRS pulling back automated audit-selection models after GAO found audit rates that did not track compliance risk.

In every case the remedy was the same, and it was unglamorous: specify requirements, curate and document data, evaluate for accuracy and fairness, deploy only with notice and human review, monitor in production, maintain through disciplined retraining, and retire with an audit trail. There is no version of this where a clever technical fix substitutes for the sequence.

Putting It to Work This Month

Priya's recovery plan was not glamorous either, and that is the point. She froze new rollouts, pulled the training data sheet (there was not one, so she reconstructed what could be reconstructed and documented the rest as unknown), ran the fairness test that had been skipped, stood the monitoring dashboard back up, and assigned one accountable name to each stage. Within six weeks the agency could finally answer the auditor's three questions: who tested it, who approved it, and how we would know if it failed. The model itself barely changed. What changed was that it now had a lifecycle around it.

You do not need to be an OMB lawyer to operate this way. You need to know which phase you are in, what artifacts that phase is supposed to produce, and which oversight body is likely to ask for them first. If you take one action from this lesson, take this: pick one AI system your office relies on and find out who is accountable for monitoring it. If the answer is a shrug, you have found your next priority.

Anti-Patterns

These are the failure shapes that show up repeatedly in lifecycle reviews. Each one looks like diligence from the inside.

  • Treating the FISMA authorization as AI coverage. A clean Authority to Operate says the confidentiality, integrity, and availability controls were assessed. It does not say the model was tested for subgroup disparity, checked for drift, or probed adversarially. Signing one and considering the AI risk closed is the most common version of this mistake.
  • Retraining in place. A model retrained on new data is a new model. Pushing it to production on a maintenance ticket, without a fresh evaluation report and an updated model card, means the deployed system is no longer the one that was authorized, and nobody in the approval chain knows.
  • Assigning a stage to a title. "The data lead owns training" survives exactly as long as the current data lead. Names in the RACI chart with a review date attached survive turnover; titles do not.
  • The artifact written for the auditor. A monitoring plan drafted the week before a review, describing a process nobody has performed, is worse than no plan. It documents an intention as though it were a practice, and the gap between the two is exactly what a reviewer is trained to find.
  • Monitoring only what the dashboard already plots. Drift is detected only in what you chose to measure. A dashboard showing throughput and uptime will look perfectly healthy while approval rates for one language group fall away underneath it.
  • The system with no off switch owner. When nobody is willing to own the decision to stop, the default is to keep running. That is not a decision to continue; it is the absence of one, and it is indistinguishable from a decision to continue right up until it is examined.

Practice Prompts

  • Pick one AI or automated decision system your office uses. For each of the seven phases, write down the artifact that would prove the phase happened, then go find out whether it exists. Mark each one found, missing, or stale.
  • Take that same system and classify it: is it rights-impacting, safety-impacting, both, or neither? Write the one-paragraph justification you would give an Inspector General, and note which minimum practices your classification triggers.
  • Draft the RACI chart for that system with real names in the Accountable column. Send it to each named person and ask them to confirm in writing that they accept the stage. Note who does not reply.
  • Write the retirement plan for a system that is currently running fine. Include the records retention schedule and the process for reviewing historical decisions it has already made. Notice which parts you cannot answer without asking someone else.
  • Identify which sector layers apply to your system's data: health, education, criminal justice, federal information security, or agency-specific direction. For each layer, name the phase where it changes what you must produce.

Reflection

Think about the last AI or analytics system your organization put into production. At which phase did the documentation trail go cold? In most agencies the answer is somewhere between deployment and monitoring, because that is where a project team disbands and no standing owner takes over. Ask yourself who would notice, today, if that system's outputs shifted for one group of applicants, and how long it would take them to notice. If you cannot name the person and estimate the interval, you have described a phase with no owner rather than a system with no problems.

Glossary

  • AI lifecycle: The sequence of phases an AI system passes through from requirements to decommissioning, each with its own duties, owner, and evidence.
  • Authority to Operate (ATO): A formal authorization, issued under FISMA, permitting a system to run in production. Written against the SP 800-53 control catalog, it addresses confidentiality, integrity, and availability, not model behavior.
  • Model card: A standardized document describing what a model was trained to do, what data it used, how it performs across subgroups, and what its known limitations are. Invalidated by retraining.
  • Drift: Degradation of model performance over time as the world it was trained on changes, often while inputs still appear normal.
  • Red-teaming: Deliberately adversarial testing that attempts to make a system fail, misbehave, or produce disparate outcomes before deployment does it for you.
  • Shadow mode: A deployment pattern in which the model produces outputs that are recorded and compared but not acted on, while humans continue to make every decision.
  • Rights-impacting and safety-impacting AI: Risk categories under OMB Memorandum M-24-10, each carrying its own set of minimum practices demonstrated at specific lifecycle phases.
  • Privacy Threshold Analysis and Privacy Impact Assessment: Sequenced privacy reviews, the first determining whether a fuller assessment is required and the second documenting the handling of personal information.
  • System of Records Notice (SORN): The published notice required under the Privacy Act when an agency maintains records retrievable by personal identifier.
  • Decommissioning plan: The documented approach to retiring a system, covering stakeholder notice, records preservation, and review of decisions the system already made.

Closing

An AI system is not finished when it launches. Launch is the moment it starts to drift, and the moment someone must start watching. Everything in this lesson follows from taking that sentence literally: the artifacts exist because a later reader will need them, the phase gates exist because each one attaches a different duty, and the named owner exists because duties assigned to no one are performed by no one.

Priya's tool is still running. The difference is that a named person receives an alert when its rejection rate moves, a dated evaluation report sits in a folder that someone can find, and there is a written answer to the question of when it gets turned off. None of that made the model more accurate. All of it made the agency able to answer for it, which in government is the harder and more consequential achievement.

Key Takeaways

  • AI is a lifecycle, not a project. Development, training, testing, deployment, monitoring, and retirement each need attention; launch is the beginning of the work, not the end.
  • A FISMA authorization does not cover AI-specific risk. Subgroup error, silent drift, adversarial manipulation, and uneven explainability each attach to a different legal duty and none are found by the traditional authorization process alone.
  • Bias enters during training. A model learns the patterns of its historical data, including past unfairness, so document where the data came from, under what authority you may use it, and what is missing.
  • Test across groups before anyone is affected. Overall accuracy can hide serious harm to a specific community; measure per group, probe adversarially, and finish the model card before deployment, not after.
  • Deploy gradually, never all at once. Shadow mode and staged rollouts catch problems while a human still owns every decision, and deployment is where notice, opt-out, human review, and Section 508 obligations become concrete.
  • Retraining re-enters the lifecycle. A retrained model is a new model; it needs a new evaluation, an updated model card, and an updated authorization, not a maintenance ticket.
  • Monitoring is the stage that fails silently. Models drift as the world changes, and drift is detected only in what you chose to watch; a named human plus a live dashboard turns month-fourteen disasters into week-two reviews.
  • Plan for retirement, including the decisions already made. Preserve the records, notify stakeholders, and review past decisions for outstanding harm; systems with no off-switch owner become unaccountable zombies.
  • Artifacts are the proof a phase happened. Oversight bodies request lifecycle documents, not recollections, and the agencies that can produce them close inquiries fastest.
  • Assign one accountable name per stage. A RACI chart with names rather than titles, crosswalked to the voluntary NIST AI Risk Management Framework and the GAO accountability principles, is the cheapest insurance your agency can buy.

Frequently Asked Questions

Is it six stages or seven phases?

Both describe the same arc. The six-stage version is how a practitioner holds it in mind: development, training, testing, deployment, monitoring, retirement. The seven-phase federal version splits maintenance and retraining out of monitoring, because retraining is a return to the earlier phases under change-management discipline rather than a routine maintenance activity. Use the seven-phase list when you are producing artifacts for oversight, because that is the vocabulary the requests arrive in.

Our AI is a commercial tool we did not build. Does the lifecycle still apply?

Yes, and the artifacts still have to exist. You may not control training, but you still own the decision to use the tool, the authority under which you use it, the evaluation on your own data, the notice to affected individuals, the monitoring, and the retirement. Lifecycle control is something the practitioner does. It is not passive observation of a vendor's process, and a vendor's documentation does not discharge your agency's duty to produce your own.

How do I know whether my system is rights-impacting?

Start from what the output does to a person rather than from how the model works. If the output influences a determination about benefits, enforcement, employment, or another consequential decision affecting an individual's rights, treat it as rights-impacting and write down the reasoning. The classification decision belongs in the requirements phase, in writing, because every later phase inherits obligations from it. If the classification is genuinely uncertain, that uncertainty is itself something to escalate rather than resolve quietly.

What is the single most common gap you would find in my agency today?

A monitoring plan with no owner, or an owner with no dashboard. The requirements and evaluation artifacts usually exist in some form because a project team produced them under deadline. Monitoring artifacts fail differently: the project team disbands after go-live, the standing owner is never named, and the documents that do exist describe a process nobody performs. Priya's fourteen-month gap was this exact shape.

Does following the lifecycle mean my system is compliant?

No. Working the lifecycle produces the evidence that lets you demonstrate what you did and did not do. It does not certify that the system is fair, accurate, or lawful, and a complete set of artifacts describing an inadequate evaluation is still an inadequate evaluation. The lifecycle makes your agency's actual practice visible and reviewable. Whether that practice is good enough is a separate judgment, and one an oversight body will make for itself.