←
AI for Government
Proficient · M20 · lesson 20 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Continuity of Operations with AI
📖
now learning

Continuity of Operations with AI

15 min

At 6:14 on a Monday morning, the AI vendor that powered a state unemployment agency's eligibility-screening tool pushed a routine update and broke its own application programming interface, the connection that let the agency's systems talk to the model. The screening tool went dark. By 9 a.m., 3,400 claims were stacked up with nowhere to go, and the agency's program director, Reuben Carter, made the discovery that ruined his week: nobody had ever written down what to do when the AI stopped working. Staff had quietly stopped reviewing claims manually months earlier because the tool was so reliable. They had lost the muscle. The outage lasted six hours. The backlog took eleven days to clear.

This lesson is about that gap. Continuity of operations, what government calls COOP, is the discipline of keeping essential functions running through disruption. Agencies have COOP plans for fires, floods and power loss. Very few have one for the moment their AI fails. As a program director or strategist, building that plan is your job, and Reuben's morning is the exact failure you are planning against.

Why AI Breaks Traditional Continuity Planning

A normal COOP plan assumes a function either works or stops. AI introduces a third state: it keeps running and starts being wrong. That is the failure mode traditional planning misses entirely. Reuben's outage was the easy case, because everyone could see the tool was down. The dangerous case is the model that drifts, quietly approving or denying claims it should not, while every dashboard stays green and nobody has any reason to look.

AI systems are also fragile in ways conventional software is not. They depend on particular computational environments, particular framework versions, particular configuration. When any of those changes unexpectedly, behaviour changes with it, and you often cannot simply restart the previous version, because the data it was trained on may be outdated, unavailable or no longer permitted for use. Recovery is therefore rarely a matter of bringing a server back. It is a matter of establishing whether the thing you brought back is still the thing you validated.

Government carries a specific burden here. Citizens cannot opt out of your service and cannot switch to a competitor when it fails. Many agencies operate continuously across time zones, so there is no window in which an outage inconveniences nobody. And a great many functions carry legal deadlines: processing an application, issuing a permit, responding to a request. An AI failure that causes you to miss a statutory timeframe does not just create a backlog. It creates liability, and it creates it on a clock that started before you noticed.

Three Operating States You Have to Plan For

Whatever the underlying cause, an AI-supported function ends up in one of three states, and a plan that covers only the first is the most common gap in government.

  • Hard failure: the system is down. A vendor outage, an expired credential, a cloud region offline. Visible and total, and the easiest kind to plan for precisely because nobody has to be persuaded it happened.
  • Degraded performance: the system runs, but slowly or partially. Latency spikes, timeouts, low-confidence answers across the board. Service continues at a quality nobody has defined as acceptable or unacceptable.
  • Silent failure: the system runs and looks fine while producing wrong outputs. Data drift, a bad model update, an upstream feed that changed format. The most dangerous state, because nothing pulls the alarm and the decisions keep going out the door.

The NIST AI Risk Management Framework, which is voluntary guidance rather than binding regulation, names this directly under its Manage function: trustworthy AI requires planning for system failure and for safe degradation. Your COOP plan is where that expectation turns into a procedure a duty officer can follow at 6:14 in the morning without needing to interpret a framework.

What Actually Fails, and Why It Matters for the Plan

The three states above are what you experience. Underneath them sit six recurring causes, and each one implies a different response, which is why a plan built only around "the system is down" tends to fail on contact.

Technical failures are the familiar ones: the model server crashes, the database is unavailable, the network drops. They resemble ordinary IT failures, with the complication that restoring or retraining a model may take time that a database restore does not. Data failures are subtler: the pipeline feeding the model stops producing, or the historical data the model depends on becomes corrupted or inaccessible. A model that predicts permit approval times cannot predict anything if the permit history it draws on is offline, even though the model itself is running perfectly.

Performance degradation is the silent state, and it has four typical causes. The world has changed, which is concept drift, so a model trained on historical employment data behaves badly during an economic shock. The data distribution has changed, so a model built on submissions from one region is now being applied to another with different characteristics. The training data has become contaminated or simply out of date. Or a system update quietly changed how data flows into the model. In every case the system appears to be working, and you often discover the problem only when downstream users notice that the decisions seem wrong.

Cascading failures occur when one system's failure triggers others: a shared data platform goes down and takes several AI systems with it, or an update to a shared machine learning platform breaks every model that depends on it. Model dependency failures come from outside altogether, when a third-party provider changes a model, changes pricing, or discontinues a service. That is not a failure in your system, it is a failure in your supply chain, and your continuity plan is the only place it can be handled.

The sixth is the one people find least comfortable. Human decision-making can degrade too. When the AI fails and staff take over, quality does not automatically return to the pre-AI baseline. People who are stressed, unpractised or overwhelmed make more mistakes, apply rules inconsistently and tire faster. A system that falls back to "humans decide everything" may be technically functional and operationally worse than what it replaced. That possibility belongs in the plan, not in the after-action report.

Classify What Actually Depends on the AI

Not every AI system needs a full continuity plan, and building one for all of them guarantees that none is maintained. Plan depth should match how essential the function is. Run every AI-supported function through a single question: if this system is unavailable for a full business day, what happens to the public? The answer sorts your inventory into three tiers.

  • Tier 1, mission-essential: people are harmed or a legal deadline is missed. Unemployment eligibility, benefits payments, emergency dispatch triage. These need a tested manual fallback that can run for days, not hours.
  • Tier 2, important: service degrades and backlogs build, but nobody is immediately harmed. A permit-routing tool, a document classifier. These need a defined workaround and a recovery timeline.
  • Tier 3, convenience: staff are inconvenienced. An internal drafting assistant. These can be switched off until they are fixed.

Reuben's eligibility tool was unmistakably Tier 1, because people's rent depended on those claims. It had a Tier 3 plan, which is to say none at all. That mismatch between the tier a system occupies and the plan it was given is the most common and most expensive error in government AI operations, and it is usually invisible until the morning it is not.

Set Recovery Objectives Before You Need Them

Two standard continuity concepts apply to AI systems as they do to any other. The recovery time objective, RTO, is how long the system can be unavailable before unacceptable harm occurs. For a critical function such as immigration processing, an RTO might be four hours, beyond which applications start missing statutory deadlines. For an analyst dashboard, twenty-four hours may be perfectly tolerable. The recovery point objective, RPO, is how much data loss is acceptable: for a transactional system that might be fifteen minutes, and that answer is what drives your backup frequency.

Set both explicitly, then design against them and test the design. If your RTO is four hours and an honest recovery takes six, you have not met the requirement, and knowing that in advance is worth considerably more than discovering it during an incident. Recovery objectives also settle arguments quickly under pressure, because they convert "how bad is this?" into a comparison against a number the agency already agreed to.

The Four Pillars of an AI Continuity Plan

A workable plan rests on four pillars. Build each one for every Tier 1 system before you need it, and for Tier 2 systems in proportion to the backlog they would create.

Detection: know that it failed

You cannot recover from a failure you have not noticed, and silent failure is the hard one. Detection means monitoring the three operating states rather than uptime alone, and setting thresholds that page a named human instead of writing to a log nobody reads. It is worth being clear about the limit of this control: monitoring surfaces the failure modes you chose to watch, on the signals you chose to watch them with. A monitored system is not a safe system, it is an observed one, and the gap between those is where most silent failures live.

Fallback: keep the function running

This is the manual or simplified process that takes over. The critical discipline, and the one Reuben missed, is that the fallback has to be kept alive. A manual review process nobody has run in a year is not a fallback, it is a hope. Staff lose the skill, the forms go stale, the rules change underneath the documentation. The fallback must be documented, staffed and exercised on a schedule, and the schedule has to survive the quarters when nothing goes wrong.

Degraded operations: run at reduced capacity

Often you cannot fully replace the AI by hand, so you run a triaged subset instead. During the outage, Reuben's team could have processed the oldest and most urgent claims first while queuing the rest and telling the public a realistic delay. Defining those triage rules in advance is what makes a slower but functioning service possible, rather than a room full of people deciding priority individually under pressure. It reduces the chaos considerably; it does not eliminate it, and the plan should say which work is being deliberately deferred.

Recovery: get back to normal safely

When the AI returns, do not simply switch it back on. A model that failed silently may still be wrong, and a model that was restored may not be the same model you validated. Recovery means testing the restored system against known cases before trusting it, clearing the backlog against a plan, and reviewing the decisions the system made in the window before you caught the failure. That back-window review is the step most often skipped, and it is the one that determines whether affected residents are ever identified.

Designing a Fallback That Is Worth Having

The core principle is that every critical AI system needs a non-AI path. That does not mean running two systems in parallel forever. It means that when the AI is unavailable, a process exists that lets government keep operating, and that the process does not have to be as sophisticated as what it replaces. It has to work.

The pattern is usually to substitute a simple rule for a learned judgment. A model that predicts which applications are high priority fails, so the fallback processes applications in the order they were received: less efficient, entirely functional. A model that scores risk fails, so the fallback applies a straightforward rule based on account age or history: less sophisticated, still usable. A model that detects potential fraud fails, so the fallback flags every transaction above a set amount for human review: lower performance, and it still catches some fraud.

A fuller worked example: a social benefits agency runs a model that decides which applicants need additional verification before approval, across ten thousand applications a day. Its documented fallback is to approve anything the AI would have approved, flag anything the AI would have sent to verification for human review, and, if the model is entirely unavailable, apply a simple rule that approves returning applicants and flags new ones. The error rate rises. Service continues. That trade is the entire point, and stating it explicitly in the plan is what stops someone improvising a worse trade during an incident.

Two cautions belong with this. First, the fallback has to be designed in from the beginning rather than bolted on, because a fallback invented after deployment tends to depend on data or staff the operation no longer has. Second, a fallback can be worse than the problem. A tax agency whose fraud model failed fell back to flagging every return above an income threshold, which sent roughly forty percent of returns for investigation instead of the five percent the model had selected. The office was overwhelmed and genuine cases were missed among the false positives. Accept that a fallback will be less efficient; do not accept one that is drastically worse. Where the trade is severe, use tiered fallbacks, so a second and simpler path exists if the first is also unavailable.

What to Monitor, and What the Numbers Should Say

Detection needs specific signals with specific thresholds. Six are worth building for any Tier 1 system. Availability: is the model responding to requests at all? Latency: how long does a prediction take, and has that suddenly jumped by an order of magnitude? Output distribution: are the outputs staying in their expected range? Consistency: do identical inputs still produce identical results? Comparison against a known-answer test set, run periodically. And upstream data quality: nulls where there should not be nulls, values outside plausible ranges, formats that changed without notice.

Output distribution is the signal that catches silent failure, and it works best when the baseline is written down. One benefit eligibility system historically classified 45 percent of applications as eligible, 35 percent as requiring additional review and 20 percent as ineligible. Its alert bands were set around those proportions, at 40 to 50 percent, 30 to 40 percent and 15 to 25 percent, and a move outside any band triggered an investigation. The bands sum correctly against the baseline, which is worth checking on your own thresholds, because an alerting scheme whose bands do not reconcile will either never fire or never stop firing.

Then make the alert reach a person. Implement alerting so that a degradation notifies someone immediately, with a named owner rather than a distribution list, and do not rely on users to report the problem. By the time a downstream user complains, the failure has already been producing decisions for some period you will have to reconstruct afterwards.

Managing Degradation and Retraining

Sometimes nothing is unavailable and the predictions have simply become less accurate. The causes are the ones listed earlier: concept drift, data drift, training data that has aged out, or a system change nobody validated against the model. Handling this is a continuity concern rather than a data science luxury, because a degraded model in a Tier 1 function is an outage that has not been declared.

Five practices manage it. Establish performance baselines, so that "good" is a written number rather than an impression; for a classification model that might be a stated accuracy, precision and recall. Monitor continuously against a validation dataset, and investigate when performance falls below the baseline. Plan for retraining before you need it, including how long it takes and whether it can happen without taking the system offline, because a three-day retraining cycle is three days of something in your continuity plan. Version your models, so that a retrained model that performs worse can be rolled back quickly. And define the retraining triggers in advance, whether that is a stated accuracy floor, a fixed interval, the arrival of significant new data, or a change to the surrounding system.

AI Continuity of Operations Plan Template

Fill one of these for each Tier 1 and Tier 2 AI system. Keep it short enough that a stressed duty officer can act from it at six in the morning. The worked column shows Reuben's eligibility tool, written the way it should have been.

FieldWhat to defineWorked example (eligibility screening)
System and tierName and essentiality tierEligibility screening tool, Tier 1
Function it supportsThe public-facing serviceUnemployment claim eligibility decisions
RTODowntime before real harm4 hours before backlog risks late payments
RPOTolerable data loss, drives backup frequencyStated per system by the data owner
Detection: hardAlert, threshold, ownerAPI uptime check every 60s; page Ops on 2 failures
Detection: silentOutput signal monitoredDaily approval rate against 30-day baseline; alert if the shift exceeds 10 percent
Upstream data checkFeed validity signalNull and range checks on the document feed each morning
Fallback processManual or simplified pathCaseworker manual review using the documented eligibility checklist
Fallback readinessLast exercised, who is trainedQuarterly drill; 12 caseworkers current on the manual checklist
Fallback load capacityCan it carry peak volumeTested against peak daily claim volume, not average
Degraded triage ruleWhat to do first at reduced capacityProcess oldest claims and hardship flags first; queue the rest
Public commsWhat you tell people, who approvesPre-drafted delay notice; public information officer approves within 1 hour
Recovery validationTest before trustingRe-run 50 known cases and compare against expected before resuming automation
Back-window reviewAudit decisions made during failureRe-check claims decided in the 48 hours before the silent failure was caught
Owner and escalationWho runs this, who they callOps lead activates; escalates to the program director at hour 2

Test It, or You Do Not Have One

A continuity plan that has never been exercised is a document, not a capability. The cheapest and most valuable test is a tabletop: gather the team, declare that the eligibility tool has gone silent and approval rates have jumped 15 percent, and walk through who does what against the plan in real time. You will surface gaps in an hour that an outage would otherwise surface over eleven days, and the gaps are usually mundane. A missing phone number. A form that no longer exists. A step that assumes someone who retired.

Run three levels of testing on a schedule. Tabletops quarterly, talking through a scenario, which is cheap and finds the obvious gaps. Fallback drills at least twice a year, actually running the manual process on live or sampled work for a few hours, which is what keeps the muscle alive. And a full failover annually, deliberately taking the system offline in a controlled window and running end to end on the fallback.

Testing has to cover more than the switch. Load-test the fallback against peak volume rather than average, because a manual path that handles a quiet Tuesday and collapses under a Monday surge has not been tested. Document the procedure in enough detail that someone unfamiliar with the system can follow it, then verify that claim by having such a person try. And before a critical system goes live, run a dry run in which the AI fails on day one: walk the recovery, time it, and find the bottlenecks. One immigration agency did exactly that with a visa prioritization model and learned that processing took about 40 percent longer without the model's guidance, which told them in advance that the fallback needed either more staff or different procedures. That is a finding worth having before launch rather than during an incident.

Reuben instituted this after his eleven-day backlog. When a vendor update broke the API again ten months later, the team detected it in nine minutes, moved to the manual checklist they now drilled every quarter, processed urgent claims by hand, and validated the model against fifty known cases before turning it back on. The public-facing impact was a three-hour delay notice and no backlog. The failure was the same; the preparation was not. It is fair to say the plan accounted for most of that difference, and fair to note that a single repeat incident is one data point rather than a controlled comparison.

Three Scenarios Worth Thinking Through

The following are drawn from different levels of government and different failure causes. Each shows a continuity strategy and the limit of that strategy.

A health ministry runs a model predicting which facilities are at risk of running out of critical medicines. It works in normal conditions. When an outbreak spikes demand, its predictions become unreliable because they rest on historical patterns that no longer apply, and the ministry keeps routing supplies to facilities with low predicted demand while others run out. The continuity strategy has four parts: an automatic flag when predictions deviate materially from actual consumption, which switches supply decisions to a conservative rule based on facility size and population served; designated staff monitoring model performance during a crisis with authority to override; a retraining plan once crisis data accumulates, with extensive testing before the new version deploys; and a paper-based backup in which regional health officers submit needs directly to the national office. Note the limit on the first part: the model flags the deviations it was configured to measure, so a crisis that changes something nobody instrumented will not trigger it, which is exactly why the human monitoring exists alongside.

A municipal permitting system uses a model trained on three years of history to route applications by complexity, sending complex permits to experienced reviewers. A major development proposal produces a tenfold increase in applications, unlike anything in the training data, and classification becomes unreliable. The strategy is a load-based fallback that switches to a simple cost-threshold rule when a surge is detected, an alert to managers that temporary reviewer capacity is needed, staged processing that queues and batches applications by submission date rather than attempting everything at once, and retraining afterwards on the combined historical and surge data so the next spike is less novel.

A regional trade agreement depends on automated document verification, and when one country's model goes down, transactions stall and supply chains are disrupted. The strategy distributes backup models across multiple countries so traffic can be routed to an alternative, escalates to human reviewers when no model is available, writes response times into the agreement so that manual review triggers automatically when automated verification runs long, and includes regular drills in which one country's system is taken offline and partners must activate their fallbacks. Two honest caveats: routing to alternative models reduces dependence on any single site but does not remove shared exposure where those models rest on the same platform, data source or supplier, which is the cascading failure pattern described earlier. And escalating to human review supports continuity at reduced throughput rather than guaranteeing it, since the reviewers have their own capacity limit and the escalation is only as good as the drill that proved it.

Anti-Patterns

  • No fallback at all. A critical system is deployed with no defined non-AI path, because designing for failure is uncomfortable and the effort feels better spent on the primary system. When the model fails, staff do not know which work to prioritise and process it effectively at random. Establish the fallback requirement before you start building, not after go-live.
  • An untested recovery procedure. The document exists and has never been exercised, because testing is disruptive and the event is hypothetical. Then the backup turns out to be corrupt, the steps are in the wrong order, and the author has retired. Test at least twice a year and treat it with the seriousness of a security audit.
  • Treating a clean test as a guarantee. Re-running fifty known cases before restoring automation is a sound control and it is evidence, not proof. It tells you the system handles those cases correctly today. Pair it with heightened monitoring for a defined period after restoration rather than declaring the incident closed at the moment the test passes.
  • Ignoring performance degradation. The model runs, nobody monitors output quality, and months of decisions accumulate under a degraded model before an analyst notices a downstream number moving. Establish baselines and alerting before deployment, and review the monitoring data periodically even when nothing has fired.
  • A fallback worse than the failure. An over-simple substitute rule floods the operation with work it cannot absorb, so genuine cases are missed among false positives. Test fallbacks under realistic load, accept reduced efficiency, and redesign anything whose error rate makes the situation worse.
  • Letting the manual skill decay. Staff stop exercising the process the tool replaced, and the documented fallback quietly becomes unusable while still appearing in the plan. Schedule drills, track who is current, and treat a lapsed roster as an open finding.
  • Assuming the humans will simply cope. Falling back to human judgment does not restore the pre-AI baseline. Fatigue, unfamiliarity and inconsistency all rise under incident conditions. Plan for supervision, rotation and spot-checking during extended fallback operation.
  • Closing the incident without the back-window review. Service is restored, the queue is cleared, and nobody examines the decisions the system made while it was silently wrong. Those decisions affected real people and are the agency's responsibility whether or not anyone looks for them.

Practice Prompts

  • Identify the critical systems. List the AI systems your organisation relies on for mission-critical functions. For each, name the failure mode with the highest impact, whether that is unavailability or unreliable output, and assign a tier.
  • Design a fallback procedure. Take one Tier 1 system and write out the fallback steps. Estimate how much longer each step takes without the model, and mark where the bottleneck appears.
  • Set RTO and RPO. For two critical systems, define how long they can be down and how much data loss is acceptable. Document both in a form your vendor contract can reference.
  • Build a monitoring plan. For one system, decide what would indicate failure or degradation, establish a baseline value for each signal, and set alert thresholds and a named recipient.
  • Run a dry run. Schedule a simulation in which one AI system becomes unavailable. Walk the recovery with the team, time it, document what broke, and update the procedure.
  • Load-test the fallback. Work out whether your manual path can carry peak volume rather than average volume, and what you would do on the day it cannot.
  • Audit the back window. Take a past incident, or a hypothetical one, and define exactly which decisions you would re-examine, how you would find them, and who would contact affected residents.

Reflection

Which AI system does your organisation most depend on, and what would actually happen if it failed tomorrow morning? Is there a documented fallback, and when was it last exercised by someone who would have to run it? How long would recovery genuinely take, as opposed to how long the plan says, and have you ever measured the difference? Then ask the question that catches most agencies: if that system were producing plausible but wrong outputs right now, what signal would tell you, how long would it take to fire, and who would receive it? If the honest answer is that a downstream complaint would be the first indication, that is the gap to close this month.

Glossary

  • Continuity of operations (COOP). The discipline of keeping essential government functions running through disruption, including the plans, fallbacks and exercises that make that possible.
  • Recovery time objective (RTO). The maximum tolerable duration of an outage before unacceptable harm occurs. Set per system, and used to judge whether a recovery design is adequate.
  • Recovery point objective (RPO). The maximum tolerable data loss, expressed as a period. It determines how frequently data must be backed up.
  • Hard failure. The system is unavailable and the failure is visible. The easiest state to detect and the one most existing plans already cover.
  • Degraded performance. The system runs but slowly or partially, delivering service at a quality nobody has classified as acceptable or not.
  • Silent failure. The system runs and appears healthy while producing wrong outputs. Detected only through output monitoring or downstream complaint.
  • Concept drift. A change in the underlying relationship the model learned, caused by the world changing rather than by the data pipeline breaking.
  • Data drift. A change in the distribution of inputs reaching the model, such as applying a model built for one population to another.
  • Cascading failure. One component's failure propagating to others through shared infrastructure, platforms or suppliers.
  • Model dependency failure. A disruption originating with a third-party provider that changes, reprices or discontinues a model you depend on. A supply chain issue rather than a system issue.
  • Degraded operations. Deliberately running a reduced but functioning service under pre-defined triage rules while the primary system is unavailable.
  • Back-window review. The audit of decisions made between the start of a failure and its detection, to identify people affected while the system was wrong.
  • Tabletop exercise. A discussion-based walkthrough of a scenario against the written plan, used to surface gaps cheaply and frequently.

This lesson sits inside the risk chapter. Enterprise AI Risk Management provides the frame that decides which systems get this treatment, and Continuous Monitoring Fundamentals builds the detection layer the whole plan depends on. AI Incident Response Planning and Crisis Management for AI Failures cover what happens once a failure is declared, including the public communication your template pre-drafts. AI Red-Teaming Fundamentals is how you find the failure modes before they find you, and Third-Party AI Risk Management together with AI Supply Chain Risk address the vendor dependency that caused Reuben's outage in the first place. For the systems most likely to be Tier 1, see AI for Mission-Critical Government Functions and Rights-Impacting and Safety-Impacting AI Safeguards.

Closing

Continuity of operations is not glamorous work. It is planning for failure, writing procedures nobody wants to read, and taking systems offline deliberately to prove that the alternative works. It is also the difference between an AI system that is an asset and one that becomes a liability the first time a vendor pushes an update on a Monday morning.

The principle is narrow enough to hold in mind: every critical AI system should be able to fail gracefully. That requires a designed fallback rather than an improvised one, regular testing rather than a document, monitoring that watches output as well as uptime, explicit recovery objectives, and the ability to roll back to a known model version. None of it prevents failure. What it buys is that the failure stays inside a boundary you chose in advance, and that the people who depend on the service find out from you rather than from the delay.

Key Takeaways

  • Plan for three operating states, not one. Hard failure stops the system, degraded failure slows it, and silent failure keeps it running while it is wrong. The last is the most dangerous and the least often planned for.
  • Know the causes underneath the states. Technical, data, degradation, cascading and vendor failures each demand a different response, and human decision quality can degrade too when people take over unprepared.
  • Match plan depth to essentiality. Classify each function by what happens to the public after a day of downtime, and give Tier 1 systems a tested manual fallback that can run for days.
  • Set RTO and RPO explicitly. If recovery takes longer than the objective you published, you have a gap, and it is far better to find it on a whiteboard than during an incident.
  • An untested fallback is not a fallback. Manual skills decay quickly, so the process must be documented, staffed, drilled and load-tested against peak volume rather than average.
  • Monitor outputs, not just uptime. Track output distribution against a written baseline so silent failure raises a human alert, and remember that monitoring only covers what you chose to watch.
  • Define degraded operations in advance. Pre-set triage rules make a slower but functioning service possible and say plainly which work is being deferred.
  • Do not let the fallback become the disaster. A substitute rule that floods the operation can do more harm than the outage. Test under realistic load and use tiered fallbacks where the trade is severe.
  • Recover deliberately. Validate the restored system against known cases, keep monitoring closely afterwards, and audit the decisions made in the window before detection.
  • Test in three tiers. Quarterly tabletops, fallback drills at least twice a year, and an annual full failover are what turn a document into a capability.

Frequently Asked Questions

Does every AI system need a continuity plan?

No, and pretending otherwise is how plans stop being maintained. Tier the inventory by what happens to the public after a business day of downtime. Tier 1 systems need a tested manual fallback, Tier 2 need a defined workaround and a recovery timeline, and Tier 3 can be switched off until they are fixed. The expensive error is not having too few plans; it is a Tier 1 system carrying a Tier 3 plan.

Our AI is a vendor service. Is continuity their responsibility?

The vendor is responsible for their service. You remain responsible for the public function, and a provider that changes, reprices or discontinues a model has not breached anything you can operate through. Put the response times and notice obligations into the contract, and build a fallback that does not depend on the vendor being available to help you use it.

How do we detect a silent failure that nobody has thought of?

Partly you cannot, and the plan should acknowledge that rather than imply otherwise. Broad-signal monitoring helps: output distribution against a baseline, consistency on repeated inputs, and periodic testing against a known-answer set will catch classes of problem rather than only the specific ones you anticipated. Beyond that, keep the human channel open, because frontline staff routinely notice that decisions look wrong before any metric moves.

Fallback drills are expensive and disruptive. How do we justify them?

Compare them against the incident they are insuring against. Reuben's untested fallback produced an eleven-day backlog from a six-hour outage; the drilled version produced a three-hour delay notice. A tabletop costs an hour and surfaces the missing phone numbers and retired authors that turn a short outage into a long one. Where budget is genuinely tight, run tabletops quarterly and reserve full failover for the systems where a missed deadline creates legal exposure.

Should we tell the public when the AI fails?

Tell them what affects them: the delay, the expected duration and what to do meanwhile. That is why the template includes a pre-drafted notice with a named approver, because the drafting is much harder during the incident. Where decisions may have been wrong during a silent failure, the back-window review determines who needs to be contacted individually, and that obligation does not go away because the system has been fixed.

What if the fallback cannot handle the volume at all?

Then say so in the plan rather than discovering it live, and make degraded operations the primary strategy for that system. Define which work is processed and which is queued, communicate a realistic timeframe, and treat the capacity gap as a documented risk for leadership to accept or fund. A plan that honestly states it can carry half the volume is far more useful than one that assumes full coverage it has never tested.