AI Project Risk Management and Contingency Planning
Nina Vasquez had managed software projects for eleven years before she took on an AI implementation, a customer churn prediction model for a telecom company she consulted with. She brought her standard risk register: budget risk, timeline risk, integration risk, stakeholder risk. Three months in, two risks she had not anticipated materialized simultaneously. The model's training data had a temporal gap, because data from the COVID period had been excluded as anomalous, which meant the model had never seen the demand patterns the company was now experiencing. And the business team that had requested the model had been reorganized, leaving nobody accountable for acting on its predictions. The model was built. It ran. Nobody used it. "I had managed every project risk I knew how to name," Nina said afterwards, "but AI adds categories that were not in my toolkit."
Why AI Projects Fail Differently
AI projects fail for all the reasons traditional technology projects fail: unclear requirements, underestimated complexity, inadequate resources, stakeholder drift. Everything a competent project manager already knows still applies. But AI projects also fail for reasons specific to how AI systems are built and deployed, and those reasons rarely appear on a conventional risk register because they have no clean analogue in deterministic software. A conventional system either meets its specification or it does not. An AI system can meet its specification precisely and still be wrong in production, still degrade quietly over months, and still be ignored by the people it was built for.
Understanding those AI-specific failure modes is the starting point for managing them, and they fall into three broad categories: data and model risks, integration and usage risks, and governance and compliance risks. Nina's project was undone by one risk from the first category and one from the second. Neither was exotic. Both were the kind of thing that becomes obvious in hindsight and invisible in planning, because nothing in a standard risk workshop prompts anyone to ask whether the training data covers the period the business is currently living in, or whether the person who will act on the model's output still has that job.
Data and Model Risks
Data availability risk is the risk that the data needed to train or run the model does not exist, is inaccessible, or is of insufficient quality. This is the most common source of AI project overruns, and the easiest to underestimate, because someone senior will almost always say the data is there. The only reliable way to assess it is to access and profile the data before committing to a delivery timeline. Assurances are not equivalent to having looked at it, counted the gaps, and confirmed that the fields you need are populated in the years you need them.
Data drift is the risk that the statistical patterns in the data the model is deployed on diverge from the patterns in the data it was trained on. This happens because the world changes. A fraud detection model trained before a new tactic emerges will miss the new pattern entirely, and will miss it confidently. A demand forecasting model trained on pre-pandemic data will not generalize to post-pandemic demand dynamics, which was precisely Nina's gap. The mitigation is to define a monitoring plan before deployment and set an explicit threshold, so that a fall in model accuracy beyond a stated amount automatically triggers a retraining cycle rather than a discussion about whether one is warranted.
Model performance risk is the risk that the model does not perform well enough in production to justify the investment. Test set accuracy is not a reliable predictor of real-world business value, and treating it as one is how projects arrive at launch with an impressive metric and an unusable system. A model that is 92% accurate may still produce more errors than a business can absorb, if it is deployed at high volume or in a high-stakes context where each error carries a cost the successes do not offset. Define minimum performance thresholds in terms of business impact rather than statistical metrics, and define them before the project begins, when the number is still a requirement rather than a negotiation.
Overfitting is the risk that the model learned the training data too precisely and performs poorly on data it has not seen. This is a technical failure mode, but its risk can be assessed without deep technical expertise, because the diagnostic questions are simple and the answers are revealing. Ask any AI vendor or internal team what cross-validation methodology was used, and what the gap is between training performance and held-out test performance. A large gap signals overfit. A team that cannot answer either question has not measured the thing you are being asked to buy, which is itself the finding.
Integration and Usage Risks
Adoption risk is the risk that the model is built and deployed and the intended users ignore it or work around it. This was Nina's second failure. The model ran, but the organizational home for acting on its outputs had dissolved during a reorganization, and nothing in the project plan noticed. AI systems require not just technical deployment but an operational workflow: who receives the output, what they do with it, how it fits into their existing process, and what authority they have to act on it. A prediction delivered to nobody in particular is indistinguishable, from the organization's point of view, from a prediction that was never made.
The mitigation is to map the decision workflow before building the model. Identify the specific human who will receive the AI output and act on it, by role and by name. Confirm that the role and its authority will be stable for the duration of the project, and treat a pending reorganization in that part of the business as a live project risk rather than as background noise. Then consider how the recommendation integrates with the tools and reporting cadences that person already uses, because an output that arrives in a channel nobody reads has the same practical effect as no output at all.
Integration risk is the risk that the model cannot connect reliably to the production systems it needs to read from or write to. API changes, data schema changes, permission changes and system version updates all create integration breakpoints, and each is owned by a team that has no reason to know your model depends on it. Systems that require real-time data are especially vulnerable, because the failure often looks like degraded predictions rather than an obvious outage.
Human-AI interaction risk is the risk that users interact with the system in ways that lead to poor outcomes: over-trusting low-confidence predictions, systematically overriding correct recommendations, or misinterpreting the output format. The risk is not simply that the model is wrong. It is that humans plus model combined perform worse than humans alone, an outcome that is invisible if you evaluate the model in isolation. Design the interface with as much care as the model: present confidence levels rather than bare point predictions, show what factors drove the recommendation, and run a limited pilot in which you observe how real users behave.
Governance and Compliance Risks
Regulatory risk is the risk that the AI system violates legal requirements, whether anti-discrimination law, data protection regulation, or sector-specific AI governance rules, either at deployment or as regulations change afterwards. AI regulation is evolving rapidly, which means a system that was compliant when deployed may become non-compliant when new rules take effect, without anything about the system having changed. The mitigation is to conduct a regulatory assessment as part of project scoping rather than as an afterthought: identify which regimes apply, including data protection, sector AI rules, anti-discrimination law and consumer protection, and build regulatory review into the deployment gate.
Reputational risk is the risk that the system produces outputs which, if surfaced publicly, would cause damage: biased decisions, embarrassing errors, or outputs that conflict with the organization's stated values. This is not always captured in technical evaluation, because technical evaluation asks whether the output is accurate on average rather than whether any individual output would be defensible in a regulator's letter. The two questions have different answers, and the second one produces the crisis.
Vendor risk in AI has a specific shape. The vendor providing an AI tool or an underlying model may change their service terms, increase pricing significantly, withdraw the model, or be acquired by someone with different priorities. Dependence on a proprietary model creates concentration risk in a way that dependence on a commodity software component does not, because the model's behavior is the product and cannot be replicated by switching to an equivalent. The mitigation is to understand the vendor's data practices and exit terms before contracting, and to assess honestly how easily you could bring the capability in-house.
The Risk Register for AI Projects
A risk register for an AI project should include the standard project risk categories and add the AI-specific ones described above. For each risk you need the same fields you would use anywhere: a description, the likelihood of occurrence, the business impact if it occurs, the current mitigation, the residual risk after that mitigation is applied, and the owner. The discipline is unremarkable. What is remarkable, in practice, is how consistently the AI-specific risks arrive at deployment with the owner field empty, because they belong to a phase of the work that the project plan treats as the end rather than the beginning.
| Risk that commonly lacks an owner | The question that names the owner |
|---|---|
| Model monitoring | Who watches performance after launch? |
| Retraining | Who decides when retraining happens, and who initiates it? |
| Incident response | Who acts if the model produces a harmful output? |
| Regulatory currency | Who tracks evolving AI regulation and responds to it? |
If these four do not have named owners in your project plan, assign them before deployment rather than after. An unowned risk is not a risk that has been accepted; it is a risk that will be discovered when something goes wrong, by whoever happens to be standing closest to the failure. The distinction matters because accepted risks are visible to governance and unowned ones are not, which means the organization believes it is carrying less exposure than it actually is. Naming an owner also forces a useful conversation about capacity.
Contingency Planning: What You Do When the Model Fails
Every AI system will eventually fail. It will perform below acceptable thresholds, encounter a distributional shift it was not designed for, be affected by a regulatory change, or experience an ordinary technical outage. Contingency planning asks the question that follows: what happens then? The answer needs to exist before the failure, because the moment of failure is the worst possible time to design a process, negotiate authority, and find out whether the people who used to do the work manually still remember how.
For each high-stakes AI system, define a fallback. In practice the fallback is one of three things: a manual process in which humans review without AI support, a simpler rule-based system such as the predecessor logic the AI replaced, or a different model, either a secondary system or an earlier version that is known to be stable. Which one is right depends on the volume of decisions involved and how long you can tolerate reduced throughput. What matters more than the choice is that the fallback is written down, that someone can invoke it, and that invoking it does not require a decision from a person who may be unreachable.
The fallback also needs to be maintained. Organizations that retire their manual process or rule-based predecessor when they deploy AI frequently discover that they have no fallback at all when the AI system is taken offline, a condition worth naming as brittleness. The savings from decommissioning the old process are booked immediately; the cost appears later, as an outage with no floor under it. The test to apply before launching any AI system is simple: if this model were offline for two weeks starting tomorrow, what would we do? If nobody knows, the system is not ready for production, whatever its accuracy on the test set.
Monitoring as a Continuous Discipline
Post-deployment monitoring is where most AI risk management breaks down. Organizations invest heavily in pre-deployment testing, treat launch as the finish line, and then deprioritize monitoring once the system is live and the project team has moved on. Data drift accumulates gradually. Performance degrades in small increments that no single week makes obvious. Nobody notices until something goes wrong visibly enough to reach someone senior, at which point the question is not what happened last week but how long this has been happening, and the honest answer is usually that nobody can tell.
An effective monitoring regime has four components. Statistical monitoring of model inputs detects data drift before performance degrades, which is the only form of early warning available. Performance monitoring against business metrics, rather than statistical accuracy alone, catches the case where the model is still technically correct and no longer commercially useful. Periodic shadow testing, running the current model and a baseline in parallel and comparing their outputs, reveals degradation that neither system reports about itself. And a defined retraining trigger converts all of that observation into action, by specifying the threshold at which a retraining cycle begins without anyone needing to argue for it.
Monitoring also requires someone responsible for it, which means adding the monitoring role to the project's RACI matrix before deployment rather than assuming the platform team will absorb it. An alert that fires and has no designated recipient is equivalent to no alert. This is the same failure that leaves the risk register with empty owner fields, and it has the same remedy: name a person, confirm they have the access and the time, and make the monitoring output part of a routine that already exists.
Anti-Patterns
- Accepting assurances that the data exists. A delivery timeline committed to before anyone profiled the data is a commitment made without information.
- Treating test set accuracy as a business case. A statistically strong model can still produce more errors than the business can absorb at production volume.
- Planning the build and not the workflow. If no named person receives the output and has authority to act on it, the model can run flawlessly and change nothing.
- Leaving the owner field empty on post-launch risks. Unowned risks surface as incidents rather than as decisions.
- Retiring the predecessor process at launch. Decommissioning the system the AI replaced books an immediate saving and removes the only fallback the organization had.
Practice Prompts
- Write the risk register for a current AI project with the six standard fields, then mark every risk whose owner you cannot name as a person rather than a team.
- For one deployed model, profile the training data against the period it is now operating in and identify any temporal gap, excluded period or population that the model has never seen.
- Write the business-impact threshold for one model as a decision rule: at what error volume and error cost would the organization stop using it? Compare that against the statistical target being reported today.
- Map the decision workflow downstream of one AI output: the named recipient, the action they take, the tool it arrives in, the authority they hold, and whether that role is currently under review.
- Write the fallback plan for your highest-stakes AI system, then answer the two-week question honestly: if it were offline from tomorrow, what would happen, who would do it, and does that capability still exist?
Reflection
Nina's account is worth sitting with, because she was not careless and her risk register was not thin. The two risks that undid her project were invisible precisely because they sat between the categories her discipline had taught her to use: one was a property of the data rather than the plan, the other a property of the organization rather than the system. Neither belongs to the project manager, the data scientist or the sponsor in any obvious way, which is how they end up belonging to nobody. The uncomfortable version of the question is what your own organization would do the week an AI system it depends on started being quietly wrong. If noticing requires a person, and nobody has been named, then the monitoring, the fallback and the retraining trigger are all still theoretical.
Glossary
- Data drift: divergence between the patterns in the data a model was trained on and the data it is now deployed on, caused by the world changing rather than the model.
- Overfitting: a model that learned its training data too precisely and performs poorly on unseen data, diagnosable through the gap between training and held-out test performance.
- Adoption risk: the risk that a deployed model is ignored or worked around by its intended users, usually because no named person owns the decision it feeds.
- Fallback: the process that runs when the AI system cannot, whether a manual review, a rule-based predecessor or a secondary model.
- Shadow testing: running a current model and a baseline in parallel on the same inputs, to detect degradation neither reports about itself.
- Retraining trigger: a pre-agreed performance threshold that automatically initiates a retraining cycle, removing the need to argue for one after degradation is noticed.
Related Lessons
Several lessons extend this material. AI Risk Taxonomy & Assessment develops the categories used here into a classification suitable for a portfolio of systems rather than a single project. Model Performance Risk Management goes deeper into threshold-setting and degradation. AI Observability and Production Monitoring treats the monitoring regime described here as an engineering discipline in its own right. Incident Response for AI Security Breaches and Crisis Communication for AI Incidents cover what happens after a contingency plan is invoked. Ongoing Vendor Management & Governance addresses concentration risk, and Data Quality Assessment gives you the profiling methods that make the data availability question answerable before you commit.
Closing
The lesson from Nina's project is not that AI projects need more risk management than other projects, but that they need it aimed at different things. The categories that matter sit outside the build: whether the data covers the world the model will operate in, whether a named person will act on the output, whether a regulator's view of the system might change, and whether anyone is watching once the launch is over. Each becomes tractable the moment it has an owner and a threshold. Left unowned, they are not risks the organization has accepted; they are incidents it has not had yet.
Key Takeaways
- AI projects add failure categories beyond standard project risks. Data drift, overfitting, adoption gaps and AI-specific regulatory exposure require dedicated assessment in the project risk register rather than being folded into generic technology risk.
- Assess actual data before committing to timelines. Assurances that the data exists are not equivalent to profiling it for availability, completeness and quality, and that assessment has to happen before scoping rather than during build.
- Map the decision workflow before building the model. Identify the specific human who receives the output and acts on it, confirm their role will be stable, and confirm the workflow fits the process they already run.
- Define business-impact thresholds, not just statistical accuracy targets. A model that is 92% accurate may be operationally unacceptable if volume is high and the consequences of each error are significant.
- Unowned risks become incidents. Assign named owners for model monitoring, retraining decisions, incident response and regulatory currency before deployment, not after something has gone wrong.
- Every AI system needs a maintained fallback. A manual process or rule-based predecessor that is retired when the AI launches leaves the organization with no option when the AI goes offline.
- Post-deployment monitoring requires an assigned owner and a defined retraining trigger. Monitoring infrastructure without a responsible person and a threshold that triggers action is not an operational control.
Frequently Asked Questions
How do we assess model quality if nobody on our side is technical? Two questions carry most of the weight: what cross-validation methodology was used, and what the gap is between training performance and held-out test performance. A large gap signals overfit. You do not need to evaluate the methodology yourself; you need to see whether the team can produce the numbers at all, because a team that cannot has not measured the system you are being asked to rely on.
The model is accurate and nobody uses it. What went wrong? Almost certainly the decision workflow, which is the failure Nina's project reached. Accuracy determines whether the output is worth acting on; it says nothing about whether a named person receives it, in a tool they use, with the authority to act. Reorganizations are the usual culprit, because they dissolve the accountable role without anybody updating the project plan. Treat the recipient's role as a project dependency.
Is a fallback really necessary if the system has high availability? Availability is only one of the ways an AI system stops being usable. It can also drift below an acceptable performance threshold, encounter conditions it was never trained for, or become non-compliant when regulation changes, and none of those show up as downtime. An organization that has decommissioned the process the AI replaced has no answer to the two-week question, which is the definition of brittleness.
We have monitoring dashboards already. Is that sufficient? Only if three things are true: someone is named as responsible for looking at them, a defined threshold triggers retraining rather than a discussion, and the monitoring covers business metrics and input drift rather than statistical accuracy alone. Dashboards without a named recipient are the same defect as an alert with no addressee, and they fail the same quiet way.
Skill.re