Algorithmic Impact Assessments
Tomas Rivera, an AI program lead at a state child-welfare agency, was handed a 60-page document by a vendor and told it was the system's "algorithmic impact assessment." His agency was about to deploy a tool that would help screen which abuse hotline calls got investigated. Tomas had asked for the assessment as a condition of the contract. Now he held it, and he realized he had no idea how to tell a real assessment from an elaborate sales brochure. The document was full of confident language. It was also missing the one thing he most needed: any honest account of how the system could be wrong, and who would be hurt when it was.
An algorithmic impact assessment, or AIA, is a structured study of what an AI system might do to the people it touches before you turn it on. It is the homework a vendor or program owner does to surface harms, bias, privacy risks, and failure modes while they are still cheap to fix. This lesson covers when to require one, what a complete one contains, how to read what comes back, and the red flags and green flags that separate a genuine assessment from a glossy evasion.
What an Impact Assessment Is Actually For
An AIA is the AI equivalent of an environmental impact statement. Before a highway is built, the law forces planners to study what it will do to the water, the air, and the neighborhood. An AIA forces the same discipline for an algorithm: before it decides who gets investigated or denied, study what it will do to accuracy, fairness, privacy, and the people on the wrong end of a mistake.
The purpose is not paperwork. It is to move the discovery of harm from after deployment, where it shows up as a lawsuit or a news story, to before deployment, where it shows up as a fixable design note. Government AI systems reach into lives directly. A benefits determination system might deny someone critical assistance. A hiring algorithm might discriminate. A security screening system might flag innocent people. Without assessments, agencies find these problems reactively, when residents complain, when a civil rights audit reveals discrimination, or when Congress investigates. With them, agencies find the same problems early enough to change something.
Be clear about what the instrument is and is not. An AIA is not perfect prediction. It is structured thinking about what could go wrong and how to mitigate it, written down by people who were obliged to look. That is genuinely valuable and it is also bounded: an assessment documents the harms you thought to examine, using the data you had, on the populations you chose to compare. What it does not do is certify that no other harm exists. For Tomas, the immediate question is narrower still, which is whether the document in his hands did any of that work at all, or simply claimed to.
When an Assessment Is Required
OMB Memorandum M-24-10, issued in 2024, requires agencies to conduct algorithmic impact assessments for high-risk AI systems. That is a policy mandate rather than a matter of local preference. The stronger practice, and the one this curriculum teaches, is to conduct an assessment for any system that makes or recommends decisions affecting the public, whether or not it lands inside a formal high-risk designation. The cost of assessing a system that turned out to be low-risk is a few weeks of documentation. The cost of skipping one that turned out not to be is measured in denied applications.
Beyond the compliance question, assessments do governance work that nothing else does. They document what you know and what you do not know about a system's risks. They establish the baseline against which later fairness and bias monitoring is measured, so that a disparity discovered in year two can be compared to something. They identify where human oversight is essential rather than decorative. And they create accountability: when something goes wrong, you will be asked whether you conducted an assessment, what it revealed, and how you addressed what it found.
That accountability runs in one direction only, and it is worth being precise about it. An agency that skipped the assessment has no answer at all. An agency that conducted a rigorous one has a record of what it examined, what it found, and what it did in response. That record is what makes the decision reviewable. It is not a clearance, and it does not convert a harm the assessment failed to look for into a harm nobody could have anticipated. The assessment shows your work; it does not grade it.
Requiring Assessments from Vendors: Build It into the Contract
The time to require an assessment is before you sign, not after you deploy. If you ask afterward, the vendor has every incentive to produce a document that justifies what they already built. Put the requirement in the solicitation and the contract, and specify what "complete" means so you are not handed a brochure. A usable contract clause covers six things:
- Scope of decisions. The assessment must describe every decision or recommendation the system makes that affects a person.
- Intended and unintended uses. What the system is for, and what it must never be used for.
- Data provenance. What data trained the model, where it came from, and whether it was lawfully obtained.
- Fairness testing. Measured performance across relevant groups, with the actual numbers, not a promise that testing occurred.
- Failure analysis. What the system gets wrong, how often, and who bears the consequences of each error type.
- Mitigations and monitoring. What safeguards are in place and how the system will be watched after launch.
Notice that this aligns with what the major frameworks already ask. The federal guidance on rights-impacting AI directs agencies to assess impacts before deploying systems that affect people's rights, and the NIST AI Risk Management Framework, a voluntary standard, builds its "Map" and "Measure" functions around essentially the same discipline. You are not inventing a new burden. You are operationalizing an expectation that already exists, at the one moment when you have the leverage to make a vendor meet it.
What a Complete Assessment Contains
M-24-10 sets out the required elements. An assessment covers the purpose and function of the AI system; its intended uses and reasonably foreseeable misuses; the categories of individuals who could be affected; potential impacts on civil rights and civil liberties; potential impacts on privacy; the steps taken to mitigate identified risks; and the human review and override mechanisms in place.
Good practice extends that list. Add data quality and provenance, covering what the system was trained on, how recent that data is, and what biases it carries. Add fairness and bias analysis showing whether the system performs differently across demographic groups. Add explainability and transparency, meaning whether you can explain decisions and whether affected people can understand why they were affected. Add security considerations, including vulnerability to adversarial attack, model theft and data poisoning. Add monitoring and ongoing assessment describing how you will detect problems after deployment. And add stakeholder engagement, recording who was consulted and, just as importantly, who is missing.
Component by Component: Weak Versions and Strong Ones
The difference between a serious assessment and a compliant-looking one shows up section by section. In each case the weak version is not false; it is simply unfalsifiable, which is what makes it useless.
Purpose and function. This should be an operational description, not marketing language. The weak version reads "AI system for benefits determination." The strong version reads: a machine learning model classifies applications for unemployment benefits eligibility, outputs a recommendation of eligible, ineligible or escalate to human review, and a human caseworker makes the final decision and can override the model recommendation. The second version tells you exactly where the machine stops and a person starts.
Intended uses and foreseeable misuses. Intended uses state what you are trying to accomplish, such as screening applications efficiently and reducing processing time. Foreseeable misuses state how the system could be used in ways you do not intend: using its output as the basis for denying a citizenship review, automating the decision without human oversight, or repurposing an employment tool for credit decisions or immigration matters. Writing the misuses down is what lets you prohibit them in the contract and detect them in an audit.
Affected populations. Not "the general public." Be specific: unemployed workers applying for benefits, and disproportionately workers from particular industries, geographic regions and demographic groups. Include the secondary population too, meaning caseworkers whose workload changes and supervisors managing reviewers, because a system that quietly shifts work onto three overloaded people fails in a way no fairness metric will catch.
Civil rights and civil liberties impacts. Four questions carry most of the weight. Discrimination risk: might the system treat different demographic groups differently? Due process: are applicants notified of decisions, and can they appeal? Transparency: do people understand why they were accepted or denied? Privacy: what data is collected and stored, and how is it protected?
Human oversight mechanisms. The assessment should say exactly where a human reviews or overrides. Strong versions are specific and statistically targeted rather than universal: every decision above a defined certainty threshold reviewed by a human rather than all decisions reviewed by nobody in particular; automatic escalation to a supervisor when fairness metrics degrade; a route for an affected person to request human review of a decision they disagree with; and an audit trail of overrides recording how often humans override and in which kinds of cases. That last item matters more than it looks. An override rate near zero usually means the review has become a formality rather than a safeguard.
The Fairness Section: Disaggregate or It Says Nothing
The single most common failure in an impact assessment is a headline accuracy figure with nothing behind it. "System achieves 94% accuracy" is compatible with the system working beautifully for most people and badly for one group, and a reader cannot tell which. The strong version breaks performance out by group and publishes the gaps.
| Breakdown | Reported accuracy |
|---|---|
| Age 18-25 | 96% |
| Age 25-40 | 95% |
| Age 40+ | 91% |
| Women | 94% |
| Men | 95% |
| White | 95% |
| Latino | 91% |
| Black | 89% |
| Asian | 96% |
Read that table the way a reviewer should. Taking the figures as reported, the spread across the racial breakdown runs from 89% to 96%, and across the age bands from 91% to 96%, and the assessment says so in public. That disclosure is what makes a serious assessment useful, and it is exactly what the weak version conceals behind a single blended figure. Note also that the age bands as reported overlap at 25, which is the kind of definitional sloppiness worth sending back for clarification rather than silently correcting yourself.
Fairness thresholds work the same way. A strong assessment states in advance what performance variance across groups the agency considers acceptable, records that decision, and names the trigger that fires when the variance is exceeded. The threshold is a commitment your agency makes to itself, written before the results arrive so it cannot be adjusted to fit them. It is not a legal test, and it does not establish that a variance below the line is lawful or that one above it is unlawful. Those are questions for counsel on the specific facts. What the recorded threshold gives you is a decision made in advance rather than under pressure.
The Monitoring Plan
An assessment is a point-in-time document; monitoring is what carries its findings forward. Without a monitoring plan, the risks the assessment identified simply go unwatched, and a system can operate for months before anyone notices its fairness metrics have degraded. A serious plan names specific metrics, a review cadence, alert thresholds, and the people responsible.
A workable pattern combines several rhythms: a monthly accuracy audit on a sample of cases; a quarterly fairness audit with disaggregated metrics; a real-time alert when fairness metrics degrade beyond the recorded threshold; an annual third-party fairness audit; continuous tracking and analysis of complaints from the public; and a performance degradation alert if accuracy drops more than 2% from the established baseline. The point of the mixture is that different failures move at different speeds. Drift is slow and shows up in the quarterly comparison; a bad model update is fast and needs the alert.
Evaluating a Submitted Assessment
When the assessment arrives, read it against one organizing question: does this document help me understand how the system could harm someone, or does it work to reassure me that it cannot? A real assessment is honest about limits. A fake one is relentlessly confident. A real impact assessment tells you how the system can fail; a sales brochure tells you it cannot.
Work through the submission section by section and ask:
- Does it name the affected populations specifically, including the most vulnerable, or speak only in generalities?
- Does it report fairness results with actual measured numbers broken out by group, or just assert that the system is "tested and fair"?
- Does it describe concrete failure modes and who they hurt, or skip straight to benefits?
- Does it explain how an affected person learns about and challenges a decision?
- Does it say what data trained the model and acknowledge gaps in that data?
- Does it commit to ongoing monitoring with named owners and thresholds, or treat launch as the finish line?
- Who wrote it? An assessment conducted only by the people who built the system carries a conflict of interest that no amount of care inside the document resolves.
Four qualities separate a thorough submission from a cursory one, and they are worth scoring separately. Specificity: is it specific or vague? Completeness: thorough or perfunctory? Credibility: does it seem honest about risks, or dismissive of them? Actionability: are the findings translated into concrete decisions to deploy, modify, or not deploy? A document can be long, specific and complete while ending in no decision at all, and that is its own kind of failure.
Red Flags and Green Flags
Tomas eventually built a simple two-column instinct for triaging any assessment. Use it as your quick screen before the deep read.
Red flags: send it back
- No fairness numbers, only the word "fair," or "bias testing complete; system is fair" with no specifics. Confidence without measurement.
- No failure analysis. A system that "doesn't make mistakes" has simply not been studied. Related: "system is always right; minimal override needed."
- Accuracy reported only as a single overall figure, hiding how it performs for specific groups, or no demographic analysis at all.
- "We will evaluate fairness after deployment." Fairness evaluation belongs before launch, or at the very latest alongside it.
- "System is fully explainable" asserted without evidence. This is rarely true, and a document that claims it casually has not tested it.
- Vague data sourcing, or refusal to disclose what the model was trained on.
- No mechanism for a person to understand or appeal a decision.
- Monitoring described as "as needed," or "we will monitor and adjust," with no owner, metric or trigger.
- No stakeholder input, with civil rights offices and affected community groups never consulted.
- The assessment was conducted only by the system's developers.
- The document reads like marketing: heavy on benefits, silent on limits.
Green flags: a serious assessment
- Specific, quantified group-level performance numbers, including where the system performs worse and why.
- A frank failure section naming error types and the people each one harms, with the mitigation attached to each.
- Specific, named affected populations and the most exposed among them, backed by documented examples rather than categories.
- A clear description of when humans override, how escalation works, and how often overrides actually happen.
- A clear, plain-language explanation and appeal path for affected individuals.
- Honest discussion of data limitations and what they mean for whom.
- Quantified fairness thresholds recorded in advance, with the acceptable variance stated rather than implied.
- A monitoring plan with named owners, specific metrics, a review cadence, and thresholds that trigger review.
- Third-party audit of the assessment itself, or at minimum a civil rights office review.
- Documented stakeholder input from the civil rights office and affected communities.
- A re-assessment schedule, annually or whenever significant changes occur.
- Budget allocated for monitoring and remediation, which is the difference between a plan and an intention.
- Recommended conditions or restrictions on use, evidence that the author took the risks seriously enough to constrain their own product.
Two Worked Cases
Screening job applications
A federal agency is considering AI to screen job applications. The fairness analysis starts from the historical record: 22% of applications from women were accepted, against 28% from men. The question the assessment must answer is whether the AI system will perpetuate that disparity or correct it. Tested on historical data, the model accepted 20% of women and 26% of men. Read the arithmetic carefully, because the source narrative calls this a slight improvement and the numbers do not support that reading: the gap was 6 percentage points before and is 6 percentage points after. Both rates fell; the distance between them did not close. The honest finding is that the disparity survived the model, and an assessment that describes it as improvement is doing the reader a disservice.
The resulting recommendation was appropriately narrow. The system should not be used for final hiring decisions, only as a first-pass screen feeding human reviewers who are trained to account for historical bias. Human oversight was specified concretely: the system outputs qualified, borderline or unqualified; 100% of borderline cases are reviewed by a human; 10% of unqualified cases are reviewed to catch system errors; 5% of qualified cases are reviewed for quality assurance; and the agency tracks how often humans override and in what kinds of case. Monitoring runs monthly on acceptance rates by demographic group with an alert if disparity emerges, quarterly as a fairness audit compared against pre-AI hiring patterns, and annually as a third-party audit of both the system and the decision process. The escalation rule is the sharpest part: if a demographic disparity emerges, use of the system pauses until the root cause is understood.
Flagging suspicious benefit claims
An AI system flags suspicious government benefit claims. Here the assessment turns on failure mode analysis, because the two failure directions harm different people. A false positive means an honest person is wrongly investigated. A false negative means fraud goes uncaught. Which is worse is a mission question, not a technical one: a fraud detection posture tolerates more false positives to catch more fraud, while a fairness posture tolerates more false negatives to avoid wrongly accusing people. The assessment has to state which posture the program has chosen and why, because the choice determines who absorbs the errors.
The assessment then fixes the operating limits in advance. It sets a false positive rate ceiling of 5%, above which the system is judged too aggressive, and raises an alert if the false negative rate exceeds 15%, which indicates the system is missing fraud. Alongside those, it commits to a monthly analysis of the demographic distribution of flagged cases, asking directly whether some groups are over-represented among the flags. Set those numbers before the system runs, record them, and treat a breach as a trigger rather than a topic for negotiation. The source text for the 5% figure is internally muddled, describing it both as an acceptable rate and as an indicator that the system is too aggressive, so read it as the ceiling and confirm the intended reading with whoever owns the program before you rely on it.
How Tomas Used the Screen
Tomas ran the vendor's 60 pages through the red-flag list and got his answer fast. The document reported one overall accuracy figure of 89% and never broke it out by race, age, or neighborhood. It had no failure section. Its data sourcing was a single sentence. By the green-flag list, it scored almost nothing.
He sent it back with a specific request: group-level performance numbers, a failure analysis naming who gets hurt by false positives and false negatives, and a monitoring plan with a named owner. The revised assessment told a different and more useful story. The system was strong overall but flagged calls from lower-income areas for investigation at a noticeably higher rate, and the vendor recommended a human review step for those cases before deployment. That single disclosure, absent from the original glossy version, changed the agency's launch plan and likely kept families from being investigated on a machine's hunch. The assessment did its job only once Tomas refused to accept the brochure.
Anti-Patterns
- Assessing after the system is built. Governance gets deferred while the build proceeds, and the assessment surfaces fairness problems at the point where changing the system is expensive and slow. Conduct an initial assessment during design, refine it as the system is developed, and produce the final version before deployment. The early draft is the cheap one.
- Letting the builders assess their own work. Internal teams are convenient and external reviewers cost money, but developers have every incentive to downplay risk and cannot see their own blind spots. The assessment misses obvious biases because nobody looked for them, and the problems surface after launch. External review is essential: a civil rights office review at minimum, a third-party bias audit where the stakes justify it.
- Reporting aggregate accuracy only. A single 94% figure hides the fact that one group sees 91% and another 97%. Aggregate metrics are simpler and disaggregated analysis takes real work, so the shortcut is tempting. The system deploys, and a civil rights audit finds the disparity you should have found first. Require disaggregated metrics for every relevant group.
- Treating the assessment as the finish line. The document feels complete, so the risks it identified go unmonitored and the system runs for months before anyone notices degradation. Every assessment must include a monitoring plan with specific metrics, a review cadence, and alert thresholds, plus the budget to run it.
- Skipping the people the system acts on. The AI team writes the assessment internally, the civil rights office later identifies a discrimination risk nobody saw, and affected communities are rightly angry they were never asked. Require consultation with the civil rights office, affected community groups and internal stakeholders before the assessment is final, and record who was consulted and who was not.
- Treating a completed assessment as a clearance. A finished AIA proves that someone examined a defined set of risks, on the data available, for the groups they chose to compare. It does not certify that the system is safe or fair, and it cannot speak to the harm nobody thought to look for. An agency that answers a complaint by pointing at the existence of an assessment has misunderstood what the assessment is for.
- Producing findings that decide nothing. A long, specific, well-sourced assessment that ends without a recommendation to deploy, modify or stop has converted a governance instrument into a filing. Every assessment should terminate in a decision and the conditions attached to it.
Practice Prompts
- Scope three assessments. You are about to require impact assessments for three proposed AI systems at your agency. For each, define who should conduct the assessment, what the key impacts to assess are, which demographic groups are most at risk, and what the monitoring plan looks like.
- Grade a real one. Read an existing assessment, real or fictional. Score each section on specificity, completeness, credibility and actionability. Then write the one-paragraph note you would send back, naming the specific missing pieces rather than asking for "more detail."
- Design the fairness component. For a system you know, design the fairness analysis section: which demographic groups matter, which metrics you would measure, what performance gaps your agency would record as acceptable in advance, and how you would detect problems after deployment.
- Design the monitoring plan. Specify what metrics you would track, how frequently, what alert thresholds would trigger escalation, who reviews the monitoring data, and what happens when a threshold is exceeded. Include who pays for it.
- Plan the consultation. Decide who should be consulted internally and externally, how you would gather their input, how you would address the concerns raised, and how the consultation would be documented in the final assessment.
Reflection
Take an AI system your agency is acquiring or building. What would the critical elements of its impact assessment be, and who should be involved in assessing impacts? Which fairness metrics matter most for this particular system, and who is qualified to interpret them? How would you monitor it after deployment, and what specifically would be your threshold for pausing the system if problems emerge?
Then ask the question that assessments are worst at. What harm would this system cause that your assessment plan would not detect, because you were not measuring for it or the affected group is not in your data? Write it down. That sentence is the most honest paragraph in any impact assessment, and it is the one most likely to be missing.
Glossary
- Algorithmic impact assessment (AIA). A systematic evaluation of an AI system's risks before deployment, identifying potential harms and mitigation strategies.
- Disaggregated metrics. Performance reported separately for different demographic groups. Essential for fairness assessment, because a blended figure can hide a large group-level gap.
- Fairness. The principle that AI systems should not systematically disadvantage protected groups. Multiple definitions exist, and the one in use must be defined explicitly for each system.
- Foreseeable misuse. A way the system could be used contrary to its intended purpose. Worth identifying precisely so it can be prohibited and detected.
- Human override. The ability for a person to make a decision different from the AI system's recommendation. Essential for high-stakes decisions, and only meaningful when it is genuinely available and its use is tracked.
- Monitoring plan. The ongoing process for detecting whether system performance degrades after deployment, including metrics, review cadence and alert thresholds.
- Rights-impacting system. A system whose decisions affect a person's legal rights, benefits, access or opportunities, and which therefore carries heightened assessment and oversight obligations.
Related Lessons
- AI Vendor Evaluation Methodology covers scoring the vendor whose assessment you are about to demand as a contract condition.
- Writing AI Requirements in RFPs and SOWs is where the six-part assessment clause becomes solicitation language.
- Federal Acquisition of AI: FAR/DFARS covers the procurement framework the contract clause has to sit inside.
- Privacy Impact Assessments for AI Systems handles the privacy analysis that runs alongside the algorithmic one.
- OMB M-24-10 Deep Dive: Full Implementation works through the memorandum's requirements in full.
Closing
An algorithmic impact assessment is where governance thinking becomes concrete. It converts "we should think about fairness" into disaggregated accuracy by named group, a variance the agency recorded as acceptable in advance, and a monitoring cadence with an owner attached. Conduct assessments early, during design. Involve the people the system will act on. Be honest about limitations. Make sure monitoring is funded and staffed before launch rather than promised in a paragraph.
A thorough assessment documents what you know, what you do not know, what could go wrong, and how you will respond. That record is how an agency demonstrates responsible AI governance and how a resident harmed by a decision can find out what the agency understood at the time. Require rigorous assessments, let their findings actually decide something, and monitor afterward to check whether the assumptions held. Tomas got a useful document only because he sent the first one back.
Key Takeaways
- An assessment moves harm discovery earlier. Like an environmental impact statement, it surfaces bias, privacy, and failure risks before deployment, when they are still cheap design notes rather than lawsuits.
- Assessments are required for high-risk systems. OMB Memorandum M-24-10, issued in 2024, mandates them for high-risk AI. The stronger practice is to assess any system that makes or recommends decisions affecting the public.
- Require it before you sign. Build the assessment into the solicitation and contract; asking after deployment only invites a document that justifies what was already built.
- Specify what "complete" means. A usable clause demands scope of decisions, intended uses and foreseeable misuses, data provenance, fairness numbers, failure analysis, and a funded monitoring plan.
- Read against one question. Does the document help you see how the system can harm someone, or work to reassure you that it cannot? Honesty about limits is the tell.
- Fairness analysis must be disaggregated. Aggregate accuracy hides disparities. Group-level numbers, published with the gaps visible, are the difference between an assessment and an assertion.
- External review prevents blind spots. Developers assessing their own system carry a conflict of interest. A civil rights office review is the minimum; a third-party audit is better.
- Human oversight must be specific and tracked. Say which cases a person reviews, how escalation works, and how often overrides actually occur. An override rate near zero means the review has become a formality.
- Monitoring plans carry the findings forward. An assessment is point-in-time; without named metrics, a cadence, alert thresholds, owners and budget, the risks it identified simply go unwatched.
- A completed assessment is a record, not a clearance. It shows what you examined, on the data you had, for the groups you compared. It cannot speak to the harm nobody thought to look for.
- Send weak assessments back. A specific request for the missing pieces often surfaces exactly the disclosure that changes your launch plan and protects the public.
Frequently Asked Questions
Who should actually write the assessment, the vendor or us?
The vendor produces the technical content because they alone hold the training data, the model and the test results, and your contract should require it. Your agency owns the assessment as a governance artifact: you define the scope, you review it, and you decide whether it is complete. Add an independent reviewer, at minimum your civil rights office, because a document written entirely by the people who built and sold the system carries a conflict of interest that no amount of internal care removes.
Our system is low-risk. Do we still need one?
The formal mandate attaches to high-risk systems, so the compliance answer depends on how your system is classified. The practical answer is that risk classification is itself a judgment made before you have studied the system, and agencies are routinely surprised by which tools turn out to touch people's rights. A short assessment for a genuinely low-risk tool costs little and gives you a documented basis for the low-risk classification, which is exactly what an auditor will ask you to produce.
What do we do when the vendor says fairness testing is impossible without demographic data we do not hold?
Treat it as a real constraint and a documented one rather than a reason to skip the section. Write down which groups you could not compare and why, because that limitation is a finding in its own right and it belongs in the assessment. Then look for what is available: geographic proxies, program-level records, or testing on your own historical case data during the pilot. What you must not do is let the absence of data become a silent gap that reads, later, as though fairness was checked.
How specific should our fairness thresholds be?
Specific enough to trigger something automatically, and recorded before the results come in. State the variance across groups your agency will treat as acceptable, name the metric it applies to, and define what happens when it is breached, whether that is escalation, a pause, or a root-cause investigation. Understand what the threshold is: a commitment your agency makes to itself, not a legal test. Whether a given disparity is lawful is a question for counsel on the specific facts, and a number you set internally neither creates nor discharges that question.
The assessment came back and it is genuinely good. Are we clear to deploy?
You are clear to make a decision, which is different. A strong assessment tells you what the system does to the groups it examined, on the data it had, and what it recommends you constrain. Deploy against those constraints, fund the monitoring, and keep the re-assessment schedule, because the finding that matters most is usually the one that arrives in month seven. Treating a good assessment as the end of the obligation is the single most common way a well-assessed system still goes wrong.
How often should an assessment be redone?
Annually as a floor, and immediately on any significant change: a model update, a change in the population served, a new use for the system, or a monitoring alert that fires. The trigger list matters more than the calendar. A system that was assessed twelve months ago and has been retrained twice since then has an assessment describing something that no longer exists.
Skill.re