←
AI for Government
Strategic · M3 · lesson 3 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Assurance Programs
📖
now learning

AI Assurance Programs

15 min

Celestine Broadwater had the audit report memorized by the time the Government Accountability Office published it. Her agency, the Social Security Administration's disability processing division, had deployed a document-classification AI eighteen months earlier. The system was flagging medical records for manual review. Nobody had checked whether it was flagging the right ones. The GAO finding was blunt: "The agency lacks a continuous monitoring program capable of detecting model drift or systematic error in production." Celestine had treated the pre-deployment audit as the finish line. The GAO made clear it was only the starting line.

Audit Versus Assurance

Think of an audit like a home inspection. A home inspector walks through once, notes what is broken or out of code, and hands you a report. That report reflects the condition of the house on that single day, judged against the code the inspector was applying, in the rooms the inspector actually entered. If a pipe bursts the following week, the inspector is not responsible. They have already left the building, and nothing in their report ever claimed to cover next week.

An assurance program is more like a building management system with sensors, alarms, and scheduled maintenance cycles. It provides ongoing evidence that the house remains safe to occupy. When something changes, a new load on the structure, a temperature spike, unexpected moisture, the system detects it and triggers a response before damage accumulates. The value is not in any single reading. It is in the fact that somebody is still reading, on a schedule, against thresholds agreed in advance by people who understood what a breach would mean.

For government AI, the distinction is consequential. A pre-deployment audit tells you whether the system met standards on the day it launched. An assurance program tells you whether it continues to meet them as data shifts, usage patterns evolve, staff turnover changes how humans interact with it, and policy changes alter what the system should be doing. Celestine's classifier had not broken on launch day. It had been quietly drifting for a year and a half while everyone pointed at an approval memo dated eighteen months earlier.

What Assurance Can and Cannot Establish

Set the expectation before you build the program, because the language of assurance invites overclaiming. An assurance program produces evidence about the indicators you chose to watch, on the population that used the system during the period observed, measured against the thresholds you set. It does not establish that the system is safe, lawful, or fit for its purpose. Those are conclusions somebody has to reach using the evidence, and they remain conclusions about the sampled behavior, not guarantees about behavior nobody sampled.

The same limit applies to certification and third-party attestation, which agencies increasingly cite as proof that an AI system can be trusted. A management-system certification attests to a conforming management system, sampled across a declared scope at a point in time. It is not proof that any given AI system is safe, lawful, or fit for the purpose your program is putting it to. Read the scope statement on any certificate before you rely on it; the scope is frequently narrower than the reassurance the certificate is being used to provide.

This is not an argument for skipping assurance. It is an argument for stating what your evidence covers whenever you pass it upward. "The classifier met its accuracy and parity thresholds on every monthly sample in the period reported" is a defensible sentence. "The classifier is validated" is not, because it hides the sample, the period, and the thresholds inside a single word. Under oversight, the second sentence collapses into the first anyway, and it collapses in public rather than in a working meeting.

Three Levels of Assurance

Federal audit and assurance practice recognizes three levels of assurance, adapted here for AI systems. The point of the ladder is to spend proportionate effort: it is entirely rational to accept thin evidence about a scheduling tool and unreasonable to accept the same evidence about a system that freezes bank accounts. Choose the level from the consequence of being wrong, not from the sophistication of the technology.

Limited assurance

Limited assurance is based on spot checks: periodic, selective tests that sample system behavior without covering every scenario. Think of it as the inspector returning every six months to check a few rooms. It is appropriate for low-risk AI systems, meaning internal tools, document drafting aids, and scheduling optimizers that do not affect citizen-facing decisions. The evidence standard is lower, the monitoring frequency is lower, and the required documentation is lighter, and all three of those should be written down so nobody later mistakes light coverage for a clean bill of health.

A typical limited assurance program for a low-risk system involves quarterly accuracy checks on a random sample of 200 to 500 outputs, an annual policy alignment review, and logging of any staff complaints or anomalies. Total staff cost runs roughly 0.25 to 0.5 full-time equivalents per system per year. The complaint log is the part that gets skipped and the part that most often catches the first sign of trouble, because staff notice a tool going wrong long before a quarterly sample does.

Reasonable assurance

Reasonable assurance is based on systematic, scheduled testing across defined risk scenarios. It is the standard applied to most citizen-facing AI systems in benefits, permitting, and eligibility programs. "Reasonable" in audit language means that a qualified professional reviewing the evidence would conclude the system is performing as intended with high confidence, though not with certainty. The words "though not with certainty" are load-bearing and belong in your reporting, not just in the definition.

A reasonable assurance program includes monthly performance reporting against defined thresholds, quarterly structured testing by staff independent of the system's operators, semi-annual equity audits checking for disparate performance across demographic groups, and annual third-party review. For a mid-size benefits AI touching 500,000 claimants per year, this typically costs $150,000 to $300,000 annually in staff time and contractor reviews. Independence in the quarterly testing is the requirement agencies quietly drop first, and dropping it converts the test into a self-report.

The annual third-party review deserves its own governance. Decide in advance what the reviewer is being engaged to examine, what evidence they will be given access to, who inside the agency receives their findings, and how disagreements between the reviewer and the program are resolved and recorded. A review whose scope the reviewed party defined, and whose findings route back only to that party, produces a document rather than an assurance. Where the reviewer is another agency or an academic partner rather than a contractor, settle the same questions plus publication rights before the work begins, not after an uncomfortable finding appears.

Continuous real-time monitoring

The highest level is often labelled absolute assurance, which practitioners more accurately call continuous real-time monitoring. It is reserved for AI systems where errors have immediate, high-stakes consequences: fraud-detection systems that freeze accounts, clinical-decision support tools in government health programs, or systems that trigger law enforcement action. "Absolute" is a shorthand and a misleading one. No assurance is truly absolute. What this level buys is near-real-time instrumentation of the system's outputs and automated alerting when performance crosses defined thresholds.

Continuous monitoring at this level requires technical instrumentation built into the system from the start, and you cannot retrofit it cheaply. The monitoring stack typically includes automated accuracy metrics computed on every output batch, statistical process control charts that detect drift before it becomes a failure, and paging alerts to a designated on-call team when thresholds are breached. Annual cost for a high-risk system runs $500,000 to $1.2 million, including the instrumentation infrastructure. That figure is why the level has to be chosen during design rather than after an incident.

Building the Monitoring Framework

Every assurance program starts with a measurement plan. After the GAO finding, Celestine's team could not answer a basic question: what should the document-classification AI's accuracy rate be, and how would they know if it fell below that? They had no baseline. Without a baseline, every subsequent number is unreadable, because any given accuracy figure is excellent news or a crisis depending entirely on what the system did when it was approved, and nobody had written that down in a form anyone could retrieve.

A measurement plan specifies the key performance indicators for the system, the baseline values established during pre-deployment testing, the acceptable operating thresholds, and the alert thresholds that trigger escalation. For a classification system, indicators typically include overall accuracy, false-positive rate meaning how often the system incorrectly flags something, false-negative rate meaning how often it misses something it should flag, and demographic parity metrics showing whether accuracy is consistent across racial, gender, age, and language groups.

Each of those indicators needs its own threshold, because a single composite score hides exactly the failure you are watching for. Add an indicator for drift itself: the GAO finding against Celestine's agency named drift detection specifically, and a metric that reports only this month's accuracy will never show you a slope. Record alongside each threshold who agreed to it and on what reasoning, so that a successor inheriting the plan can tell a deliberate tolerance from a placeholder nobody ever revisited.

Setting Thresholds With Program Staff

Engineers can tell you what the system's accuracy rate is. Program staff, the people who know what happens when the system makes a mistake, can tell you what accuracy rate is acceptable. These are different conversations and they need different people in the room. Celestine learned that a two-percentage-point drop in the classifier's recall rate, meaning the percentage of relevant documents it correctly identifies, translated to roughly 3,000 additional manual reviews per month for frontline staff. That was a workload spike nobody had budgeted for and nobody had connected to a metric.

Threshold-setting therefore requires a joint session with engineering, program operations, and the equity officer. Set separate thresholds for overall performance and for performance on demographic subgroups, and record who agreed to each one. A system that performs at 94 percent accuracy overall but drops to 78 percent for non-English speakers does not meet a reasonable assurance standard, even though the headline number looks acceptable and will be the number that reaches leadership if you let a single figure carry the report.

Managing Evidence

An assurance program generates evidence: test logs, accuracy reports, equity audits, remediation records. That evidence serves three purposes at once. It tells you whether the system is working. It satisfies inspector general and GAO reviewers who will ask for it on their own schedule rather than yours. And it creates an accountability record if something goes wrong, which is the purpose nobody values until the week they need it and discover the logs rotated out of retention long ago.

Store evidence in a structured, version-controlled repository accessible to the inspector general's office on request. Under the Freedom of Information Act, which gives the public the right to request government records, summary monitoring reports may be releasable. Design your evidence management so that operationally sensitive details, such as the exact detection thresholds a bad actor could game, can be withheld without losing the public accountability record. That separation is a design decision made once, not a redaction argument conducted repeatedly under a statutory clock.

Use a simple readiness test on your own program. If the GAO asked you today to show continuous monitoring evidence for your three highest-risk AI systems, how long would it take you to produce it? If the answer is more than two business days, your evidence management is a finding waiting to happen. The exercise is worth running as a drill, because the gap it exposes is almost never the monitoring itself. It is that the monitoring output lives in four systems and one person's spreadsheet.

When Things Go Wrong

Thresholds are only useful if they trigger real responses. Define your escalation chain before you need it: who receives the alert, who decides whether to pause the system, who notifies affected citizens, and who briefs the agency head. A well-run assurance program treats a threshold breach the way an air traffic control system treats a proximity alert, as an action signal requiring an immediate decision, rather than as information to be evaluated at the next scheduled review.

Celestine's agency now has a documented AI incident response protocol that mirrors its IT incident response process. A threshold breach at the reasonable-assurance level triggers a 24-hour review window. If the problem cannot be diagnosed and mitigated within that window, the system is paused and manual processing resumes. This is a deliberate, pre-authorized decision made at the program design stage, not a panicked improvisation in the middle of a crisis, and the difference shows in how quickly anyone is willing to make it.

Sustaining the Program

Assurance programs decay quietly. The monitoring stays switched on, the reports keep generating, and gradually nobody reads them, the thresholds stop matching the current policy, and the alert routing points at a distribution list of people who have left. Build a review of the program itself into the cycle: at least annually, confirm that each threshold still reflects what program staff consider acceptable, that each alert reaches a named current employee, and that the metrics still correspond to the decisions the system is actually making today.

Treat the program as something you build capability for rather than something you install. Decide which skills have to exist in house, sampling design, equity analysis, and reading a monitoring chart against a policy question, and which you can reasonably contract. Design the recurring processes so they run the same way each cycle regardless of who is on duty, and write down how the outputs are reviewed. Then build in a way to improve: a standing habit of asking, after each cycle, which threshold turned out to be wrong, which alert was noise, and which failure the program missed entirely.

Ask the funding question early and repeatedly. An assurance program funded from a project budget disappears when the project closes, which is usually the moment the system enters the long production life the program exists to cover. Ask the embedding question too: is assurance part of standard operations with a named owner and a line in the base budget, or is it sustained by one conscientious person? And ask what happens to the institutional knowledge, the baselines, the threshold rationales, the incident history, when that person moves to another agency.

Anti-Patterns

  • Treating the pre-deployment audit as the finish line. Approval evidence describes launch day, not drift, changed usage, staff turnover, or amended policy. Celestine's classifier had a clean approval memo and eighteen months of unexamined production behavior behind it.
  • Citing a certificate as proof the system is trustworthy. A management-system certification attests to a conforming management system, sampled across a declared scope at a point in time. It is not proof that a given AI system is safe, lawful, or fit for your use. Read the scope before you rely on it.
  • Reporting a single headline metric. Overall accuracy conceals subgroup failure by construction. Publish disaggregated figures alongside the aggregate, or the number reaching leadership is the one hiding the problem.
  • Letting the independent test stop being independent. When the quarterly structured test is run by the team that operates the system, it becomes a self-report with an audit heading. Independence is the property being purchased; without it, the cost buys nothing.
  • Setting thresholds without program staff. Engineers can say what the metric is. Only the people who absorb the consequences can say what value is acceptable, and the workload effect of a small metric change is routinely invisible from the engineering side.
  • Alerting into a void. Thresholds with no named recipient, no pause authority, and no pre-authorized decision produce a documented breach and an undocumented response. Define the chain, then drill it.
  • Funding assurance from the project budget. The money ends when the project closes and the production life begins. Put assurance in the base budget with a named owner, or plan on losing it at the worst moment.

Practice Prompts

  • Assess where you stand. Write an assessment covering current monitoring capability, gaps, organizational readiness, stakeholder alignment, and resource constraints. Identify which deployed systems have no retrievable baseline.
  • Assign the levels. List your agency's AI systems and assign each one limited, reasonable, or continuous monitoring, justified by the consequence of being wrong rather than by the technology used. Note every system whose current coverage is below its assigned level.
  • Write one measurement plan. For a single system, specify the indicators, the baselines, the acceptable operating thresholds, and the alert thresholds. Mark every value you had to invent because no baseline was ever recorded.
  • Run the two-day drill. Ask your team to produce continuous monitoring evidence for the three highest-risk systems as if the GAO had requested it. Time it. Document what was slow and why.
  • Rehearse the escalation. Walk a threshold breach through the chain: alert recipient, pause authority, citizen notification, agency head briefing. Confirm each named person still works there.
  • Plan the sustainment. Document how the program is funded, whether that funding survives project closure, who owns it, and how baselines, threshold rationales, and incident history survive that owner's departure.

Reflection

Take twenty minutes and pick the AI system whose failure would harm the most people. Write down what you know about how it is performing this month, then write down where that knowledge came from. If the answer is an approval document, a vendor dashboard, or the absence of complaints, you have a monitoring gap rather than a healthy system. The absence of complaints is the most seductive of the three: a system failing quietly for people who do not know they were failed generates none.

Then examine the sentence your agency uses when it describes that system as working. Does it say what was measured, over what period, on which population, against which threshold? If not, someone downstream is treating a bounded observation as a general guarantee, and that person is probably deciding whether human review is still needed. Rewriting the sentence is cheap; discovering the gap it concealed in a hearing is not.

Glossary

  • Audit. A point-in-time examination of a system against stated criteria, producing conclusions about what was examined during the period examined.
  • Assurance program. An ongoing process producing evidence that a system continues to meet defined standards in production, against thresholds agreed in advance.
  • Limited assurance. Periodic spot checks that sample behavior without covering every scenario, appropriate for low-risk systems that do not affect citizen-facing decisions.
  • Reasonable assurance. Systematic scheduled testing across defined risk scenarios, supporting a conclusion held with high confidence but not with certainty.
  • Continuous monitoring. Near-real-time instrumentation of outputs with automated alerting on threshold breaches. Sometimes labelled absolute assurance, which no assurance is.
  • Measurement plan. The document specifying indicators, baselines, acceptable operating thresholds, and alert thresholds for a system before it goes live.
  • Baseline. The performance values established during pre-deployment testing, without which later measurements cannot be interpreted as good or bad.
  • Model drift. Gradual degradation in a model's performance as the data or the world it operates in diverges from the conditions it was built for.
  • Demographic parity metric. A measure of whether performance is consistent across groups such as race, gender, age, and primary language, reported alongside rather than inside the aggregate.

Closing

Assurance is the discipline of continuing to look. The technical parts are the easy parts: metrics can be computed, charts drawn, alerts routed. The hard parts are organizational. Someone has to decide what value is unacceptable and defend that number. Someone has to hold authority to pause a system a program depends on. Someone has to keep the funding line alive after the project that built the system closed.

Celestine's finding was not that her agency had built a bad classifier. It was that her agency had stopped asking. An approval memo eighteen months old had become the answer to a question about today, and no one had noticed the substitution because nothing had visibly broken. Build the program that keeps asking, keep its evidence where you can retrieve it in two days, and describe what it shows in language that survives being read by someone who was not in the room.

Key Takeaways

  • Audits end; assurance continues. A pre-deployment audit tells you whether the system was safe to launch, not whether it remains safe as data, usage, staffing, and policy evolve.
  • Assurance evidence is bounded. It covers the indicators you chose, the population observed, and the period measured. It does not establish that a system is safe, lawful, or fit for purpose, and neither does a certificate whose declared scope you have not read.
  • Match the level to the consequence. Limited spot checks for internal tools, reasonable systematic testing for citizen-facing systems, continuous monitoring for high-stakes automated decisions. Choose from the harm of being wrong, not from the technology.
  • Start with a measurement plan. Indicators, baselines, acceptable thresholds, and alert thresholds belong in place before launch. Without a baseline, no later number can be read as good or bad.
  • Set thresholds jointly. Engineers establish what can be measured; program staff and the equity officer establish what is acceptable. Both conversations are required, and only the second one knows the workload cost of a two-point metric move.
  • Disaggregate your metrics. A 94 percent overall accuracy rate concealing 78 percent accuracy for non-English speakers does not meet a reasonable assurance standard.
  • Manage evidence for retrieval. Keep monitoring records in a structured, version-controlled repository and be able to produce evidence for your highest-risk systems within two business days of a request.
  • Pre-authorize the response. Define the escalation chain, the pause decision, and the citizen-notification path before you need them, and re-check annually that the named people still hold the roles.

Frequently Asked Questions

We passed our pre-deployment audit. Why do we need this? Because the audit described a specific day: what the auditors sampled, against the criteria they applied, on the version that existed then. Everything an assurance program watches for happens afterward. The input distribution shifts, a policy changes what a correct answer is, staff learn to override the system in ways nobody recorded. Celestine's agency had a clean approval and a classifier drifting for eighteen months.

Our vendor provides a monitoring dashboard. Is that sufficient? It is a useful input, not an assurance program. A vendor dashboard reports the indicators the vendor chose, usually availability and throughput rather than the equity and accuracy measures your program is accountable for, and it is produced by the party with the least interest in finding a problem. Use it, but set your own thresholds and keep your own evidence.

How do we set a threshold when we have no baseline? Establish one deliberately rather than guessing. Run a structured sample against a ground truth your program staff agree on, record the result as the baseline with its date and sample described, and set provisional thresholds you revisit next cycle. Document that the baseline was reconstructed after deployment rather than measured before it. An inspector general would far rather find that honest record than a confident number with no origin.

Who should own the pause decision? Someone senior enough to absorb the operational consequence and outside the team whose delivery schedule the pause disrupts. Pre-authorize the decision at design time with the criteria written down, because the moment you need it is the moment everyone is arguing about whether the breach is real. Celestine's agency fixed the review window at 24 hours so that argument runs against a clock somebody already agreed to.