←
AI for Government
Proficient · M1 · lesson 1 of 50 · in progress
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Audit Preparation
📖
now learning

AI Audit Preparation

15 min

Claudette Beaumont-Hazel had managed IT governance for the Illinois Department of Human Services for seven years when the Office of Inspector General notified her in August that it would conduct a performance audit of the department's AI systems, beginning in sixty days. The notification listed fourteen document categories the OIG would request. Claudette read the list twice. She could produce nine of the fourteen categories within a week. Three categories, fairness testing results, demographic performance analysis, and human override tracking, she could not produce at all, because the processes those documents were supposed to document had not been implemented. Two categories, incident response records and monitoring reports, existed on paper but reflected procedures that staff had stopped following eight months ago. She had sixty days to either produce the documentation or explain its absence. This lesson is about building the documentation architecture before the auditors send that letter.

Audit preparation is not paperwork for its own sake. It is the evidence that your systems deserve the trust the public is being asked to place in them. An agency that is unprepared for an audit turns a routine review into a crisis. An agency that is prepared turns the same review into a demonstration of competence, and usually into a shorter report.

Who Audits Government AI, and What They Look For

Government AI systems face several kinds of external scrutiny, each with a different focus, authority, and timeline. Knowing which body is asking changes what a good answer looks like.

Offices of Inspector General conduct performance audits of agency programs, including AI systems, focused on whether programs are operating as intended, achieving their stated objectives, and complying with applicable policy. OIG audits are authorized by the Inspector General Act and cannot be blocked by agency leadership. They produce public reports. An unfavorable OIG finding about an AI system, phrased as "the agency's AI screening tool lacks documented fairness testing and operates without human oversight as required by OMB guidance," is a significant reputational and political problem well beyond the program office.

The Government Accountability Office conducts audits at congressional request, typically focused on cross-agency issues or high-profile programs. GAO AI reviews have increased significantly since 2022. GAO published the AI Accountability Framework in 2021, which organizes federal oversight expectations around four principles: governance, data, performance, and monitoring. GAO reviews are slower than OIG audits but have greater reach, and they typically produce more detailed recommendations that a committee will later ask about again.

Civil rights offices and oversight bodies, including the Equal Employment Opportunity Commission for employment-related AI, the Department of Housing and Urban Development for housing-related AI, and state civil rights agencies, focus specifically on whether AI systems produce discriminatory outcomes. These reviews are triggered by complaints or patterns of outcomes, not by a schedule. They concentrate almost exclusively on fairness: do similarly situated individuals from different demographic groups receive different treatment from the AI system?

Agency internal auditors are the fourth group and the one agencies most often overlook in their planning. They review on the agency's own schedule, and their findings frequently become the first draft of what an external body examines later. Treating an internal review as a lower-stakes rehearsal is reasonable. Treating it as unimportant is how a small documented gap becomes a large public one.

Evidence, Not Assurances

What every auditor shares is a preference for evidence over assertion. "We believe our system is fair" does not satisfy an auditor, because belief is not reviewable. Compare it with this: "We conducted demographic performance analysis in March, comparing outcomes across race, gender, age, and primary language; the results are documented in this report; we found no significant disparity above our threshold of 2 percentage points; we repeat this analysis quarterly." That answer names the date, the dimensions, the artifact, the threshold, and the cadence. Every one of those is independently checkable, which is exactly the property an auditor is looking for.

This distinction is worth internalizing before the questions start, because it changes what you build during the year rather than what you say during the review. The documentation architecture described below exists so that the checkable version of every answer already exists in a folder. Nothing about audit preparation requires making a system look better than it is. It requires being able to show what the system actually does.

What Auditors Ask About

Auditors examining AI systems work through a consistent set of areas. The specific wording varies, but the areas do not, and they map directly to the documents you are expected to produce. Reading down this list and answering each question honestly is the cheapest self-assessment available to you.

AreaThe questions that get asked
System designHow was the system built? What methodology was used? Was it tested before deployment? Who reviewed and approved it?
Accuracy and performanceWhat accuracy was achieved? How does it compare to the manual process? Has accuracy been validated? How is accuracy monitored?
Fairness and biasHas the system been tested for bias? Are outcomes equal across demographic groups? Is there any evidence of discrimination? How is fairness monitored?
Data qualityWhat data is the system trained on? Is data quality documented? Are there known data issues? How is data quality assured?
Human oversightDo humans review decisions? Can humans override the system? How are overrides tracked? How is override effectiveness monitored?
ComplianceWhich laws and regulations are relevant? Is the system compliant? How is compliance ensured? Who monitors it?
Incident responseHave there been incidents? How were they handled? What was learned? What changed as a result?

The Seven Document Categories

The fourteen categories on Claudette's OIG notice can be organized into seven buckets. Every government AI system in production should have documentation in all seven, and the ones that are missing are almost always missing for the same reason: the process behind them was never stood up.

1. System design and approval records

This bucket covers what was built, why, by whom, and who approved it. It holds a system design document describing what the AI does, what data it uses, what outputs it produces, and what decisions it supports, plus an approved business case or project charter, a testing and validation plan, the results of testing, and a deployment approval signed by an appropriate authority. The deployment approval is often the missing piece: the system was tested, found to work adequately, and deployed without a formal authorization document. Fix this retroactively for existing systems by having the responsible executive sign a one-page authorization describing the system, its purpose, and the oversight conditions under which it operates.

2. Data documentation

What data was the system trained on? Where did it come from? What quality checks were performed? Are there known limitations or biases in the training data? Is the training data still representative of current conditions, or has the real-world situation changed since the model was trained? A data lineage document for each AI system, typically two to three pages, answers these questions and provides the baseline for any fairness analysis. You cannot test for demographic bias if you cannot describe the demographic composition of the training data, which is why this bucket has to be filled before bucket four is even possible.

3. Performance baseline and ongoing monitoring reports

Before an AI system is deployed, the agency should measure the current-state performance of the process it is replacing or augmenting. "The manual process approved 78% of applications within 14 days" is the baseline against which AI-assisted performance is measured, and without it you have no way to answer whether the system helped. Post-deployment, quarterly monitoring reports compare current performance to baseline and to target. These reports should be brief, one to two pages, and should be on file for the entire operational life of the system. Claudette's monitoring reports existed for the first two quarters and then stopped. The gap is the problem. Auditors will ask what happened in the quarters with no reports.

4. Demographic performance analysis

This is the single most common documentation gap in government AI audits. Demographic performance analysis answers one question: does the AI system produce materially different outcomes for different demographic groups? For benefits systems, are approval rates, processing times, or denial rates significantly different across race, gender, age, or primary language? For enforcement systems, are audit selection rates, citation rates, or penalty amounts significantly different across demographic groups? The analysis does not need to be statistically sophisticated. A table showing the relevant outcome metric broken down by group, with a clear definition of the threshold for significant disparity, is sufficient for most audits. It needs to exist, and it needs to be repeated at least annually.

5. Human oversight documentation

Every government AI system should have a documented human oversight process: which outputs are reviewed by a human before action is taken, what the review consists of, who performs it, and how reviews are recorded. For rights-impacting systems such as benefits determinations, enforcement decisions, or hiring recommendations, the documentation should also include override tracking, meaning records of how often a human reviewer overrode the AI recommendation and in which direction. Override records serve two purposes. They give you a measurable signal about the review process, and they reveal whether the system has systematic tendencies that reviewers are quietly correcting case by case.

Be precise about what an override log can and cannot establish. A log records that a reviewer entered a decision. It cannot record that anyone read the underlying case file, weighed the evidence, or considered overriding and declined. A queue that is being rubber-stamped produces a log that is indistinguishable from a scrupulous one, and in some respects looks better, because the override rate is low and the throughput is high. Override tracking is evidence worth having and worth reading closely for anomalies. It is not proof that human review is substantive, and presenting it that way to an auditor invites exactly the follow-up question you do not want.

6. Incident records

Every AI error that affected a citizen or an employee should be documented: what happened, when it was discovered, what the immediate response was, what investigation was conducted, what was found, and what was changed to prevent recurrence. An agency with no incident records for a system that has been in production for two years is not demonstrating excellent performance. It is demonstrating inadequate detection, and auditors know this. The expected question is "what is your mechanism for discovering AI errors?" If the answer is "we wait for complaints," the next question is "how many complaints have you received, and how were they resolved?"

7. Compliance records

What legal and regulatory requirements apply to this system? For federal systems that includes OMB Memorandum M-24-10, the Privacy Act, and the civil rights statutes relevant to the program. For state systems it includes applicable state AI law, state privacy law, and the federal requirements attached to federally funded programs. How is compliance with each requirement ensured? Who is responsible for monitoring it? When was it last verified? A compliance matrix, a table listing each applicable requirement, the agency's compliance approach, and the evidence of compliance, can be prepared in a few hours and makes an auditor's verification work far faster. It does not change what the evidence shows; it changes how long it takes to find.

Two Categories Agencies Forget

Two more document sets sit outside the seven buckets and go missing with striking regularity. The first is the risk classification and its justification: the written record of whether this system was assessed as rights-impacting or safety-impacting, who made that call, and on what reasoning. Auditors ask for this early, because it determines which obligations attach to everything downstream. An agency that cannot produce a classification memo has effectively told the auditor that no one decided which rules applied.

The second is training and change management. How were staff trained to use the system? What were the training completion rates? What was the change management approach, and what do the adoption metrics show? These records answer a question auditors increasingly ask: whether the human oversight described on paper is being performed by people who were actually taught what to look for. A well-documented oversight procedure staffed by untrained reviewers is a finding waiting to be written, and the training records are what distinguish the two situations.

The Audit Readiness Checklist

Before auditors arrive, verify each item below. The value of the list is not that it is comprehensive; it is that working it forces you to confirm the artifact exists rather than assume it does. Run it against one system at a time.

  • Documentation: system design documented; testing methodology documented; test results documented; fairness testing completed and documented; incident procedures documented; governance structure documented.
  • Performance data: baseline measurements taken before deployment; current performance measured and compared; demographic performance analysis completed; success criteria assessed against what was actually achieved.
  • Fairness and bias: fairness testing completed; either no significant unexplained disparities, or disparities documented and justified in writing; monitoring in place for fairness drift over time.
  • Human oversight: human review procedures in place; override tracking in place; staff trained on the procedures; override rates monitored and read, not merely collected.
  • Incidents: incident procedures in place; any incidents documented; incidents investigated and resolved; changes made that prevent recurrence rather than merely acknowledging it.
  • Data: data sources documented; data quality assessment completed; known data issues disclosed rather than omitted; data governance in place.

The Mock Audit

The most effective audit preparation technique is a mock audit conducted six months before you expect a real one. Assign a colleague who did not work on the AI system to ask the questions an auditor would ask, working from the areas and buckets above, and to attempt to gather the documents an auditor would request. The point of using someone unconnected to the build is that they cannot fill a gap from memory, which is precisely what the real auditor also cannot do. The gaps revealed by the mock audit are the gaps you have time to close. The gaps revealed by a real audit are the ones that appear in the public report.

The gaps that turn up are remarkably consistent. No deployment authorization document, which is easy to create retroactively if the responsible executive will sign it. Monitoring reports that exist for some quarters and not others, which you reconstruct from system logs where possible and document honestly as gaps where the logs do not exist. Demographic analysis conducted at deployment but never repeated, so the most recent analysis is stale and the repeat needs scheduling before the audit. Human override tracking that was designed but never implemented, which you implement now and date from the day tracking began. Do not backfill data you do not have.

Two further patterns appear often enough to name. Assumptions that were documented at design time but that the data never supported, which is worse than an undocumented assumption because it looks like diligence. And procedures that exist on paper but that staff stopped following, which is the shape of Claudette's incident response and monitoring records. Both are found by asking the person doing the work rather than reading the procedure, which is why a mock audit that only reviews documents finds half of what a real one will.

Audit Communication

How an agency communicates during an audit matters as much as what it can produce. Auditors develop their assessments from a combination of documents and interviews. Unclear answers in interviews raise questions that lead to additional document requests. Clear, consistent answers narrow the scope. That is the entire mechanism, and it means communication discipline is not about presentation. It is about not accidentally expanding the audit.

Three principles govern audit interviews. First, answer the question asked, not the question you wish had been asked; volunteering information that was not requested creates new lines of inquiry. Second, when you do not know an answer, say so and commit to a specific date for follow-up rather than speculating, because a speculative answer that later proves wrong costs you credibility on every other answer. Third, if documentation is missing, say so directly and explain what was in place instead. Auditors treat honesty about gaps differently than they treat attempts to obscure them.

The underlying posture is easy to state and hard to sustain under pressure. Be honest: if there are problems, disclose them, and do not make excuses. Be transparent: show the documentation, explain the decision-making, answer questions directly. Be prepared: have documents ready and know your system well enough to explain why it was built the way it was. Be proactive: disclose known issues, show what you have already done to address them, and explain the improvements made. An agency that says "we did not conduct demographic analysis in the third quarter because our analyst was on extended leave; we have since restored the quarterly schedule and the current analysis is attached" demonstrates accountability. An agency that produces a last-minute analysis dated the week before the audit demonstrates panic.

Anti-Patterns

Each of these is a natural response to audit pressure, and each one makes the outcome worse.

  • Hiding problems. Auditors will find them anyway, and the concealment becomes a separate and more serious finding than the original gap. A documented problem with a documented response is a manageable finding. The same problem discovered by the auditor is a credibility question about everything else you produced.
  • Blaming the vendor. You are responsible for your systems. A vendor's failure to deliver fairness testing is your failure to require and verify it, and an auditor will record it against the agency that deployed the system, not the company that sold it.
  • Treating the override log as proof of oversight. The log establishes that decisions were entered. It cannot establish that the underlying case was read. Offering it as evidence that human review is meaningful invites the auditor to test that claim directly, usually by interviewing reviewers about specific cases.
  • Making promises you cannot keep. Committing to a corrective action timeline you have not costed produces a second finding at the follow-up review. Be realistic in the interview and the follow-up is a formality.
  • Backfilling documentation. An analysis dated the week before the audit, or override data reconstructed for a period when no tracking existed, is detectable and reads as fabrication rather than diligence. Start tracking now, date it honestly, and describe the gap.
  • Defensive responses. Treating the auditor as an adversary lengthens the audit. A cooperative posture narrows scope; a defensive one signals that there is something behind the defensiveness worth looking for.
  • No documentation at all, then a scramble. Document as you go. Everything in this lesson is cheap when it is a byproduct of running the system and expensive when it is reconstructed under a sixty-day deadline.

Practice Prompts

  • Develop the audit documentation checklist for one AI system you are responsible for. Work the seven buckets plus risk classification and training records, and mark each artifact as present, stale, or absent.
  • Prepare for a mock audit by gathering the documents an auditor would request for that system. Time yourself. The elapsed time is itself a finding.
  • Write a one-page summary of the AI system suitable for an auditor's review: what it does, what data it uses, what decisions it supports, who approved it, and how it is monitored.
  • Build the compliance matrix: one row per applicable requirement, with the agency's compliance approach and the specific evidence in the third column. Note every row where the evidence column is empty.
  • Take your most recent demographic performance analysis and check its date. If it is older than a year, draft the schedule that will keep it current and identify who owns the schedule by name.
  • Rehearse three interview answers out loud: one where you know the answer, one where you do not, and one where the documentation is missing. Notice how strong the urge is to explain more than was asked.

Reflection

Think about an AI or automated system your office runs today. If the notification letter arrived this morning with a sixty-day clock, which of the seven buckets would you fill within a week, which would take the full sixty days, and which would you have to explain rather than produce? Then ask a harder question: for the buckets you would have to explain, is the documentation missing because nobody wrote it down, or because the underlying process was never performed? Those are different problems with different remedies, and only one of them can be fixed in sixty days.

Glossary

  • Performance audit: A review of whether a program is operating as intended, achieving its objectives, and complying with applicable policy. The standard OIG format for examining an AI system.
  • Data lineage document: A short record of where a system's training data came from, what quality checks it received, and what limitations are known, providing the baseline for any fairness analysis.
  • Performance baseline: The measured performance of the process before AI was introduced, against which post-deployment performance is compared. Without it, improvement claims are unfalsifiable.
  • Demographic performance analysis: A breakdown of outcome metrics by demographic group, with a stated threshold for what counts as significant disparity. The most common documentation gap in government AI audits.
  • Override tracking: Records of how often, and in which direction, human reviewers change an AI recommendation. Evidence about the review process, not proof that review is substantive.
  • Compliance matrix: A table pairing each applicable legal or policy requirement with the agency's compliance approach and the specific evidence supporting it.
  • Risk classification: The documented determination of whether a system is rights-impacting or safety-impacting, which governs which obligations attach downstream.
  • Mock audit: A rehearsal in which someone unconnected to the system asks an auditor's questions and attempts to gather an auditor's documents, conducted early enough that the gaps found can still be closed.

Closing

Audit preparation is not about making a system look better than it is. It is about documenting what a system actually does, so that the evidence, rather than the absence of evidence, tells the story. Every bucket in this lesson exists because some agency could not answer a question it should have been able to answer, and the resulting finding was published.

Claudette's sixty days were survivable for nine of fourteen categories and not for the rest, and the difference was never about writing speed. It was that five processes had either never been implemented or had quietly stopped. Documentation cannot be produced for work that was not done, which is the real lesson of every audit scramble. Treat the audit as the occasion to find out which of your processes are real, and you will get something more valuable out of it than a clean report.

Key Takeaways

  • Several kinds of auditor examine government AI: OIG, GAO, civil rights oversight bodies, and your own internal auditors. Each has different authority, focus, and timeline, but all require documented evidence rather than verbal assurances. An OIG finding is public. A GAO recommendation goes to Congress. A civil rights finding can trigger enforcement action.
  • Seven documentation buckets must be in place before auditors arrive. System design and deployment authorization; data lineage; performance baselines and monitoring reports; demographic performance analysis; human oversight and override tracking; incident records; and compliance records. Missing any one is a finding waiting to happen.
  • Risk classification and training records are the two documents agencies forget. The classification memo determines which obligations apply to everything else; training completion and adoption metrics show whether the oversight described on paper is performed by people who were taught what to look for.
  • Demographic performance analysis is the most common gap. Conduct it at deployment and at least annually thereafter. A table of the outcome metric by race, gender, age, and primary language, with a defined disparity threshold, is sufficient. It must exist.
  • An override log is evidence, not proof. It records that a reviewer entered a decision; it cannot record that anyone read the case file. A rubber-stamped queue produces a log indistinguishable from a scrupulous one, so read override records for anomalies and never offer them as proof that review is substantive.
  • A mock audit six months before the real one reveals closeable gaps. Assign a colleague who did not work on the system to gather the documents an auditor would request, and have them interview the people doing the work, not only read the procedures.
  • Honesty about gaps produces better audit outcomes than obscuring them. An agency that documents a gap, explains the cause, and demonstrates corrective action is treated differently from one whose inconsistencies suggest concealment. Never backfill data you do not have.
  • Incident records should exist for every production system. An AI system with no incident records after two years of operation is demonstrating inadequate error detection, not flawless performance. Document your detection mechanism and produce records that show it working.
  • Use the audit as a learning opportunity rather than a compliance exercise. The gaps it exposes are usually process gaps, and a process that only exists on paper will fail the next review too, whoever conducts it.

Frequently Asked Questions

We bought the system from a vendor. Is the documentation their responsibility?

The vendor may hold the source material, but the obligation to produce it is yours. Auditors examine the agency that deployed the system and made decisions with it, and "the vendor did not give us that" is recorded as a procurement and oversight failure rather than an excuse. Write the documentation deliverables into the contract before award, and verify on receipt that what arrived is what an auditor would accept, not a marketing summary.

How far back should our monitoring reports go?

For the entire operational life of the system. The concern is not volume but continuity: a run of quarterly reports that stops and restarts invites the question of what happened in between, and that question is harder to answer than the reports were to write. Where reports are genuinely missing, reconstruct from system logs if the logs exist, and where they do not, document the gap and its cause rather than leaving it unexplained.

What if our demographic analysis shows a disparity?

Document it, investigate the cause, and record what you did about it. A documented disparity with a documented response is a manageable finding. The unmanageable version is a disparity the agency did not know about because the analysis was never run, or knew about and did not record. Set the threshold for what counts as significant in advance and in writing, so the judgment is not being made after you already know the result.

Can we prepare for an audit in sixty days?

You can prepare documentation in sixty days. You cannot retroactively perform a process that never ran, and the difference is what separates Claudette's nine recoverable categories from her five. Sixty days is enough to assemble records, write a compliance matrix, obtain a retroactive deployment authorization, and refresh a stale analysis. It is not enough to manufacture a year of monitoring or a history of override tracking, and attempting either creates a worse problem than the gap.

Is a completed readiness checklist enough to pass an audit?

No. The checklist confirms that artifacts exist; the audit examines what they say. A complete file documenting a weak fairness test is still a weak fairness test, and a checked box has never once made a shaky system sound. Use the checklist to find what is missing, then read what you found with the same skepticism the auditor will bring to it.