←
AI for Government
Strategic · M4 · lesson 4 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Audit Methodology
📖
now learning

AI Audit Methodology

15 min

Tobias Castellano had been a senior evaluator in the Department of Labor's Office of Inspector General for eleven years when her division chief handed her an assignment she had never prepared for: lead the agency's first structured audit of an AI system. The system was a machine-learning model the Employment and Training Administration had deployed eighteen months earlier to prioritize job-seeker referrals at American Job Centers across six states. Nobody in the OIG had audited an AI system before. No standard procedure, checklist, or interview protocol existed. Tobias had a $280,000 budget, a two-person team, and six months before the report had to reach the House Education and the Workforce Committee. She built the methodology from scratch.

What an Audit Establishes

Start with the limit, because every decision that follows is shaped by it. An audit examines what the auditor sampled, against the criteria the auditor stated, over the period the auditor selected. It does not establish that a system is sound. A report with no adverse findings means that the tests performed, on the sample drawn, against those criteria, did not surface a reportable condition. Tobias examined 24 centers out of roughly 340 and twelve months out of eighteen. Everything outside those boundaries is untested, and the report has to say so.

This matters because of what happens to an audit report downstream. A program office under pressure will quote a clean finding as a clearance, and a vendor will quote it in the next procurement as proof the model works. Neither use is supported by the document. The defense against both is in the report itself: state the scope, the sample, the period, and the criteria in the same breath as the conclusion, and write the summary so that lifting a sentence out of it does not produce a stronger claim than the audit supports.

The corollary applies to your own team. A procedure that returns no exceptions is evidence about that procedure. If the compliance matrix showed no gaps, the honest statement is that the requirements on the matrix were satisfied by the evidence obtained, which is a narrower claim than compliance. The gap between those two sentences is where audits get into trouble later, usually in front of people who have had time to read the methodology section closely.

The Governing Framework: GAO Standards

Tobias's first move was to anchor the audit in existing professional standards rather than inventing a framework for AI. The Government Accountability Office, the independent congressional watchdog, publishes Government Auditing Standards, commonly called the Yellow Book. The Yellow Book applies to federal performance and compliance audits, including audits of systems that happen to use machine learning. Anchoring there mattered practically as well as procedurally: it gave her a body of accepted practice to cite when the auditee questioned a method, which they did.

A performance audit asks whether a program is achieving its intended results. For an AI system: is the referral prioritization model actually placing more job-seekers in employment than the previous manual process did? Performance audits require the auditor to establish a baseline, measure current outcomes, and attribute the difference to the program being examined. Attribution is the hard part, because a labor market that improved on its own will happily take credit for a model that did nothing.

A compliance audit asks whether a program is operating within applicable legal and regulatory boundaries. For an AI system: is the model consistent with OMB Circular A-130, the Privacy Act, Title VI of the Civil Rights Act, and agency-specific AI use policies? Compliance audits are checklist-driven, because the legal requirements are the standard and the auditor is not at liberty to substitute a better one. The judgment lies in deciding which requirements apply, not in deciding what compliance means.

Most government AI audits blend both types. Tobias's mandate was primarily a performance question, does the model work, but any compliance violations discovered had to be reported regardless. The Yellow Book requires it. Which type you are conducting shapes every subsequent decision: what evidence you collect, how you analyze it, how you frame findings, and what you are entitled to conclude. Decide it explicitly and write it in the plan, because an audit that drifts between the two ends up defensible as neither.

Planning Scope Before You Collect Anything

Think of an AI audit as a structural inspection of a building that is still occupied. You cannot examine everything at once. You must decide which load-bearing elements to inspect first, which systems to sample, and which finishes to skip entirely. Scope planning makes that triage explicit and defensible instead of accidental. Tobias spent the first three weeks, before requesting a single document, defining four scope boundaries and writing down the reasoning behind each one.

System boundary. Tobias included the model, its training data, its deployment configuration, and the required human review step before referral. She excluded downstream case management at the Job Centers. Boundary decisions must be documented, because they matter enormously when the auditee argues that a finding falls outside the audit scope, and a boundary defended from memory is a boundary that moves.

Time period. The model had been live for eighteen months. Tobias selected the most recent twelve: long enough to capture seasonal variation in a labor program, short enough to analyze within a six-month timeline. Recording why she chose twelve rather than eighteen turned out to matter when the agency later suggested that the excluded early period would have looked better.

Geographic sampling. Six states, roughly 340 American Job Centers. Auditing all of them was not feasible on a $280,000 budget. Tobias selected 24 centers across four states, stratifying the sample to include high-volume urban centers, lower-volume rural centers, and centers serving high proportions of non-English-speaking job-seekers. The stratification was designed to surface equity patterns that a simple random sample would very likely have diluted below visibility.

Key questions. Tobias's audit plan listed five. Is the model producing more employment placements than the pre-model baseline? Are outcomes distributed equitably across race, gender, primary language, and disability status? Is human oversight operating as designed rather than as described? Does the system comply with applicable data privacy requirements? Are model performance metrics being reported to agency leadership? Each question is phrased so that a specific body of evidence could answer it, which is what stops a plan question from becoming a topic heading that fieldwork wanders around inside.

The written plan, which ran fourteen pages, governs everything that follows. Evidence or analysis not traceable to one of those questions should not appear in the final report, and a question the plan did not ask should not acquire an answer partway through fieldwork. If the audit needs to add a question, amend the plan, date the amendment, and say why. That record is what separates an audit that followed the evidence from one that followed whatever turned out to be interesting.

Designing Test Procedures

Test procedures must be designed before evidence collection begins. Designing them afterward invites confirmation bias, because you will build tests that validate what you already suspect and you will not notice yourself doing it. Design first, then collect. Where a procedure needs a threshold, fix the threshold in the plan and record where it came from, so that the number is defensible as a standard rather than as a convenient line drawn after the data was visible.

Performance procedure

Tobias used a difference-in-differences design. She compared placement rates at the 24 sampled centers during the audit period against rates at the same centers in the twelve months before deployment, controlling for macroeconomic conditions using the state unemployment rate as a covariate. Comparing locations to themselves over time is more defensible than comparing model-using centers to non-model-using centers, which may differ in staffing, local industry mix, and caseload in ways no covariate captures.

Equity procedure

She computed disparity ratios between the highest- and lowest-performing demographic groups. Her reportable-finding threshold was a ratio above 1.25, the point at which the lower group's placement rate falls below four-fifths of the top group's rate. The threshold is drawn from the four-fifths rule used in employment discrimination analysis at the Equal Employment Opportunity Commission, and it was documented in the audit plan before any data was reviewed. Both the value and its provenance belong in the plan; a threshold with a published pedigree is much harder to argue with than one an auditor selected.

Compliance procedure

She built a 23-row compliance matrix listing each applicable requirement, the confirming evidence needed, and the source of the requirement. It was used as a literal checklist throughout document review, with each row marked satisfied, not satisfied, or not evidenced. That third state matters: a requirement for which no evidence was produced is not the same as a requirement that was met, and collapsing the two is one of the quieter ways an audit overstates what it found.

Collecting Evidence

Documentary evidence, meaning system design documents, testing reports, approval records, and monitoring logs, establishes what the agency said it would do. Tobias submitted a 31-item request in week four. She received 22 of those items within the 15-day return window her request set, and three never arrived at all. Missing documents are themselves a finding: the Yellow Book requires the auditor to note unreturned requests and to assess whether the resulting gap affects the audit's conclusions rather than quietly working around it.

Testimonial evidence, meaning interviews with program staff, data scientists, and contractors, surfaces the undocumented practices that make up most of how a system is actually operated. Tobias conducted 17 structured interviews using a standardized protocol so that answers could be compared across sites rather than read as anecdotes. She recorded each interview in a structured notes form that the interviewee reviewed and signed. Signed notes are far more defensible than reconstructed summaries when the auditee later disputes what was said.

Analytical evidence, meaning the auditor's own analysis of program data, is often the most consequential and the most contested. Tobias ran her analysis in a documented, reproducible script with every step logged, from the raw extract through each filter and join to the final figure. The auditee's technical staff will scrutinize your methods, and they will be better at your tooling than you are. Show exactly what you did and why, and treat any step you cannot reproduce on demand as a step you cannot rely on.

The general rule behind all three: in an AI audit, the strength of a finding depends not just on what the evidence shows but on whether a skeptical technical expert can follow every step from raw data to conclusion without asking you a question you cannot answer from the workpapers. Document the chain while you are building it. Reconstructing provenance under a response deadline is how findings get withdrawn.

Interpreting Results and Writing Findings

Tobias's analysis produced results she had not expected. The model improved aggregate placement rates by 8.4 percentage points, a statistically significant gain, and on the headline performance question the program looked like a success. But the disparity analysis showed that job-seekers whose primary language was not English experienced placement rates 31 percent lower than English-speaking job-seekers with similar qualification profiles. That gap put the disparity ratio well past her pre-registered 1.25 threshold, and it was invisible in the aggregate number the program had been reporting upward.

A finding under Yellow Book standards has four required components: criteria, meaning the applicable standard; condition, meaning what the auditor found; cause, meaning why the gap exists; and effect, meaning the consequences. Tobias's equity finding had all four. The criterion was equitable service delivery under Title VI. The condition was the 31 percent placement gap for non-English-speaking job-seekers. The cause was that 68 percent of training records predated the agency's language-access expansion in 2019, leaving that population underrepresented in the data the model learned from.

The effect was an estimated 2,200 job-seekers per year receiving lower-priority referrals than their qualifications warranted. Note the word estimated, and note that the basis for the estimate belongs in the workpapers where a reader can check it. A quantified effect is what moves a finding from an observation to a decision, and it is also the number most likely to be challenged, quoted out of context, and repeated in a hearing. Every factual assertion must cite the specific document, interview, or analysis that supports it. A finding that cannot be traced to its evidentiary basis will not survive the auditee's response, and it should not.

Working With the Auditee

The agency you are auditing is preparing for you, and it helps to understand what that preparation looks like from their side. AI Audit Preparation covers the auditee's half of this process in detail, including how a program assembles and organizes what an auditor will ask for. From the auditor's chair, the useful consequence is that a well-prepared auditee produces evidence quickly and a poorly prepared one produces delay that you must then interpret: is the documentation missing, disorganized, or being withheld? Those are three different findings.

Keep the relationship procedural rather than adversarial, and keep it on the record. Send document requests in writing with a return date. Confirm interview notes in writing. When the auditee objects to a method, hear the objection, note it, and respond in the workpapers rather than in a hallway. Inspectors general have statutory access to agency records, and you rarely need to invoke it; what you need is a paper trail showing that access was requested clearly and that any shortfall was documented at the time rather than reconstructed afterward.

Presenting to Oversight Bodies

Tobias presented the draft report to the Employment and Training Administration before release, as the Yellow Book requires. The agency disputed two findings and accepted three. Both disputed findings were retained with the agency's response attached, which is standard practice and better practice than negotiating a finding away: the reader can see the disagreement and judge it. The 47-page report posted publicly on the OIG website within five days of the response period closing.

The Committee hearing came six weeks later. Congressional staff focus on two things: the number that captures the problem, which here was 2,200 job-seekers per year, and the agency's accountability for fixing it, which took the form of a corrective action plan due in 90 days. Tobias spent more time preparing for questions than drafting her statement. Know your key number before you sit down at the witness table, and know exactly how it was calculated, because the follow-up question is always how you got it.

Be equally ready for the question about what you did not examine. It arrives in every serious hearing, and the wrong answer is a defensive one. The right answer is the scope statement you wrote in week three: these centers, these months, these criteria, chosen for these reasons, with these areas excluded and this effect on what the report can conclude. An auditor who can recite that fluently sounds rigorous. An auditor who improvises it sounds like someone who did not think about it until asked.

Anti-Patterns

  • Letting a clean report become a clearance. A report with no adverse findings says the tests performed, on the sample drawn, against the stated criteria, surfaced nothing reportable. It does not establish that the system is sound. Write the scope into the summary so that a sentence lifted out of it does not overclaim.
  • Choosing thresholds after seeing the data. A disparity ratio or significance level selected once the results are visible is advocacy wearing audit clothing. Fix the value in the plan, record its provenance, and accept the finding it produces.
  • Treating not evidenced as satisfied. A compliance row with no supporting document is an open question, not a pass. Collapsing the two is the quietest way an audit overstates what it actually found.
  • Collecting evidence before designing procedures. Tests built after the data is in hand tend to confirm what the auditor already suspected, and the auditor will not notice. Design first, then collect.
  • Working around missing documents. Unreturned requests are a finding in their own right and must be noted with an assessment of how the gap affects conclusions. Quietly substituting a weaker source hides a governance problem inside a methodological one.
  • Reporting only the aggregate. An 8.4 percentage point headline gain coexisted with a 31 percent gap for non-English speakers. Disaggregate before you conclude, or the number you report will be the one that conceals the finding.
  • Stating an effect without its basis. The quantified effect is the number that will be quoted, challenged, and repeated in a hearing. If its derivation is not in the workpapers, it will not survive contact with the auditee's analysts.
  • Negotiating a disputed finding away. Retain it and attach the response. A finding withdrawn under pressure disappears from the record; a finding published with a rebuttal leaves the reader able to judge both.

Practice Prompts

  • Assess your capability. For AI audit work in your office, write an assessment covering current capability, gaps, organizational readiness, stakeholder alignment, and resource constraints. Name the skill you would have to contract for on your next AI engagement.
  • Write a scope statement. Pick one deployed AI system in your agency and draft the four boundaries: system, time period, sampling, and key questions. For each, write the sentence you would use in a hearing to justify what you excluded.
  • Design one procedure. Choose a plan question and write the test that answers it, including the threshold, its provenance, and what result would and would not constitute a finding. Do this before looking at any data.
  • Build the matrix. List the requirements applicable to one system, the confirming evidence each needs, and its source. Mark honestly which rows you could evidence today.
  • Draft a finding. Take a known weakness and write it up with criteria, condition, cause, and effect. Quantify the effect, then write down exactly how you would defend that number under questioning.
  • Plan the presentation. Decide the single number that captures your most significant finding, and rehearse the answer to the question about what you did not examine.

Reflection

Take twenty minutes with the most consequential AI system your office is likely to audit next. Write down the four scope boundaries you would set and, next to each, what a critic could reasonably say you were avoiding. Then write the conclusion sentence you expect to end up with, and check whether it claims more than the sample would support. Most overclaiming in audit reports is not dishonesty. It is a summary written for readability by someone who knew the caveats and assumed the reader would supply them.

Then consider the thresholds. Which numbers in your plan came from a published standard, which from precedent in your office, and which from your own judgment? The last category is legitimate, but it must be labelled as judgment and defended on its merits. A threshold presented as a rule, when the auditee discovers it was a preference, damages every other number in the report.

Glossary

  • Yellow Book. Informal name for Government Auditing Standards, published by the Government Accountability Office and governing federal performance and compliance audits.
  • Performance audit. An examination of whether a program is achieving its intended results, requiring a baseline, a measurement, and an attribution argument.
  • Compliance audit. An examination of whether a program operates within applicable legal and regulatory boundaries, driven by the requirements rather than by the auditor's judgment of what is adequate.
  • Scope boundary. A documented limit on what the audit examined: the system elements, the time period, the sample, and the questions asked.
  • Difference-in-differences. A design that compares the same locations before and after a change, rather than comparing adopters to non-adopters who may differ in unrelated ways.
  • Disparity ratio. The ratio between the highest- and lowest-performing groups on an outcome measure, used here with a reportable threshold of 1.25.
  • Four-fifths rule. The convention, drawn from employment discrimination analysis, that a lower group's rate falling below four-fifths of the top group's rate warrants examination.
  • Compliance matrix. A row-per-requirement checklist recording each applicable requirement, the evidence that would confirm it, its source, and whether it was satisfied, not satisfied, or not evidenced.
  • Criteria, condition, cause, effect. The four required components of a Yellow Book finding: the standard, what was found, why, and the consequence.
  • Workpapers. The auditor's documented record of evidence, analysis, and reasoning, sufficient for a skeptical reader to retrace every step from raw data to conclusion.

Closing

Auditing an AI system is not a new discipline so much as an old one applied to an unfamiliar object. The standards already exist, the finding structure already exists, and the obligations to the auditee and to the public already exist. What is new is that the object under examination produces its outputs by a process nobody can fully narrate, which puts unusual weight on scope, sampling, and reproducibility. Get those three right and the rest of the methodology carries over intact.

Tobias built her methodology from scratch because no procedure existed, and the durable part of what she produced was not the analysis. It was the discipline of writing down what she would examine, why, and against what standard, before she looked at anything. That is what let her report a 31 percent equity gap inside a program with a genuine 8.4 percentage point success, and defend both numbers in the same hearing. An audit is only as strong as the boundaries it was honest about.

Key Takeaways

  • An audit examines what was sampled against stated criteria. It does not establish that a system is sound. Tobias covered 24 of roughly 340 centers and twelve of eighteen months; everything else is untested and the report must say so.
  • Anchor the audit in the Yellow Book from day one. Government Auditing Standards govern performance and compliance audits of AI systems, and knowing which type you are running determines your evidence requirements, analysis, and how you frame findings.
  • Write the plan before collecting a single document. System boundary, time period, sampling strategy, and key questions must be documented first. A written plan is what protects a finding when the auditee argues it falls outside scope.
  • Fix thresholds and their provenance in advance. A disparity threshold of 1.25, tied to the four-fifths rule, is defensible because it was set and sourced before the data was seen. A threshold chosen afterward is advocacy.
  • Treat missing documents as a finding. Note every unreturned item and assess its effect on your conclusions. Record not evidenced separately from satisfied.
  • Build every finding on criteria, condition, cause, and effect. Quantify the effect and put its derivation in the workpapers; that number is what motivates corrective action and what will be challenged first.
  • Disaggregate before concluding. An 8.4 percentage point aggregate gain concealed a 31 percent placement gap for non-English-speaking job-seekers.
  • Know your key number and your exclusions before the hearing. Staff will remember one figure, and they will ask what you did not examine. Both answers should come from documents you wrote months earlier.

Frequently Asked Questions

Do we need machine learning expertise on the audit team? You need enough to design defensible procedures and to recognize an evasive explanation, plus access to more when the analysis gets technical. Tobias ran a two-person audit and leaned on the Yellow Book's structure rather than evaluating the model's internals. The questions that mattered most, does it work, is it equitable, is oversight real, are answered with outcome data and interviews.

What if the agency says our sample is too small? Answer with the sampling design rather than the size. Tobias stratified 24 centers to include urban, rural, and high non-English-speaking populations precisely so the sample would surface patterns a random draw would dilute. Explain the stratification, state what the sample does not support, and note the budget constraint that set the size. A documented sampling rationale survives that challenge; an undocumented one does not.

How much of the audit should the auditee see before publication? The draft report, as the Yellow Book requires, and enough of the method for the response to be substantive rather than procedural. Tobias's agency disputed two findings and accepted three, and both disputes were published with the responses attached. Showing the method invites informed pushback, which is uncomfortable and produces a stronger final report than pushback based on guesses about what you did.

Can we reuse this methodology for the next AI system? The structure carries over: anchor in the standards, write the plan, design procedures with sourced thresholds, collect three kinds of evidence, build findings on the four components. The specifics do not. A different system has different applicable requirements, a different population, and a different meaning of working. Reusing a scope statement without rederiving it is how an audit rigorously answers the previous system's questions.