Congressional and IG Reporting on AI
Ronan Donnelly is the Chief AI Officer at a federal regulatory agency with about 3,200 employees, and he has been in the role for 14 months. In that time he has submitted two required AI use case inventories to OMB (the Office of Management and Budget, which coordinates federal AI governance), responded to one data call from the Government Accountability Office (GAO, the investigative arm of Congress) about the agency's algorithmic accountability practices, and received formal notification from the agency's Office of Inspector General (OIG, the independent watchdog office within each federal agency) that AI governance had been added to the OIG's annual audit plan. He knew these obligations were coming. He was less prepared for the specific questions they asked, the documentation they expected to find, and the distance between what he believed the agency was doing and what the record could actually demonstrate. The OIG's preliminary findings were not damning. They were a useful education in what "documented governance" means as against "we have a governance process."
Documented Governance Versus Having a Process
The OIG audit standard is documentary. Findings rest on what the record shows, not on what agency officials say happened. An official can give entirely credible oral testimony that the agency followed a governance process, and if that process left no documentary trace the finding will be that the process cannot be verified. For practical purposes that is the same finding as one saying it did not occur, because an auditor has no method for distinguishing an undocumented process that ran from an undocumented process that did not.
Officials tend to hear that as an accusation of bad faith, and it is not one. It is a statement about what an audit can establish. The consequence for how you run a governance function is direct: every step that matters has to produce an artifact with a date on it, filed somewhere a stranger could find it. A decision made in a meeting and communicated verbally happened, in the sense that it changed what the agency did, and did not happen, in the sense that survives an audit. The gap between those two senses is where Ronan spent his first audit.
The Three Reporting Channels
Federal agencies face three distinct categories of AI reporting obligation, and they differ in audience, in the standard applied to a response, and in what happens when a response is inadequate. Treating them as one compliance function is the first mistake, because a document that satisfies one channel frequently fails another. The table below sets out the differences, and the paragraphs after it fill in what each channel actually asks for.
| Channel | What it asks for | Standard applied | What follows a poor response |
|---|---|---|---|
| OMB and the White House | Annual AI use case inventory, certification that minimum risk management practices were applied to high-impact systems, progress reporting on governance implementation. | Certification against stated federal requirements. | Primarily reputational and programmatic. |
| Congress, proactively | Annual reports required by statute, submitted to authorizing and appropriations committees. | Accurate to the technical reality and comprehensible to a non-technical reader at the same time. | Follow-up inquiry, and an authorizing or appropriations record that is harder to correct than to write. |
| Congress, reactively | Responses to data calls, letters and subpoenas from oversight committees or individual members, usually on short notice. | Same dual standard, under time pressure and without the chance to organise first. | Escalation from a data call to a letter to a hearing. |
| Inspector General | The documentary record of governance for each system: authorizations, assessments, protocols, monitoring, incidents. | Documentary. What the record shows, not what officials state. | An audit finding, published, that the agency then has to remediate and report against. |
Reports to OMB and the White House are grounded in OMB Memoranda, particularly M-24-10, OMB's 2024 guidance on advancing governance, innovation and risk management for federal AI, and in legislation such as the AI in Government Act. The consequences of failure in this channel are primarily reputational and programmatic: failure to certify compliance can result in OMB withholding authorization for certain AI acquisitions and can generate negative attention in the President's Budget. Confirm the current consequence structure with your own general counsel and OMB liaison before relying on it in a briefing, since guidance in this area is revised.
Congressional reporting takes two forms and they demand different preparation. Proactive reporting means annual reports required by statute, going to authorizing and appropriations committees, on a schedule you can plan against. Reactive reporting means responses to data calls, letters and subpoenas, on a schedule you cannot. Congressional staff frequently have limited technical background and ask questions shaped by policy concerns rather than technical precision, which means a response has to be simultaneously accurate to the technical reality and comprehensible to a non-technical reader. Those two requirements pull against each other in every sentence.
Inspector General oversight is the most operationally consequential of the three. Every federal agency has an OIG that operates independently of agency leadership, auditing agency programs for efficiency, effectiveness and compliance. That independence is the point: the OIG is not a stakeholder to be managed, and treating an audit as a communications exercise is the fastest route from a routine finding to a serious one. Where the OIG's specific statutory authorities and access rights are in question, that is a matter for your agency counsel and your OIG liaison, not for a governance handbook.
What Submitting a Report Does Not Discharge
The most common misreading of this whole area is that filing closes something. It does not. A submitted use case inventory is a snapshot of what the agency believed on the day it was compiled. Submitting it does not discharge the duty to keep it accurate, and an inventory that was correct at submission and wrong two months later is a governance failure whether or not the filing was timely. The same holds for a certification: certifying that minimum risk management practices were applied is a statement about a moment, and it says nothing about the system as modified since.
Nor does a completed report discharge the underlying obligation it reports on. A progress report describing governance implementation is evidence that a report was written. If the practices it describes are not running, the report is a specific, dated, signed statement by the agency that will be read back to it. Certification is the highest-stakes version of this, because it converts an operational shortfall into a documented representation. If the agency's practices do not meet the standard, the answer is to improve the practices, not to certify and plan to catch up.
The same discipline applies to a briefing. Appearing before staff, walking a committee through a system, or answering a data call fully and promptly are all good things and none of them closes an oversight matter. Oversight bodies decide when a matter is closed, and they do so in writing. Treat every interaction as adding to a record you will be held to, rather than as discharging a duty you can then stop tracking.
Preparing the IG Documentation Package
When an OIG announces that AI governance is on the audit plan, the agency typically has 60 to 90 days before the formal audit begins. That window is not for creating governance after the fact. It is for organising, reviewing and completing the documentary record of governance that has already occurred. The distinction is not a technicality, and the paragraph below on retroactive documentation explains why treating it as one is the single most dangerous move available to you during those two to three months.
Ronan's agency had governance processes in place. What it lacked was organised documentation. The AI use case inventory was accurate but stored across three different spreadsheets in two divisions. Risk assessments for high-impact systems had been completed, but the completed assessments sat in individual program managers' email inboxes rather than in a central repository. The human-in-the-loop review protocols existed in a SharePoint document that had last been updated ten months earlier and did not reflect changes made during a system update in the intervening period. None of that is misconduct. All of it produces findings.
For each active AI system, an OIG expects to find six things in the record:
- The system's authorization record: when it was approved, by whom, under what governance framework, and with what documented conditions
- The pre-deployment risk assessment, including the determination of whether the system is a rights-impacting or safety-impacting AI system as those terms are defined under OMB M-24-10
- Evidence that minimum risk management practices were applied, including the specific checklist or framework used and the date of completion
- The current human-in-the-loop protocol and the date it was last reviewed
- Monitoring reports covering the period of production use, with dates, methodology and findings
- Any incident reports for the system, with dates, descriptions and the corrective actions taken
If any of these documents do not exist, the honest answer to the OIG is that they do not exist. It is not that they will be created to cover the gap. Retroactive documentation produced after an audit announcement is a significant integrity problem and may compound an audit finding into something considerably more serious than the finding it was meant to avoid. A missing monitoring report is a control weakness. A monitoring report dated last spring and written last week is a different category of problem entirely, and it is one that attaches to the people who produced it rather than to the program.
Congressional Data Calls
Congressional data calls, meaning formal requests for information from oversight committees, typically arrive with two to three weeks of lead time and ask for documentation the agency may never have organised in the form requested. Five requests recur in AI oversight: a list of all AI systems in production use, the vendor or development source for each, the decisions or determinations each system assists with, whether the system has been subject to bias testing, and whether affected parties have been notified that AI is used in decisions affecting them.
Read that list as a readiness checklist rather than as a description of one committee's curiosity, because the five questions are largely the same ones an OIG asks in a different register. An agency that maintains the six-item documentation package can answer all five inside the lead time. An agency that does not will spend the two to three weeks assembling facts under deadline, which is precisely the condition under which people round, guess and reconcile inconsistent spreadsheets by picking one. The last of the five questions is the one agencies most often cannot answer, because notification practice is usually set at the program level and never aggregated.
Testifying in Outcome Language
Preparing testimony for a hearing on AI requires translation between technical and policy frames. The members asking questions are not typically AI practitioners. They are legislators who want to know whether their constituents are being served fairly, whether the agency is spending technology funds prudently, and whether there are risks that warrant legislative action. Congressional oversight of AI is, in the end, oversight of outcomes for constituents, and an agency that can speak fluently about how a system affects people rather than how it works internally is a more credible witness.
An answer that says "we use ensemble methods with regularization to prevent overfitting in our claims processing model" will prompt follow-up questions and skepticism. An answer built on outcomes lands: the system reviews about 4,000 claims per month and recommends approval or denial to a human case manager who makes the final decision; it has been bias-tested; and the case manager's final decision differs from the recommendation in roughly 8% of cases, which on that volume is on the order of 320 cases a month.
Two things in that answer need stating precisely, and both are places where a witness can accidentally overclaim. The 8% figure measures how often the human departs from the recommendation. That is an override rate, not an error rate. It does not tell you how often the system was wrong, because a case manager who agrees with a wrong recommendation produces no divergence at all, and a case manager who overrides a correct one produces divergence without any system error. Report it as an override rate, say what it is measured against, and resist the pull toward calling it accuracy.
The bias claim needs the same care. What testing supports is that no statistically significant difference in recommendation rates was found across the demographic groups that were measured, on the data available, at the statistical power that data provided. Say it in those terms. "We found no bias" is a broader claim than any test delivers, and a committee that later learns a group was not measured, or that the sample could not have detected a difference of the relevant size, will treat the original phrasing as an attempt to mislead rather than as a simplification.
Voluntary Disclosure
Some AI governance failures are found internally before any oversight body finds them. In those cases proactive disclosure to OMB, the OIG or the relevant oversight committee is almost always better than hoping the problem stays buried. Disclosure is not an admission of catastrophic failure. It is the same good faith that makes a governance function credible in the first place, exercised at the moment it costs something.
Be precise about what disclosure achieves, because the standard framing of this overpromises. Disclosing removes the concealment question, preserves the credibility of everything else the agency reports, and lets the agency present the scope of the problem and its remediation plan rather than having both characterised by someone else. It does not erase the underlying failure, cure the obligation that was missed, or bind any oversight body to respond in a particular way. Disclose because concealment is the larger failure and because the record is better with it, not because disclosure is expected to buy a lighter outcome.
Ronan made one voluntary disclosure during his tenure. A bias monitoring report for a medium-risk system was six weeks overdue when the compliance coordinator discovered it had never been produced. The disclosure went to the OIG's office as a one-page summary of what happened, a corrective action plan with named owners and dates, and a request for confirmation that the disclosure had been received. The OIG acknowledged receipt and made no further inquiry. Note what that acknowledgment was: a receipt, not a determination, and nothing about it prevented the office from returning to the matter later. The same failure surfaced during the audit instead would likely have been at least what Ronan's agency classified as a Level 2 compliance finding.
Anti-Patterns
- Treating a filed report as a discharged obligation. A submitted inventory is a snapshot of a belief on a date. An inventory accurate at filing and wrong two months later is a governance failure regardless of timeliness. Tie inventory updates to the change-control trigger for the systems themselves.
- Certifying ahead of practice. Certifying that minimum risk management practices were applied when they were only partly applied converts an operational shortfall into a documented representation by the agency. If the practices do not meet the standard, fix the practices.
- Producing documentation after the audit announcement. The 60 to 90 day window is for organising and completing the record of governance that occurred, not for manufacturing a record of governance that did not. A backdated artifact converts a control weakness into an integrity problem that attaches to individuals.
- Calling an override rate an error rate. How often a human departs from a recommendation measures divergence, not accuracy. Agreement with a wrong recommendation is invisible in that number and an override of a correct one inflates it. Name the metric by what it measures, in testimony and in the monitoring report both.
- Saying "we found no bias." What a test supports is no statistically significant difference in the groups measured, on the data available, at the power that data provided. The unqualified version is broader than any test delivers, and the qualification is what protects you when someone asks which groups were measured.
- Governance records in personal storage. Assessments in individual inboxes, protocols in a stale shared document, an inventory split across spreadsheets in two divisions. Each was real work that registers as an absence. A central indexed repository per system is the minimum requirement, not an aspiration.
- Reporting a briefing as an outcome. Answering a data call fully, testifying well or walking staff through a system are additions to the record. None closes an oversight matter. Oversight bodies close matters in writing, and until they do the matter is open.
- A portfolio with no incident reports. An active AI portfolio with no documented incidents in a year is far more likely to have an undocumented incident history than a genuinely clean one, and auditors read it that way. An incident report with a corrective action is evidence of a working governance function, not an admission.
Practice Prompts
- Assemble the six-item documentation package for your highest-impact AI system today: authorization record, pre-deployment risk assessment with its rights-impacting or safety-impacting determination, evidence of minimum risk management practices with dates, current human-in-the-loop protocol with its last review date, monitoring reports, incident reports. Note where each lives and how long it took to find.
- Answer the five recurring congressional data call questions for your full AI portfolio, without asking anyone else for help, and time yourself. Whichever question you cannot answer is the one that will arrive with a two-week deadline.
- Take one metric you report externally and write down exactly what it measures, what it does not measure, and how someone could misread it in your favour. Do this first for anything you currently describe as an error rate or an accuracy figure.
- Audit one governance artifact for currency. Find the authorization document for a system that has been updated since deployment and check whether it reflects the system as it runs now. If it does not, you have found the shape of your next audit finding.
- Draft the voluntary disclosure you would send if you found a materially overdue compliance obligation tomorrow: a one-page summary, a corrective action plan with named owners and dates, a request for confirmation of receipt. Having the template written removes the delay that turns a disclosure into a discovery.
Reflection
Take twenty minutes with these questions. For your highest-impact AI system, which of the six documentation items could you produce within an hour, and which do not exist? Where does your agency rely on someone's recollection of a decision rather than a dated artifact? Which external claims about your systems are stated more broadly than the evidence supports? If a compliance obligation were missed today, how long would it take to find out, and who would decide whether to disclose it?
Glossary
- AI use case inventory. The required listing of an agency's AI systems submitted to OMB. A snapshot of a belief on a date, not a standing description of the portfolio.
- Rights-impacting and safety-impacting AI. Risk categories used to determine which systems carry heightened requirements, as those terms are defined under OMB M-24-10.
- Data call. A formal request for information from an oversight committee, typically arriving with two to three weeks of lead time and asking for records in a form the agency has not organised.
- Override rate. How often a human decision-maker departs from a system's recommendation. Distinct from an error rate, which it neither measures nor bounds.
- Voluntary disclosure. Proactively reporting an internally discovered governance failure to an oversight body, with documented scope and a remediation plan, before that body finds it.
- Documentary standard. The audit principle that a finding rests on what the record shows rather than on what officials state, under which an undocumented process is treated as unverifiable.
Related Lessons
- Oversight Mechanisms: IG, GAO, Congress maps how the three channels interact and where one fact ends up in all of them.
- AI Audit Preparation covers the operational work inside the window before an audit begins.
- AI Use Case Inventory and Documentation (OMB M-24-10) is the detailed treatment of the inventory obligation itself.
- Briefing Lawmakers on AI covers the oral counterpart to the written reporting in this lesson.
- Risk Classification: Safety-Impacting vs. Rights-Impacting covers the determination that anchors the documentation package.
- Minimum Risk Management Practices covers what the certification actually certifies.
- AI Incident Documentation and Response covers the incident record auditors expect to find, and are suspicious to find empty.
- GAO AI Accountability: Four Principles in Practice covers the framework behind many GAO data calls.
Closing
Almost everything in this lesson reduces to one habit: produce a dated artifact for every governance step that matters, and keep it somewhere a stranger could find it without your help. Ronan's agency was not badly governed. It was governed in email, in meetings and in the heads of program managers, which is how most agencies are governed, and which is invisible to every one of the three reporting channels. The audit did not find bad practice. It found unverifiable practice, and the remedy for that is filing, not reform.
The second habit is precision about claims. Say what a metric measures, say what a test found and in which groups, say what a submission is a snapshot of. Every overstatement in a report to OMB, a response to a data call, or an answer in a hearing is a hostage handed to whoever reads it next. Understating is inconvenient once. Overstating is a finding, and unlike a finding about a missing document, it attaches to the person who said it.
Key Takeaways
- IG findings rest on the documentary record. Credible oral testimony that a governance process ran does not survive an absence of documentation. Every governance process must produce a durable, organised, dated record that a stranger can locate and review.
- Central repositories are a prerequisite, not an improvement. Records spread across inboxes, personal folders and disconnected spreadsheets cannot be audited. One indexed location per active system is the minimum.
- The three channels differ in audience, standard and consequence. OMB certifies against federal requirements, Congress demands technical accuracy and non-technical comprehensibility at once, and the IG applies a documentary standard. A document that satisfies one frequently fails another.
- Filing does not discharge. An inventory is a snapshot, a certification is a statement about a moment, and a briefing is an addition to a record. Oversight bodies close matters, in writing; nothing you submit closes one for you.
- The audit window is for organising, never for manufacturing. Retroactive documentation created after an announcement is an integrity problem that compounds the finding it was meant to prevent. If a document does not exist, say so.
- Name metrics by what they measure. Divergence between a human decision and a recommendation is an override rate. A test that found no significant difference found it in the groups measured, on the data available, at the power available. State both that way every time.
- Update documentation when systems change. An authorization record reflecting a system as deployed 18 months ago, with two updates since, is not current documentation. Tie the governance record to the same change trigger as the system.
- Disclose early, and for the right reason. Voluntary disclosure with documented scope and a corrective action plan removes the concealment question and preserves credibility. It does not erase the failure or bind any oversight body's response.
- Incident reports are a governance strength. An active portfolio with no documented incidents is more likely to have an undocumented incident history than a genuinely clean one, and auditors read it that way.
Frequently Asked Questions
What do we do if the audit window opens and a required document genuinely does not exist? Say so, in writing, with the date you established it was missing and what you are doing about it going forward. Create the missing control from today rather than a record of it from the past. A gap disclosed by the agency is a finding you can remediate; a gap papered over is a different investigation.
Should a data call response go through general counsel? Route it through counsel and through whoever owns your relationship with the requesting body, every time, and build that into the two to three week timeline rather than discovering it on day twelve. The point is not review for tone. It is that a response to Congress is a representation by the agency and needs to be treated as one.
How much detail belongs in a written report to OMB? Enough to support every claim you make and no claim you cannot support. The failure mode in these reports is not insufficient detail but confident summary: a sentence that describes the intended process rather than the process that ran. Write what happened, with dates, and let the summary follow from it.
Skill.re