AI in Financial Regulation
Priya Nair is the head of market surveillance at a state financial regulator with jurisdiction over 2,400 licensed firms and a team of 31 examiners. On a Tuesday, her surveillance system generated 4,100 alerts. Her team can investigate maybe 60 a day. The other 4,040 sat in a queue, and somewhere in that queue, she suspected, was the next enforcement case that would otherwise surface only after a retiree lost their savings. The problem was not that her agency lacked data. It was drowning in it. The alerts were 95 percent noise, and the noise was hiding the signal she was paid to find.
Financial regulation is where AI's promise and peril are sharpest, because the people you regulate are using AI faster than you are. A trading firm deploys a model that can move markets in milliseconds; you have a quarterly examination cycle and a spreadsheet. Closing that asymmetry, what people call RegTech, is the real subject here. Used well, AI lets a small public team supervise a fast, complex market. Used badly, it produces confident accusations a court will throw out, and it does so at a volume no examiner corps can review.
The regulatory stack you are operating inside
United States financial regulation is a patchwork, and knowing which piece you are standing on determines which statute constrains your model. The Securities and Exchange Commission regulates securities markets and investment advisers under the Securities Act of 1933, the Securities Exchange Act of 1934, the Investment Company Act of 1940, and the Investment Advisers Act of 1940. The Commodity Futures Trading Commission regulates derivatives under the Commodity Exchange Act. The Federal Deposit Insurance Corporation, the Office of the Comptroller of the Currency, the Federal Reserve Board, and the National Credit Union Administration regulate depository institutions under their respective banking statutes.
The rest of the map matters just as much. The Consumer Financial Protection Bureau regulates consumer financial products under Title X of the Dodd-Frank Act. FinCEN enforces the Bank Secrecy Act and anti-money-laundering rules. The Office of Financial Research at Treasury supplies analytics to the Financial Stability Oversight Council. The Federal Financial Institutions Examination Council coordinates prudential regulation across the bank agencies. The Federal Housing Finance Agency regulates the housing government-sponsored enterprises and the Federal Home Loan Banks. Every one of these bodies is deploying AI internally for examination and surveillance, and every one is also supervising AI inside the entities it regulates.
That dual role is the structural fact of this lesson and it has no analogue in most of government. You both use AI and regulate others' AI. The same fairness standard you enforce on a lender's model, you must meet in your own surveillance model. The same model documentation you demand of a supervised institution, an inspector general or the Government Accountability Office can demand of you. Holding firms to a standard you violate is the fastest way to lose a case, and it is a failure mode that is invisible from inside the agency until someone external looks.
Fraud detection and market surveillance
This is Priya's problem, and it is the clearest win. Traditional surveillance uses fixed rules: flag any trade over a threshold, any account with a sudden spike. Rules are simple to defend but generate the 95 percent noise, because a rule cannot tell a legitimate large trade from a suspicious one. It only sees size. Machine learning changes the question from whether something crossed a line to whether it looks unlike normal behavior for this account, in this market, at this time. A model that learns each account's baseline can rank the 4,100 alerts so the 60 most anomalous rise to the top.
At federal scale the same approach runs on far larger data. The Securities and Exchange Commission's Consolidated Audit Trail aggregates every order, cancellation, and execution in United States equity and options markets, running to trillions of events. Its Market Information Data Analytics System handles public trade and quote data. Enforcement staff apply machine learning across both to identify spoofing, layering, marking the close, and insider trading patterns that no human eye finds in that volume. The National Exam Analytics Tool helps examination staff triage adviser books for suspicious trading. Derivatives surveillance combines futures exchange data with swap data repositories to detect manipulation across linked markets, and FINRA, the self-regulatory organization overseen by the Commission, runs cross-market equities supervision with models that flag manipulation spanning venues.
The governance problems are equally real and worth naming before the capability is procured. Model drift produces false positives that flood investigators or false negatives that miss manipulation, and drift in a market surveillance model is guaranteed because markets change by design. Explainability is hard, because a judge will want to know why the model flagged a specific trader. Adversarial actors probe the system, and unlike most government AI, the population under observation includes sophisticated parties with the resources and motive to learn what triggers an alert.
The non-negotiable rule follows directly: an alert is a lead, not a charge. Securities enforcement staff are explicit that AI outputs are leads and not evidence, and that a human investigator builds the case using traditional methods. The model tells an examiner where to look. The examiner builds the case with evidence a human can explain to a judge. If you cannot articulate, in plain English, why a trade is suspicious without saying the algorithm said so, you do not have a case. You have a liability.
The explainability requirement is not optional here
In most government AI uses, explainability is good practice. In financial enforcement it is a due-process requirement. When you sanction a firm, it has the right to know and contest the basis. A black-box model that cannot show its reasoning fails that test. This is why regulators favor models whose outputs a human can interrogate, and why every AI-surfaced lead must be re-grounded in documented evidence before it becomes an action.
Be careful about what satisfies that requirement, because vendors will offer you an explainability feature and treat the box as checked. An explanation generated after the fact, which sounds plausible but does not reflect what actually drove the output, is worse than no explanation at all. It manufactures confidence in a proceeding where confidence is exactly what is being tested. The consumer-protection guidance discussed later in this lesson says this outright about credit denials, and the principle is identical in enforcement: the reason given has to be the actual reason, demonstrably, not a narrative reconstructed around a score.
Anti-money-laundering and Bank Secrecy Act compliance
FinCEN, under the Bank Secrecy Act, requires financial institutions to file Suspicious Activity Reports and Currency Transaction Reports. The volume is enormous, exceeding four million such reports per year, and legacy rules-based monitoring systems produce false-positive rates often above 95 percent. Machine learning and network analysis can cut false positives while surfacing genuine networks that transaction-level rules never see, because the signal in money laundering is frequently relational rather than individual.
The supervisory expectations here are unusually specific. The FFIEC Bank Secrecy Act and anti-money-laundering examination manual, updated in 2024, explicitly addresses automated monitoring systems and requires institutions to validate models, document tuning decisions, and provide audit trails. FinCEN has signaled through advisory guidance and subsequent innovation statements that responsible AI adoption is welcomed, provided sound risk management accompanies it. Federal examination teams now expect institutions to explain their monitoring models, document false-positive rates, show evidence of ongoing tuning, and demonstrate a clear path from alert to filing decision.
The permanent cautionary tale is the Danske Bank Estonian branch case, in which roughly $220 billion in suspicious transactions moved before anti-money-laundering monitoring caught up. The lesson for AI is not that better models would have prevented it. It is that no model overcomes gaps in data coverage, failures of governance, or an institutional unwillingness to act on the alerts the system already produced. A supervisor evaluating an institution's AI monitoring should therefore spend as much attention on what happens after an alert as on how the alert was generated.
One narrowing is worth stating explicitly, because the pitch in this space is always about volume reduction. Cutting false positives is genuinely valuable and it is not free. Tuning that reduces alert volume can also reduce true positives, and the true positives you lose are invisible by construction. Any claim that a model reduced alerts by some proportion is incomplete until it is paired with evidence about what happened to detection of known cases. Ask for both numbers or treat the first one as a cost claim rather than a performance claim.
Prudential supervision and model risk management
Prudential supervisors deploy AI to triage examination resources, score institutions on dimensions resembling the standard supervisory ratings, build early warning indicators, and stress-test portfolios. Innovation offices at the banking agencies were among the first parts of the federal supervisory apparatus to formalize AI experiments. The policy foundation underneath all of it is Supervisory Letter SR 11-7, issued by the Federal Reserve in 2011, together with the paired Office of the Comptroller of the Currency Bulletin 2011-12, the supervisory guidance on model risk management.
Those documents require institutions and their supervisors to manage models across the full life cycle: development, implementation, use, validation, governance, and controls. SR 11-7 predates modern machine learning, and its principles scale better than people expect: independent validation, ongoing monitoring, documentation, and board reporting are technique-agnostic requirements. What does not scale automatically is the validation practice built on top of those principles, which generally assumes a model you can inspect and training data your institution controls. Neither assumption survives contact with a third-party foundation model, which is where the framework has to be extended rather than merely applied.
Vendor models attract a rule that supervisors state without qualification and that you should carry into your own agency's practice. When a bank uses a vendor's machine learning fraud model, the bank remains responsible. Outsourcing the model does not outsource the accountability. The identical principle applies when an agency buys a vendor's model for supervisory analytics: the agency is responsible for the model choices it made, and the Government Accountability Office will ask the agency to explain them. Build contractual access to model logic and error metrics into the procurement, and be clear-eyed that contractual access is not the same thing as institutional capability to evaluate what you receive.
Where modern machine learning stretches the framework
The model risk framework was written for classical models: regression, structural credit models, value-at-risk. Modern machine learning, deep learning, and foundation models stretch it in three ways the source sets out. First, training data is massive and frequently third-party, so validation teams need access to data lineage and not only to model weights, and examination expectations now reach documentation of data provenance for AI systems. Software supply chain requirements, including the Office of Management and Budget memo M-22-18 on secure software development practices and the NIST Secure Software Development Framework it points to, apply to machine learning pipelines as much as to traditional software.
Second, foundation models are not built inside the institution. An institution using a general-purpose model for internal research still owns the model risk when the outputs affect a regulated decision. The workable governance pattern treats the foundation model as one layer of a stack, with each layer carrying its own validation plan and its own statement of what it is and is not relied upon for. Third, drift and adversarial robustness become first-class concerns rather than periodic checks, and ongoing monitoring under the Measure and Manage functions of the risk framework becomes continuous in a way classical models never required. A 2024 Federal Reserve review of model risk management identified large language models as a supervisory priority.
Consumer protection and fair lending
On the protection side, AI helps a regulator see harm at scale. It can cluster thousands of consumer complaints to reveal that 600 separate complaints about a lender are actually one pattern of illegal fee-stacking. It can scan marketing materials for deceptive claims. It can monitor for the newer frontier of harm: AI-driven discrimination in lending, where a model denies credit in ways that correlate with race even when race was never an input. That last point cuts both ways, and it is the part to internalize. The firms you regulate are deploying AI for credit, pricing, and underwriting, and your job is increasingly to audit their models.
The legal architecture here is dense and every element of it predates AI. The Equal Credit Opportunity Act at 15 USC 1691 and following, implemented through Regulation B, prohibits discrimination in credit. The Fair Credit Reporting Act at 15 USC 1681 governs consumer report data. The Fair Housing Act reaches housing-related credit. Unfair, deceptive, or abusive acts and practices authorities cover conduct the specific statutes do not. Together these create a broad fairness regime that AI systems must satisfy, and fair-lending model risk accordingly requires disparate-impact testing, review of alternative data, and documentation that survives an examination.
The Consumer Financial Protection Bureau's Circulars 2022-03 and 2023-03 are the specific guidance on adverse action and AI, and their content is unambiguous. Creditors using AI must still comply with the Equal Credit Opportunity Act and must provide specific, accurate reasons in adverse action notices even when the decision came from a complex model. Notices must specify the principal reasons for denial even where the denial was AI-driven, and post-hoc explanations that do not reflect the actual drivers of the decision are non-compliant. The model is too complex to explain is not an acceptable adverse action reason. Carry those statements exactly as written when you brief your own staff, because they are the operative standard and paraphrase loses the part that binds.
What you should ask a supervised lender follows directly from that standard. Show me your model's approval rates by protected class. Show me your fair-lending testing and what it covered. Show me what your model uses as a proxy for income and what else that proxy correlates with. Show me a sample of adverse action notices and demonstrate that the stated reasons are derived from the decision rather than assembled around it. A regulator who does not understand AI cannot supervise an industry that runs on it, and these four requests are the minimum literacy the job now requires.
Consumer-facing AI harms also reach beyond the financial regulators proper. The Federal Trade Commission polices unfair and deceptive acts under Section 5 of the FTC Act, including AI-driven dark patterns and deceptive AI marketing, and its actions concerning biased hiring algorithms and deceptive conversational agents supply useful precedent even for supervisors outside its jurisdiction. Fair-lending examination is coordinated across the banking agencies, so an institution's exposure is rarely confined to one supervisor's view of it.
The frameworks that govern your own use
The NIST AI Risk Management Framework, with its Govern, Map, Measure, and Manage functions, gives you the structure. It is voluntary and non-binding. Its Measure function is where financial regulators must be rigorous: you must document your model's false-positive rate, meaning legitimate trades wrongly flagged, and its false-negative rate, meaning real manipulation missed. For a screening tool that only ranks where examiners look, a high false-positive rate is tolerable. For anything that triggers an action, it is not, and the acceptable threshold should be set explicitly per consequence level rather than inherited from whatever the model happened to produce.
The Office of Management and Budget memo M-24-10, issued in 2024, classifies AI affecting access to financial services or imposing penalties as rights-impacting. A surveillance ranker is comparatively low-risk. A system that automatically suspends a license or freezes an account is rights-impacting and demands an impact assessment, human decision, and an appeal path. Narrow the low-risk framing slightly as you apply it: a ranker still determines who receives examination attention, which is a distributive decision even when no individual output is an adverse action, and it deserves fairness monitoring for that reason alone.
RegTech beyond surveillance
Beyond enforcement, AI takes administrative weight off a thin examiner corps. It can read a firm's filing and check it against requirements, draft a first-pass examination summary, and triage incoming filings by risk so the high-risk ones reach a human first. A regulator that automates the completeness check on filings can redirect substantial examiner time from clerical review to actual investigation, and this category deserves to be funded first for the same reason it deserves less governance overhead: nothing in it decides anything about anyone.
Guard the boundary of that category deliberately. A completeness checker that begins generating a risk score has entered a different class. A first-pass examination summary that becomes the examination summary has quietly relocated a judgment from an examiner to a model. Write down where the administrative category ends, and revisit it whenever a vendor proposes an enhancement, because the enhancement is usually the crossing.
International alignment
Institutions under your supervision are increasingly multinational, which means foreign requirements land on your desk whether or not you invited them. The European Union's AI Act, in force from August 2024, classifies high-risk AI and imposes pre-market and post-market obligations, and it treats credit scoring for natural persons as high-risk, requiring conformity assessment, data governance, human oversight, and post-market monitoring. Institutions operating in the European Union face this extraterritorially. The Bank of England supervisory statement SS1/23 modernizes United Kingdom model risk management expectations, and the Monetary Authority of Singapore's FEAT principles, covering fairness, ethics, accountability, and transparency, are an influential global reference.
The Financial Stability Board and the International Organization of Securities Commissions have each published principles and recommendations for AI in financial services, and your agency's counterparts meet in those fora. The positions taken there shape the global standard of practice, and inconsistency between jurisdictions costs institutions money and creates opportunities for regulatory arbitrage. That is a reason to participate rather than a reason to defer: a standard set without your input still applies to the firms you supervise.
Enforcement case studies and what went wrong
The pattern across enforcement actions involving AI is remarkably consistent. Inadequate governance of model choices, insufficient fairness testing, and poor adverse action documentation produce penalties regardless of how sophisticated the underlying model was. Model quality has never been the defense. Documentation and governance have been.
In 2024 the Securities and Exchange Commission charged two investment advisers with making false and misleading statements about their use of AI, having claimed AI-driven investment processes they did not actually have. The case establishes that AI claims are material and actionable under Section 206 of the Investment Advisers Act, and it is worth reading as a warning about marketing language rather than about technology. The Office of the Comptroller of the Currency has taken multiple fair-lending actions where machine learning models or automated underwriting contributed to disparate outcomes, with consent orders typically requiring model retraining, new monitoring, and board reporting. FinCEN has imposed substantial penalties on institutions for inadequate transaction monitoring, frequently including model governance deficiencies among the findings.
One consumer case is worth stating in full because of what the institution argued. A large national bank settled redlining allegations in 2023 after an examination found that its mortgage marketing algorithm effectively steered offers away from majority-minority census tracts. The bank argued that the algorithm had optimized for revenue. The regulator held that disparate impact is disparate impact whether or not the intent is algorithmic. Carry that holding verbatim into your own supervisory practice: optimization for a legitimate business objective is not a defense to a discriminatory outcome, and an institution that offers it as one has told you how its model governance works.
Two cases from outside financial services are cited so routinely by financial supervisors that they belong here. Michigan's Integrated Data Automated System, an unemployment fraud detection system whose acronym unhelpfully matches the securities market data system named earlier and which is entirely unrelated to it, falsely accused tens of thousands of people of fraud. It is the standing cautionary tale for any AI system making adverse decisions at scale, and supervisors cite it alongside the COMPAS criminal-justice risk assessment case. Settlements involving third-party facial recognition data sourcing illustrate a further exposure: financial institutions ingesting third-party identity intelligence inherit the regulatory liability attached to how that data was collected.
A surveillance AI governance checklist
Run any AI surveillance or enforcement tool through these eight checks before deployment, and require written answers rather than assurances.
- Lead, not verdict. Confirm the tool ranks and surfaces; humans build and own every case.
- Explainable basis. Confirm an examiner can state the reason for suspicion without citing the algorithm, and that the stated reason is derived from the decision rather than reconstructed around it.
- Documented error rates. Know and record false-positive and false-negative rates, and set an acceptable threshold for each consequence level in advance.
- Fairness self-audit. Test whether your tool's attention or actions skew by firm size, geography, or protected class, and remember that an audit tests only the disparities you thought to measure.
- Rights-impacting classification. Sort each tool in writing; apply impact assessments and appeal paths to anything that penalizes.
- Evidence re-grounding. Require every AI-surfaced lead to be independently documented before any enforcement step.
- Drift monitoring. Markets change by design; schedule retraining and revalidation so the model does not silently decay between examinations.
- Vendor transparency. If you buy the tool, the contract must give you access to its logic and error metrics rather than a score alone, and you must have someone able to evaluate what that access delivers.
What Priya builds first
Priya does not need to automate enforcement. She needs her 4,100 daily alerts ranked so her 60 investigations land on the 60 most anomalous rather than the 60 oldest. That is a low-risk re-ranking layer on a system she already runs. She can pilot it for one quarter, measure whether it surfaces real cases earlier than the old queue, and report the lift to her board, all without ever letting the model touch an enforcement decision. The model finds the needle. Her examiners still thread it.
The measurement design matters more than the model here, and it is the part most agencies skip. Before the pilot starts, decide what counts as a real case, decide how you will detect the ones the ranker pushed down, and decide what result would cause you to stop. Re-ranking looks harmless because nothing is automated, which is exactly why nobody demands evidence that it worked. A ranker that reliably deprioritizes a category of misconduct is doing real harm quietly, and the only thing that surfaces it is a measurement somebody committed to in advance.
Anti-patterns to watch for
Accepting an explainability feature as an explanation. The standard in this domain is that the stated reason must be the actual reason. A narrative generated after the fact, however fluent, is the specific thing consumer-protection guidance calls non-compliant when it does not reflect the drivers of the decision. Test the claim rather than the feature: change an input and see whether the explanation changes with it. If it does not, you have a storytelling layer sitting on top of a black box.
Reporting alert reduction as performance. Every monitoring vendor leads with the proportion by which they cut false positives, and every tuning decision that reduces alerts can also reduce detection. The true positives lost are invisible by construction, which is why the claim has to be paired with evidence about detection of known cases. A reduction figure standing alone is a description of cost savings dressed as a description of accuracy.
Treating a ranker as consequence-free. A tool that only decides where examiners look has still decided who gets examined and who does not, and that distribution can skew by firm size, geography, or the characteristics of a customer base. The low-risk classification is right about the safeguards it triggers and wrong if it is read as a reason to skip fairness monitoring. Monitor concentration in your leads the way you would monitor concentration in your actions.
Outsourcing the model and assuming you outsourced the risk. Supervisors state without qualification that the institution using a vendor's model remains responsible, and the identical logic applies to your agency when it buys supervisory analytics. Contractual access to logic and error metrics is necessary and insufficient; someone on your side has to be capable of evaluating what arrives. A contract clause you cannot exercise is a paper control.
Applying classical validation practice to foundation models unchanged. The model risk principles scale. The practices built on them assume an inspectable model and controlled training data, and a third-party general-purpose model offers neither. Extend the framework explicitly with layer-by-layer validation and a written statement of what each layer is relied upon for, rather than performing a validation that the model's structure has quietly made meaningless.
Enforcing a standard on lenders that your own models would fail. This is the dual-role trap, and it is invisible from inside the agency. Run the fairness testing you demand of supervised institutions against your own surveillance and triage models, on the same schedule, documented the same way. The first time an inspector general or opposing counsel asks the question is a bad time to discover the answer.
Practice prompts
- Map your agency's AI portfolio against the model risk management controls in SR 11-7 and OCC Bulletin 2011-12. Identify each gap and write the remediation you would propose, including who owns it.
- Draft adverse action notice language for an AI credit model that would satisfy the specific-and-accurate-reasons standard, then draft the test you would run to demonstrate that the stated reasons reflect the actual drivers.
- Design a joint supervisory approach with a foreign counterpart for a dual-authorized institution using a foundation model in credit decisions. Identify where the two regimes' requirements conflict and how you would sequence them.
- Present a model risk memo to a mock board committee explaining an anti-money-laundering monitoring model, including its false-positive rate, its tuning history, and what happens to an alert after it is generated.
- Write the supervisory playbook section specifying the evidence your examiners must collect from a regulated entity using AI: which artifacts, which tests, which documentation, and what an inadequate answer looks like.
Reflection
These are worth working through with your examination staff rather than alone, because the gap between how a supervisory tool is described and how it is used is where most of the exposure in this lesson actually sits.
- Which of your own models would fail the fairness testing you require of the institutions you supervise?
- If a firm challenged an examination that began with an AI-generated lead, what could you put in the record about how that lead was produced?
- Where has a tool in your agency crossed from administrative triage into shaping a judgment, without anyone reclassifying it?
- Do you know what your surveillance model deprioritizes, and how would you find out?
- How much of your team's AI literacy is sufficient to audit a supervised institution's underwriting model, and how much is sufficient only to be impressed by one?
Glossary
- RegTech. The use of technology, increasingly AI, to perform regulatory and supervisory functions, closing the capability gap between a public supervisor and the firms it oversees.
- Model risk management. The life-cycle discipline of developing, validating, monitoring, documenting, and governing models, set out in the paired 2011 supervisory guidance and extended with difficulty to modern machine learning.
- Adverse action notice. The notice a creditor must provide stating the principal, specific, and accurate reasons for a denial, a requirement that complexity of the model does not excuse.
- Disparate impact. A discriminatory outcome regardless of intent. Optimization for a legitimate business objective is not a defense against it.
- Significant non-compliance. In supervisory analytics, the elevated-risk classification that triage models are typically trained to predict.
- Suspicious Activity Report. The filing institutions make under the Bank Secrecy Act, produced in volumes exceeding four million per year, and the output end of most anti-money-laundering monitoring.
- Model drift. Degradation of a model's performance as the environment moves away from its training conditions. Structural in market surveillance, because markets change continuously.
- Data lineage. The documented history of training data, including origin, transformations, and control, which validation of a machine learning model requires and inspection of model weights does not supply.
- Rights-impacting AI. A system affecting access to financial services or imposing penalties, requiring impact assessment, a human decision, and an appeal path.
Related lessons
- AI Regulatory Design covers the design of regulatory regimes for AI itself, which is the other half of the dual role described here.
- Enterprise Risk Frameworks for AI generalizes the model risk life cycle beyond the financial supervisory guidance it originated in.
- Third-Party AI Risk Management develops the vendor-model accountability rule that this lesson states and does not elaborate.
- Algorithmic Fairness in Government supplies the disparate-impact testing methods behind both the fair-lending supervision and the self-audit obligation.
- Rights-Impacting and Safety-Impacting AI Safeguards covers what the rights-impacting classification obliges once a supervisory tool starts imposing consequences.
- Continuous Monitoring Fundamentals addresses the drift monitoring that market surveillance makes mandatory rather than periodic.
- Evaluating AI Vendor Claims is the right companion for the alert-reduction and explainability claims described in the anti-patterns above.
- AI for Mission-Critical Government Functions places financial supervision in the broader class of high-consequence government deployments.
Closing
The asymmetry is the whole problem. The firms you oversee already run on AI, and a regulator without AI literacy is bringing a quarterly spreadsheet to a millisecond market. Closing that gap is not optional and it is not primarily a technology project. It is a governance project, because the constraint on a public supervisor is never raw capability. It is the requirement that every action be explainable, documented, contestable, and defensible in front of a decision-maker who is entitled to disagree.
That constraint is also the advantage, if you use it. An institution that cannot explain its own model to you has told you something important about its governance. An enforcement case built on evidence rather than on a score survives a challenge that a score-based case would not. And an agency that holds its own models to the standard it enforces on others arrives at every hearing with the one thing that cannot be bought late: a record that shows the work. Priya's ranker is worth building. What makes it worth trusting is everything she committed to measuring before she turned it on.
Key takeaways
- An alert is a lead, not a charge. Enforcement staff are explicit that AI outputs are leads and not evidence; a human examiner must build every case on evidence a judge can follow.
- Explainability is a due-process requirement. A firm you sanction has the right to contest the basis, and the stated reason must be the actual reason rather than a plausible narrative assembled around a score.
- You both use and regulate AI. Meet the same fairness standard in your surveillance model that you enforce on a lender's, and run the testing on the same schedule.
- Adverse action standards do not bend for model complexity. Notices must give specific, accurate, principal reasons even for AI-driven denials, and the model is too complex to explain is not an acceptable reason.
- Disparate impact is disparate impact. Optimization for a legitimate business objective is not a defense to a discriminatory outcome, whether or not the intent was algorithmic.
- Document both error rates and pair reduction claims with detection evidence. High false positives are tolerable in a screening ranker and unacceptable in anything that penalizes, and alert reduction alone says nothing about what stopped being caught.
- Outsourcing the model does not outsource the risk. The institution, or the agency, remains responsible for the model choices it made and must be able to explain them.
- Governance failures produce the penalties, not model quality. Across enforcement actions the recurring findings are weak model governance, insufficient fairness testing, and poor adverse action documentation.
- Start by re-ranking, not automating. Put your best investigations on the most anomalous alerts, commit to the measurement before the pilot, and keep the model away from enforcement decisions.
Frequently Asked Questions
Can an AI alert be the basis for an enforcement action?
It can be the basis for deciding where to look and nothing more. Enforcement staff at the federal level are explicit that AI outputs are leads and not evidence, and that a human investigator builds the case using traditional methods. Every AI-surfaced lead has to be independently documented before any enforcement step. If you cannot state, in plain English, why the conduct is suspicious without referring to the model, you do not have a case that will survive a challenge.
What makes an explanation good enough for an adverse action notice?
It has to give the principal reasons, specifically and accurately, and those reasons have to reflect what actually drove the decision. Guidance is explicit that post-hoc explanations which do not reflect the actual drivers are non-compliant, and that model complexity is not an acceptable reason. The practical test is causal: change an input, and see whether the stated reason changes with it. An explanation layer that produces the same narrative regardless is a compliance risk rather than a compliance control.
Does using a vendor's model reduce my institution's or agency's exposure?
No. Supervisors state without qualification that an institution using a vendor's model remains responsible for it, and the same principle applies to an agency buying supervisory analytics. Negotiate contractual access to model logic and error metrics rather than a score alone, and then make sure someone on your side can actually evaluate what that access delivers. A contract clause nobody has the capability to exercise is a paper control with a signature on it.
How does model risk guidance written in 2011 apply to foundation models?
The principles carry over well: independent validation, ongoing monitoring, documentation, and board reporting are technique-agnostic. The practices built on them do not, because they assume a model you can inspect and training data you control. Extend the framework rather than reapplying it: demand data lineage and not only model weights, treat the foundation model as one layer of a stack with its own validation plan, and make drift and adversarial robustness continuous concerns rather than periodic checks.
Is a tool that only ranks alerts genuinely low-risk?
It is low-risk in the sense that matters for classification: it imposes no penalty and makes no adverse decision, so it does not carry the impact assessment and appeal obligations that a penalizing system does. It is not consequence-free. It determines who receives examination attention, and that distribution can skew by firm size, geography, or customer base. Monitor concentration in your leads, and measure what the ranker pushes down as deliberately as you measure what it raises.
Why do enforcement actions involving AI keep finding the same things?
Because the findings are about governance rather than technology. Across cases the recurring defects are inadequate governance of model choices, insufficient fairness testing, and poor adverse action documentation, and these produce penalties regardless of how sophisticated the model was. Sophistication has never been a defense. What has protected institutions is a documented record showing that the model choices were deliberate, tested, monitored, and explainable to someone entitled to disagree with them.
Skill.re