AI in Public Safety
Chief Deputy Aisha Morrow runs emergency management for a county of 1.4 million people that spans coastline, wildland, and a dense urban core. During a wildfire two summers ago, her team made evacuation calls off a patchwork of phone tips and a map updated by hand. People in one canyon were told to leave hours after the fire had already cut their only road out. After the fire, a vendor offered her an AI platform that fused satellite, weather, and traffic data into live evacuation recommendations. It was genuinely impressive. It was also the kind of tool that, in the wrong configuration, could send 40,000 people the wrong direction with total confidence. Aisha's challenge was not whether to use AI. It was how to use it without surrendering judgment to it.
Public safety is the most ethically charged arena for government AI, because the same technology that speeds a fire response can, applied to policing, encode decades of bias into who gets stopped, watched, or arrested. That is why the framing here is deliberately split: AI is a powerful ally in response and a tool requiring extreme restraint in enforcement. A leader who blurs that line will eventually face a lawsuit, a wrongful outcome, or both. This lesson gives you the evidence base to hold the line and the vocabulary to defend it in public.
Why the stakes here are concrete and asymmetric
Public safety is the highest-stakes domain for government AI because it combines coercive state power, meaning arrest, detention, surveillance, border enforcement and emergency response prioritization, with life-safety consequence, meaning false arrest, denied disaster assistance and missed rescue. It does this in systems affecting people who rarely have the ability to choose another provider or walk away. A dissatisfied customer switches vendors. A person subject to a face-recognition search does not have that option, and neither does a household waiting on a damage assessment.
The asymmetry runs in both directions and both directions hurt. When a damage-assessment model undercounts destruction in a census tract after a hurricane, federal assistance can be misrouted for years. When a predictive-policing model over-weights historically over-policed census blocks, the feedback loop is not theoretical, it is arrests of real people. Agencies deploying AI in policing, corrections, homeland security, emergency management, fire services, maritime safety, border security, transportation security and public health emergency response operate under a dense constellation of authorities, and each of them places boundaries on what an AI system may do, who is accountable when it fails, and how an affected individual obtains redress.
The source list of those authorities is worth reading aloud in a governance meeting, because it makes the point that public-safety AI is not governed by AI policy alone. It names the Fourth and Fourteenth Amendments of the United States Constitution, the Civil Rights Act of 1964, the Posse Comitatus Act, the Stafford Act for disaster response, the Homeland Security Act, the Aviation and Transportation Security Act, the Support Anti-Terrorism by Fostering Effective Technologies Act of 2002, the First Step Act at the federal corrections level, state criminal procedure codes, state and local consent decrees, and, for federal law enforcement, the rights-impacting provisions of the Office of Management and Budget memo M-24-10.
Emergency management: AI's strongest case
This is Aisha's domain and where AI saves lives with the least controversy, because the decision is about routing resources rather than judging a person. AI fuses satellite imagery, weather models, sensor data and live traffic into a single picture faster than any human team. It can predict a wildfire's spread, model a flood's path block by block, and recommend which neighborhoods to evacuate in what order so that roads do not gridlock. It can scan post-earthquake satellite imagery to find collapsed buildings and direct search teams. Federal and state emergency management have moved toward this kind of data fusion precisely because minutes matter and no human team integrates that many feeds in real time.
The federal version is further along than most county emergency managers realise. FEMA runs post-disaster imagery triage, including damage proxy models built with Oak Ridge National Laboratory, integrated into National Response Framework workflows. FEMA's Office of Response and Recovery has reported that imagery-triage tools shorten initial damage estimates after hurricanes from weeks to under seventy-two hours. That is a genuine and large gain. It is also where the asymmetry bites, because errors in that pipeline translate directly into denied Individual Assistance applications and delayed rebuild grants that can cascade through a household's finances for years. Speed and accuracy are not the same virtue, and the second one is the one people live inside.
Hazard prediction is the adjacent win. Satellite plus machine learning pipelines are used across the United States Geological Survey, the National Oceanic and Atmospheric Administration and CAL FIRE, and AI supports the Wildfire Crisis Strategy run by the U.S. Forest Service. The discipline that keeps all of this safe is simple to state and hard to hold: AI recommends, the incident commander decides, and uncertainty travels with every recommendation. Aisha's tool should never trigger an evacuation order on its own. It should produce a ranked recommendation with a confidence range, and a human who knows the canyon makes the call.
Fire services, 911 and dispatch
Beyond wildfire, AI predicts which buildings face the highest fire risk so inspectors can prioritize, optimizes station placement and crew dispatch, and analyzes sensor data for early detection. These are resource-allocation uses with low civil-rights exposure and they ship comfortably. Emergency communications is a harder call than it first appears. Machine learning triage of incoming 911 calls by likely severity, natural-language transcription, and drone-as-first-responder programs of the kind operated in Chula Vista and Brookhaven all sit closer to the line, because dispatch prioritization decides who gets help first.
That is not a rhetorical point. Emergency-response dispatch prioritization is named in M-24-10 alongside law-enforcement AI as a category carrying minimum practices. If your county is scoping a computer-aided dispatch model, scope it as a rights-impacting and safety-impacting system from day one, define the monitoring regime before the pilot, and write down the rollback trigger. Expect the workforce question too: a triage model is frequently read by staff as a headcount argument, and if you have not answered that question honestly you will be answering it in a grievance instead.
Law enforcement: the oversight-first frontier
Here the framing inverts. In policing, the default posture is restraint, and the leader's primary job is oversight rather than adoption. The reason is documented rather than ideological. Predictive policing models forecast where crime will occur or who is likely to offend, and they have a well-established failure mode. Trained on historical arrest data, such a model learns where police made arrests before, not where crime actually happened. Because over-policed neighborhoods generated more arrests, the model sends more police there, which generates more arrests, in a feedback loop that launders historical bias as objective prediction.
The record is specific. The RAND Corporation's 2016 evaluation of Chicago's Strategic Subject List documented the pattern, and the Chicago Office of Inspector General flagged it again in a 2020 follow-up report. Los Angeles terminated its PredPol contract. Santa Cruz abandoned predictive policing in 2020 and surfaced the ethical tradeoffs publicly while doing so. New Orleans ran a predictive-policing engagement with Palantir that became a case study in what happens when a program is not disclosed to the city council that would have had to authorize it. Products in this space have included PredPol, later Geolitica, and HunchLab. Several cities abandoned these systems after exactly the outcome the theory predicts.
None of this means the technology is uniformly banned. It means a public-safety leader holds these uses to the strictest standard: never as the sole basis for any action against a person, always with human judgment and documented independent evidence, always with transparency to the public and an audit trail, and frequently with the conclusion that a given use should not be deployed at all. The right posture is to put the burden of proof on the tool rather than on the citizen. Public-safety AI deployed without community input has a documented failure pattern, running from the Los Angeles contract termination through Santa Cruz to the controversies around the Allegheny County Family Screening Tool in Pittsburgh.
Facial recognition: what the record actually shows
Facial recognition technology carries documented accuracy gaps across demographic groups and has already produced wrongful arrests of innocent people. The measurement comes from the government's own laboratory. The NIST Face Recognition Vendor Test Part 3, on demographic effects, published as NIST IR 8280 in December 2019, found false-positive differentials of up to two orders of magnitude across demographic groups on some algorithms. That is a hundredfold difference in how often the system falsely matches a person, depending on who the person is.
The consequences have names. Robert Williams was wrongfully arrested by the Detroit Police Department in January 2020, in what the source records as the first publicly documented United States arrest based on a false face-recognition match, and at least six additional publicly known false-match arrests have followed, disproportionately affecting Black men. Nijeer Parks in New Jersey and Randal Reid in Louisiana and Georgia in 2022 are two of them. Together they establish that face-recognition false positives are not edge cases. A leader at this level should be able to name these cases in public, say specifically what went wrong in each, and explain how the proposed governance avoids the same failure mode.
Deployment reaches well beyond municipal policing. Face recognition is used through the FBI's Next Generation Identification Interstate Photo System, in Customs and Border Protection's Biometric Entry/Exit program, and in state-level systems. Transportation Security Administration face-verification pilots at more than twenty-five airports and the CBP program have both drawn sustained congressional and Government Accountability Office scrutiny, including GAO-22-106154, for oversight gaps. Jurisdictions have answered differently: San Francisco banned the technology for city use in 2019, Portland, Oregon followed in 2020, and Vermont enacted its own restriction, while Washington State chose an accountability-report regime instead.
Be precise about what the controls buy you. A required human review step, an independent corroboration rule and an in-situ accuracy test are all worth having, and none of them is a guarantee against misidentification. Every wrongful arrest named above happened with human officers in the loop. Testing bounds the error rate you have agreed to live with; it does not remove the error. Treat these controls as reducing the frequency and improving the recoverability of a wrong outcome, and design the appeal and disclosure path on the assumption that a wrong outcome will occur anyway.
Risk assessment, benefits enforcement and the due-process baseline
Risk scoring tools reach the courtroom and the benefits office, and the case record there is older and more developed than the face-recognition record. The COMPAS pretrial risk assessment tool was examined by ProPublica in 2016 and defended by Northpointe, now Equivant. Its enduring lesson is mathematical rather than partisan: competing fairness definitions, predictive parity on one hand and equal false-positive rates on the other, cannot all be satisfied at once. The choice among them is a policy decision that belongs with elected officials, not with a vendor's default configuration. Risk assessment instruments including COMPAS have been challenged on due-process, equal-protection and transparency grounds, notably in State v. Loomis, decided in Wisconsin in 2016.
Michigan's MIDAS unemployment fraud system, deployed from 2013 to 2015, is the canonical United States cautionary tale for automated decisions of this kind. The state later acknowledged that the system wrongly accused tens of thousands of claimants of fraud, at rates later reported above 90 percent error on the contested determinations, with minimal human review. It is not, strictly, a public-safety system. Every principle it violated, meaning inadequate testing, no meaningful human in the loop, and no appeal pathway calibrated to the stakes, is exactly what M-24-10 now requires for public-safety AI.
The European record is where the jurisprudence is most developed. The Netherlands' SyRI welfare risk-scoring system was struck down by the District Court of The Hague in February 2020 under Article 8 of the European Convention on Human Rights, and the Dutch childcare benefits scandal, the toeslagenaffaire, brought down the Rutte III cabinet in January 2021. Both stand for the proposition that algorithmic risk scoring in an enforcement context demands a dramatically higher evidentiary standard than a commercial setting would require. Closer to home, the Houston Independent School District's value-added teacher evaluation system was invalidated by a federal court in 2017 after the Houston Federation of Teachers suit, and the Ohio Bureau of Motor Vehicles' pandemic-era automated license suspensions underscored the same baseline: automated is not a defense to constitutional process.
Gunshot detection and the money that is already spent
Gunshot-detection systems illustrate the hardest version of the governance problem, which is a system that is already installed. ShotSpotter, now SoundThinking, has been deployed in Chicago, New York and other jurisdictions and has been the subject of ongoing litigation and Office of Inspector General reports questioning accuracy and community impact. The MacArthur Justice Center published a 2021 Chicago analysis with false-positive rate findings, and the City of Chicago decided in 2024 to let its contract lapse. Other jurisdictions point to counter-evidence from their own deployments.
The scenario a senior leader actually faces is narrower than the public debate. Suppose a mid-sized department procured a gunshot-detection contract in 2019 using asset-forfeiture funds, and the question on your desk is whether to condition the next Byrne Justice Assistance Grant award on demonstrating M-24-10 minimum practices. You have to weigh the published analyses against the counter-evidence, and you have to do it inside the existing authorities of that grant program rather than the authorities you wish it had. That constraint, not the merits, is usually what decides these cases, which is a reason to insert governance review into the procurement pathway rather than the renewal pathway.
The policy stack that governs public-safety AI
Leaders at this level operate fluently across seven overlapping instruments, and knowing which one a given objection lives in is most of the skill. Executive Order 14110, signed on October 30, 2023, on the safe, secure and trustworthy development and use of artificial intelligence, directed federal departments to issue guidance for AI in law enforcement. The Office of Management and Budget memo M-24-10, issued in 2024, designates law-enforcement AI, emergency-response dispatch prioritization and critical-infrastructure monitoring as both safety-impacting and rights-impacting. That classification triggers minimum practices including pre-deployment testing, impact assessment, ongoing monitoring, public notice, and the right to a human alternative and appeal. Waivers require Chief AI Officer approval with notice to OMB.
The remaining five fill in procurement, method and disclosure. OMB memo M-24-18 layers vendor disclosure, testing evidence and post-award monitoring obligations onto acquisition. The NIST AI Risk Management Framework 1.0, published in January 2023, together with its Generative AI Profile, NIST AI 600-1 of July 2024, structures the Govern, Map, Measure and Manage functions and is the default reference for agency AI governance boards. It is voluntary guidance rather than binding regulation, which is exactly why agencies adopt it by policy and then have to enforce it themselves.
Three more instruments handle disclosure and method. The DHS AI Use Case Inventory, published annually under Section 7225 of the fiscal year 2021 National Defense Authorization Act and OMB guidance, forces disclosure of every AI use case and has become a primary source for oversight reporting. 28 C.F.R. Part 23 continues to govern federally funded criminal intelligence systems and is interpreted by the Department of Justice's Office of Justice Programs to reach AI-augmented intelligence products. For biometrics specifically, the 2023 Department of Justice interim guidance on facial recognition and the NIST FRVT benchmarks establish the accuracy and operational-use expectations that plaintiffs, defense counsel and inspectors general will cite in court and in audits.
The practical implication is a sequence rather than a checklist. Never deploy rights-impacting public-safety AI without a completed AI Impact Assessment documenting statutory authority, use-case scope, affected populations, disparate-impact testing tied to the NIST AI RMF MEASURE function and Department of Justice guidance, and a human-in-the-loop protocol wherever facial recognition, predictive assessment or automated triage informs coercive action. Meeting those requirements is a floor for deployment, not a certificate of fairness. A system that cannot meet them is not ready; a system that meets them can still be wrong about a person, which is why the appeal path matters as much as the assessment.
State, local and comparative frameworks
State-level frameworks vary widely enough that a multi-jurisdiction program needs a map rather than a policy. Washington State's face-recognition law at RCW 43.386 requires accountability reports. Illinois's Biometric Information Privacy Act governs private-sector biometrics and reaches government vendors downstream, including Clearview AI. California's AB 1215 imposed a moratorium on face recognition in police body cameras and expired in 2023. New York City's Local Law 144 and the city's AI playbook set procurement-side disclosure requirements. Seattle's Surveillance Ordinance is the reference model for proactive community engagement rather than minimum legal compliance.
For agencies with international partners there is a comparative yardstick worth understanding. The European Union AI Act, which entered into force in August 2024, designates law enforcement, migration and border control AI as high-risk or prohibited under specific conditions. It does not bind a United States agency, and it is not a retrospective standard against which older domestic deployments should be judged. It is useful as a statement of where a large peer jurisdiction drew its lines, and as a preview of the questions an international partner will ask before sharing data with you.
Institutional architecture: who does what
A credible public-safety AI program needs a clear division of labor, because diffuse accountability is how these programs fail quietly. The Chief AI Officer, required under M-24-10 for CFO Act agencies and strongly advisable for state and local governments operating at scale, owns the governance program and signs waivers. The AI Governance Board, typically chaired by that officer with general counsel, the chief information officer, the chief information security officer, the privacy official, a civil-rights or equity lead, the mission owner, and a line attorney from the inspector general's office as an observer, reviews new use cases, sets minimum testing and monitoring standards, and approves waivers. Independent red-teaming, consistent with NIST AI 600-1, tests both model behaviour and operator workflow, which is the part agencies skip.
Outside the agency, the counterweights are what make the system honest. Inspectors general, operating at federal agencies under the Inspector General Act of 1978 as amended and at state and local levels where equivalents exist, audit against agency-defined standards and produce the public record that Congress, courts and journalists rely on. GAO reports are the external audit trail: GAO-21-518SP set out an AI accountability framework, GAO-22-106154 examined CBP biometrics, and GAO-23-105923 addressed federal law-enforcement facial recognition. Civil-society organizations including the ACLU, the Electronic Frontier Foundation, the Electronic Privacy Information Center, the NAACP Legal Defense Fund, the Leadership Conference on Civil and Human Rights, Upturn, the AI Now Institute and the Center for Democracy and Technology supply the outside pressure.
The vendor ecosystem is neither enemy nor ally. Companies including Axon, Motorola Solutions, Palantir, Clearview AI, IDEMIA, Thomson Reuters CLEAR, NEC, Dataminr and SoundThinking are counterparties, to be procured against under FAR Part 39 and OMB M-24-18 with exit rights, delivered testing evidence and FedRAMP Moderate authorization where software as a service is involved. Treating a vendor as a partner in governance is a category error. Treating one as an adversary is equally unhelpful. Write the obligations into the contract, because that is the only place your leverage is durable.
Five choices that define the program
Five decisions define a public-safety AI program, and each is political rather than technical. First, deployment posture: ban certain uses outright, impose a moratorium with a sunset, or permit with oversight. Each is defensible and each has costs. Second, disclosure posture: minimum legal compliance, a full use-case inventory plus impact assessments as the M-24-10 baseline, or proactive community engagement on the Seattle model. Third, testing regime: vendor-provided evidence only, independent pre-deployment red-teaming, or continuous monitoring with statistical early-warning triggers.
Fourth, human oversight design, where the vocabulary matters in court. Reviewer-on-the-loop means a human reviews flagged outputs. Reviewer-in-the-loop means a human must concur before action is taken. Reviewer-after-the-loop means a human reviews only a sample after the fact. For facial recognition and other rights-impacting uses, the M-24-10 default is reviewer-in-the-loop with documented independent corroboration before any enforcement action. Fifth, sunset and reauthorization: no sunset, periodic reauthorization by the legislative body, or performance-contingent continuation. Resist the temptation to treat any of these as a technical question. Each requires legitimacy built through stakeholder engagement rather than imposed by fiat.
Scenarios you will actually face
Three more situations recur often enough to rehearse. In the first, a damage-assessment pipeline trained on pre-event and post-event imagery misses entire hollows in a mountainous county because cloud cover defeated the classifier and manual teams could not reach the area. Your questions are how to adjust the rule that triggers human re-inspection, and how to communicate that adjustment without destroying field trust in a tool that is otherwise working. The answer that holds is to publish the change and the reason, because a silent tightening reads as a cover-up when it eventually surfaces.
In the second, a state fusion center asks for access to a face-recognition search against driver's license photos to identify a suspect in a violent-crime investigation, in a state with no explicit enabling legislation for the technology. The federal interim guidance sets a minimum bar and leaves judgment on specific queries to agencies. You need a written decision procedure before the request arrives: who signs off, what is logged, what is retained, and what is disclosed to the defense under Brady obligations. In the third, a dispatch vendor proposes a severity-triage model and the union reads it as a staffing reduction. Scope the pilot under the rights-impacting requirements, define the minimum acceptable monitoring regime, and name the rollback trigger in the pilot charter rather than after the first bad call.
A public-safety AI deployment gate
| Use | Category | Posture | Required control |
|---|---|---|---|
| Disaster modeling and evacuation routing | Safety-impacting | Adopt with discipline | Recommends only; commander decides; uncertainty shown |
| Damage assessment from imagery | Safety-impacting | Adopt with a re-inspection rule | Documented trigger for human re-inspection; appeal path for denied assistance |
| Fire-risk prioritization and dispatch optimization | Resource allocation | Adopt | Human confirms; ships comfortably |
| Search-and-rescue imagery analysis | Safety-impacting | Adopt | Directs teams; humans verify |
| 911 and dispatch severity triage | Safety-impacting and rights-impacting | Pilot under full minimum practices | Monitoring regime and rollback trigger defined before launch |
| Predictive policing and place-based forecasting | Rights-impacting | Strong restraint; frequently decline | Bias-tested; never sole basis; public transparency; sunset date |
| Facial recognition for enforcement | Rights-impacting | Strongest restraint; frequently decline | Never sole basis; independent corroboration; audit trail; disclosure to defense |
| Pretrial or risk-assessment scoring | Rights-impacting | Strongest restraint | Fairness definition chosen by elected officials; explainable in court; contestable |
What Aisha does
Aisha adopts the evacuation tool on her terms. It produces ranked recommendations with confidence ranges, it never issues an order automatically, and she runs a tabletop exercise in which her commanders practice overriding it when local knowledge conflicts with the model. She documents the configuration so an after-action review can reconstruct every recommendation the system made and what was done with it. She writes down, in the pilot charter, the performance level at which the tool comes out of the loop.
On the policing side of her county's broader AI conversation she advocates the opposite posture, treating any enforcement AI as unproven until tested and audited, because the cost of a wrong evacuation route and the cost of a wrongful arrest are different kinds of harm and only one of them takes away a person's liberty. She also does the unglamorous work early: she identifies the three parties most likely to sue or investigate her, usually a civil-rights organization, the inspector general and a legislative oversight committee, and briefs them before they read about the system in the press. That sequencing is the difference between an oversight relationship and an oversight event.
Anti-patterns
- The vendor demo trap. A compelling demonstration in ideal conditions, with good lighting, curated imagery and a balanced dataset, becomes the basis for procurement. Operational performance on body-worn camera footage in urban night conditions is far worse, as the NIST FRVT demographic testing and independent academic work consistently find. Require the vendor to submit NIST FRVT results for the specific algorithm version and require in-situ testing on your own data before deployment.
- The "it is just a tip" defense. Officials claim the AI output is only a lead and therefore needs no oversight comparable to evidence. This collapses at trial and under Brady and Giglio obligations. In the Detroit case and in the New Orleans predictive-policing engagement, the lead in practice became the primary basis for action.
- Treating human review as a guarantee. A required concurrence step is a control, not a shield. Every documented wrongful face-recognition arrest occurred with officers in the loop. Human review reduces the frequency of a wrong outcome and improves your ability to catch it; it does not prevent one, and any briefing that says it does is setting up the next incident.
- The procurement shortcut. Using asset-forfeiture funds, exigent sole-source authorities or pilot carve-outs to avoid competitive procurement and governance review. The money is real and the exemption from scrutiny is temporary.
- The federated fig leaf. Calling a system federated or de-identified, without cryptographic or statistical backing, in order to justify reduced oversight. A label is not an architecture.
- The measurement trap. Tracking alerts generated or matches returned rather than end-to-end outcomes: lawful arrests, convictions, errors avoided, and demographic disparity in who the system touches. Output volume is the easiest metric to collect and the least informative one to report.
- The public commitment trap. A political leader announces a system before the governance review, and everyone below is then told to make it work. The defense is structural: insert governance review into the announcement pathway rather than after it.
- Buying the fairness definition from the vendor. Predictive parity and equal false-positive rates cannot both be satisfied. If nobody in your agency chose which one the system optimises, a vendor's default chose it for you, and the choice is a policy decision that belongs to elected officials.
Practice prompts
- List every public-safety AI system your agency operates or contracts for, then cross-check the list against your jurisdiction's published AI use case inventory or the federal equivalent. The gap between the two lists is your first finding.
- For each system, classify it as safety-impacting, rights-impacting, both or neither, and write down the reasoning. Assume you will defend that reasoning to an inspector general who has already read the vendor's marketing.
- For each safety-impacting or rights-impacting system, document the minimum practices: impact assessment, pre-deployment testing evidence, ongoing monitoring, public notice, and the human alternative and appeal path. Identify which are missing and build a remediation plan with dates against each.
- Draft the public-facing notice a citizen would see if they were subject to your highest-stakes system, in plain language, and test it on people who do not work in government. If they cannot say what happened to them and what to do about it, it is not notice.
- Identify the stakeholders most likely to sue or investigate you and brief them before they learn about the system from the press. Then set a reauthorization date and an early-warning metric, with the rule that if the metric breaches, the system stops until governance reapproves it.
Reflection
Consider the last time your organization decided that a control was sufficient. Somebody said the human review step would catch it, or that the corroboration policy meant a match could never be the sole basis for an arrest, and the room moved on. Now go and find out what that control looks like in the middle of a busy night shift, when the reviewer is under time pressure and the system has already produced a name with a confidence score attached. The distance between the policy and that moment is the actual risk you carry.
Then ask the legitimacy question, which is harder than the technical one. Every posture in this lesson, from an outright ban to permissive adoption with oversight, is defensible on the evidence. What is not defensible is choosing one without telling anyone. The programs that survived contact with the public in this record were not the most accurate ones; they were the ones whose owners could explain, in advance and in plain language, what the system did, who it could affect, and how a person harmed by it would get that harm undone.
Glossary
- Safety-impacting AI. Under OMB M-24-10, AI whose output could meaningfully affect human safety, including public-safety dispatch, emergency alerts and critical-infrastructure monitoring.
- Rights-impacting AI. Under OMB M-24-10, AI whose output could meaningfully affect civil rights, civil liberties or access to critical services.
- Minimum practices. The floor of testing, monitoring, notice and appeal required for safety-impacting or rights-impacting AI, waivable only by the Chief AI Officer with notice to OMB.
- NIST AI Risk Management Framework. Voluntary federal guidance structured around the Govern, Map, Measure and Manage functions. It is a reference, not a regulation, and agencies make it binding on themselves by policy.
- FRVT. The NIST Face Recognition Vendor Test, the authoritative accuracy and demographic-differential benchmark, whose Part 3 report on demographic effects is the standard citation for false-positive disparities.
- 28 C.F.R. Part 23. The federal regulation governing operation of federally funded criminal intelligence systems, now read to reach AI-augmented intelligence products.
- FedRAMP. The Federal Risk and Authorization Management Program. FedRAMP Moderate is the default baseline for software as a service handling controlled unclassified information.
- AI use case inventory. The periodic public disclosure of an agency's AI uses, which has become a primary source for oversight and journalism.
- Reviewer-in-the-loop, on-the-loop, after-the-loop. Three distinct oversight postures with different evidentiary weight in court and in audit: concurrence before action, review of flagged outputs, and post-hoc sampling respectively.
- Predictive parity and equal false-positive rates. Two competing fairness criteria for a risk-scoring system that cannot be satisfied simultaneously, making the choice between them a policy decision.
Related lessons
- AI for Mission-Critical Government Functions sets the wider frame for high-consequence deployments.
- AI in Defense and National Security covers the adjacent domain where coercive power and classification interact.
- AI in Infrastructure covers the critical-infrastructure monitoring category named alongside law enforcement in M-24-10.
- Judicial and Legal Implications develops the due-process and evidentiary questions raised by risk assessment and biometric identification.
- Oversight Mechanisms: IG, GAO, Congress covers the external audit relationships described here.
- Rights-Impacting and Safety-Impacting AI Safeguards covers the minimum practices in operational detail.
- Public Reporting and Algorithmic Transparency covers the disclosure posture and the use case inventory.
- Community Engagement in Government AI addresses the legitimacy work that determines whether any of this holds.
Closing
The split this lesson opened with is the thing to carry away. In disaster response, AI's job is to route trucks and people faster than a human can, and the discipline is to keep uncertainty attached to every recommendation and a commander at the point of decision. In enforcement, AI's job is to be doubted, audited and kept far away from any decision to deprive a person of liberty. Governing both with one policy produces either an emergency management program strangled by process or a policing program with none.
The constitutional floor sits above all of it and does not move. Due process and equal protection are not waivable by a model's confidence score, and a tool that cannot be explained to a court cannot be the basis for depriving someone of liberty. Every case in this lesson, from Detroit to Chicago to The Hague, is a version of an institution discovering that principle after the fact. The work available to you is to discover it beforehand, in writing, with a date on it and a name attached.
Key takeaways
- Split response from enforcement. AI is a strong ally in disaster and emergency response and a tool demanding extreme restraint in policing. Never govern them the same way.
- In response, AI recommends and a commander decides. Evacuation, damage assessment and dispatch tools route resources fast, but they must carry uncertainty and must never issue an order on their own.
- Speed is not accuracy. Cutting damage assessment from weeks to under seventy-two hours is a real gain, and an error in that pipeline still becomes a denied assistance application, so build the re-inspection and appeal path alongside the model.
- Predictive policing launders bias. Trained on arrest history, it sends police where police already were. The RAND evaluation of Chicago's Strategic Subject List and the 2020 inspector general follow-up documented the loop that several cities then abandoned.
- Facial recognition has caused wrongful arrests. NIST measured false-positive differentials of up to two orders of magnitude across demographic groups, and Robert Williams' January 2020 arrest was followed by at least six more publicly known false-match arrests, disproportionately of Black men.
- A control is not a guarantee. Human review, corroboration rules and in-situ testing reduce and bound error. Every wrongful arrest in this record happened with a human in the loop, so design for the wrong outcome you will still get.
- Law-enforcement AI is the highest-scrutiny category. M-24-10 makes it both safety-impacting and rights-impacting, triggering impact assessment, testing, monitoring, public notice and a human alternative with appeal.
- Somebody has to choose the fairness definition. Predictive parity and equal false-positive rates cannot both hold. If elected officials did not choose, a vendor default did.
- The constitution outranks the confidence score. Due process and equal protection cannot be waived by a model, and a tool that cannot be explained cannot deprive someone of liberty.
- Put the burden of proof on the tool. In enforcement, assume an AI is unfair until tested and audited. Declining to deploy is a legitimate and frequently correct outcome.
Frequently Asked Questions
Is facial recognition banned for government use?
No, and the picture is jurisdictional rather than national. San Francisco banned it for city use in 2019, Portland, Oregon followed in 2020, and Vermont enacted its own restriction, while Washington State chose an accountability-report regime under RCW 43.386 and California's body-camera moratorium under AB 1215 expired in 2023. Federal use continues in programs including CBP's Biometric Entry/Exit and TSA verification pilots, under congressional and GAO scrutiny. If you operate across jurisdictions, you need a map of these regimes rather than a single policy.
Our vendor says its algorithm tested well. What should we ask for?
Ask for NIST FRVT results on the specific algorithm version you would deploy, not on the product family or an earlier build, and then require in-situ testing on your own data under your own operating conditions. Demonstrations run in good lighting on curated imagery, and operational footage does not resemble that. Add contractual audit rights, published model cards and data sheets, and delivery of disparate-impact results as a deliverable rather than a request.
If a human officer reviews every match, is that sufficient oversight?
It is necessary and not sufficient. The documented wrongful arrests in this record all involved human officers. What raises the standard is documented independent corroboration before enforcement action, which is the M-24-10 default for rights-impacting uses, together with logging of the query, retention limits, disclosure to the defense under Brady obligations, and a review of whether the corroboration is genuinely independent rather than a second look at the same lead.
How does an emergency-management use get classified as rights-impacting?
Through the consequence, not the technology. Emergency-response dispatch prioritization is named in M-24-10 alongside law-enforcement AI, and a damage-assessment model whose output determines who receives individual assistance is deciding access to a critical service. The test is whether the output could meaningfully affect civil rights, civil liberties or access to critical services. Write the classification down with the reasoning, because you will be asked to defend the reasoning rather than the conclusion.
The system was bought before any of this governance existed. What now?
Legacy systems are the common case rather than the exception, and the leverage points are renewal, grant conditions and incident response rather than the original award. Bring the system into the use case inventory, run the impact assessment retrospectively, publish the notice, and set a reauthorization date. Where funding flows through a federal grant program, examine what that program's existing authorities actually permit you to condition, because the answer is usually narrower than the policy conversation assumes.
Does the EU AI Act apply to a United States agency?
Not directly. It entered into force in August 2024 and designates law enforcement, migration and border control AI as high-risk or prohibited under specific conditions. Treat it as a comparative yardstick and as an indication of what an international partner will ask before sharing data with you, not as a standard against which to judge domestic deployments that predate it.
Skill.re