AI in Healthcare Delivery
Dr. James Okafor runs clinical informatics for a regional Veterans Affairs health network serving 380,000 enrolled veterans across 14 facilities. A vendor pitched him a sepsis early-warning model that promised to alert nurses hours before a patient crashed. The pilot looked great in the demo. Six weeks into live use, his charge nurses had quietly started ignoring it, because it fired forty alerts a shift and thirty-eight were false alarms. The model was not wrong, exactly. It was untuned for his patients, and nobody had budgeted for the work of tuning it. The result was worse than no tool: alert fatigue had trained his best clinicians to dismiss the very warnings that might save someone.
Government healthcare is the highest-stakes arena for AI in the public sector, because the agencies here both deliver care and pay for it for more than 150 million Americans. A miscalibrated model does not just waste money. It can deny a sick person coverage or numb a nurse to a real emergency. As a leader, your job is to know which uses are ready, which are clinical decisions a machine must never make alone, and how to deploy without the alert-fatigue trap that ate James's pilot. Everything that follows is organised around that judgment.
The federal health AI landscape you are operating inside
Federal involvement in health AI spans payers, providers, regulators, researchers and public-health authorities, and each arm carries a different authority and a different risk profile. The Centers for Medicare and Medicaid Services operates Medicare, Medicaid, the Children's Health Insurance Program and the Marketplace, touching nearly half the United States population through payment and coverage decisions. The Veterans Health Administration is the largest integrated health system in the country, with more than 170 medical centers and over 1,000 outpatient clinics. Knowing which of these you are standing in tells you which rules bind you first.
The regulators and researchers sit alongside them. The Food and Drug Administration regulates medical devices, including software. The Centers for Disease Control and Prevention conducts surveillance through its Center for Forecasting and Outbreak Analytics and the National Syndromic Surveillance Program. The National Institutes of Health fund biomedical AI research through Bridge2AI, launched in 2022, the All of Us Research Program, and AIM-AHEAD for health equity. The Health Resources and Services Administration runs community health programs, the Indian Health Service serves tribal populations, and the Substance Abuse and Mental Health Services Administration covers mental health and substance use.
Two offices set rules that catch almost every project. The Office of the National Coordinator for Health IT sets electronic health record standards, and its HTI-1 Final Rule, issued in January 2024, includes algorithm transparency requirements for certified health IT. The Department of Health and Human Services Office for Civil Rights enforces HIPAA and Section 1557 nondiscrimination in federally funded health programs, and its 2024 Final Rule reaches patient care decision support tools directly. Coordination runs through the HHS AI Council chaired by the department's Chief AI Officer, and HHS published its first AI use case inventory in 2024.
Three tiers, sorted by what happens if it is wrong
The cleanest way to govern healthcare AI is to sort every use by the cost of an error. That sort tells you how much human oversight you owe, how much validation evidence you should demand before go-live, and how fast you are allowed to move. It is also the sort that survives contact with an inspector general, because it forces you to write down, in advance, what you thought was at stake.
Administrative AI: start here
The lowest-risk, highest-volume win is paperwork, and government health programs drown in it. AI can draft prior-authorization responses, summarize a patient's record before a visit, transcribe a clinician's spoken notes, and triage the mountain of claims and disability filings that federal health and benefits agencies process. The VA and CMS handle tens of millions of claims a year, so shaving even 20 percent off the clerical time on each is enormous capacity returned to care. If an administrative summary is wrong, a human reviewing it catches it before harm reaches a patient. That is why this tier ships first.
Diagnostic AI: human-confirmed only
AI that reads a chest X-ray, flags a suspicious retinal scan, or, like James's tool, predicts a patient deteriorating, is genuinely powerful. Some imaging models match specialist accuracy on narrow tasks. But these are clinical decisions, and the rule is absolute: the model assists, a licensed clinician decides. A diagnostic model has two ways to be useless. Too many false alarms and clinicians stop trusting it. Too few and it misses real cases. Both happen when you buy a model trained on someone else's patient population and never tune it to yours.
The fix is not a better model. It is budgeting for calibration, measuring the false-alarm rate against what nurses will actually tolerate, and tuning until an alert means something again. Notice what that implies for procurement: the tuning work is a line item, not a favour you ask of an already stretched informatics team after go-live. Every deployment plan that treats calibration as a post-launch activity is describing James's pilot before it happened.
Coverage and eligibility AI: the rights-impacting tier
When AI touches who gets Medicare or Medicaid coverage, or how a disability claim is decided, you have entered the most regulated and most dangerous tier. A model that recommends denying coverage, or that sets a care-hour limit, directly affects a person's access to care. These uses are squarely rights-impacting, and they have already produced litigation when agencies and insurers let algorithms override clinician judgment on coverage. Never let a model deny care alone. The person on the other end of that decision needs a named human to appeal to, and you need a record showing that human actually looked.
FDA regulation of AI and machine learning devices
The FDA regulates medical devices under the Federal Food, Drug, and Cosmetic Act, and software qualifies as a device when it is intended for a medical purpose. That category is called Software as a Medical Device, or SaMD. AI raises a genuinely novel regulatory problem here, because models are not static objects: retraining, fine-tuning and distributional shift change a device's behaviour after it has been cleared. A rule written for a scalpel does not obviously fit a thing that quietly becomes different next quarter.
The agency has worked the problem in public. Its April 2019 discussion paper on a proposed regulatory framework for modifications to AI and machine learning based SaMD introduced the Predetermined Change Control Plan concept, and a January 2021 Action Plan updated that thinking. In October 2021 the FDA published Good Machine Learning Practice Guiding Principles jointly with Health Canada and the United Kingdom's Medicines and Healthcare products Regulatory Agency. April 2023 guidance then operationalised Predetermined Change Control Plans under Section 3060 of the 21st Century Cures Act.
The Good Machine Learning Practice principles are worth reading in full, because they read like a procurement checklist. The source describes ten foundational principles and names these: multi-disciplinary expertise applied throughout the life cycle; sound software engineering and security practices; representative clinical study participants and datasets; independence of training and test data; deliberate consideration of the human in the loop; testing that demonstrates device performance under clinically relevant conditions; clear information for users; performance monitoring of deployed models; and clear communication of updates. As of 2024, more than 900 AI-enabled medical devices had been authorized through the 510(k), De Novo or premarket approval pathways. If you are procuring one, confirm the authorization, understand the predicate device it was cleared against, and read the labeling for intended use and intended population, because that population may not be yours.
HIPAA and the three places your AI touches patient data
HIPAA has three rules that matter most for AI. The Privacy Rule governs uses and disclosures of protected health information. The Security Rule specifies administrative, physical and technical safeguards for electronic PHI. The Breach Notification Rule governs what happens after an incident. The useful mental model is that an AI project touches PHI in three distinct places, and each one creates its own obligation: the training data, the inference inputs, and the outputs that become part of the medical record.
For training on PHI, the Privacy Rule permits uses for health care operations broadly, while research uses typically require institutional review board approval with a waiver or authorization, or de-identification first. De-identification can follow either the Safe Harbor method or Expert Determination, and both have real limitations for AI, because high-dimensional health data frequently remains re-identifiable. Treat de-identification as risk reduction, not as a guarantee that a record can never be traced back to a person. For inference, any vendor processing PHI needs a Business Associate Agreement under 45 CFR 164.504(e). For outputs that land in the record, model cards and intended-use documentation should travel with them.
Two more layers sit on top. The 2024 Office for Civil Rights Final Rule under Section 1557 extended nondiscrimination requirements to patient care decision support tools, including AI, requiring covered entities to make reasonable efforts to identify and mitigate discrimination risk. And state law adds complexity that federal programs routinely underestimate: the California Confidentiality of Medical Information Act, the Texas Medical Records Privacy Act and the New York SHIELD Act each impose requirements beyond HIPAA. Substance use disorder records carry stricter re-disclosure rules under 42 CFR Part 2.
Case: the sepsis model that did not travel
James's problem is not unique to James, and the best-documented version of it is the Epic Sepsis Model, deployed at hundreds of hospitals in the United States. Epic marketed the model with a reported area under the receiver operating characteristic curve of 0.76 to 0.83 for predicting sepsis. In June 2021, Wong and colleagues published a study in JAMA Internal Medicine evaluating the model at the University of Michigan and found substantially lower performance in practice: an AUROC of 0.63, sensitivity of 33 percent at typical operating thresholds, and a numbers-needed-to-evaluate ratio that produced significant alert fatigue.
Read that sensitivity figure carefully, because this is where healthcare AI briefings go wrong. Sensitivity of 33 percent means the model identified roughly a third of the actual sepsis cases at those thresholds. It is not a false-alarm statistic, and it does not become one when someone subtracts it from 100. Keep sensitivity, specificity and positive predictive value labelled with their own definitions in every slide you show a clinical committee, because the moment one of them is renamed, the committee is reasoning about a different model than the one you bought.
The performance gap relative to vendor claims was attributed to differences between the development and deployment populations, concept drift, and changes in clinical workflow. Epic subsequently worked with institutions to adjust the model and released improvements. Five lessons carry over to federal program management. Vendor-reported metrics must be independently validated before deployment. Performance depends heavily on local population, record configuration and workflow, not on the model alone. Alert fatigue is a safety concern, not a usability complaint. Ongoing monitoring with published real-world performance is the standard of care. And systems that run their own clinical decision support pipelines, including the VA and the Military Health System on MHS GENESIS, should never assume a vendor model is production-ready without local validation.
Case: what happens when the training data is not real
IBM Watson for Oncology was marketed as an AI system to assist oncologists in treatment planning, developed in collaboration with Memorial Sloan Kettering Cancer Center and deployed at MD Anderson Cancer Center beginning in 2013. Internal documents reported by STAT in 2017 and 2018 indicated the system had been trained on hypothetical cases rather than real patient data, and physicians at multiple sites reported unsafe and incorrect treatment recommendations. MD Anderson terminated the relationship in 2017 after spending more than $60 million. IBM subsequently wound the product down, and it is cited here as history, not as a tool you could buy.
The transferable lessons are unglamorous and expensive. Data provenance is not optional, and training on synthetic or hypothetical data without rigorous external validation is a design defect rather than a shortcut. Marketing claims deserve scrutiny against clinical evidence, every time. Expensive failures damage institutional trust long after the product is retired, which is a cost that never appears in the contract. Ordinary procurement safeguards, including pilot evaluation, independent validation and staged deployment, would likely have surfaced the problem far earlier. The Department of Veterans Affairs explicitly cited this failure when designing the validation requirements for its National AI Institute.
Case: the proxy label that under-served an entire population
This is the single most important governance point in healthcare AI, and it is not optional. Clinical AI trained on historical data inherits historical inequity, and the canonical demonstration is precise enough to teach from. In October 2019, Obermeyer, Powers, Vogeli and Mullainathan published an analysis in Science of a widely used chronic care management algorithm affecting more than 200 million Americans. The algorithm, attributed to Optum, ranked patients for high-risk care management by predicted future healthcare costs.
Because Black patients had historically received less spending for the same level of medical need, a product of systemic inequity in access and utilization, using cost as a proxy for need systematically understated those patients' needs. The model was not malicious. Its label was wrong, and nobody had tested for it. The authors estimated that changing the label from cost to actual illness burden would have more than doubled the share of Black patients automatically identified for additional care. Optum and the authors subsequently collaborated on remediation.
So before any clinical or coverage model goes live, demand stratified validation and prove the accuracy holds across race, ethnicity, sex, age, language, disability, geography and the intersections of those groups. A model that is 94 percent accurate overall but 80 percent accurate for one group is not a 94 percent accurate model. It is a model that fails a specific community, and you are accountable to that community. Risk stratification algorithms at CMS, in state Medicaid programs and at the VA should be audited specifically against this pattern, because Section 1557 of the Affordable Care Act applies to algorithmic discrimination in federally funded health programs and the 2024 Office for Civil Rights Final Rule codifies that expectation.
REACH VET: what a disciplined deployment looks like
It is worth studying a federal deployment that worked, because the contrast with James's pilot is the whole lesson. REACH VET, which stands for Recovery Engagement and Coordination for Health, Veterans Enhanced Treatment, was launched nationally by the VA in 2017. It is a statistical model that uses electronic health record data to identify veterans at elevated risk of suicide, overdose and other adverse events. The top 0.1 percent of veterans by risk score are flagged for proactive outreach by local clinicians. Academic evaluation, including work by Kessler and by Bossarte and colleagues, found reductions in healthcare utilization associated with the intervention.
Look at the design choices rather than the headline. It was developed by VA researchers using in-house data and validated in house. It operates at a very high precision threshold instead of a low-threshold alert, which is the exact opposite of the forty-alerts-a-shift configuration that broke trust in James's network. It is paired with a specific clinical workflow, outreach by a provider who already knows the veteran, rather than left as standalone automation. It is monitored through performance reporting and has been evaluated in peer-reviewed publications. Treat it as a case study in responsible deployment, not as a template to copy, because its strength comes from those choices and not from the algorithm.
Payment integrity, fraud detection and the provider on the other side
CMS uses AI and advanced analytics extensively. The Fraud Prevention System, operated through the Center for Program Integrity, screens Medicare fee-for-service claims for fraud indicators. The Predictive Learning Analytics Tracking Outcomes system and related tools support risk adjustment for Medicare Advantage. The Government Accountability Office has audited this territory repeatedly, including reports in 2019 and 2023 on Medicare Advantage risk adjustment errors, noting both the benefits and the concerns. The CMS Innovation Center has separately piloted AI for care quality measurement and beneficiary outreach.
Four issues recur in payment integrity AI and they are worth naming before a program office discovers them the hard way. Disparate impact: algorithmic selection can concentrate audits on particular geographic or provider populations. Due process: providers must have a genuine opportunity to contest. Adversarial adaptation: providers adjust billing practices in response to detection algorithms, so a model's measured performance decays for reasons that have nothing to do with the code. Data quality and concept drift: claim coding changes over time. Under the Office of Management and Budget memo M-24-10, issued in 2024, CMS is responsible for classifying its AI uses and applying the minimum practices where they are rights-impacting or safety-impacting.
Public health surveillance and drug safety
Two more uses deserve their own note. In drug safety, AI can scan adverse-event reports across millions of records to detect a dangerous drug interaction far faster than manual review, which is exactly the kind of signal the FDA hunts for. In public health, models can forecast disease spread and target outbreak response. Both are screening tools: they surface a signal that a human epidemiologist or pharmacologist then confirms. The risk in both is overconfidence in a model that may be reacting to a data artifact rather than a real signal, so uncertainty must travel with every output rather than being stripped out for the briefing slide.
The CDC's Center for Forecasting and Outbreak Analytics, established in 2021 with $200 million in American Rescue Plan funding and a $50 million annual appropriation, integrates machine learning with traditional epidemiology for respiratory disease forecasting, wastewater surveillance analysis and outbreak detection. The National Syndromic Surveillance Program aggregates de-identified emergency department data from hospitals nationwide, supporting anomaly detection, genomic epidemiology such as SARS-CoV-2 variant tracking, and vaccine effectiveness monitoring. Persistent challenges include data-sharing limits with state and local public health, variable record quality, privacy frameworks under HIPAA and state public health exceptions, and the need to adapt models quickly during a novel outbreak.
The design lessons that COVID-19 pushed into that center's architecture are the ones worth borrowing: durable data pipelines rather than emergency improvisation, academic partnerships through Insight Net, publicly available code and evaluation, and integration with the established communication pipeline of the Morbidity and Mortality Weekly Report. If you are building surveillance AI, the published materials are a better starting point than a vendor architecture diagram.
Health equity and access as a design requirement
Federal healthcare AI has to serve every population it touches, and the obligations here are specific rather than aspirational. Section 1557 of the Affordable Care Act prohibits discrimination on the basis of race, color, national origin, sex, age or disability in federally funded health programs, and the 2024 Office for Civil Rights Final Rule extends that explicitly to patient care decision support tools. The NIH AIM-AHEAD program funds AI and health equity research at minority-serving institutions.
Particular populations require particular attention at design time. The Indian Health Service and tribal epidemiology centers serve Native American populations with distinct data sovereignty considerations grounded in tribal sovereignty principles and the CARE Principles for Indigenous Data Governance. Rural health, reached through HRSA programs and Medicare rural adjustments, requires deliberate inclusion in training and validation populations rather than an assumption that a national model generalises. Populations with limited English proficiency are covered under Executive Order 13166 and need translation and culturally appropriate design. Accessibility obligations under the Americans with Disabilities Act and under Section 508 apply to AI interfaces. Specify equity evaluation criteria up front, and gather community input through federally qualified health centers, tribal governments and patient advocacy organizations while the design is still changeable.
Clinical governance: the committee that actually holds
Responsible healthcare AI needs governance structures beyond generic agency AI governance, because the questions are clinical. A Clinical AI Oversight Committee should include medical leadership, informatics leadership, quality and safety, nursing leadership, pharmacy, patient advocacy, bioethics, legal, privacy, security and data science and engineering. That is eleven seats, counted from the source list, and each of them exists because a real deployment failed for the lack of it. Physician informaticists should serve as embedded owners for each clinical AI system, so that every model has a name attached to it.
The committee's pre-deployment review should demand seven artifacts: intended use, clinical validation, a fairness audit across relevant subgroups, an implementation plan, a monitoring plan, an incident response plan, and deactivation criteria. That last one is the one agencies forget. If you cannot say in advance what performance would make you turn the system off, you have not made a decision, you have made a purchase. After deployment the committee should receive periodic performance reports and incident reports on a schedule, not on request.
Patient advisory participation should be structural rather than token, and the whole apparatus needs to align with CMS Conditions of Participation, Joint Commission standards and institutional bylaws. Give explicit attention to alert fatigue, documentation burden, and liability allocation between the algorithm developer and the clinician, which is the question your general counsel will ask first. The anti-information-blocking rules of the 21st Century Cures Act apply, and state medical board guidance increasingly addresses AI liability and scope of practice.
Procurement: where the governance is either won or lost
Healthcare AI procurement in federal practice has to integrate the Federal Acquisition Regulation, HIPAA Business Associate Agreements, FedRAMP for cloud services, FDA authorization where applicable, and clinical acceptance criteria that a clinician wrote. Common vehicles include the VA's FedRAMP Moderate cloud environment, Indian Health Service and CMS enterprise agreements, GSA Multiple Award Schedule special item numbers for AI, and Defense Health Agency vehicles for military medicine.
Seven contract terms do most of the protective work. Data rights: state who owns fine-tuned weights trained on your PHI. Training data restrictions: the vendor may not use agency PHI to train shared models without specific authorization. Performance evidence: require delivery of real-world performance metrics stratified by subgroup, not a single headline number. Change control: require notification and re-validation on any material model update. Exit: require data return and destruction at contract conclusion. Incident response: bind notification obligations to HIPAA Breach Notification timelines. Transparency: require model cards, data statements and intended-use documentation as deliverables. The Department of Defense Business Associate Agreement template and the Office for Civil Rights sample BAA are reasonable starting points. Research uses may implicate Certificates of Confidentiality under 42 U.S.C. 241(d), and FDA-regulated devices carry additional Quality System Regulation requirements under 21 CFR 820.
A clinical AI deployment rubric
| Gate | Question | Pass condition |
|---|---|---|
| Tier | Administrative, diagnostic, or coverage? | Classified in writing; oversight scaled to consequence |
| Human authority | Who makes the final clinical or coverage call? | Named licensed human; the model never decides alone |
| Local validation | Was it validated on the population it will serve? | Local evidence, not vendor-supplied metrics alone |
| Calibration | What is the false-alarm rate against clinician tolerance? | Tuned to your population; alert fatigue measured |
| Equity | Does accuracy hold across all relevant subgroups? | Stratified validation passed and documented |
| Label integrity | Is the predicted label clinically grounded? | No cost or utilization proxy standing in for need |
| Privacy | Is PHI protected across training, inference and output? | Business Associate Agreement; approved systems only |
| Regulatory | Is it an FDA-regulated device? Rights-impacting? | Authorized as required; impact assessment completed |
| Appeal | Can a patient contest an AI-influenced decision? | Documented human appeal path, published in plain language |
| Deactivation | What performance turns this off? | Written criteria and an owner empowered to use them |
What James fixes
James does not need a new vendor. He needs to treat his sepsis tool as the medium-stakes diagnostic it is: pull it back to two pilot units, fund a clinical informaticist to tune the alert threshold to his own patient population, measure the false-alarm rate weekly, and re-expand only once nurses trust the alerts again. The REACH VET pattern gives him the argument to make to his leadership, because a high-precision threshold paired with a named clinician's workflow is a defensible design and a forty-alert shift is not.
Meanwhile he ships the safe administrative wins today. Note summaries and prior-authorization drafting return clinician hours immediately, they carry a human check by construction, and they buy him the institutional credibility to insist that the high-stakes tool gets the calibration it always needed. He also writes down, before anything expands, the performance level at which the sepsis model gets switched off. That single paragraph is what separates a governed deployment from a purchase with hope attached.
Anti-patterns
- Deploy and forget. Procurement ends at go-live, with no monitoring plan, no incident response and no owner. Every failure in this lesson had this component. Monitoring detects only what you chose to watch, so choose deliberately and write it down.
- Trusting the vendor's metric. Deploying on vendor-published performance without local validation. A metric measured on one population is evidence about that population and nothing more.
- Demographic blindness. Reporting a single overall accuracy number and never stratifying by race, ethnicity, disability, language, age, geography or their intersections. An aggregate number can hide a group-specific failure indefinitely.
- Cost as a proxy for need. Predicting spending, utilization or contact volume when access inequities exist, then treating the output as a measure of sickness. The label is the model.
- Alert fatigue by configuration. Setting a low alert threshold so nothing is missed, which trains clinicians to dismiss everything. A high recall setting that nobody reads has a real-world sensitivity of zero.
- Missing Business Associate Agreement. A vendor touches PHI without one. This is an enforcement matter, not a paperwork lapse.
- Shadow AI. Clinicians or administrators pasting PHI into unapproved consumer AI tools because the approved path is slower. If the sanctioned tool is unusable, you have created this problem yourself.
- Human review as a rubber stamp. A required confirmation click is not oversight. If the reviewer sees only the model's conclusion, has no time budgeted, and faces no consequence for agreeing, the human in the loop is a signature, not a safeguard.
- No deactivation criteria. No written standard for turning the system off when it underperforms, which means it will stay on through its own failure.
Practice prompts
- Take one clinical or coverage AI system your organization operates and classify it into the three tiers. Write the reasoning in a paragraph you would be willing to defend to an inspector general, then have a clinician who was not involved read it and tell you where it is wishful.
- Pull the performance evidence for that same system and identify, for each headline number, which metric it actually is. Sensitivity, specificity, positive predictive value and overall accuracy answer different questions. Flag every place the definition has drifted between the vendor's document and your own briefing materials.
- Write the deactivation criteria for your highest-stakes deployed model in one page: the metric, the threshold, who watches it, how often, and who has the authority to switch the system off without asking permission.
- Audit one risk stratification or prioritization model for label integrity. What is it actually predicting? If the answer involves cost, utilization or prior contact with the system, work out who is under-counted by that proxy and what it would take to test the gap.
- Draft the plain-language notice a patient would receive if an AI system influenced a decision about their care or coverage, including how to reach a human and how to appeal. Test it on people who do not work in healthcare and see whether they can say what happens next.
Reflection
Think about the last clinical or administrative tool your organization adopted and ask who owned it once the launch team moved on. In most agencies the honest answer is nobody, because the project team disbanded at go-live and the operational owner inherited a system they did not choose and cannot evaluate. That gap is where alert fatigue accumulates, where drift goes unnoticed, and where a model that has quietly stopped working keeps producing outputs that people keep acting on.
Then ask a harder question about your own incentives. Every failure described in this lesson was visible to somebody inside the organization before it became public. What would it cost, in your setting, for a nurse or a caseworker to say out loud that the tool is not working? If the answer is that it would cost them something, your monitoring plan is decorative, because the earliest and best signal you have is a frontline clinician's judgment, and you have priced it out of reach.
Glossary
- Software as a Medical Device (SaMD). Software intended for a medical purpose, regulated by the FDA as a device in its own right rather than as a component of hardware.
- Predetermined Change Control Plan. A pathway allowing a manufacturer to specify in advance how a cleared AI device may be modified, so that anticipated retraining does not require a new submission.
- Protected health information (PHI). Individually identifiable health information governed by HIPAA, which an AI project touches in training data, inference inputs and outputs that enter the record.
- Business Associate Agreement (BAA). The contract required under HIPAA before a vendor may process PHI on a covered entity's behalf, setting safeguard and breach obligations.
- De-identification. Removal or transformation of identifiers under the Safe Harbor or Expert Determination methods. It reduces re-identification risk; it does not eliminate it, particularly for high-dimensional health data.
- AUROC. Area under the receiver operating characteristic curve, a single summary of a model's ability to rank cases above non-cases. It says nothing on its own about performance at the threshold you actually run.
- Sensitivity. The share of actual positive cases the model identifies. Distinct from specificity, from positive predictive value, and from overall accuracy.
- Proxy label. A measurable variable used to stand in for the thing you actually care about. When the proxy is generated by an inequitable process, the model inherits that inequity.
- Alert fatigue. The degradation of clinician response caused by a high volume of low-value alerts. It is a patient safety issue, not a user experience complaint.
- Rights-impacting AI. Under OMB M-24-10, AI whose output could meaningfully affect civil rights, civil liberties or access to critical services, triggering minimum practices including impact assessment, monitoring, notice and appeal.
- Stratified validation. Measuring and reporting model performance separately for each relevant subgroup rather than reporting a single pooled figure.
Related lessons
- AI in Social Services applies the same eligibility and rights-impacting patterns to human services programs.
- AI and Equity: Reaching All Communities develops the stratified validation and access obligations sketched here.
- Rights-Impacting and Safety-Impacting AI Safeguards covers the minimum practices in operational detail.
- Privacy Impact Assessments for AI Systems is the companion for the HIPAA and PHI questions.
- Evaluating AI Vendor Claims is the direct remedy for the Epic Sepsis and Watson failure modes.
- Human-in-the-Loop: Design and Implementation covers how to keep clinical review from decaying into a rubber stamp.
- AI for Mission-Critical Government Functions sets the wider frame for high-consequence public sector deployments.
Closing
Healthcare is where the general principles of government AI stop being abstract. A rights-impacting classification is a real person's coverage. A monitoring plan is whether anyone notices that a model has drifted away from the patients it was built for. Stratified validation is whether an entire community gets referred to the care it needs. The frameworks matter here precisely because the consequences are individual and often irreversible.
The pattern across every case in this lesson is the same and it is not technical. Someone accepted a number without asking what population produced it, what label it predicted, or what would happen at the threshold they actually intended to run. Ask those three questions of every model your organization deploys, write the answers down, name the human who owns each one, and decide in advance what would make you switch it off. That discipline is available to you today, costs almost nothing, and would have prevented most of what went wrong here.
Key takeaways
- Sort every use by the cost of an error. Administrative AI ships first, diagnostic AI is human-confirmed, and coverage and eligibility AI is the rights-impacting tier carrying the heaviest oversight.
- The model assists; a licensed human decides. No model denies care or makes a clinical call alone, and a patient must always be able to appeal to a person who actually looked.
- Validate locally or do not deploy. Epic Sepsis performed far below its marketed AUROC at a single independent site. A metric measured on one population is evidence about that population.
- Keep metric definitions straight. Sensitivity, specificity, positive predictive value and accuracy answer different questions, and a renamed metric produces a decision about a model you did not buy.
- Interrogate the label. The Optum chronic care algorithm predicted cost and was read as predicting need, which under-referred Black patients at national scale. Ask what your model actually predicts.
- Calibration is a safety issue. A tool that cries wolf trains clinicians to ignore it. REACH VET works partly because it runs at a very high precision threshold with a defined clinical follow-up.
- Stratified validation is mandatory. Prove accuracy holds across race, ethnicity, sex, age, language, disability and geography, and treat Section 1557 as binding on algorithmic outcomes.
- HIPAA governs three places, not one. Training data, inference inputs and outputs in the record each carry obligations, and any vendor touching PHI needs a Business Associate Agreement.
- Write the deactivation criteria before launch. If no stated performance level turns the system off, nothing will.
Frequently Asked Questions
Does FDA authorization mean a model is safe for my patients?
No. Authorization means the device met the agency's requirements for its stated intended use and intended population, evaluated against a predicate or a de novo pathway. It is not a statement that the device performs equivalently in your setting. Read the labeling for the intended population, compare it to yours, and plan local validation regardless. More than 900 AI-enabled devices had been authorized as of 2024, across a very wide range of evidence quality.
If our vendor signs a Business Associate Agreement, are we covered on privacy?
A BAA is necessary and not sufficient. It binds the vendor to safeguard obligations and breach notification, but it does not tell you whether your PHI is being used to train shared models, whether you own the resulting weights, or what happens to the data at contract end. Those are separate contract terms and you have to write them. Section 1557, state medical privacy laws and 42 CFR Part 2 obligations sit on top of HIPAA and are not addressed by a BAA at all.
Our model was validated in a peer-reviewed study. Is that enough?
It depends entirely on where the study was run. The Epic Sepsis Model carried vendor-reported AUROC of 0.76 to 0.83 and measured 0.63 at the University of Michigan, with 33 percent sensitivity at typical thresholds. Published evidence tells you the model can work somewhere. Local validation tells you whether it works here, on your population, in your workflow, with your record configuration.
How do we test for bias when we do not collect complete demographic data?
Incomplete demographic data is a finding, not an excuse. Document what you can measure, state plainly what you cannot, and treat the gap as a risk to be closed rather than a reason to skip the analysis. In practice you will often find proxies for the missing dimension in geography or program enrollment, and you should be equally careful that those proxies do not themselves encode the pattern you are trying to detect.
Where should an agency with limited capacity start?
Start in the administrative tier, where a human already reviews the output before it affects anyone. Document summarization, prior-authorization drafting and queue triage return real capacity, carry a natural human check, and build the institutional muscle for governance before the stakes rise. Use the credibility earned there to insist on proper validation budgets when a diagnostic or coverage tool arrives.
Who should own a clinical AI system after launch?
A named physician informaticist or equivalent clinical owner, reporting into a standing oversight committee that receives scheduled performance and incident reports. Ownership that lives with the project team disappears at go-live. Ownership that lives with a clinician who uses the system stays attached to the thing that matters, which is whether the tool is still helping the people in front of them.
Skill.re