←
AI for Government
Visionary · M14 · lesson 14 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI in Social Services
📖
now learning

AI in Social Services

15 min

Marcus Delacroix runs human services for a county of 1.1 million people. He has 280,000 active benefits cases, a caseworker vacancy rate of 22 percent, and a child welfare hotline that received 41,000 calls last year. In a budget hearing, a county commissioner asked him a simple question: can artificial intelligence help us do more with the staff we have? Marcus said yes, and he meant it. But a second commissioner asked a different question: what happens when the computer is wrong about a family? Marcus did not have a good answer that day. Social services is the part of government where a wrong automated decision does not cost a customer a refund. It can cost a child a safe home, or an elderly woman her heating in January.

As an agency leader, Marcus is responsible for both the efficiency gain and the human cost, and he cannot trade one against the other in the way a commercial operator can. This lesson is about holding those two things at once: moving fast enough to relieve a caseload that is genuinely hurting people through delay, while refusing to move fast in the places where being wrong is not recoverable.

The pressure Marcus is under is real

It is worth taking the efficiency argument seriously before taking it apart, because a leader who treats it as merely a temptation will lose the room and deserve to. Marcus has 280,000 active cases and a caseworker vacancy rate of 22 percent, which is not an abstraction to the households in those cases. Every week a benefits application sits in a queue is a week somebody is deciding between rent and medicine. Every hour a screener spends assembling a file by hand is an hour not spent on judgment. The commissioner who asked whether AI could help do more with the staff on hand was asking a legitimate question about a real harm.

The 41,000 hotline calls make the same point from the other direction. A volume like that is not handled by hiring alone in any budget Marcus is going to get, and the screeners working it are making consequential decisions under time pressure with incomplete information. Anything that puts relevant history in front of a screener faster is straightforwardly good. What is not straightforwardly good is anything that converts that history into a number and puts the number where the judgment used to be. Holding that distinction is most of the job, and it is easiest to hold when you have already conceded that the pressure driving the other direction is genuine.

The stakes are not symmetric

Most AI business cases assume errors balance out. A few false positives, a few false negatives, and the net benefit is positive. Social services breaks that assumption, because the two kinds of error fall on different people and carry wildly different costs. The arithmetic that works in a commercial setting, where the same customer experiences both error types and the aggregate is what matters, does not describe a benefits program at all.

Consider child welfare screening. A false negative, meaning the model fails to flag a genuinely at-risk child, can end in a death that makes the front page and a federal civil rights investigation. A false positive, meaning the model flags a safe family for investigation, subjects parents to a traumatic and intrusive intervention they did not deserve, and disproportionately falls on poor and minority households. Neither error is acceptable, and you cannot minimize both at once. Every threshold you set is a moral choice about which family bears the risk.

That is worth saying plainly to your own leadership, because the threshold conversation is usually held as a technical one. Somebody proposes a cutoff, somebody else asks what it does to the alert volume, and the meeting moves on. What actually happened in that meeting is that a group of people decided how many safe families would be investigated in order to reduce how many at-risk children were missed. If that decision was not made by someone accountable for it, and recorded as a decision rather than as a configuration value, it was still made. It was just made by whoever set the default.

Why this is not a routine IT procurement

The National Institute of Standards and Technology AI Risk Management Framework, a voluntary federal standard for governing AI risk, calls these high-impact or rights-impacting systems. Voluntary is the operative word: it is a reference an agency adopts by policy and then has to enforce on itself, which means the framework only ever works as well as the person inside the organization who is willing to invoke it.

Federal AI policy in the Office of Management and Budget memo M-24-10, issued in 2024, goes further than a reference. For rights-impacting uses, agencies must run impact assessments, test for disparate outcomes, provide human review of decisions, and give people a way to appeal. Benefits eligibility and child welfare sit squarely on that list. The important structural point is that these obligations attach to the consequence rather than to the technology. A simple rules engine that decides eligibility carries them. A sophisticated model that only sorts a queue may not. Classify by what the output does to a person, write the reasoning down, and expect to defend the reasoning rather than the label.

Treat those four requirements as a floor rather than a certificate. An agency can complete an impact assessment, run a disparate-outcome test, staff a review step and publish an appeal path, and still harm a family, because each of those controls has a failure mode of its own. The assessment covers the harms someone thought of. The disparate-outcome test measures the groups you had data for. The human review is only as real as the time and authority the reviewer actually has. The appeal path only works for people who receive the notice, understand it, and can act on it. Meeting the minimum is the beginning of the safety argument, not the end of it.

Sequence by risk, not by ambition

Marcus's instinct to start with the riskiest use case, child welfare prediction, was wrong, and it is the most common instinct in the field because that is where the moral urgency is. The smart sequence starts where AI reduces friction without making consequential decisions about people, and earns trust before touching anything rights-impacting. This is not timidity. It is how an organization builds the governance muscle, the monitoring habits and the political credibility it will need when a genuinely high-stakes proposal arrives.

There is a second reason to sequence this way, which is about your own staff. Caseworkers and screeners will decide very early whether the agency's AI program is something being done to them or something built with them, and that judgment is remarkably durable. A first project that visibly removes drudgery makes the second project possible. A first project that arrives as a scoring system pointed at their professional judgment makes every subsequent one a negotiation.

Four domains where AI actually helps

Benefits administration: lower risk, high payoff

Most benefits backlogs are not decision problems. They are paperwork problems. AI that reads uploaded documents, extracts income figures, checks for missing pages and routes complete applications to the right queue can cut processing time dramatically without ever deciding who is eligible. The key design rule is simple enough to put in a procurement document: the AI assembles and triages; a human decides. A document-classification model that says an application is missing a pay stub is a clerical assistant. A model that says to deny a household is a decision-maker, and must be governed as one.

Hold that line at the level of the actual workflow rather than the org chart, because the line moves quietly. A triage model that sorts applications into likely-approve and likely-deny queues has not decided anything on paper, and has substantially decided things in practice if the second queue is worked last by the most overloaded staff. When you review a system, follow a case through the process as it is actually staffed, and ask where a human exercised judgment that could have gone the other way.

The upside of getting this tier right is larger than it looks on a slide. Document handling and completeness checking are exactly the work that consumes the caseworker capacity Marcus does not have, they fail in ways a person catches in the ordinary course of reviewing an application, and they produce a measurable change in processing time that a commissioner can understand without a briefing on machine learning. They also generate the operational evidence, the monitoring habits and the incident history that the agency will need the first time a genuinely rights-impacting proposal reaches the table.

Child welfare: highest risk, assist and never decide

Predictive risk models in child welfare are the most scrutinized AI in government, and for good reason. The defensible posture is narrow: use AI to help a trained screener gather and surface relevant prior history faster, and never to generate a risk score that drives the screening decision on its own. The human screener must be able to see why a case was surfaced, override it, and have that override recorded.

Recording overrides is where the design becomes real. An override capability that nobody uses is not oversight, it is a disclaimer, and the override rate is the single most informative number your program will produce. If it is near zero, either the model is extraordinary or your screeners have concluded that disagreeing with it is not worth the trouble, and you should find out which before a case goes wrong. If it is high and consistently in one direction, that is a finding about the model rather than about the screeners.

Housing and homelessness

Coordinated-entry systems prioritize scarce housing slots, and AI can help match a person's documented needs to available units and flag people who have fallen out of contact. Both of those are genuinely useful. But prioritization formulas encode value judgments about who counts as most vulnerable, and those judgments belong to elected and appointed officials rather than to a vendor's default model. If your jurisdiction has never had that debate in public, adopting a tool settles it privately, and the settlement will be discovered later by whoever was ranked lower.

The flagging use is more comfortable than the ranking use and deserves separating from it in your governance. Identifying someone who has fallen out of contact prompts an outreach attempt, and the cost of being wrong is a phone call to a person who did not need one. Ordering a queue for scarce units allocates a resource that somebody will not receive. Those two functions frequently arrive inside the same product, and a program that has approved the first has often not noticed that it approved the second.

Food security

For programs like the Supplemental Nutrition Assistance Program, AI is best used to detect likely errors in the agency's own work, meaning duplicate records, data-entry mistakes and cases stalled in a queue, rather than to detect fraud in applicants. That distinction sounds like a preference and is closer to a rule. Fraud-detection models point suspicion at the public and have a track record of false accusations against vulnerable people. Error-detection models point the same analytical capability at the agency, where the organization can absorb being wrong, and where the correction improves service instead of triggering an accusation a household has to defend itself against.

Notice and appeal, as they actually work

The appeal requirement is the control most often satisfied on paper and least often in practice, so it deserves its own treatment. A household can only contest a decision it knows was made, understands the basis of, and can act on within its own circumstances. That means the notice has to say plainly that an automated system was involved, what it did, and what the person can do next, in language and in a channel that reaches them. A policy that grants a right of appeal while the letter that goes out is unreadable has created a right that only the well-resourced can exercise, which is the opposite of what the requirement is for.

The second half of the problem is on your side of the counter. An appeal path is only real if the person handling the appeal can actually reach a different conclusion, which means they need to see the reasons behind the original output rather than only the output, and they need to be somewhere other than the queue that produced it. Track how many appeals you receive and how many succeed, and treat a very low volume as a question rather than as evidence that the system is working. In a program serving households under this much pressure, silence is at least as likely to mean the notice never landed as it is to mean the decision was right.

The bias problem, in plain terms

The most common failure in social services AI is not a software bug. It is the model faithfully learning patterns from biased historical data. If your county investigated poor neighborhoods more often in the past, a model trained on that history will recommend investigating them more often in the future, and will present that recommendation as objective. Nothing malfunctioned. The system did exactly what it was built to do, which was to reproduce the pattern in the record.

This is why "the model is not biased, it just reflects the data" is not a defense. Reflecting the data is the mechanism of harm, not an exoneration of it. And it is why bias work cannot be delegated entirely to a technical team: the question of whether past investigation patterns represent risk or represent enforcement history is a substantive question about your county's own conduct, and the people who can answer it are caseworkers, community organizations and the families themselves, not the people holding the training set.

Three questions to ask of any model

Marcus's team needs to ask three things of any vendor or in-house model, in plain language, and to keep asking until the answers are documents rather than assurances. The questions are deliberately non-technical, because the point is not to out-argue a data scientist. It is to establish whether anyone in the room can describe what the system learned from, how it performs for the communities the agency serves, and what a caseworker is able to do about an output they believe is wrong. A vendor who cannot answer those three in writing has told you something important about how the product was built.

  • What did it learn from? If the training data reflects past human decisions, it will inherit past human bias. Ask to see how the training population breaks down by race, income, geography and disability status, and treat an inability to answer as an answer.
  • Does it perform equally across groups? Demand error rates broken out by protected class, not a single overall accuracy number. A model that is 90 percent accurate overall can be 70 percent accurate for one community, and the aggregate figure will never reveal that.
  • Can a caseworker understand and override it? If the model is a black box that staff cannot question, it is not decision support. It is an unaccountable decision-maker wearing a badge it did not earn.

Be precise about what a passing answer buys you. Disparate-impact testing measures the gaps you thought to look for, in the groups your data lets you see, at the time you ran it. It does not certify that the system is fair, and it says nothing about a group your records do not distinguish. Run the test on a fixed schedule rather than once at launch, publish what you found, and treat every category you cannot measure as an open risk rather than as a clean result.

A usable artifact: the social services AI deployment rubric

Before any social services AI goes live, run it through this rubric. Any red answer means the system is not ready for the public. A row of green is a precondition for deploying, not a warrant that the system is safe, and the rubric is most useful in the argument it forces a program team to have out loud before anything ships.

DimensionGreen (deploy)Yellow (deploy with controls)Red (do not deploy)
Decision authorityAI assists; human decides every caseAI scores; human reviews flagged casesAI decides automatically, no human in the loop
ReversibilityWrong output is easily correctedCorrection requires escalationWrong output causes irreversible harm before review
Disparate impact testingTested across protected groups, gaps documented and closedTested, small gaps with a mitigation planNot tested, or large unexplained gaps
Explainability to staffCaseworker can see and explain the reasonReason available on requestBlack box, no rationale
Appeal path for the publicClear notice and human appeal existAppeal exists but is unclearNo way for a person to contest the result
Override loggingEvery override recorded and reviewedOverrides allowed, partially loggedNo override possible

What Marcus did

Marcus took this rubric into his next commissioner hearing, which changed the character of the conversation because he was no longer answering a question about technology. He was presenting a decision procedure. His benefits document-triage tool scored green across the board and went live, cutting average application processing from 18 days to 6. His proposed child welfare risk score scored red on decision authority and explainability. He paused it, restructured it as a screener's research aid, and committed to publishing its disparate-impact testing before any pilot.

That is the answer he did not have on day one: the agency deploys where the system can only help, and holds back where it could harm a family without anyone able to explain why. Notice what made the answer credible in a public hearing. It was not the framework and it was not the vendor's evidence. It was that he had applied the same rubric to the project he wanted and to the project he was proud of, and that the rubric had stopped one of them.

The restructuring is worth studying too, because pausing is not the only option available when a system scores red. The child welfare proposal did not die. It became a screener's research aid, which is the same underlying capability pointed at assembling history rather than at producing a score, and it kept the part of the original idea that was genuinely useful to an overloaded hotline. Red on the rubric is a statement about how a capability is being applied, not always a verdict on the capability itself, and the most productive response is usually to ask what version of this would score green.

Governance that survives an audit

The Government Accountability Office's AI Accountability Framework expects four things from agencies using AI: clear governance, sound data, documented performance and ongoing monitoring. For social services specifically that translates into concrete habits. Name an accountable human owner for each model. Keep a register of every model in production and its risk tier. Run disparate-impact tests on a fixed schedule rather than only at launch. And publish, in plain language, what the system does and how to appeal it.

An agency that can hand an auditor those four artifacts on request is in a fundamentally different position from one scrambling after a child is harmed, and the difference is not mainly about the audit. It is that each artifact is a standing question somebody has to answer, which is what keeps a deployed system connected to a person who is thinking about it. Ongoing monitoring detects only what you chose to watch, so the register and the schedule are where you decide, in advance and in writing, what would count as the system going wrong.

Anti-patterns

  • Starting with the highest-stakes use case. Beginning where the moral urgency is greatest rather than where the error is most recoverable. It spends institutional credibility before the agency has learned how to govern anything, and it makes every subsequent project a negotiation with staff.
  • Setting the threshold as a configuration value. Choosing an operating point in a technical meeting without naming it as a decision about which families bear which risk. Whoever set the default made the policy.
  • The triage tool that quietly decides. A model that sorts applications into likely-approve and likely-deny queues has decided nothing on paper and a great deal in practice, if the second queue is worked last by the most overloaded staff.
  • Treating a completed control as proof of safety. An impact assessment covers the harms someone thought of, a disparate-impact test measures the groups your data lets you see, and a review step is only as real as the reviewer's time and authority. A completed checklist has never once made a wrong decision right.
  • The override nobody uses. An override capability with a near-zero usage rate is a disclaimer rather than oversight. Watch the override rate as a first-class metric and investigate both extremes.
  • Buying the prioritization formula. Adopting a vendor's default ranking for scarce housing or services settles a public value judgment privately, and the settlement is discovered later by the people who were ranked lower.
  • Pointing fraud detection at the public. Using the agency's best analytical capability to find suspected wrongdoing in applicants rather than errors in its own work, which places the burden of being wrong on the household least able to carry it.
  • Testing once, at launch. Running disparate-impact analysis as a launch gate rather than on a schedule, so that drift, policy changes and caseload shifts go unmeasured for years.
  • Publishing notice nobody can act on. An appeal path that exists in the policy and not in the letter, written at a reading level and in a language the affected household does not use.

Practice prompts

  • Take one automated or semi-automated process in your agency and follow a single case through it as it is actually staffed today, not as the process document describes it. Mark every point where a human exercised judgment that could have gone the other way. If there are none, you have found an automated decision that nobody classified as one.
  • Write down the operating threshold for your highest-stakes screening or scoring tool, then write the sentence that describes what it does in human terms: how many families in which situation bear which risk. Take that sentence to the person who is actually accountable and ask them to own it.
  • Pull the override rate for any tool where staff can disagree with an output. Investigate it whichever way it points, and treat a near-zero rate as a finding rather than as a success.
  • Run the deployment rubric against a system you already have in production rather than one you are considering. Legacy systems rarely get scored, and they are where the unclassified rights-impacting decisions usually live.
  • Draft the plain-language notice a household would receive when an automated system affected their case, including what happened, what it means, and how to reach a person. Test it with people outside government and revise until they can say what to do next without help.

Reflection

Think about the last time your organization described a control as sufficient. Somebody said a caseworker reviews every flagged case, or that an appeal path exists, and the room accepted it. Now go and look at what that control is at the end of a shift, with a vacancy rate, a queue and a supervisor asking about throughput. The distance between the policy and that moment is the actual risk your agency carries, and it is not visible in any document you would show an auditor.

Then ask Marcus's second commissioner's question about a system you already run: what happens when it is wrong about a family? Follow it all the way through. Who notices, how long does it take, what does the household have to do, what does it cost them in the meantime, and who inside the organization is told. If the answer breaks down at any step, that is where your next piece of work is, and it will do more good than any improvement to the model.

Glossary

  • Rights-impacting AI. A system whose output could meaningfully affect a person's civil rights, civil liberties or access to critical services, which triggers requirements for impact assessment, disparate-outcome testing, human review and appeal.
  • False negative. A case the system fails to flag that should have been flagged. In child welfare screening, a genuinely at-risk child who is not surfaced.
  • False positive. A case the system flags that should not have been. In child welfare screening, a safe family subjected to an intrusive investigation.
  • Threshold. The operating point at which a score becomes an action. Because the two error types fall on different people, it is a policy choice and not a technical setting.
  • Disparate impact testing. Measuring error rates separately by protected group rather than reporting one overall accuracy figure. It measures the gaps you thought to look for, in the groups your data can distinguish.
  • Coordinated entry. The process by which a jurisdiction prioritizes scarce housing resources, whose ranking formula encodes a public judgment about who is most vulnerable.
  • Override rate. How often staff disagree with a system's output. It is the clearest available signal of whether human review is real, and both a very low and a very high rate are findings.
  • Model register. A maintained list of every model in production with its risk tier and accountable owner, which is what allows anyone to answer what the agency is running.

Closing

In commercial AI you optimize for the average case. In social services you are judged by your worst case, which is the one family the system failed. That is not a reason to avoid the technology, because the delay and the backlog in an understaffed benefits program are themselves a harm falling on people who cannot absorb it. It is a reason to be deliberate about where the technology is allowed to touch a decision and where it is only allowed to touch the paperwork.

The discipline that holds all of this together is unglamorous and entirely available to Marcus today. Sequence by recoverability of error. Keep consequential decisions with accountable people who can see the reasoning and disagree with it. Make the threshold a decision somebody owns. Test across groups on a schedule and publish what you find. Write the appeal in language a person can act on. None of that requires a better model, and all of it is what stands between a routine error and a family harmed by a system nobody could explain.

Key takeaways

  • Errors are not symmetric. A false negative and a false positive harm different people in very different ways, so every threshold is a moral choice rather than a technical setting.
  • Somebody owns the threshold. If it was not chosen by an accountable person and recorded as a decision, it was chosen by a default, and the policy was made anyway.
  • Sequence by risk, not by ambition. Start with paperwork and triage uses that assemble information, and earn trust before letting any model touch rights-impacting decisions.
  • The AI assembles; the human decides. Keep consequential decisions about real families with accountable people who can see the reasoning and override it, and check where judgment is exercised in the workflow as actually staffed.
  • Bias is usually inherited, not coded. Models learn from past human decisions, so demand error rates broken out by protected group rather than a single accuracy number.
  • A control is not a guarantee. Impact assessments cover the harms someone imagined, disparate-impact tests measure the groups your data can see, and a review step is only as real as the reviewer's time and authority.
  • Watch the override rate. An override capability nobody exercises is a disclaimer, and both a near-zero and a lopsided rate are findings worth investigating.
  • Point analytics at your own work first. Detecting the agency's errors improves service, while detecting suspected fraud in applicants places the cost of being wrong on the household least able to carry it.
  • Governance is your insurance. A model register, named owners, scheduled bias testing and plain-language public notice are what stand between a routine error and a governance scandal.

Frequently Asked Questions

How do we know whether a system counts as rights-impacting?

Classify by what the output does to a person, not by how sophisticated the technology is. A simple rules engine that determines eligibility affects access to a critical service and carries the obligations. A large model that only sorts a work queue may not. Write down the reasoning behind your classification, because the reasoning is what you will be asked to defend, and a classification with no recorded rationale reads as a conclusion someone wanted.

Our vendor says the model is not biased because it does not use race as an input.

Excluding a protected characteristic from the inputs does not remove it from the model. Geography, prior contact with the agency, income and household composition all carry information about who a person is, and a model trained on past decisions inherits the pattern in those decisions regardless of which columns it was shown. The only way to know is to measure outcomes by group after the fact, which is why disparate-impact testing is an ongoing measurement rather than a design claim.

Is a human review step enough to keep us out of trouble?

It is necessary and not sufficient. A review step becomes a rubber stamp when the reviewer sees only the model's conclusion, has no time budgeted for the task, and faces no consequence for agreeing. Design it so the reviewer sees the underlying reasons rather than only the output, ensure the override is genuinely available, log every override, and review the override rate. If nobody ever disagrees, you do not have oversight, you have a signature.

What should we do about a legacy system that was never classified?

Score it against the rubric now, exactly as if it were a new proposal, and accept that this is where the uncomfortable findings usually are. Legacy systems predate the governance and are frequently making decisions that nobody has ever labelled as decisions. Bring the system into your model register, name an owner, run the disparate-impact test, and publish the notice. If it scores red on decision authority or explainability, the fact that it is already running is not an argument for leaving it alone.

We are badly understaffed. Is it responsible to wait?

Delay is itself a harm, and pretending otherwise is its own failure of judgment. The resolution is not to move slowly everywhere, it is to move quickly in the places where an error is recoverable. Document handling, completeness checking and queue routing return real capacity, carry a natural human check, and can ship while the higher-stakes proposals go through proper review. The sequencing exists precisely so that an understaffed agency can act.

How often should disparate-impact testing be repeated?

On a fixed schedule set in advance and recorded in the model register, rather than in response to a complaint. The reason is that the things which shift a system's behaviour, meaning caseload composition, policy changes, data pipeline changes and model updates, all happen without announcing themselves. A test run only at launch describes a system that no longer exists, and the gap between that description and reality grows silently.