←
AI for Government
Visionary · M3 · lesson 3 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI and Equity: Reaching All Communities
📖
now learning

AI and Equity: Reaching All Communities

15 min

When the state rolled out its new AI-powered benefits chatbot, the dashboard looked like a triumph. Applications rose 30 percent in the first quarter, average processing time dropped, and the press release wrote itself. Then Lila Moreno, the agency's director of program integrity, pulled the numbers apart by county. In the three counties with the highest poverty and the lowest broadband access, applications had actually fallen. The chatbot worked beautifully for people with a smartphone, fast internet, and English fluency. For everyone else, it had quietly become a new wall. The agency had spent $2.4 million making it easier for the already-served to get served, and harder for the people the program existed to reach.

This is the central trap of AI in government. AI optimizes for the average user, and the average user is rarely the person who needs public services most. Reaching all communities is not a feel-good add-on to an AI deployment. It is the test of whether the deployment did its job at all, because a service that does not reach the underserved is a service that failed its mandate. This lesson follows Lila as she turns the chatbot from a wall back into a door, and builds an equity practice any leader can apply.

The Exclusion Patterns, Named

Government serves every citizen; that is the constitutional commitment. AI threatens and reinforces that commitment at the same time, and the threat arrives in a small number of recognizable shapes. Knowing them by name is the difference between discovering an exclusion in an audit and designing it out in a review. Each of the patterns below has produced real harm to real people in real programs.

  • Biometric verification that performs unevenly. When a face verification service cannot recognize darker skin tones reliably, it excludes Black Americans from a service they are entitled to. Federal evaluations have documented accuracy variation across demographic groups.
  • Language coverage that stops at two. An AI voice assistant supporting only English and Spanish excludes speakers of Mandarin, Vietnamese, Hmong, Arabic, Somali, Navajo, and dozens of other languages spoken in households across the country.
  • Mobile-first assumptions. An application flow that assumes smartphone ownership and broadband access excludes seniors, rural residents, and low-income households in one design decision.
  • Historical data as ground truth. A scoring tool trained on historically biased enforcement records reproduces those disparities at scale and calls the output objective.
  • Poverty used as a risk proxy. A screening algorithm that treats receipt of public benefits as a risk signal is penalizing poverty and reporting it as risk.

The affected populations are not an abstraction either. Communities of color, tribal nations, rural populations, seniors, people with disabilities, LGBTQ+ Americans, people with limited English proficiency, immigrants, and low-income households each encounter these patterns differently, and a control that addresses one of them often does nothing for another. That is why "make it accessible" is not a plan.

Notice what these patterns have in common. None of them requires anyone to intend exclusion, none announces itself in an error log, and every one of them is compatible with a system that is performing well on its stated objective. That is what makes them a governance problem rather than an engineering problem. An engineer asked to make a chatbot answer questions accurately can succeed completely and still ship every pattern on this list, because none of them is a failure of the thing that was measured.

The Four Divides AI Can Widen

Lila's diagnosis identified four distinct gaps, each of which an AI system can either bridge or deepen. Naming them separately matters, because each has a different fix.

  • The access divide. Broadband, devices, and digital literacy. An AI tool that assumes a smartphone and fast internet excludes the people least likely to have both.
  • The geographic divide. Rural communities with thin connectivity and few in-person offices. The chatbot's biggest drop was in the most rural counties.
  • The disability divide. AI interfaces that fail with screen readers, voice control, or cognitive accessibility needs.
  • The language divide. Tools built and tested in English that degrade or mislead in other languages, including the machine translations they fall back on.

An AI service that works for the average citizen and fails the marginalized one has not improved the program. It has automated its blind spots, and it has done so behind a metric that says the opposite. That is the quiet part of the problem: exclusion produced this way is invisible to the people who caused it, because their dashboard is genuinely improving while their program is genuinely failing.

Equity in government AI is not only a values commitment. A substantial body of civil rights law already applies to these systems, and it applied before anyone deployed a model. Your general counsel and civil rights officer own the analysis; the leadership point is that these obligations attach to the outcome, not to the technology, and no system architecture exempts a program from them.

Title VI of the Civil Rights Act of 1964 prohibits discrimination on the basis of race, color, or national origin in programs receiving federal financial assistance. Title VII covers employment. The Americans with Disabilities Act and Section 504 and Section 508 of the Rehabilitation Act carry accessibility obligations, with Section 508 the standard most often invoked for electronic and information technology. Executive Order 13166 requires agencies to take reasonable steps to provide meaningful access for persons with limited English proficiency. Executive Order 13175 requires consultation with tribal governments.

Layered on top are the AI-specific instruments. OMB Memorandum M-24-10, issued in 2024, identifies rights-impacting AI and attaches pre-deployment testing, ongoing monitoring, and public notice to it; a citizen-facing benefits system of the kind Lila runs sits squarely in that category. The Blueprint for an AI Bill of Rights articulates an Algorithmic Discrimination Protections principle alongside Notice and Explanation; it is a non-binding blueprint published by the Office of Science and Technology Policy, not a statute or a regulation, and it should be cited as a standard your agency has chosen rather than one imposed on it. Executive Order 13985 on racial equity and Executive Order 14110 on AI shaped agency practice during their periods of effect and are best cited historically; confirm the current status of any executive order before relying on it.

Enforcement is real and already active in this space. The Equal Employment Opportunity Commission has published technical assistance addressing AI hiring tools under Title VII and the ADA. The Justice Department's Civil Rights Division has opened investigations into AI-related discrimination under the ADA and Title VI. The practical implication for a program leader is straightforward: the question an investigator asks is not whether you intended to exclude anyone. It is what your system did, to whom, and what you knew about it.

The Programs That Fund the Other Side of the Divide

Accessibility of a service and availability of the connection are different problems, and a leader who solves only the first has solved the easier one. Several federal programs have addressed the second, and their status is worth checking before you rely on any of them in a plan. The Justice40 Initiative committed that at least 40 percent of the benefits of certain federal investments would flow to disadvantaged communities, and agencies deploying AI for benefits, licensing, or enforcement were expected to measure against it; cite it historically and verify what replaced it in your own program area.

The Digital Equity Act of 2021 funded state digital equity planning. The Broadband Equity, Access, and Deployment program funds broadband infrastructure in unserved and underserved areas, and its coverage determinations rest on federal broadband maps that your own service data can usefully contradict. Household connectivity subsidies have come and gone, and a design that assumed a subsidy still exists is a design that stopped working the day the subsidy ended. The durable lesson is architectural rather than programmatic: do not let a digital divide become an AI divide by assuming the connectivity problem is being solved somewhere else.

Why "Benefits Should Flow to All" Is Harder Than It Sounds

The reflex answer is "make the chatbot accessible." But Lila found the deeper issue. The harm was not only in the chatbot's design; it was in how success was measured. The agency tracked total applications, an aggregate that rises even as specific communities fall. Equity is invisible to aggregate metrics by construction. You cannot manage what you only measure in totals.

Her first structural fix, therefore, was not technical. It was to mandate disaggregated measurement on every consequential AI system: performance broken out by income, geography, language, and disability status. This mirrors the GAO AI Accountability Framework's emphasis on assessing performance across the populations a system affects, not just in aggregate. The day the agency started reporting county-level and language-level outcomes, the equity gap stopped being a discovery and became a standing management number.

Subgroup Evaluation and the Limits of Fairness Metrics

Disaggregation raises an immediate technical question: measured how? Three fairness measures dominate practice and they answer different questions. Demographic parity asks whether selection rates are similar across groups. Equalized odds asks whether error rates are similar across groups. Calibration asks whether predicted probabilities match actual rates within each group. The MEASURE function of the voluntary NIST AI Risk Management Framework is the usual governance anchor for running these evaluations and recording what they showed.

Here is the fact that trips up most programs: no single metric captures all fairness concerns, and the metrics can trade off against each other mathematically, so satisfying one may make another worse. That is not a defect in the field. It is a property of the problem, and it means the choice of metric is a policy judgment about which kind of error the program is least willing to impose on people. Select the measures that fit your specific decision context, document why you chose them, and test empirically rather than assuming.

The four-fifths rule deserves particular care because it is so often misused. It is a screening rule of thumb for flagging adverse impact in selection procedures, not a safe harbor. Clearing it does not establish that a selection procedure is lawful, and failing it does not by itself establish a violation. Treat a passing ratio as a reason to keep looking rather than a reason to stop, and never let a single ratio stand in for the disparate impact analysis your counsel would actually perform.

Design for the Edge, Not the Average

Lila reframed the design goal. Instead of optimizing for the typical user, the team designed for the hardest-to-reach user and let everyone else benefit from the robustness. Three moves carried the chatbot from wall to door.

First, multi-channel by default. The AI assistant gained a phone-based voice version and an SMS path that works on a basic phone with no app and minimal data. The same intelligence, reachable without broadband. Applications in the rural counties recovered within two quarters.

Second, real language access. Rather than trusting raw machine translation, the agency had its top five non-English languages reviewed by human translators and ran the chatbot through native-speaker testing before launch. Machine translation drafts; humans verify the high-stakes paths. A mistranslated eligibility question is not a typo; it is a wrongful denial. Meaningful access for people with limited English proficiency runs wider than a chatbot in any case: it takes translation of vital documents, qualified interpreters, and multilingual digital services, and a translated interface on top of English-only notices and English-only phone support is not language access.

Third, accessibility as a gate, not a patch. The interface was tested against Section 508 standards, the federal requirement that electronic and information technology be usable by people with disabilities, with actual assistive-technology users in the loop. Accessibility that is verified by checklist alone routinely passes audits and still fails real users. Conformance is generally expressed against the Web Content Accessibility Guidelines at level AA, and the curriculum points to WCAG 2.1 AA as the working target; confirm the version your own agency's standards incorporate, and pair whichever you use with assistive technology testing and a staffed human alternative for the cases where the AI fails a disabled user outright.

Community Engagement That Is Not a Checkbox

Every fix above is a design decision made by people who are not the ones being excluded, which is exactly the condition that produced the problem. Engagement is how you close that loop, and it is not a checkbox. Each form of it has its own requirements and its own failure mode.

Tribal consultation under Executive Order 13175 requires government-to-government engagement, not a stakeholder meeting with tribal members present. Those are different things, and conflating them is a substantive failure, not a procedural one. Language access requires the translation, interpretation, and multilingual service described above. Disability access requires conformance, assistive technology testing, and human alternatives. Rural access requires attention to actual broadband availability rather than to the coverage a map asserts. LGBTQ+ inclusion requires attention to how systems handle names, pronouns, and gender markers, which is where automated matching and verification most often fail people. Immigrant inclusion requires attention to fear of government interaction and to alternative documentation pathways for people who cannot produce what the default flow demands.

Community advisory boards are the standing mechanism, and they work only when they have three things: real authority over something, a budget, and term limits with a succession plan. A board with none of those provides cover rather than counsel. Be honest about what engagement discharges, too. Consulting a community does not establish that a system is fair, a translated notice does not establish that language access was achieved, and a completed equity assessment does not establish that a system is non-discriminatory. Each of those is evidence you looked. The outcome data is the evidence of what you found.

What the Case Record Shows

The patterns are not hypothetical, and the cases are worth knowing precisely because leaders reach for them as precedent. In 2022 the Internal Revenue Service adopted a facial recognition service for taxpayer authentication; after civil rights objections from privacy organizations, members of Congress, and advocacy groups, the Treasury Department reversed course. The concerns were accuracy differences across skin tones documented by federal evaluations, exclusion of taxpayers without smartphones, and inadequate alternative pathways.

Michigan's MiDAS system issued tens of thousands of false unemployment fraud determinations. Working-class families, disproportionately Black, were devastated. The system had no adequate appeals pathway and minimal human oversight, and a class action settlement followed. The absence of the appeals pathway is the part to sit with: an automated determination without a real route to contest it converts a model error directly into a life outcome.

The COMPAS recidivism tool was analyzed by ProPublica in 2016, which found it was twice as likely to falsely flag Black defendants as future criminals compared with white defendants. That finding was contested on the grounds that the tool was calibrated across groups, which is precisely the tradeoff described earlier: error-rate parity and calibration can pull against each other, and a vendor and a critic can both be arithmetically correct while disagreeing about what fairness means. The case is the standing illustration that the metric choice is the policy choice.

The Houston Federation of Teachers litigation over an EVAAS-based teacher evaluation model in 2017 established that opaque proprietary scoring used in consequential employment decisions raises serious due process problems, with disproportionate effect on teachers of color. The Dutch childcare benefits scandal, which came to a head in 2021, showed a fraud risk model targeting dual-nationality families, overwhelmingly Moroccan and Turkish Dutch, driving twenty-six thousand families into wrongful fraud accusations and bringing down the national government.

One case runs the other way and is the more useful template. The Allegheny County child welfare screening tool took a different path: published validation, community advisory boards, and a shadow-mode period before production use. Criticism of it remains, and that is the point rather than a caveat. A contested system with published evidence and a real engagement structure is a governable system; an uncontested system with no published evidence is usually just one nobody has examined yet.

An AI Equity Impact Review

Lila institutionalized her lessons as a required review before any citizen-facing AI system launches or materially changes. It is short on purpose, so it actually gets done. Every item demands evidence, not assertion.

  1. Who is the hardest-to-reach user? Name the specific community most likely to be excluded by this system.
  2. Channel coverage. Can someone with a basic phone, no broadband, and no app complete the task? If not, what is the non-digital path?
  3. Language coverage. Which languages are supported, which high-stakes paths were verified by a human translator rather than machine-translated, and are the vital documents and the phone support covered too?
  4. Disability access. Was it tested with real assistive-technology users against Section 508 standards, not just an automated checker, and is there a human alternative when it fails?
  5. Subgroup evaluation. Which fairness measures were run, why those, and what did they show? Record the tradeoff you accepted.
  6. Disaggregated metrics. Will outcomes be reported by income, geography, language, and disability, with thresholds that trigger review?
  7. Accessible appeals. Is there a contest pathway with plain-language notice and explanation, reachable through every channel the system uses, and does anyone actually staff it?
  8. Fallback to a human. When the AI cannot serve someone, is there a clear, staffed path to a person?
  9. Affected-community input. Did members of the underserved community test the system before launch, and did anything change as a result?
  10. Public transparency. Is the system documented in the agency's public AI inventory, with the equity findings available rather than filed internally?
  11. Sunset trigger. What level of measured exclusion would require pausing or redesigning the system?

The Leadership Shift

The hardest part for Lila was not engineering. It was resetting what "success" meant for her leadership team. As long as the agency celebrated the 30 percent aggregate rise, no one had reason to look harder. The shift was to declare that an AI deployment is not successful until it works for the community that needs it most, and to put the disaggregated numbers on the same dashboard as the headline metric. Equity moved from a values statement in the strategic plan to a number the agency reports every quarter. That is what reaching all communities actually requires: not better intentions, but better measurement, multi-channel design, and the discipline to design for the edge of the population rather than its center.

Anti-Patterns to Refuse

  • The assessment as absolution. Completing an equity impact assessment and treating the obligation as discharged. The assessment records that you looked. Whether the system discriminates is answered by the outcome data, and an assessment filed before launch cannot answer a question about what happened after it.
  • The translated notice as language access. Translating the interface and leaving vital documents, phone support, and the appeals path in English. A person who can start an application in their language and cannot contest a denial in it has not been given access; they have been given an entrance.
  • The community meeting as consent. Holding a listening session and treating the community as having endorsed the system. Consultation informs a decision your agency still owns and is still accountable for, and a meeting where nothing changed afterward is a record of attendance.
  • Stakeholder engagement in place of tribal consultation. Inviting tribal members to a general stakeholder session and calling it consultation. Consultation under the tribal consultation executive order is government-to-government, and substituting the easier thing is a substantive failure.
  • The four-fifths pass as a clean bill. Treating a passing ratio as proof that a selection procedure is lawful. It is a screening rule of thumb, not a safe harbor, and it addresses one narrow question about one kind of decision.
  • The automated accessibility checker as accessibility. Running a scanner, collecting a green result, and shipping. Checkers find a subset of defects and cannot tell you whether a person using a screen reader can finish the task.
  • The aggregate as evidence. Reporting a rising total as proof the system serves everyone. Aggregates conceal subgroup collapse by construction, which is exactly how Lila's agency shipped a wall and celebrated it.
  • Automated denial with no route back. Deploying a consequential determination without a staffed, reachable, plain-language contest pathway. The absence of the appeal is what turns an error into a harm, and it is the single most consistent feature of the worst cases in this field.

Practice Prompts

  1. Take the headline success metric for one AI system you own and break it out by geography, language, and disability status. Do not adjust anything yet. Just look at whether any subgroup moved in the opposite direction from the total.
  2. Trace the full contest pathway for an adverse determination in one of your programs, end to end, in a language other than English and on a device without an app. Note every step that breaks.
  3. List the languages your program actually encounters, from your own contact records rather than from census assumptions. Compare that list to the languages your AI tooling supports, and identify which high-stakes paths in the gap have never been reviewed by a human translator.
  4. Choose the fairness measure you would defend for one consequential model, write down why that measure and not the others, and name the tradeoff you are accepting and who bears it.
  5. Audit one community advisory structure in your agency against three questions: what does it have authority over, what is its budget, and what are its terms. If all three answers are vague, you have cover rather than counsel.
  6. Draft the sunset trigger for one live system: the specific measured level of exclusion at which you would pause it. Then check whether you currently collect the data that would let you notice you had crossed it.

Reflection

Think about the last AI system your organization launched and ask who was in the room when its success criteria were set. Were any of them people the program exists to serve rather than people who run it? If the answer is no, the criteria almost certainly encode the experience of the people who were present, and the resulting metric will keep reporting good news for exactly as long as nobody disaggregates it.

Then consider what your agency does when equity evidence arrives late. Lila found her gap because she went looking; nothing in the system told her. If the same gap opened tomorrow in a different program, would anything surface it, and would surfacing it be a career-safe thing for the person who noticed to do? The measurement problem and the culture problem are the same problem, and only one of them can be fixed with a dashboard.

Glossary

  • Disaggregated measurement. Reporting outcomes broken out by population characteristics such as income, geography, language, and disability, rather than only in aggregate.
  • Disparate impact. A facially neutral practice that produces significantly different outcomes for a protected group, which can raise civil rights exposure regardless of intent.
  • Demographic parity. A fairness measure asking whether selection rates are similar across groups.
  • Equalized odds. A fairness measure asking whether error rates are similar across groups.
  • Calibration. A fairness measure asking whether predicted probabilities match actual rates within each group.
  • Four-fifths rule. A screening rule of thumb for flagging adverse impact in selection procedures. It is not a safe harbor, and clearing it does not establish lawfulness.
  • Limited English proficiency. The status of a person who does not speak, read, write, or understand English well enough to interact with a program, triggering meaningful-access obligations.
  • Tribal consultation. Government-to-government engagement with tribal nations, which is distinct from and not satisfied by general stakeholder outreach.
  • Section 508. The federal accessibility requirement, part of the Rehabilitation Act, that electronic and information technology be usable by people with disabilities.
  • Rights-impacting AI. The category, defined in federal AI policy, covering systems whose outputs affect a person's rights, opportunities, or access to critical services, and which attract heightened testing, monitoring, and notice obligations.
  • Equity impact assessment. A pre-deployment review documenting who could be excluded by a system, what was tested, and what will be monitored. It records that you looked; it does not establish that the system is fair.

Closing

Every case in this lesson has the same structure. A system was deployed, it worked for most people, its aggregate numbers improved, and a specific community was quietly moved further from a service it was entitled to. Nobody intended it. In several instances nobody noticed until an outside party looked, and in the worst of them the people harmed had no realistic way to contest what had been decided about them.

The work that prevents this is not exotic. Measure by subgroup and put the result next to the headline number. Build more than one channel, and make the non-digital one real. Verify the high-stakes language paths with humans. Test accessibility with people who use assistive technology. Consult affected communities in the form the obligation actually requires, and let the consultation change something. Staff an appeals path that reaches everyone the system reaches. And decide in advance what level of measured exclusion would make you stop. Do that and equity stops being an aspiration in the strategic plan and becomes what it should have been all along: a number on the dashboard, reported every quarter, that leadership is accountable for.

Key Takeaways

  • AI optimizes for the average user. The average user is rarely the person who needs public services most, so AI can quietly exclude the underserved.
  • Aggregate metrics hide inequity by design. Total counts can rise while specific communities fall; only disaggregated measurement reveals it.
  • Name the four divides. Access, geography, disability, and language each require a distinct fix, not a single "make it accessible" gesture.
  • Civil rights law already applies. Title VI, Title VII, the ADA, and Sections 504 and 508 of the Rehabilitation Act attach to outcomes; no architecture exempts a program from them.
  • Meaningful access is wider than an interface. Limited English proficiency obligations reach vital documents, qualified interpreters, and multilingual services, not just a translated screen.
  • Tribal consultation is government-to-government. A stakeholder meeting with tribal members present is not consultation, and substituting it is a substantive failure.
  • Design for the hardest-to-reach user. Multi-channel by default, including voice and SMS, makes the system robust for everyone.
  • Verify high-stakes language paths with humans. Machine translation may draft, but a mistranslated eligibility question is a wrongful denial, not a typo.
  • Treat accessibility as a gate. Test against Section 508 with real assistive-technology users; passing an automated checker is not the same as working.
  • The fairness metric choice is a policy choice. Parity, equalized odds, and calibration can trade off against each other, so document which you chose and who bears the tradeoff.
  • An assessment is evidence you looked, not proof you are fair. The same is true of a translated notice and a community meeting; outcome data is what answers the question.
  • No consequential determination without a real appeal. The missing contest pathway is what turns a model error into a life outcome, and it recurs in every worst case.

Frequently Asked Questions

We ran an equity impact assessment and it came back clean. Are we covered? No. The assessment documents what you examined before launch, using the assumptions you held at the time. Whether the system excludes anyone is answered by disaggregated outcome data after deployment, and the gap between a clean assessment and a failing system is precisely the space Lila's chatbot occupied for a full quarter. Keep the assessment, then measure.

Our vendor says the model was tested for bias. Is that sufficient? Ask which fairness measure they used, on which population, and what tradeoff they accepted, because a model can satisfy one measure while failing another and both results are honest. Then ask whether the test population resembles yours. Vendor testing is an input to your evaluation, not a substitute for it, and the obligation stays with the agency that deploys the system.

We cannot support every language spoken by our applicants. What is the minimum? There is no universal number, and the requirement is meaningful access rather than a fixed language count, which means the answer depends on your program, your population, and the guidance your agency operates under. What is clear is where to spend first: the paths where a translation error becomes a wrongful denial, the contest and appeals pathway, and the vital documents. Ask your civil rights officer for the analysis your program requires.

Does an AI system need a human alternative, or is a good AI enough? Provide the human alternative. Any system will encounter people it cannot serve, including users whose assistive technology it fails, whose documentation does not match the expected pattern, or whose situation falls outside its training. A staffed path to a person is what keeps those cases from becoming denials, and it needs to be reachable through the same channels the AI uses rather than through a separate one nobody finds.

Is the Blueprint for an AI Bill of Rights something we have to comply with? It is a non-binding blueprint published by the Office of Science and Technology Policy, not a statute or a regulation. Its principles, including Algorithmic Discrimination Protections and Notice and Explanation, are a strong and widely adopted standard, and many agencies have chosen to hold themselves to them. Keep your documentation clear about which of your commitments are legal obligations and which are adopted standards, because conflating them creates confusion in exactly the audit where clarity matters.

Disaggregating by race feels legally risky. Should we do it? This is a question for your civil rights officer and general counsel, and the answer depends on your program's authorities and the data you lawfully hold. What you should not do is treat the discomfort as a reason to measure nothing: geography, language, connectivity, and disability-related access are often available where demographic data is not, and they surface most of the same exclusions. A program that measures no subgroup at all is a program that will learn about its gaps from an investigator.