←
AI for Government
Capable · M24 · lesson 24 of 42 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Measuring AI Impact
📖
now learning

Measuring AI Impact

15 min

Nadia Osei-Mensah spent three years building the Ohio Department of Job and Family Services' AI-assisted benefits-eligibility screener. In the spring budget hearing, a state senator held up a printout and asked her one question: "How do you know it actually helped anyone?" She had deployment logs, model accuracy scores, and a vendor dashboard. None of it answered the question. She left the hearing without her supplemental funding. That story has a thousand variations across state agencies, federal programs, and city halls right now. Government has accelerated AI adoption, but measurement has not kept pace, and without credible evidence of impact an AI project carries the same political vulnerability as any other spending line. It deserves to.

What Measurement Buys You, and What It Cannot

Government is accountable for how it spends money and how it deploys systems. Measurement is how an agency shows that a system works, that public money was well spent, and that constituents benefit. It also drives improvement, because it tells you which parts are working and which need adjustment. Without it you cannot demonstrate return on investment, cannot improve the system on evidence, and cannot make a defensible decision about whether to scale it.

Be careful with the word that usually attaches to this. Measurement does not prove that a system delivered the value it promised. A before-and-after comparison is an attribution made under assumptions, and it is only as strong as the assumptions you can document and defend. What good measurement gives you is evidence of a size and shape a reviewer can check, together with an honest account of what else might explain the result. That is a far more durable position in front of an auditor than a claim of proof, and it is the difference between a program that survives a hostile question and one that does not.

Think of the whole discipline as a before-and-after photograph. The image only means something if you took the "before" shot before the change happened, under the same lighting, with the same camera. Everything that follows flows from that analogy.

Defining KPIs Before You Deploy

A key performance indicator, or KPI, is a measurable metric tied to a specific outcome your agency is accountable for. Not "AI accuracy." Not "model confidence." Those are vendor metrics. Your KPIs live in your program statute and in your oversight body's expectations.

For an eligibility screener, the KPIs might be average processing time per application, with a target of under 3 business days against a baseline of 11; error rate on initial determinations, targeted under 4 percent and measured by appeals overturned; applicant drop-off rate before submission, as a proxy for system friction; and the equity gap, meaning the approval rate difference across demographic groups, targeted under 2 percentage points. Notice what those are not. They are not technical metrics. They are program outcomes. The AI is a means to an end, and your KPIs must reflect the end. Notice also that each target is a number your agency chooses, writes down, and can defend, not a threshold anyone else imposed on you.

The five categories of KPI

The curriculum groups indicators into five families. Most agencies instrument the first two and forget the rest, which is exactly how a project ends up efficient, accurate, inequitable, and unused all at once.

CategoryRepresentative metricsQuestion it answers
EfficiencyHours of manual work eliminated per case; cost per decision before and after; cases processed per day; output per staff memberDid this save money or capacity?
QualityPercentage of correct decisions; consistency of outcomes on similar cases; rework rate; downstream quality of outcomesAre decisions more accurate and more consistent?
EquityDemographic parity in outcomes; disparate impact analysis; equity in access; equity in speed of serviceAre all populations served, and served equally well?
AdoptionPercentage of cases actually routed through the system; staff confidence in it; training completion; whether it runs without constant central supportIs anyone using it, and will it survive the project team?
BusinessReturn on investment; timeliness of decisions; stakeholder satisfaction; problems caught early that would otherwise have escalatedWas the investment worth making?

Adoption metrics deserve particular attention because they are the ones that quietly explain a disappointing result. If only a third of eligible cases actually go through the system, an unimpressive efficiency number is not evidence that the tool does not work. It is evidence that the tool is not being used, which is a different problem with a different fix.

Lock your KPIs before go-live

If you define metrics afterward you will, consciously or not, select the ones that make the tool look good. A Government Accountability Office auditor will ask when the metrics were set. So will an Inspector General. So will a journalist with a Freedom of Information Act request. Date the document, circulate it, and keep the version that predates deployment.

Lead and lag indicators

Some outcomes take months to appear. A fraud-detection model's effect on improper payments may not show up in quarterly data at all. Use lead indicators, early signals that predict later outcomes, alongside the lag indicators that oversight actually cares about. For fraud detection, a lead indicator might be the percentage of flagged cases that pass manual review within 30 days; the lag indicator is the annual improper payment rate in your program's report to the Office of Management and Budget.

Taking the Before Photograph

Baseline measurement means collecting data under the same conditions you will use after deployment. Start 6 to 12 months before go-live. Waiting until two weeks before launch is common and almost always fatal to the analysis, because you end up with a single noisy snapshot rather than a distribution.

The method is straightforward. Define exactly what you are measuring, for example the average time to process a case. Measure it over a defined period rather than at a moment, four weeks of manual processing being a reasonable minimum. Collect multiple measurement points rather than one, so that you can see variance and not just a central value. And document any unusual circumstances that affected the baseline window, such as a hiring freeze, a system outage, or a seasonal surge, because you will need that note when someone challenges the comparison two years later.

Your baseline should cover five dimensions. Volume and throughput: how many cases, requests, or decisions per month, and what the distribution looks like including peaks, troughs, and seasonal spikes. Processing time: end to end, not just the step the AI will touch, because a two-hour gain means nothing if it moves a bottleneck into a two-week manual review queue. Error and rework rates: what percentage of outputs require correction, appeal, or manual override, broken down by case type and worker cohort. Cost per transaction: staff hours multiplied by loaded salary cost, plus system costs, which becomes the denominator for any return-on-investment calculation. And equity distribution: outcomes broken down by race, income, geography, language preference, and disability status wherever your program has the data and the obligation.

That last one is not optional analysis for a public agency. If you do not have an equity baseline, you cannot prove the AI did not worsen disparities, and the absence of the data is not a defense. Store the whole baseline in a versioned location outside the live system. Operational databases get migrated, purged, and reformatted. You need data integrity across a 24-month window at minimum.

The Before and After Comparison, Worked

After deployment, allow time for adoption before you measure anything. Two to four months is the usual window. Measuring at two weeks, when adoption is low and staff are still learning the interface, produces a disappointing result that says nothing about the system. Then measure the same metrics, the same way, over the same kind of window, and account explicitly for differences in staffing, volume, and case mix.

A worked example from the curriculum shows what a complete comparison looks like. Baseline: average processing time without AI of 45 minutes per case, cost per case of $18, accuracy rate of 87 percent, and processing capacity of 50 cases per day. Post-deployment: average processing time of 18 minutes per case, a 60 percent reduction; cost per case of $7, a 61 percent reduction; accuracy of 91 percent, a 4 point improvement; and capacity of 120 cases per day, a 140 percent increase. Each of those four derived figures follows from its stated inputs, which is exactly the property a reviewer will test.

The efficiency arithmetic runs on from there. At 50 cases a day at $18 a case, total daily cost is $900. At 120 cases a day at $7 a case, total daily cost is $840. Net savings of $60 a day, with more output, and roughly $15,000 a year assuming 250 work days. Note the shape of that result: the per-case cost fell by 61 percent while the total daily cost fell by far less, because output rose at the same time. Both figures are in front of you, $900 against $840, so a reviewer can see the whole picture. Presenting only the per-case number to an appropriations committee would be technically true and substantively misleading, and the committee staffer who works it out will not describe it charitably.

Isolating AI Impact From Other Factors

This is the hardest part, and the part most AI vendors quietly skip in their case studies. When processing times drop 40 percent after a deployment, at least three things might explain it: the AI, a new intake form that launched the same month, and three experienced staff you hired six months earlier who are now fully productive. Separating them is your analytical job. A legislator's staff researcher will try to separate them for you if you do not, and they will not be charitable either.

Four confounders account for most of the trouble. Volume changes, because economies of scale can produce time savings on their own. Staffing changes, because new hires and departures move throughput independently of any tool. Concurrent process improvements, because agencies rarely change one thing at a time. And seasonal factors, because much government work has natural annual rhythm.

Phased rollout with a control group

Deploy to half your offices or case types first, and measure both groups for 90 days. This is the closest a government program usually gets to a randomized controlled trial without a formal research protocol. Attribute the difference between the groups to the tool and the movement common to both to external factors: if the comparison offices also improved, that drift was not caused by your deployment. The attribution holds only to the extent the two groups were genuinely comparable to begin with, so document how you matched them. A phased rollout also gives you a fallback if the tool fails, which is a bureaucratic advantage on top of the analytical one.

Difference-in-differences

Compare the change in outcomes for the group using AI against the change for a comparable group that is not. If processing times fell 5 percent nationally, from staffing improvements or seasonal patterns, and fell 22 percent in your AI-using offices, the attributable effect is roughly 17 percentage points. Your agency's data office or a university partner can usually run this analysis for under $30,000, which is inexpensive relative to defending a $2 million procurement with no evidence behind it.

Interrupted time series

If a control group is impossible, plot the outcome metric monthly for 24 months, mark the deployment date, and fit a regression line to each segment. A visible, statistically significant break in trend at the deployment date, with no other major program change at the same time, is defensible evidence. It is not proof of causation, but it is credible attribution when documented honestly.

When you have none of the above

Sometimes there is no control group, no comparison jurisdiction, and no clean 24-month series. The fallback is not to abandon attribution but to constrain it. Document every other change you made in the same window. Account for those changes explicitly in the write-up. Make conservative claims about the AI's contribution, deliberately smaller than the raw difference. And be transparent about the confounders you could not rule out. A modest claim that survives scrutiny is worth more than an ambitious one that collapses under it. You are not trying to prove the AI was the only factor. You are trying to show, with documented evidence, that it was a meaningful one and that you looked hard for alternative explanations.

Measuring Efficiency

Efficiency metrics answer the budget question: did this save money or capacity? The usual measures are processing time in hours per case before and after, cost per decision including labor, system cost, and overhead, throughput per full-time equivalent staff member, and rework rate.

Converting those to dollars is where measurement work most often falls apart. Suppose a tool saves 4 staff-hours per day across 12 offices at a loaded labor rate of $35 per hour. Those are your three inputs, and the annual labor value is their product multiplied by your agency's number of paid working days. Carry all four numbers into the appendix so a reviewer can reproduce the calculation themselves. An earlier version of this material stated a specific annual figure that does not follow from those inputs; it has been removed rather than adjusted, because a number a reviewer cannot reproduce does more damage in a hearing than no number at all.

One further discipline. Do not count headcount reduction as a gain unless it actually happened through attrition or redeployment. Claiming savings from positions you did not eliminate is precisely the kind of projection that attracts an audit and then fails it.

Measuring Quality

Quality metrics answer the program integrity question: are decisions more accurate and more consistent? Track accuracy, consistency of outcomes across similar cases, rework and overturn rates, appeals, and complaint volumes. If your agency runs a quality assurance review process, as most federal social services programs do under regulation, compare its error rates before and after deployment. A system that speeds up decisions while increasing error rates is a liability, not an asset.

Quality is genuinely harder to measure than efficiency because it requires knowing the true right answer, which the operational record does not contain. Four techniques get you there: spot checks of decisions, expert review of a sample, tracking downstream outcomes to see whether the decision held up in the real world, and monitoring complaints and appeals as an independent signal. None of these is free, and the temptation to skip them because processing time is sitting right there in the log file is the reason so many programs can describe their speed and not their correctness.

Measuring Equity

Equity metrics answer the civil rights question: did the system treat people differently based on protected characteristics? Measure approval rates, processing times, and error rates disaggregated by race, national origin, and disability status at minimum. Four specific measures are worth naming: demographic parity, meaning whether approval rates are equal across groups; error-rate comparison, sometimes discussed as equalized odds, which asks whether error rates such as the false positive rate match across groups; wait time equity; and outcome quality by group. All four require linking decisions to demographic characteristics, which is an infrastructure investment you have to make deliberately.

Do not assume fairness. Verify it. If the system was trained on historical data that reflected past discrimination, it may replicate those patterns, and you need the data to find out, as well as to defend yourself when a civil rights organization believes disparities exist and you believe they do not.

Two cautions on how to read the results. First, any threshold you apply is a threshold your agency sets for itself and records in advance, and a measured disparity is a signal that triggers investigation and legal review rather than a legal conclusion about discrimination. Statistical analysis identifies where to look; it does not decide the question. Second, the legal landscape here is layered and it moves. Title VI of the Civil Rights Act and state-level equity statutes may apply to your program. Executive orders on federal equity assessment have been issued, revoked, and replaced across administrations, so confirm what is currently in force for your agency rather than citing an order from memory.

Reporting to Oversight Bodies and the Public

You will have at least four distinct audiences, and they need different things from the same underlying data.

  • Legislators and appropriations committees want dollar savings, error rate changes, and whether the project came in on budget and on time. Lead with numbers and keep the methodology in an appendix, but make sure the appendix exists.
  • Inspector General and Government Accountability Office reviewers want methodology first and will question every assumption. Prepare a written measurement plan dated before deployment, plus a data dictionary. Document every deviation from the plan and why it happened.
  • The public and press want plain-language summaries, concrete examples, and honest disclosure of limitations. A two-page public summary and a published FAQ substantially reduce the risk that someone else's Freedom of Information Act request sets the narrative first.
  • Advocacy organizations and affected communities want equity data and a route to raise concerns. Proactive publication of disaggregated outcomes with a named point of contact is standard practice in mature AI governance programs.

Annual reporting aligned to the budget calendar is the practical minimum. For high-stakes systems, those affecting benefits, enforcement, child welfare, or housing, quarterly reporting is increasingly the expectation. Report the areas needing improvement alongside the successes; a report with no bad news in it invites a reader to go looking for the bad news themselves.

Common Measurement Pitfalls

  • Succeeding at the wrong metric. You measure processing time, it improves, and the actual goal was accuracy. Define success metrics upfront as part of requirements, and measure what matters rather than what is easy to instrument.
  • Measuring too early. Results at two weeks reflect low adoption and a learning curve, not the system. Allow two to four months before declaring success or failure.
  • Ignoring concurrent changes. You hired five new staff the same quarter and attributed everything to the tool. Document what else changed, use a control group where possible, and make conservative claims.
  • Metric gaming. Staff learn what is measured and optimize for it, so processing time improves because cases are being rushed. Measure quality alongside efficiency and never incentivize a single metric in isolation.
  • Skipping equity because it is hard. Processing time is easy to measure and fairness is not, so fairness gets dropped. Invest in the measurement infrastructure for the metrics that matter, including the difficult ones.
  • Measuring what the vendor tracks. Vendor dashboards show model performance. Your oversight body needs program outcomes. These are rarely the same thing.
  • Announcing projected savings before measuring actual ones. A press release claiming $4 million in projected savings becomes a liability when an audit finds $800,000 realized two years later.
  • Calling automation AI without disclosure. Rules-based automation and machine learning behave differently under novel conditions, and auditors increasingly know the difference.
  • Ignoring spillover. A tool that speeds one step may create a downstream bottleneck, increase call center volume, or shift burden onto applicants. Measure the whole process.
  • No measurement owner after go-live. Deployment teams move on. If nobody owns the plan 18 months later, the data will not exist when the oversight request arrives.
  • Conflating uptime with performance. A system available 99.9 percent of the time and producing wrong answers is not performing well. Track operational and programmatic metrics separately.

Anti-Patterns

  • Calling a before-and-after comparison proof. It is attribution under stated assumptions. Say which assumptions, and say what else could explain the result.
  • Setting the metrics after you see the data. Any KPI defined post-deployment is vulnerable to selection bias, and the first question an auditor asks is when it was set.
  • Reporting the per-unit gain and burying the total. A 61 percent fall in cost per case alongside a much smaller fall in total daily cost is one result, not two, and presenting half of it is how credibility is lost. Publish both, with the daily totals behind them.
  • Publishing a figure a reviewer cannot reproduce. If the inputs do not yield the number, drop the number and print the inputs.
  • Treating a statistical disparity as a legal verdict, in either direction. Finding a gap does not establish discrimination, and finding none does not establish compliance.
  • Attributing everything in the window to the deployment. Every uncontrolled change in the same period is an alternative explanation until you have addressed it in writing.
  • Skipping adoption metrics. Without usage rates, a weak result is uninterpretable: you cannot tell a tool that does not work from a tool nobody runs.
  • Letting the vendor define success. Model accuracy is their metric. Program outcomes are yours, and only one of the two appears in your statute.

Practice Prompts

  • Design a KPI framework for an AI hiring assistance system. What efficiency gains would you measure, what quality metrics matter, which equity metrics are critical, how would you establish the baseline, and what is your post-deployment measurement plan?
  • You deploy a system and observe processing time down 30 percent, accuracy up 5 percent, and diversity of hires up 15 percent. Write the list of questions you would need answered before attributing any of those to the system.
  • Design a quarterly impact dashboard for leadership. Which metrics from all five categories appear, and how would you present them to make the case for continued investment without overstating attribution?
  • Take a system your agency already runs. Write down today what you would need to have measured before it launched, and note which of those baselines you can still reconstruct from archived data.
  • Draft the two-page public summary for a deployed system, including its limitations. Then decide whether you would be comfortable if a journalist obtained the full underlying data through a records request.

Reflection

Think of a change you have implemented, at work or personally, that you believed worked. How would you actually measure whether it did? What would you compare against, and what did you know about the before state at the time you made the change? Most of us discover at this point that we never took the before photograph, which is precisely the position Nadia was in when the senator asked his question.

Then ask the uncomfortable version. If the measurement came back flat, or negative, what would you do with it? A measurement plan that only has a path for good news is not a measurement plan, and the agencies that get into trouble are rarely the ones whose systems underperformed. They are the ones that could not say so before somebody else did.

Glossary

  • KPI (key performance indicator). A measurable metric indicating whether an initiative is delivering its intended value, tied to a program outcome rather than to a technical property of the model.
  • Baseline. Measurement of the current state before an intervention, collected over a defined window with multiple points, used for comparison afterward.
  • Lead indicator. An early signal that predicts a later outcome, used when the outcome of record takes months or years to appear.
  • Lag indicator. The outcome of record itself, typically reported annually and typically the one oversight bodies act on.
  • Control group. A population similar to the group receiving the intervention but not receiving it, used to separate the effect of the intervention from other changes.
  • Difference-in-differences. An attribution method comparing the change in outcomes for the treated group against the change for a comparable untreated group.
  • Interrupted time series. An attribution method that plots an outcome over a long period, marks the intervention date, and tests for a break in trend.
  • Disparate impact. A statistical difference in outcomes for protected groups that may indicate discrimination even without intent. A signal that triggers review, not a legal finding on its own.
  • Demographic parity. An equity measure asking whether outcome rates, such as approval rates, are equal across groups.
  • Metric gaming. Behavior in which performance is optimized for the metric being measured at the expense of the true objective, distorting the result.
  • Loaded rate. The full hourly cost of a staff member including benefits and overhead, used as the multiplier when converting saved hours into dollars.

Closing

Impact measurement is what moves a program from "it seems to be working" to "here is the evidence, here is how we collected it, and here is what we could not rule out." It does three things at once: it demonstrates value, it drives improvement by showing what needs to get better, and it builds trust because an agency that publishes its own limitations is harder to discredit than one that publishes only its wins.

The senator's question was not hostile. It was the right question, and Nadia should have been able to answer it with a dated measurement plan, a baseline taken a year earlier, a comparison group, and a candid paragraph about the intake form that launched the same month. None of that is technically difficult. All of it has to be decided before go-live, by a named person who is still responsible for it two years later, which is why measurement is a governance problem rather than an analytics problem.

Key Takeaways

  • Define KPIs before deployment and date the document. Metrics set after go-live are vulnerable to selection bias and will not survive Inspector General or Government Accountability Office scrutiny.
  • Instrument all five KPI categories. Efficiency, quality, equity, adoption, and business. Skipping adoption makes a weak result impossible to interpret; skipping equity is not an option for a public agency.
  • Collect baselines 6 to 12 months early, over a window rather than at a moment. Multiple measurement points, documented anomalies, and storage outside the live system with 24-month integrity.
  • Allow two to four months for adoption before measuring impact. A two-week reading measures the learning curve, not the system.
  • Isolate the AI effect with a phased rollout, difference-in-differences, or an interrupted time series. Where none is possible, document concurrent changes and make deliberately conservative claims.
  • Measurement produces evidence, not proof. Attribution rests on assumptions; state them, and state what you could not rule out.
  • Publish figures a reviewer can reproduce. Carry the inputs into the appendix. If a computed result does not follow from its inputs, drop the result and keep the inputs.
  • Measure efficiency, quality, and equity separately. Each answers a different oversight question and carries a different legal obligation, and quality requires knowing the true answer, which the operational record does not contain.
  • Equity thresholds are yours to set and record in advance. A measured disparity triggers investigation and legal review; it does not settle the legal question in either direction.
  • Translate the same data for four audiences. Dollars for legislators, methodology for auditors, plain language for the public, disaggregated outcomes for affected communities.
  • Assign a named measurement owner for at least 24 months after deployment and document every deviation from the plan. Undocumented deviations look like cover-ups even when they are not.

Frequently Asked Questions

When should we start measuring? Before you deploy. Baseline collection should begin 6 to 12 months ahead of go-live, over a defined window with multiple measurement points, and the KPI document should be dated and circulated before the system is live.

How soon after launch can we report results? Allow two to four months for adoption first. Measuring at two weeks captures low usage and an unfinished learning curve, and a disappointing reading at that point tells you nothing about whether the system works.

Model accuracy is up. Is that impact? No. Model accuracy is a vendor metric. Impact is measured in program outcomes: processing time, error rates, appeals overturned, equity gaps, cost per transaction. A more accurate model that nobody uses has produced no impact at all.

We have no control group. Can we still claim impact? Yes, with constraints. Use an interrupted time series if you have a long enough data history, document every other change in the window, account for confounders explicitly, and make claims deliberately smaller than the raw difference. A modest claim that survives scrutiny beats an ambitious one that does not.

How do we convert time savings into dollars? Multiply the hours saved by the number of sites, by your loaded hourly labor rate, by your agency's number of paid working days, and publish all four inputs so a reviewer can redo the calculation. Do not count headcount reduction unless the positions were actually eliminated through attrition or redeployment.

Our equity analysis found a gap. Does that mean the system is discriminating? Not by itself. A statistical disparity is a signal that triggers investigation and legal review. It identifies where to look. The reverse holds too: finding no gap in your analysis does not establish compliance.

Processing time improved but staff say quality dropped. Which is right? Possibly both, and that pattern is the classic signature of metric gaming. When a single metric is measured and incentivized, behavior shifts toward it. Measure quality alongside efficiency, using spot checks, expert review of samples, downstream outcomes, and appeal rates.

Who owns measurement after the project team disbands? Somebody named, in writing, for at least 24 months past deployment. Without that, the deployment team moves on, collection lapses quietly, and the data is missing on the day the oversight request arrives.