←
AI for Government
Proficient · M5 · lesson 5 of 50 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Metrics and KPIs for Government
📖
now learning

AI Metrics and KPIs for Government

15 min

Carmen Iglesias is a program director at a city permitting department. Her team deployed an AI assistant to help review building permit applications, and six months in, her boss asked the question every program director dreads: "Is it working?" Carmen had a number ready. The AI was processing applications 40 percent faster, and she was proud of it. Then a councilmember at the public meeting asked a follow up. "Faster is great. But are the permits you approve any good? Are you turning down the same share of applicants as before? Are residents happier, or just getting wrong answers faster?" Carmen did not have those numbers. She had measured the one thing that was easy to measure and missed everything that decided whether the city was better off.

The trap of the easy metric

Speed is the easiest thing to measure in any AI deployment, so it is what most agencies measure, and it is almost never the thing that matters most. A system that processes applications 40 percent faster but approves 40 percent more bad permits is not a success. It is a faster way to fail. Government exists to deliver outcomes such as safe buildings, fair benefits and served residents, not to be busy. Measurement is the hinge between promise and accountability, and an agency that cannot show measurable outcomes has no defense when its budget, its audits or its complaints arrive.

The controlling analogy is the doctor's vital signs. A doctor never judges a patient on one number. Pulse alone tells you the heart is beating, not whether the patient is healthy. You need pulse, blood pressure, temperature and oxygen together to see the real picture, and a single number out of range is a warning rather than a verdict. AI metrics work the same way. You need a balanced panel, read together, or you will declare a sick system healthy because its pulse looked good. A dashboard that shows only speed tells you the AI is busy. Whether the public is better off is a different question, and only that one is your job.

The failure cases are documented and they repeat. A widely deployed sepsis prediction model ran for years at most sites without local measurement, which allowed underperformance to persist unnoticed. A large chronic care algorithm went unmeasured for equity until academic researchers intervened. The Michigan MIDAS unemployment system operated without the outcome measurement that would have surfaced what the record describes as a 93 percent false positive rate. Against those, the Veterans Affairs REACH VET program offers the template: it published its evaluation through peer review, updated it over time and engaged academic partners, which is how a federal AI program earns durable credibility.

Six families of metrics, not one number

Carmen rebuilt her dashboard around four dimensions that a city program director can own: efficiency, quality, satisfaction and compliance. Federal practice widens that into six families, and the two views nest rather than compete. The six families are model performance, fairness and equity, operational, cost and efficiency, trust and transparency, and mission outcome. Carmen's efficiency dimension sits inside cost and efficiency, her quality dimension spans model performance and fairness, her satisfaction dimension sits inside trust, and her compliance dimension cuts across all of them. Whichever framing you adopt, the rule is the same: report the families together or you are reporting a fragment.

Efficiency, in its place

This is the easy dimension and it does belong on the panel. Processing time, throughput, cost per transaction and staff hours saved all live here, including Carmen's 40 percent speed gain. It is useful and it is never sufficient on its own. Pair every efficiency number with a quality number in the same report, on the same page, so that nobody can claim victory by trading correctness for speed. An efficiency figure reported alone is not a lie, but it is an invitation to be read as one.

Quality, the dimension agencies skip

For Carmen, quality means the share of AI assisted permit reviews later found correct, how often a human reviewer overturned the AI, and whether approval and denial rates hold steady across neighborhoods and demographic groups. A quality dashboard that does not break results out by group is hiding the question most likely to end up in litigation. This is also the dimension where federal guidance is most explicit: fairness and equity measurement is expected for rights impacting AI under OMB Memorandum M-24-10, issued in 2024, and civil rights obligations attach independently through Section 1557 of the Affordable Care Act for federally funded health programs and through Title VI, Title VII, Title IX and the Americans with Disabilities Act for other protected classes.

Satisfaction, from the resident's side of the counter

Government's customers are residents, and many of them cannot take their business elsewhere. Measure their experience directly: satisfaction scores, complaint volume, appeal rates, appeal outcome distribution and time to resolution counted from the resident's point of view rather than the agency's. A faster internal process that produces more confused residents and more appeals is a net loss, and only this dimension will catch it. Time to first successful use and adoption rate belong here too, because a tool nobody can complete a task with is not serving anyone regardless of how fast it runs.

Compliance, which is close to unique to the public sector

Track human review rates for consequential decisions, audit log completeness, accessibility conformance under Section 508, privacy incidents and override frequency. A compliance metric trending the wrong way is an early warning of the kind of failure that draws an inspector general or a courtroom. Compliance measurement is also where the oversight relationship lives: audit and inspector general finding closure rates, completeness of the agency's AI use case inventory, and whether model cards and data statements are current rather than written once and abandoned.

Model performance, calibration and the operating point

Model performance metrics vary by task, and picking the wrong one for the task is a common way to be confidently wrong. Binary classification uses accuracy, precision, recall, F1, area under the ROC curve, area under the precision recall curve, and balanced accuracy when classes are imbalanced. Multi class work uses per class precision and recall with macro or micro F1. Ranking uses nDCG, MAP and MRR. Regression uses MAE, RMSE, MAPE and R squared. Generative systems use BLEU, ROUGE, BERTScore, model graded evaluation and human preference, plus faithfulness and grounding metrics where retrieval is involved.

Calibration should always accompany accuracy, through the Brier score, expected calibration error and reliability diagrams. A model can rank cases well and still be badly wrong about how confident it should be, which matters enormously when a score is handed to a caseworker as a number. For safety impacting systems the operating threshold matters more than an optimal F1 score, because the threshold is what actually decides cases. Publish the operating point you chose and the workflow that sits downstream of it, so a reader can see what the system does rather than what it could do at some other setting.

Name the metric exactly, because the definitions invert

More measurement disputes come from mislabeled metrics than from bad models. Precision asks, of the cases the system flagged, what share were truly positive. Recall asks, of the truly positive cases, what share the system flagged. The false positive rate asks, of the truly negative cases, what share were wrongly flagged. The false negative rate asks, of the truly positive cases, what share were missed. Subtracting a false negative rate from one gives you recall, not precision and not specificity. Any figure quoted without its denominator is unusable, which is exactly why the MIDAS number above is carried in the form the record states it rather than converted into some other rate.

Pre-register the evaluation, and protect the test set

Federal programs should write the evaluation plan before the results exist. That plan names the held out test sets, the strata for subgroup analysis and the acceptance criteria that will decide go or no go. Guard against benchmark contamination: a vendor that has trained on your test set can show you any number you like. Published academic benchmarks and public evaluation programs are useful context and they do not replace local evaluation, because real world deployment performance can differ substantially from benchmark performance. The sepsis model case is the standing proof of that gap.

Fairness and equity measurement

The core fairness metrics are demographic parity, meaning equal positive rates across groups; equalized odds, meaning equal true positive and false positive rates across groups; calibration by group; and group conditional precision and recall. The important technical result, from the work of Kleinberg and colleagues and of Chouldechova, is that demographic parity, equalized odds and calibration cannot all be satisfied at once when base rates differ between groups. You must choose which property to prioritize, and you must write down the choice and the reason. An agency that has not made that choice explicitly has still made it, implicitly, and cannot explain it later.

Intersectional analysis matters because a system can be fair across race and fair across sex when each is examined alone, and unfair for a specific combination such as Black women. Report subgroup sample sizes alongside subgroup results, because conclusions drawn on small subgroups are unreliable and reporting them without the counts invites both false alarms and false comfort. Equity measurement also depends on upstream data. Demographic labels have to exist and be accurate, which is its own difficulty under federal confidentiality rules, and proxy methods such as Bayesian surname and geocoding estimation are imperfect substitutes that carry their own error.

Stratified measurement is the practical discipline underneath all of this. Aggregate accuracy is the number that hides disparate impact, so build the dashboard to show the strata by default rather than on request. Publish equity results alongside aggregate performance rather than in a separate document nobody opens. Open toolkits for group fairness analysis exist and are a reasonable starting point; the standard reference texts in algorithmic fairness are worth reading before you pick a metric, because the choice of metric encodes a policy position whether or not you intended one.

Operational and cost metrics

Operational metrics describe system behavior in production. They include availability and service level attainment, latency at the median and at the 95th and 99th percentiles, throughput, error rates covering timeouts and server and model errors, drift indicators for input and output distributions, retraining frequency, rollback count, incident count by severity, mean time to detect, mean time to remediate, and change failure rate. Generative systems add hallucination rate measured by sampling and review, refusal rate, jailbreak success rate carried over from red team work, and prompt injection detection rate. Agentic systems add task success rate, tool call correctness and human override rate.

Data quality metrics sit underneath everything else: completeness, consistency, timeliness, freshness and validity. A model performance number computed on stale or incomplete data is a measurement of your pipeline, not of your model. These operational measures feed the MEASURE function of the NIST AI Risk Management Framework, a voluntary and non binding framework that many agencies have adopted, and they support the continuous monitoring expectation in M-24-10 minimum practices. Design the dashboard around thresholds and alerting rather than expecting a human to notice a trend by looking at a chart.

Cost metrics exist because the money is the public's. Total cost of ownership should include compute, whether accelerator hours or per request inference charges; data costs covering collection, labeling, storage and lineage; platform costs for operations, monitoring and governance tooling; workforce costs for both federal staff and contractor hours; and governance overhead. Express unit economics as cost per decision, cost per case processed or cost per hour of analyst time released. Compare return against the status quo baseline rather than a hypothetical optimum, and count only realized savings, because hypothetical savings do not survive federal reporting.

Two cost traps are worth naming in advance. The first is vendor cost opacity: demand itemised billing and understand the pricing levers, which typically include input tokens, output tokens, context length and model tier. The second is scale. What is cheap at pilot volume can be prohibitive at agency volume, and a cost curve that was never modeled beyond the pilot is a budget surprise waiting for the next appropriation cycle. Federal benefit cost methodology and performance budget integration guidance already exist in OMB Circulars A-94 and A-11; AI programs do not get their own accounting rules.

Trust, transparency and mission outcomes

Trust metrics ask whether the people who use the system, the people subject to it and the people who oversee it actually trust it. On the user side, adoption rate, time to first successful use, satisfaction and override rate. On the subject side, complaint rate, appeal rate and the distribution of appeal outcomes. On the oversight side, audit and inspector general finding closure rates and inventory completeness. Transparency metrics count coverage: what proportion of use cases have a published model card, a stated intended use, a fairness report and a red team summary, and how quickly affected parties receive a response through the reporting channel.

Agencies are tempted to conceal trust problems and it reliably backfires when oversight bodies or the press find them first. Mature agencies publish their problem statistics along with the improvement trend. Opacity reads as an audit flag, not as discretion.

Mission outcome metrics are what the agency exists to deliver, and they are specific to the mission rather than to the technology. Benefits administration measures claim accuracy and processing time. Health payers measure improper payments and beneficiary outcomes. Tax administration measures taxpayer service, fraud reduction and dispute resolution time. Immigration services measure processing time and approval consistency. Suicide prevention outreach measures outcomes in the identified cohort. Tie each AI initiative to a mission objective from the agency strategic plan rather than inventing an outcome measure after deployment to justify what was already built.

The hard part of outcome measurement is the counterfactual. Would the outcome have improved anyway, without the AI? Randomised pilots, stepped wedge rollouts and quasi experimental designs are the honest answers to that question, and federal evaluation guidance already addresses evaluation design. Post hoc storytelling is the dishonest answer, and it is the one that collapses under audit. Pre register the design, then share the result whichever way it comes out.

Leading and lagging indicators

Carmen also learned the difference between metrics that tell you what already happened and metrics that warn you before it does. Lagging indicators, such as last quarter's appeal rate or an audit finding, confirm a problem after it has cost you something. Leading indicators, such as a rising rate of human reviewers overturning the AI, a creeping gap in approval rates between neighborhoods, or a drop in confidence on particular case types, flag trouble while a fix is still cheap. A mature dashboard watches both and invests in the leading set, because they are the smoke detector rather than the fire report. Drift indicators and data quality checks belong in the leading set for the same reason.

A usable artifact: the government AI KPI dashboard

This is the balanced panel Carmen built. Adapt the specific metrics to your program, but keep every dimension represented, and give each metric a target, a type and a named owner. Read the target column carefully: every figure in it is a number the agency chose for itself and recorded in advance, so that performance could be judged against a commitment rather than against whatever the system happened to produce. None of them is an objective standard or a legal test, and the example values below are illustrative of Carmen's program rather than benchmarks to copy.

DimensionExample metricTarget set in advanceTypeOwner
EfficiencyAverage processing time against baseline40 percent reductionLaggingOps lead
Cost per transactionDown quarter over quarterLaggingFinance
QualityHuman override rateBelow 10 percent and stableLeadingQA lead
Approval rate gap across groupsWithin the tolerance recorded before launchLeadingEquity officer
SatisfactionResident satisfaction scoreUp against baselineLaggingService lead
Appeal and complaint volumeFlat or downLeadingService lead
ComplianceHuman review rate on consequential decisions100 percent where requiredLeadingCompliance
Section 508 accessibility conformanceFull conformanceLaggingAccessibility lead

A dashboard is only half of the measurement infrastructure. The other half is automated alerting against those thresholds, an incident tracker that records what went wrong and what was done about it, and an executive reporting rhythm that puts the same numbers in front of leadership on a schedule. Without the alerting, nobody sees the threshold breach until the quarterly review. Without the incident tracker, the same failure gets rediscovered. Without the reporting rhythm, the dashboard becomes an artifact that a small team maintains and no decision maker reads.

The frameworks that shape federal measurement

Four instruments structure how federal agencies measure. The GPRA Modernization Act of 2010 requires agencies to publish performance plans with measurable indicators, and the Foundations for Evidence-Based Policymaking Act of 2018 strengthened that through learning agendas and chief data officer requirements. OMB Memorandum M-24-10, issued in 2024, requires performance information in the annual AI use case inventory where appropriate. The NIST AI Risk Management Framework contributes its MEASURE function, which is voluntary rather than binding. Together they mean that AI measurement is not a new obligation bolted on; it is the existing performance regime applied to a new class of system.

The fourth instrument is the GAO Artificial Intelligence Accountability Framework, published in June 2021 as GAO-21-519SP, which organizes federal AI accountability into four principles: Governance, Data, Performance and Monitoring. Governance covers structures, roles and escalation paths. Data covers the quality, reliability and representativeness of the data used to build the system. Performance covers validating and monitoring against goals while watching for unintended effects. Monitoring covers continuous assessment across the life cycle. Each principle carries key practices, and auditors use the framework during audits. Treat it as a measurement design reference rather than as a checklist you complete the week before an audit.

The MEASURE function is worth reading against your own metric set, because it enumerates the trustworthiness characteristics an agency is expected to have something to say about. Counting the characteristics named in that source list gives seven: validity and reliability, safety, security and resilience, accountability and transparency, explainability and interpretability, privacy, and managed bias. If your dashboard has nothing at all to say about one of them, that gap is a finding you have made about yourself, and it is much better found internally than in an audit letter.

Public reporting and the AI use case inventory

The annual AI use case inventory is the primary public artifact. The format has called for each use case to record intended use, development stage, expected benefits, whether the use case is rights impacting or safety impacting, and, where appropriate, performance indicators. Inventories are published on agency open data pages, and the range of quality across agencies is wide: some publish detailed quantitative information, others publish the minimum required fields. The format has been iterated and will continue to be. Treat the inventory as a public performance report rather than as a compliance form, because that is how oversight bodies and journalists read it.

Public reporting creates a genuine tension between transparency and operational security. Publishing the precise thresholds of a fraud detection algorithm could help someone evade it. The standard resolution is to publish the intended use, the general performance characteristics and the governance arrangements, while withholding operational specifics, and to handle records requests through standard practice rather than inventing an AI exception. Reporting to Congress happens through statutory performance reports, responses to oversight letters and testimony. Match the tone to the evidence. Over promising AI capability is the reliable route to losing credibility later, when the same committee has your earlier statement in front of it.

Three rules for honest metrics

Carmen added three guardrails so the dashboard would tell the truth rather than flatter the program. First, always pair efficiency with quality, so no speed gain is ever reported without the accuracy and fairness numbers beside it. Second, keep a pre AI baseline, because a claim of 40 percent faster means nothing without the old number; if you did not measure before deployment you cannot prove improvement, so capture the baseline before you launch, every time. Third, disaggregate by group, because an overall average can look fine while a specific community is being harmed.

When Carmen returned to the council with the rebuilt dashboard, the picture was more honest and considerably more persuasive. Processing was 40 percent faster. Override rates were low and steady. Approval rates were consistent across neighborhoods within the tolerance she had recorded before launch. Resident satisfaction was up six points and appeals were flat. She could answer the councilmember's follow up with evidence across every dimension, and when a small gap did appear later in one district's approval rate, the leading indicator caught it in weeks rather than in a lawsuit.

Anti-patterns

Federal measurement fails in recognizable ways. A mature program audits itself against this list on a schedule, because auditors and inspectors general are looking for exactly these.

  • Aggregate only. Reporting overall accuracy while the stratified performance stays in a drawer. This is the anti-pattern that hides disparate impact.
  • Vanity metrics. Counting pilots, press mentions and slide decks as though activity were impact.
  • Benchmark tunneling. Optimising for public benchmarks that do not resemble your production distribution.
  • Pilot inflation. Endless pilots that raise activity statistics without ever producing mission impact.
  • Missing counterfactual. Claiming causation from correlation with no baseline or comparison design.
  • Measurement free go live. Deploying without acceptance criteria specified in advance.
  • Ghost KPIs. Indicators written into the plan document and never updated again.
  • Post hoc framing. Inventing the success metric after deployment so the deployment looks successful.
  • Opaque vendor metrics. Accepting vendor claimed performance without independent verification on your own data.
  • Under reporting incidents. Reclassifying problems as noise so the dashboard stays green.
  • Dashboard theater. A beautiful dashboard that is connected to no actual decision.
  • Measurement without action. The metrics exist, the threshold breaches, and nobody does anything.
  • Inverted metric definitions. Describing a false negative rate as though it were a false positive rate, or reporting a rate without its denominator. This is not pedantry; it silently reverses what the number means about the people the system missed.
  • Targets presented as standards. Reporting an internally chosen threshold as though it were an objective or legal test. Record who set the target, when and why.

Practice prompts

  1. Take your program's current status report and mark every number as efficiency, model performance, fairness, operational, cost, trust or mission outcome. Which families have no entry at all?
  2. For one deployed system, write down the operating threshold in use and the workflow downstream of it. If nobody can tell you the threshold, that is your first finding.
  3. Pick your headline quality metric and write its denominator in one sentence. Then check whether the way it is described in your last report matches that denominator.
  4. Choose between demographic parity, equalized odds and calibration by group for one rights impacting system, write down which you prioritized and why, and put the memo in the record.
  5. Draft the counterfactual paragraph for your most successful AI initiative: what evidence do you have that the outcome would not have improved anyway?
  6. Build the alerting layer for one metric on your dashboard, including who is paged and what they are expected to do.

Reflection

Carmen's first dashboard was not dishonest. It measured something real, it measured it accurately, and it was still misleading, because the selection of what to measure carried more meaning than the measurement itself. Ask yourself which number your program would report if it were allowed to report only one, and then ask who would be invisible in that number. Ask what you would have to see, on which metric, at what level, to recommend pausing a system that leadership is proud of. If you cannot name that number and that level in advance, your measurement program is a reporting function rather than an accountability function, and the difference will only become visible on the day it matters.

Glossary

  • Leading indicator. A metric that moves before the outcome does, giving you time to intervene while a fix is still cheap.
  • Lagging indicator. A metric that confirms what already happened, useful for accountability and useless for prevention.
  • Precision. Of the cases the system flagged, the share that were truly positive.
  • Recall. Of the truly positive cases, the share the system flagged.
  • Calibration. Whether the confidence a system reports matches how often it is actually right at that confidence level.
  • Stratified measurement. Reporting performance separately for defined subgroups rather than only in aggregate, so disparate impact is visible.
  • Equalized odds. A fairness criterion requiring equal true positive and false positive rates across groups.
  • Demographic parity. A fairness criterion requiring equal positive rates across groups.
  • Drift. Change over time in the distribution of inputs, outputs or labels, which degrades performance without any change to the model.
  • Counterfactual. The estimate of what would have happened without the intervention, which is what separates a measured effect from a coincidence.
  • AI use case inventory. The annual public list of an agency's AI uses, with intended use, stage, impact designation and, where appropriate, performance indicators.

Closing

The councilmember's question was not hostile. It was the right question, asked by someone whose job is to ask it, and Carmen could not answer it because she had built a dashboard for herself rather than for the public. The rebuilt version is harder to maintain, less flattering in places, and considerably more useful, because it can survive contact with an auditor, a committee and a resident who believes they were treated unfairly. Measurement in government is not a scorecard you present when things go well. It is the evidence you will need on the day things go badly, which is why it has to be built before that day and not after it.

Key takeaways

  • Busy is not the same as good. Speed is the easiest metric and the least sufficient, and a system that fails faster is still failing.
  • Measure the families together. Model performance, fairness and equity, operational, cost and efficiency, trust and transparency, and mission outcome form the panel; a fragment of it is not a measurement program.
  • Quality and compliance are the public sector difference. Fairness across groups, human review rates and accessibility conformance are not optional extras in government.
  • Name the metric exactly. Precision, recall, false positive rate and false negative rate answer different questions, and a rate without its denominator cannot be interpreted at all.
  • You cannot have every fairness property at once. When base rates differ across groups, parity, equalized odds and calibration cannot all hold, so choose deliberately and document the choice.
  • Pre register the evaluation and protect the test set. Acceptance criteria written after the results exist are not acceptance criteria.
  • Watch leading indicators. Rising override rates and creeping group level gaps warn you while a fix is still cheap.
  • Capture a baseline before launch. Improvement claims are meaningless without the pre deployment number.
  • Targets are commitments you set, not standards you discovered. Record who chose each threshold and when, and report it that way.
  • Publish, including the problems. Opacity about performance is read as an audit flag, and over promising is what costs a program its credibility later.

Frequently Asked Questions

We are a small agency with one AI system. Do we really need six metric families?

You need something in each family, not a large program in each. For a single system that can be one number per family plus the stratified version of the quality number. The families are a completeness check rather than a workload target. The failure they prevent is the common one, where a program measures the dimension it finds easy and has nothing at all to say about the dimension that eventually produces the complaint.

What if we never captured a baseline before deployment?

Say so, plainly, and stop making improvement claims that depend on a number you do not have. You can often reconstruct a partial baseline from historical case records, and where you can, document how you built it and what its limitations are. What you cannot do is present a post deployment figure as an improvement without stating what it is being compared against. That is the claim that collapses first under audit.

How do we measure fairness if we do not collect demographic data?

This is a real constraint and it does not have a clean answer. Demographic labels have to exist and be accurate for group metrics to mean anything, and federal confidentiality rules genuinely restrict collection in many contexts. Proxy estimation methods based on surname and geography exist and are imperfect, with error that varies by group in ways that can distort exactly the comparison you are trying to make. Document what you can measure, what you cannot, and what the gap means for the conclusions you are drawing.

Should we publish our performance numbers when they are bad?

The record on concealment is not encouraging. Problems that oversight bodies or reporters find first cost more than problems an agency discloses along with its remediation plan, and opacity itself tends to be read as a finding. The genuine exception is operational specifics that would help someone game a system, such as the exact thresholds of a fraud detector. The standard resolution is to publish intended use, general performance characteristics and governance arrangements while withholding those specifics.

How do we know the AI caused the improvement?

Only by designing for that question before deployment. Randomised pilots, stepped wedge rollouts and quasi experimental comparisons give you a defensible counterfactual; a before and after comparison in a changing environment does not. If the design was never put in place, the honest report says the outcome improved during the period the system was in use and that the contribution of the system has not been isolated. That sentence is much easier to defend than a causal claim you cannot support.

Who owns AI metrics, the technical team or the program?

Every metric on the panel needs one named owner, and they will not all sit in the same place. Operational and model performance metrics usually belong to the technical team. Fairness metrics usually need an equity or civil rights owner. Cost belongs with finance, satisfaction with the service owner, and compliance with the compliance function. The point of naming the owner in the dashboard itself is that an unowned metric is the one nobody acts on when it moves.