←
AI for Government
Visionary · M23 · lesson 23 of 46 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Cross-Cutting: Government Efficiency and Citizen Experience
📖
now learning

Cross-Cutting: Government Efficiency and Citizen Experience

15 min

Commissioner Dawit Reyes inherited a state Department of Motor Vehicles with a satisfaction problem so lopsided it was easier to state backwards: 94 percent of the citizens it surveyed were dissatisfied. His predecessor had bought a chatbot. It deflected calls beautifully and made everything worse, because people who could not get a real answer simply drove to a branch and waited three hours, angrier than when they started. The agency had optimized a metric, calls deflected, while degrading the mission, which is that a person renews a license without losing a day of work. Reyes decided his job was not to buy more AI. It was to redesign the journey so that efficiency and experience pulled the same way.

That is the problem at the top of an agency. You are not deploying a tool; you are orchestrating a portfolio of changes across many programs so that the public's experience improves and the cost of delivering it falls at the same time. Those two goals are usually treated as a trade-off, and the trade-off is where most modernization programs quietly go wrong. The transformative move is to stop treating them as opposites and to start hunting for the places where the cheap change and the kind change turn out to be the same change.

Orchestrating a portfolio is a different job from running a deployment, and it fails in two directions. Leave every program to pursue its own modernization and you get a hundred local optima that never meet, with each unit measuring itself on something the others do not recognize. Centralize every decision and you get a queue in front of one team, which is how a transformation office becomes the slowest part of the agency it was created to speed up. The work at the top is choosing the small number of things that must be common and then leaving the rest genuinely delegated.

The False Trade-Off and Why It Is False

Most efficiency programs cut cost by pushing work onto the citizen. Fill out the form yourself. Find the right office yourself. Call back during business hours. Costs drop on the agency's ledger and rise on the public's, which is not a saving so much as a transfer that nobody recorded. Satisfaction craters, complaints climb, and the calls come back through the legislature, usually attached to the budget hearing where you have to explain why service got worse in the year you spent money to improve it.

The reframe that makes the trade-off dissolve is easy to state and hard to act on: most citizen frustration is itself waste. It is duplicated work and failure demand sitting inside your own operation, wearing the costume of a customer service problem. Wherever a citizen is frustrated because something went wrong, the agency has usually paid to handle the same case twice. Find those places and you cut cost and frustration in one motion. Citizen frustration is not the price of efficiency; most of it is the symptom of waste you have not found yet.

It helps to understand why a predecessor reaches for deflection in the first place, because you will face the same pull. Calls avoided is cheap to instrument, moves quickly, needs no cooperation from another unit, and fits in a quarterly report to people who will not read a second page. A citizen-outcome measure is slower, crosses boundaries you do not fully control, and often gets worse before it gets better. The incentive gradient runs toward the wrong number, which is why banning it has to be a governance decision rather than a matter of individual judgment.

Rework Is the Hinge

Reyes had a number that made the argument for him. His DMV ran a 31 percent rework rate, meaning nearly a third of applications were rejected for a fixable error and resubmitted. That figure holds up against its own description, since 31 percent is indeed close to a third. Every one of those rejections was a frustrated citizen and a duplicated staff transaction at the same moment. The agency was paying twice to produce an outcome the citizen experienced as a failure. Nothing in that arrangement is efficient, and nothing in it is kind.

Read the rate carefully before you set a target against it. It counts applications rather than people, so a citizen rejected twice appears twice in the numerator, while a citizen who gave up entirely appears neither there nor in your satisfaction survey. That second group is the one you cannot see, and it tends to be made up of the people with the least time and the least capacity to try again. Attacking rework is still the clearest transformative win available to you. Just do not mistake the measured share for the whole of the harm.

Map the Journey, Not the Org Chart

Agencies are organized by function: intake, processing, payments, appeals. Citizens do not experience functions. They experience a journey with a goal attached to it. Renew my license. Get unemployment benefits while I look for work. Register my small business. The journey crosses every silo you have, and the worst pain almost always lives in the handoffs between them, which is precisely the territory that no single unit owns, no single manager is measured on, and no single system records from end to end.

So Reyes mapped the renewal journey the way a citizen travels it and counted the steps. There were eleven. Four of them asked for information the DMV already held on file. Three existed only to compensate for a handoff failure earlier in the chain. The source does not say whether those two groups overlap, so do not add them together; treat them as two separate readings of one map. AI did not appear on his first map at all, and that was the point. You map the journey, find the waste, and only then ask where a machine genuinely helps.

The mapping discipline is what converts a vague mandate to modernize into a list of specific, arguable decisions. A step that exists only because two systems cannot pass a record between them is a candidate for integration, not for automation. A step that asks for data you already hold is a candidate for pre-filling. A step that exists because a statute requires a signature is not a candidate for anything until the statute changes. Sorting the eleven steps into those categories is most of the analytical work, and it is work no vendor can do for you.

Where AI Earns Its Place in a Journey

Once the map exists, the candidates stop being a wish list and start being specific. Four patterns recur across journeys, and each attacks a waste point rather than a cost center.

  • Front-door triage. Understanding what a citizen actually needs in plain language and routing them correctly the first time, instead of requiring them to learn your org chart before they can ask a question.
  • Pre-filling and validation. Using data the agency already holds to complete forms and to catch errors before submission, which attacks the rework rate directly and at its source.
  • Drafting and summarizing for staff. Letting a caseworker handle three times the volume by drafting responses and summarizing case files, while the human keeps the decision.
  • Proactive service. Telling a citizen their license expires in 60 days and renewing it in two taps, so the transaction never gets the chance to become a problem.

None of these is a plan to replace the call center with a bot, and the difference is not stylistic. Each names a point in a mapped journey where effort is currently wasted, which means each can be tested against whether that waste actually fell. A general-purpose assistant bolted onto the front door of an unmapped process has no such test available to it, which is exactly why it survives so long on numbers that describe its own activity.

Notice the boundary drawn in the third pattern: the caseworker keeps the decision. That line is doing more work than it appears to. Drafting and summarizing change how quickly a human reaches a judgment; they do not transfer the judgment. The moment a summary becomes the only thing a decision-maker reads, the decision has effectively moved to the system that wrote the summary, whatever the process document says. Preserve the ability to open the underlying record, make it normal rather than exceptional to do so, and check periodically that someone still does.

Throughput Is Not Experience

The three-times-volume figure above is the one to handle carefully, because it is the shape that fools senior leaders most often. Handling three times the volume is a throughput measure. It describes what the office processed, not what the citizen received. A caseworker clearing three times as many files can coexist with longer end-to-end waits, more appeals and more people giving up, if the additional throughput went into work that was never the bottleneck. An efficiency gain becomes an experience gain only when you measure the citizen's end of the transaction and find that it improved too.

The same warning applies to every deflection metric, which is how Reyes's predecessor came unstuck. Calls avoided rose while service collapsed, because the calls did not stop existing. They turned into three-hour branch visits that no dashboard counted. And an aggregate improvement can hide a worse experience for the people who most needed help: a mean handling time that falls while the hardest cases, the ones with language barriers or contested facts or missing documents, get slower and less visible. Report gains by population as well as in total, or you will not see it happen.

The Accessibility and Equity Floor

A government service is not optional for the citizen, so it has to work for everyone: people with disabilities, people with limited English, people with no smartphone, people with low digital literacy. An AI front door that works only for confident English-speaking smartphone users has not improved service. It has built a faster lane for the people who needed help least, and it has usually done so with money appropriated on the argument that service would improve for everybody. That is the failure you personally will be asked to account for.

Accessibility is a legal obligation as well as a mission one. Section 508 requires that federal information technology be usable by people with disabilities. Which accessibility authority binds a particular agency depends on its jurisdiction, and a state or local program should not assume the federal citation transfers to it, so confirm the applicable requirement with your counsel rather than reasoning from someone else's example. What does not vary is the discipline underneath: every AI-touched journey keeps a fully staffed human path, gets tested with assistive technology, and is measured for whether outcomes are equal across populations rather than only on average.

Reyes added one rule to every project charter, and it did more work than any policy document. If it does not work for our hardest-to-serve resident, it is not done. The value of a rule stated that plainly is that it survives translation into a solicitation, a sprint review and a status report without being softened by anyone. It gives a program manager something to fail against early, when failing is cheap, rather than at launch, when the only options left are to defend the thing or to withdraw it.

Prioritizing the Portfolio

You cannot transform every journey at once, and choosing which to take first is the highest-leverage decision you make. Reyes ranks candidate journeys on a simple grid before he commits budget. Score each journey from 1 to 5 on the rows below, and pursue the high totals first. The grid is not a formula that decides for you. It is a device that forces the equity and waste questions to be asked at the same table, on the same day, as the volume and feasibility questions, which is where they usually go missing.

Scoring dimensionQuestion1 (low)5 (high)
Citizen painHow much frustration or lost time does this cause?Minor annoyanceDays lost, benefits delayed
VolumeHow many people travel this journey?Hundreds per yearHundreds of thousands per year
Waste and reworkHow much duplicated effort exists today?Clean processHigh rejection or rework rate
Equity stakesWho is harmed if it fails, and how vulnerable are they?Low-stakes serviceVulnerable residents, essential benefit
FeasibilityIs the data available and the change deliverable?Data missing, legally blockedData on hand, clear authority

The renewal journey scored high on volume, waste and feasibility but low on equity stakes. Unemployment benefits scored high on everything, equity included. Reyes sequenced unemployment first despite it being the harder build, because the grid told him that is where transformation mattered most rather than where it was easiest. Notice what that decision costs. The easier win would have produced a faster success story and better early numbers. Choosing the harder journey means accepting a slower start and defending it, which is a leadership problem rather than an analytical one.

Be honest about what the total does. Adding five unweighted scores treats a point of feasibility as equal to a point of equity stakes, which is a value judgment the arithmetic hides rather than resolves. Use the total to sort the shortlist and to force a conversation, then read the individual rows before you commit money. A journey that scores low on feasibility because the data sits in another agency is telling you to go and negotiate for the data, not to abandon the journey. The grid surfaces the argument; people still have to have it.

Govern the Transformation, Do Not Just Launch It

At your level the failure mode is not a bad pilot. It is a hundred uncoordinated pilots that never add up to a transformed agency, each defensible on its own terms and none of them changing what a citizen experiences. Three governance moves keep a cross-cutting program coherent, and all three work by constraining what an individual project is allowed to optimize for.

  • One outcome scorecard, agency-wide. Pick a small set of citizen-outcome measures, such as end-to-end completion time, first-contact resolution, rework rate and satisfaction broken out across populations, then hold every project to the same scorecard. Ban deflection-only metrics.
  • A single intake and review gate. Every AI-touched journey passes one charter review covering value, the equity floor, accessibility and risk. This is the control that stops the next call-deflecting chatbot before it is funded rather than after it is launched.
  • Reusable platforms over one-off builds. A shared identity service, a shared document-understanding capability and a shared human-fallback pattern let each new journey move faster and behave consistently, and they mean an accessibility fix made once is a fix made everywhere.

A gate is a control, not a guarantee. A charter review catches the projects that arrive at it and describe themselves accurately, which is most of them and not all of them. Work funded as a minor enhancement, procured as a feature of something already bought, or run as an experiment that nobody classified as a deployment will route around the gate without anyone intending to. Sample what actually shipped, periodically and without warning, and compare it against what came through the review. The gap is your real governance coverage.

The scorecard has the same limitation, which is worth saying to the executive team out loud. A measure you chose is a measure someone can manage toward. First-contact resolution improves if staff are pressed to close cases at the first touch whether or not the citizen's problem is finished, and completion time improves if the hardest applications are routed into a queue that the clock does not see. Pair every headline measure with one that would move in the wrong direction if the number were being gamed, and read them together.

Anti-Patterns to Avoid

  • Buying a front door for a broken journey. A conversational layer over an unmapped process moves the frustration rather than removing it. The citizen who cannot get an answer does not disappear; they arrive in person, later, angrier, and in a channel that costs more to serve. Map first, then decide what to buy.
  • Counting deflection as service. Calls avoided is a measure of what your agency stopped doing. It rises when a citizen gives up, which is the worst outcome in the set. Any metric that improves when service fails belongs off the scorecard entirely, not merely lower down it.
  • Reporting throughput as citizen experience. Files cleared, documents processed and cases handled per officer describe internal capacity. They become an experience claim only if the end-to-end time a citizen waits, or the share who complete their goal, moved in the same direction.
  • Reading the average and stopping there. An aggregate improvement can hide a worse experience for the people who most needed help. Break every headline number out by population, channel and case complexity before you present it, because the exception is where the political and legal risk lives.
  • Treating the staffed human path as something to fund later. The fallback is not a phase two. It is the thing that makes the automated path safe to launch, and an agency that ships the automation while the human route is still understaffed has chosen its most vulnerable residents as the group that absorbs the risk.

Practice Prompts

  • Map one journey end to end. Pick your highest-volume citizen transaction and count the steps as the citizen experiences them, not as your org chart describes them. Mark each step as duplicated data you already hold, a handoff repair, or a genuine requirement. Do not name a technology until the map is finished.
  • Find your rework rate. For one service, calculate the share of submissions rejected and resubmitted for a fixable error. Then ask who is missing from that figure, specifically how many people started and never came back, and what it would take to measure them.
  • Audit your scorecard for deflection. List every measure your AI and modernization projects currently report against. Circle each one that would improve if a citizen gave up. Propose the citizen-outcome measure that should replace it, and identify who has to approve the change.
  • Run the prioritization grid. Score several candidate journeys on citizen pain, volume, waste, equity stakes and feasibility. Where the highest total is not the easiest build, write the paragraph you would use to defend that sequencing to a budget committee.
  • Test the floor. Take one live AI-touched service and walk it as a resident with no smartphone, then as a resident using assistive technology, then as a resident with limited English. Record where each attempt stops. Bring the three failure points, not a summary of them, to your next charter review.

Reflection

Which of your current efficiency measures would improve if a citizen gave up partway through? If your agency's average handling time fell next quarter, what would you look at to find out whether the hardest cases got slower while it fell? Who in your organization is accountable for a journey rather than for a function, and if the answer is nobody, what would it take to create that role? And if your hardest-to-serve resident tried your newest service tomorrow, at what step would they stop, and do you know that or are you assuming it?

Glossary

  • Rework rate. The share of submissions rejected for a fixable error and resubmitted. It counts applications rather than people, so it understates the harm to anyone who abandoned the process.
  • Failure demand. Contact generated because something went wrong earlier, rather than because the citizen has a new need. It is waste that arrives disguised as customer service volume.
  • Deflection metric. Any measure of contacts avoided. It improves both when a problem is genuinely resolved and when a citizen gives up, which makes it unusable as a service measure on its own.
  • Citizen journey map. A record of the steps a person takes to reach a goal, ordered as they experience them and crossing whatever organizational boundaries lie in the way.
  • End-to-end completion time. Elapsed time from a citizen's first attempt to the point their goal is achieved, including waits, handoffs and resubmissions rather than only the moments an officer was working.
  • Equity floor. The set of conditions a service must satisfy for every population before it is considered done, including a staffed human path, assistive-technology testing, and outcomes measured across groups.

Closing

Reyes did not solve his DMV by buying a better chatbot. He solved the framing first, which let him see that his rework and most of his complaint volume were the same problem seen from two directions. The eleven-step map, the rework rate and the charter rule about the hardest-to-serve resident are unremarkable instruments individually. Together they made it impossible for a project to claim success while the citizen's day got worse, and that is the whole of the discipline.

Your version of this will not look like his, because your journeys and your constraints are different. What transfers is the order of operations. Map before you buy. Attack the waste that is also the frustration. Hold the floor for the people the service was hardest for. Sequence by where transformation matters, not by where it is easy. And measure the citizen's end of the transaction, because every metric that measures only your end will eventually reward you for a service the public cannot use.

Key Takeaways

  • Efficiency and experience are not a trade-off. Most citizen frustration is hidden waste, which means the cheap fix and the kind fix are usually the same fix.
  • Map the journey, not the org chart. Citizens experience goals that cross every silo, the worst pain lives in the handoffs, and that is where AI earns its place.
  • Attack rework first, and read the rate honestly. Pre-filling and validation cut cost and frustration together, but a rate counted in applications says nothing about the people who gave up.
  • Throughput is not experience. Volume handled, files cleared and calls avoided describe your operation. Only a citizen-side measure can tell you the service improved.
  • An aggregate gain can hide a worse experience for the people who most needed help, so break every headline number out by population, channel and complexity.
  • Keep an accessibility and equity floor. A staffed human path, assistive-technology testing and outcomes measured across populations, with the applicable accessibility authority confirmed for your jurisdiction.
  • Govern with one scorecard and one gate, and audit both. A gate catches what arrives at it and a metric can be managed toward, so sample what shipped and pair each measure with one that would expose gaming.

Frequently Asked Questions

Our chatbot cut call volume sharply. Is that not a real saving?

It is a real change in call volume, which is not the same claim. Calls stop for two very different reasons: the citizen got what they needed, or the citizen gave up and went somewhere more expensive, such as a counter, a legislator's office or an appeal. Before you book the saving, look for the displaced demand in your branch traffic, your complaint log and your abandonment data. If it is not there, you have a saving. If it is there, you have a transfer.

How do we know whether an efficiency gain reached the citizen?

Measure the transaction from the citizen's first attempt to the moment their goal is achieved, including the waits and the resubmissions, and compare that figure before and after. Internal cycle times, files cleared per officer and queue depths all improve without that number moving. The test is not whether the office got faster; it is whether the person waiting on the outcome got their result sooner and with fewer attempts.

Does Section 508 apply to my state or local agency?

Do not assume it does. Section 508 requires that federal information technology be usable by people with disabilities, and the accessibility authority that binds a particular program depends on its jurisdiction. This lesson deliberately does not tell you which one applies to you, because getting that wrong is a legal exposure rather than a documentation error. Ask your counsel to identify the governing requirement in writing before you set an accessibility standard in a solicitation.

Should we start with the easiest journey to build momentum?

Sometimes, but decide it explicitly rather than by default. Scoring candidates on pain, volume, waste, equity stakes and feasibility often puts the hardest journey at the top, as unemployment benefits outranked license renewal here. If you take the easier one first, say plainly that you are buying delivery confidence at the cost of deferring the journey where transformation matters most, and set a date for the harder one.

We passed the charter review. Are we safe to launch?

The review is a control, not a clearance. It examines what was presented to it, on the day it was presented, by people describing their own project. Work reclassified as an enhancement, bought as a feature of an existing contract, or run as an experiment can reach production without ever meeting the gate. Treat a passed review as evidence that the known risks were considered, and back it with unannounced sampling of what actually shipped.