Communicating AI Success Stories
Dolores Ramirez ran communications for a county Department of Social Services that had just shipped its first real AI win. A document-triage model had cut the wait for food-assistance decisions from 19 days to 6. She wrote a press release she was proud of: "County deploys cutting-edge artificial intelligence to revolutionize benefits processing." It ran on a Tuesday. By Thursday a local reporter had a different headline: "County uses secret algorithm to decide who eats." The story was wrong on the facts, but Dolores had handed the reporter the framing on a plate. She had told a technology story to an audience that only cares about a trust story.
This lesson is about telling AI success stories that build trust instead of inviting suspicion. The win was genuine. The communication turned it into a liability. As a program director or senior manager you will face the same fork, and a great deal of the difference between the two outcomes sits in how the story is told rather than in what the system did.
Why Communication Decides What Happens Next
Government AI is contested ground. Stories of bias, discrimination and failure dominate the coverage, and they dominate it for reasons that are frequently justified. Genuine successes exist alongside them: systems that improve services, reach people who were being missed, and free staff for work that needs judgment. Those stories do not tell themselves, and in the absence of your account of what a system does, the public will reasonably assume the worst available explanation.
This is not a vanity concern. Public perception shapes political support, and political support shapes policy. If the public distrusts government AI, the policy response tightens, sometimes past the point where useful systems can be deployed at all. If the public understands a system and has reason to trust the people running it, there is room to keep improving it. Communication is therefore part of the operating model of the program, not a wrapper applied after delivery.
Scale is what makes this different from corporate communication. A management approach that works for a fifty-person organization does not transfer to a service accountable to millions of residents across diverse populations and contexts. Your audience includes people who cannot opt out of the service, cannot switch to a competitor, and did not choose to interact with a model. That is the burden the story has to carry.
One Win, Two Very Different Audiences
Every AI success story has to land with two groups at once, and they want opposite things. Your internal audience, meaning the agency head, the budget office and peer divisions, wants evidence the investment paid off so it will fund the next one. Your external audience, meaning residents, advocates, journalists and oversight bodies, wants evidence that the public was protected rather than processed. Lead with technology and you reassure neither: the internal crowd cannot repeat a story about embeddings, and the external crowd hears "secret algorithm."
The fix is to make the human the subject of the sentence and the AI the tool in the predicate. Not "we deployed AI to process applications" but "families now get a food-assistance decision in 6 days instead of 19, because a tool helps caseworkers sort documents faster." Same fact, different subject. The first version is about the agency's cleverness. The second is about the resident's life, with the technology in its proper supporting role, and it is the version a reporter can quote without reframing.
What Makes a Story Worth Telling
Not every improvement is a story, and publishing weak ones spends credibility you will need later. Five properties separate a story that survives contact with a skeptical reader from one that does not: a concrete outcome rather than an adjective; an effect on identifiable people rather than an aggregate; lessons learned, including the ones that were uncomfortable; a plain explanation of how the system works; and an honest acknowledgment of its limitations.
The last two are the ones agencies drop under time pressure, and they are the two doing most of the work. A story with a number and no mechanism reads as a claim. A story with a mechanism and no limitation reads as marketing. Both invite the reader to supply the missing part themselves, and the version they supply is almost never the flattering one.
What a Success Story Can and Cannot Demonstrate
This is the part most communications plans get wrong, and it is worth being precise because the error is a credibility risk rather than a stylistic one. A success story is evidence. It is not proof, and the difference will be raised by exactly the audiences you most need to convince.
A pilot result shows what happened in the pilot, under the conditions of the pilot, with the population the pilot covered. It supports a claim that the approach is worth extending, and it does not establish that the system works, because working is a claim about a wider population and a longer period than a pilot observes. Say what the pilot showed and where it ran. Then say what you still expect to learn at scale, which in most agencies means different document quality, different languages and legacy integrations the pilot never touched.
A metric that improved after deployment is a before-and-after comparison, and a before-and-after comparison establishes attribution only to the extent that nothing else changed at the same time. Dolores's 19 days to 6 arrived alongside a staffing change and the end of a seasonal surge. That does not make the improvement fake. It means the honest sentence is that the wait fell from 19 days to 6 in the period after the tool was introduced, and that the agency's estimate of the tool's contribution rests on stated assumptions. Where you can compare against a group that did not get the tool, say so, and say how comparable the groups were.
A testimonial from a resident or a caseworker is a real account of one experience. It shows that the improvement is possible and what it feels like, which is worth a great deal in a hearing room. It does not show frequency, and a story presented as though it did will be answered by the first counter-story a reporter finds. Pair every testimonial with the aggregate figure that shows how typical it is, and if the aggregate is not flattering, publish it anyway rather than letting someone else discover it.
Why Transparency Is the Whole Story
In government, a success story that hides how the system works reads as a confession rather than a celebration. Dolores's word "secret" did the damage. The federal direction here is clear and worth borrowing at any level of government: the Office of Management and Budget's memorandum M-24-10, Advancing Governance, Innovation, and Risk Management for Agencies' Use of Artificial Intelligence, issued in 2024, requires covered agencies to publicly inventory their AI uses and to apply transparency and protection practices to systems that affect the public. The principle beneath it is simple. If a resident could be affected by a model, they have a right to understand it in plain terms.
So transparency is not a risk to be managed around in your story. It is the story. The strongest line Dolores could have written was the one she left out: the tool only sorts and flags documents, every benefits decision is still made by a trained caseworker, and any resident can ask how their case was handled. That sentence answers the most common form of the "secret algorithm" fear before a reporter raises it, though it does not settle the question for a reader who has other reasons to doubt the agency, and it should not be expected to.
The governance version of the same principle is to publish the decisions, not only the outcomes. Say why the system was deployed, what safeguards are in place, how it will be monitored, and who is accountable for it. A story that cannot survive a hostile reporter was never finished. Write the skeptic's question into your own announcement, and answer it in your own words while you still control the framing.
A Five-Beat Structure That Builds Trust
Use a fixed structure for any external AI success story. It keeps the human first and disarms suspicion in order.
- The person. Open on a resident or a frontline staffer, not the agency. "A single parent applying for food assistance used to wait nearly three weeks for an answer."
- The change. The concrete before and after, with a real number, and the period it covers. 19 days to 6.
- The role of the tool. Exactly what the AI does and, crucially, what it does not do. It sorts; it does not decide.
- The guardrail. The human in the loop, the accuracy checks, and the recourse path. Name them, and name who is accountable.
- The honest limit. One thing you are still watching or improving. This is the beat that earns belief.
Beat five feels counterintuitive and does more than the rest combined. Saying that you audit the tool monthly for bias because these systems can drift makes the other claims easier to believe, because it demonstrates that the agency is looking rather than asserting. Communications that read as flawless read as spin, and a reader who suspects spin discounts the numbers first.
The AI Success Story Builder
Draft any announcement by filling these rows before you write a single polished sentence. If a row is blank, you are not ready to publish; you are ready to investigate.
| Beat | Question to answer | Worked example (benefits triage) |
|---|---|---|
| Person | Whose life changed? | A parent applying for food assistance |
| Change | Before and after, real numbers, stated period | Decision time went from 19 days to 6 |
| What the tool does | One plain sentence | Sorts and flags case documents so caseworkers find what they need faster |
| What it does NOT do | The line that prevents fear | It does not approve, deny, or score any application |
| Human guardrail | Who stays accountable | A trained caseworker makes every decision and signs it |
| Recourse | How a resident pushes back | Any applicant can request a review and an explanation |
| Attribution basis | Why you credit the tool | What else changed in the same period, and what comparison you have |
| Honest limit | What you still watch | Monthly bias and accuracy audit; results published yearly |
| Internal version | The value line for leadership, with its rate | 13 days faster per case across roughly 2,400 cases a year, priced at the agency's own loaded cost per case-day |
Note the last row and what it does not contain. The internal cut has to carry a money figure, and that figure has to arrive with the rate it was computed from. Thirteen days faster across about 2,400 cases a year is a defensible input pair; a rounded dollar total with no stated cost per case-day is not, because no reader can rebuild it and no auditor will accept it. Publish the rate next to the total, or publish the days and the caseload and let the finance office price them.
The same win has an internal cut and an external cut drawn from the same facts. To the agency head you lead with the cleared backlog and the priced time. To the public you lead with the parent and the 6 days. You never contradict yourself across the two; you choose which true sentence goes first.
Your Staff Are the Channel You Forgot
Most of the explaining will not be done by you. It will be done by a caseworker at a counter, a call-center agent, or a supervisor at a community meeting, answering a question nobody drafted talking points for. Government staff need to understand the system well enough to explain it to residents in their own words, which means they need training on what it does and what it does not do before the press release goes out, not after.
Handled well, staff are the most credible ambassadors the program has, because they are the ones the public actually encounters. Handled badly, they become the source of the first damaging quote, and the quote will be accurate about their confusion. The internal communication plan is therefore not the budget deck. It is the plain-language brief that lets the person at the counter answer the question without guessing.
Match the Story to the Channel
Where you tell it shapes how much room you have, and different formats reach genuinely different audiences. A short press release cannot carry the guardrail beat properly, so pair it with a plain-language web page that can. Blog posts, short videos, infographics and town halls each reach people the others miss, and each has different requirements for length, reading level and review.
- Press release: the person and the change, pointing to the detail page for the guardrails.
- Public web page or AI inventory entry: the full five beats, written for a general reading level. This is your defense against the word "secret."
- Internal memo or budget deck: the value cut with its rate shown, and the public-trust framing attached so leadership tells the same story you do.
- Community briefing or town hall: the honest limit and the recourse beats first, because the room you most need to win is the skeptical one and it will not wait politely for beat five.
Timing matters as much as channel. Announce the guardrails before or alongside the win, never after the criticism. Dolores published the trust details only once the bad story had run, which made them read as defensive. The same words, published on the Tuesday, would have read as confident. You cannot retrofit transparency; you can only schedule it.
Three Practices That Build Standing Trust
Individual announcements work better when they land on ground you have already prepared. Three practices do that preparation, and each one is a way of being believed before you need to be.
The first is a public AI transparency dashboard. Several governments publish a single page listing which agencies use AI, what each system does, how it is performing against key metrics, how to appeal a decision, what safeguards are in place, recent issues and their resolutions, and where to send a question. The point is not the design. It is that citizens and advocates can check for themselves rather than take your word, and that when an issue arises they learn about it from you rather than from a news report.
The second is proactive disclosure when something goes wrong. When a system made systematic errors affecting a large number of residents, one government disclosed immediately, stated how many people were affected and how, described what went wrong and why, explained the remedy, and set out what it had changed to prevent a repeat. That is a harder document to write than a success story and it does more for standing trust than several good announcements, because it is the case where silence was the available alternative.
The third is a community advisory board with genuinely mixed membership: affected community members, civil rights advocates, ethicists, technologists and agency staff. Boards of this kind review systems before deployment, comment on design and governance, surface risks, recommend safeguards, and keep watching after launch. Where they work, they tend to improve the design and give decisions a legitimacy the agency cannot manufacture on its own. Where they are convened to approve a decision already made, participants notice quickly and the arrangement costs more credibility than it buys.
Measure Whether the Story Worked
A communications push is itself a program with outcomes you can track. Watch the share of coverage that repeats your human-first framing against the share that repeats the suspicion framing. Watch the volume and tone of resident questions after an announcement. Watch whether the next budget cycle references the win you communicated. If coverage keeps drifting toward suspicion, the problem is usually upstream in the telling, and the builder above is where you fix it.
Dolores rewrote her approach for the next milestone, a year later, when the same model expanded to housing assistance. She opened with a named applicant's shorter wait, stated plainly that caseworkers still made every decision, and published the audit schedule the same morning. That time the local paper ran her framing closely. One outlet on one story is not proof that the method works, and she treated it as a single encouraging data point rather than a validated formula. The underlying win was no bigger than the first one. The telling was better.
Anti-Patterns
- Presenting a pilot as proof. A pilot result describes the pilot. Announcing that a system "works" on the strength of one bounded test invites a reasonable challenge you cannot answer, and the challenge lands after you have already committed publicly. Say what the pilot showed, where, and over what period.
- Treating an improvement as an attribution. A metric that moved after deployment is not the same as a metric that moved because of it. Name what else changed in the same window, and state the comparison your attribution rests on.
- Using a testimonial as evidence of frequency. One resident's experience shows what is possible, not what is typical. Unaccompanied by the aggregate, it will be answered with a counter-story, and the counter-story will be equally true.
- Publishing a dollar figure without its rate. An internal value claim that no reader can rebuild from stated inputs is an assertion. Show the caseload, the time saved and the cost basis, or show the operational figures alone.
- Sequencing guardrails as a rebuttal. Transparency details published after criticism read as damage control regardless of their content. The same sentences published alongside the announcement read as confidence.
- Treating communication as optional. The work gets deferred because it feels like a nice-to-have next to delivery. Meanwhile suspicion accumulates, and it is far more expensive to answer once it has hardened than to prevent.
- Copying another jurisdiction's story wholesale. A framing that worked elsewhere was built for a different population, a different history with the agency, and a different press environment. Use it as a guide, not a template.
- Publishing with no owner. If nobody is named as accountable for the story after it goes out, questions go unanswered, corrections do not happen, and the second-day coverage is written without you. Assign the owner and the review cadence before you publish.
Practice Prompts
- Assess your current state. What has your agency published about its AI use? Could a resident find it, and would they understand it without help?
- Identify the gaps. Compare what you publish now with the five beats. Which beat is missing most often, and what would it take to fill it?
- Run the builder on a real win. Take an actual result from your program and complete every row, including the attribution basis and the honest limit. Note which row you could not fill.
- Write the hostile headline. Draft the worst fair headline a reporter could write from your announcement, then revise the announcement until that headline no longer follows from it.
- Build the two cuts. Write the internal version and the external version of the same win. Check that no sentence in either would embarrass you if quoted next to the other.
- Brief the counter. Write the half-page a frontline staff member would need to answer a resident's question about the system, in their own words, without a script.
- Set the success metrics. Decide in advance what coverage, question volume and question tone would tell you the communication worked, and who reviews them afterwards.
Reflection
Think about the last AI-related announcement your agency made, or the one you are about to make. Which of the five beats did it carry, and which did it skip? If a reporter had asked what the system does not do, would the answer have been in the release or in someone's head? Now consider the harder question: what claim in that announcement is doing more work than the evidence behind it can support, and what would the honest version of that sentence say instead? Write that sentence. It is almost always shorter, and it is the one you would rather be quoted on.
Glossary
- AI use case inventory. A published list of the AI systems an agency uses, what each does, and how it is governed. Under M-24-10, covered federal agencies are expected to maintain and publish one.
- Attribution. The claim that an observed change was caused by a specific intervention. It requires a comparison or a stated set of assumptions, not just a before-and-after pair.
- Human in the loop. A design in which a person makes or reviews the consequential decision and is accountable for it, with the model supporting rather than deciding.
- Recourse. The route by which a person affected by a decision can question it, request an explanation, and have it reviewed. In a success story it is a beat, not a footnote.
- Proactive disclosure. Publishing a problem, its scope and its remedy before anyone else reports it. The practice that builds the most standing trust and is hardest to do on the day.
- Community advisory board. A standing group of affected residents, advocates, subject experts and agency staff that reviews AI systems before and after deployment and advises on safeguards.
- Transparency dashboard. A public page consolidating which systems are in use, how they perform, what safeguards apply, how to appeal, and where to ask questions.
- Internal cut and external cut. The two versions of one true story: the value framing for leadership and the human framing for the public, drawn from the same facts and never contradicting each other.
Related Lessons
The measurement behind any story you tell is built in Measuring AI Impact and AI Metrics and KPIs for Government, which is where the figures in your release should come from rather than being assembled for the release. Communicating AI Projects to Leadership covers the internal cut in depth, and Stakeholder Management and Communication covers the audiences beyond press and leadership. Building Trust Through Transparency and Transparency: Citizens' Right to Know develop the disclosure obligations this lesson borrows, and Moving from Pilot to Production matters because most premature success stories are told about systems that have not yet made that transition. When the story you have to tell is a failure, Crisis Management for AI Failures is the companion piece.
Closing
Communicating a government AI success is not decoration on top of the delivery. It determines whether the next system gets funded, whether staff can defend the current one at a counter, and whether residents encounter your account of what happened or someone else's. The discipline is narrow enough to remember: make the person the subject, say what the tool does and does not do, name who stays accountable, state what you are still watching, and claim only what your evidence actually supports.
Dolores's second announcement was not more skilful than her first in any literary sense. It was more honest about the mechanism and more careful about the claim, and it was published in the right order. That is generally all the difference there is between a story that builds trust and one that hands away the framing.
Key Takeaways
- Make the human the subject. The resident or frontline worker is the story; the AI belongs in the predicate, never the headline.
- Claim only what the evidence supports. A pilot shows what happened in the pilot, an improvement shows a change rather than a cause, and a testimonial shows one experience rather than a frequency.
- Treat transparency as the story, not a risk. Stating what the tool does, what it does not do and who stays accountable removes the easiest version of the "secret algorithm" framing before it forms.
- Write the skeptic's question into your own announcement. A story that cannot survive a hostile reporter is unfinished, and the hostile version is cheap to draft yourself.
- Always include the honest limit. Naming what you still audit makes every other claim easier to believe than a release that sounds flawless.
- Keep one internal cut and one external cut of the same facts. Lead with value for leadership and with the human outcome for the public, and never let the two contradict each other.
- Show the rate behind any money figure. A value claim a reader cannot rebuild from stated inputs will not survive review, however sound the underlying work was.
- Brief your staff before the public. The people at the counter do most of the explaining, and they are either your most credible ambassadors or the source of the first damaging quote.
- Sequence guardrails first. Publish the human-in-the-loop and recourse details alongside the win, because the same words read as defensive once criticism has landed.
- Build standing trust before you need it. A public inventory or dashboard, proactive disclosure of problems, and a genuine advisory board make individual announcements land on prepared ground.
Frequently Asked Questions
We only have pilot results. Can we announce anything?
Yes, provided the announcement describes a pilot. State where it ran, over what period, with which population, and what you expect to learn at scale. The problem is never that a pilot is small; it is that a pilot described as a finished system creates a claim you cannot defend the moment anyone asks about the wider population.
Legal review keeps stripping the honest limit out of our releases. What do we do?
Bring them the alternative rather than the principle. The limit is what makes the rest of the release defensible, and a release that reads as flawless generates the follow-up questions that create real exposure. In practice the disagreement is usually about wording, and a limit stated as a monitoring commitment tends to survive review where one stated as an admission does not.
How do we credit the AI when other things changed at the same time?
Describe the change and the period, name the other factors, and state what your estimate of the tool's contribution assumes. If you have any group that did not receive the tool, use it and say how comparable it was. Overclaiming here is what turns a genuine improvement into a correction later.
Do we have to publish problems as well as successes?
Treat it as the practice with the highest return rather than as an obligation to litigate. Residents learn about failures either way, and the version they encounter first sets the frame permanently. Where agencies have disclosed early with scope, cause and remedy, the outcome has generally been public understanding rather than scandal.
What reading level should the public page use?
Plain language, tested on someone outside the program. The specific grade level matters less than a simple check: can a person who has never heard of your project read the page once and correctly say what the system does, what it does not do, and how to challenge a decision? If not, the page is not finished, however accurate it is.
Who should own the story after it is published?
A named person, with a stated review point. Announcements generate questions, corrections and second-day coverage, and each of those needs someone with the authority to answer. An unowned story stops being yours the moment it leaves the building.
Skill.re