Communicating AI Projects to Leadership
Marcus Bell, a project lead at a state Department of Motor Vehicles, had eleven minutes on the agenda. He walked into the deputy director's quarterly review armed with forty-three slides about his license-renewal chatbot: model architecture, token costs, an F1 score of 0.91. Four minutes in, the deputy director put down her pen and asked one question. "Marcus, will this shorten the line at the Riverside office, or not?" He did not have a clean answer. The slides had everything except the one number she could carry into her own budget meeting.
The project was real, the work was good, and he still walked out without the funding he came for. That gap is the whole subject of this lesson. Most AI projects in government do not fail in the lab. They fail in the conference room, where a technically sound effort gets translated into the wrong language for the person holding the money. Leadership does not approve, fund or defend what it does not understand, and communicating an AI project to executives is a distinct skill you can learn the way Marcus eventually did.
Why an AI briefing is not an IT briefing
Leaders already know how to hear a traditional IT project. You ask for a server, a license, or a software upgrade. The cost is fixed, the outcome is binary because it either works or it does not, and the risk is mostly about delivery dates. AI breaks all three assumptions, and if you brief it like ordinary IT you will confuse the room and lose the decision without ever being contradicted.
Three differences matter most. First, AI is probabilistic. It is right most of the time rather than all of the time, which leaders interpret as "broken" unless you frame it. Second, AI changes after launch. The data shifts, the model drifts, and the performance you reported in March may not hold in September. Third, AI carries public-trust risk that a payroll upgrade never does. A wrong answer from a benefits-eligibility model can become a news story. Your briefing has to acknowledge all three without drowning the listener in them.
The audience constraint compounds all of that. Government executives typically have limited technical background, they think in mission, budget and risk terms, and they decide on constrained timelines. If communication is too technical, leadership is lost and cannot decide. If it is oversimplified, leadership decides on incomplete information. The target is communication that is accurate, understandable and actionable at the same time, which is harder than any of the three alone and is the actual skill being taught here.
Know which of three people you are briefing
"Leadership" is not one person. In a typical agency you are speaking to some mix of three roles, and each wants a different sentence first. Getting this wrong is not a style error. It determines whether your opening bought attention or spent it.
- The mission owner, often a division director, wants to know what changes for the public and for staff. Lead with the outcome.
- The resource owner, a CFO or budget officer, wants the cost, the savings, and when the two cross. Lead with the money.
- The risk owner, a CIO, general counsel or chief privacy officer, wants to know how this fails and who is accountable. Lead with the guardrails.
Marcus had built a CFO deck and walked into a mission-owner meeting. The deputy director did not care about token costs yet. She cared about the line at Riverside. Match your opening sentence to the chair, not to your comfort zone, and if you genuinely do not know who is chairing, that is a question to ask before the meeting rather than a risk to absorb during it.
Behind those three sits a fourth audience you are always writing for even when it is not in the room. Government has multiple layers of oversight: executive sponsors, Congress, the inspector general, the Government Accountability Office and civil rights offices all review AI projects. Your ability to describe the work in compliance and accountability terms is what determines whether it survives scrutiny later. A briefing that would embarrass you in an IG review is a briefing to rewrite now, while the only cost is your own time.
Translate the metric, do not recite it
An F1 score of 0.91 means nothing to a deputy director, and it is your job to translate it rather than her job to decode it. Translation means converting a model number into a mission number she can repeat upward, and it starts with picking the measure that answers her actual question. Marcus quoted an F1 score, which combines precision and recall into a single index. It is not the share of answers that are correct, and presenting it as though it were is one of the most common ways a translation goes quietly wrong before anyone notices.
Watch the move done properly. The chatbot answers renewal questions correctly 91 percent of the time. The Riverside office handles roughly 1,200 renewal walk-ins a week, and call-center data shows about 60 percent of those are simple status or document questions the bot can handle. So the bot can deflect around 720 questions a week, get 655 of them right, and route the rest to a human. That frees an estimated 12 staff-hours a week at one office. Across nine offices that is roughly 108 hours, about three full-time positions' worth of time, redirected from answering "what documents do I bring" to clearing the actual line.
Leaders do not fund accuracy. They fund the shorter line that accuracy buys. Every step of that chain is checkable, which is exactly what makes it persuasive: a skeptical CFO can test the walk-in count, the 60 percent share and the hours estimate independently, and each one that survives strengthens the next. State which links are measured and which are estimated, because a chain presented as uniformly solid falls apart entirely when one link turns out to be a guess.
The same discipline applies to every technical phrase you are tempted to say out loud.
| What the technical team says | What the executive needs to hear |
|---|---|
| Supervised learning with logistic regression | We train the system on historical decisions so it can predict future ones |
| 89% accuracy, AUC 0.92, F1 score 0.87 | Out of 100 decisions the system gets 89 correct, against 87% for our current manual process |
| Demographic parity metrics show equalized odds at under 5% disparity | Approval rate is 80% for Group A and 78% for Group B, so the treatment is even across populations |
| False positive rate under 2%, false negative rate under 5% | The system finds more than 95% of the cases it should flag, and wrongly flags under 2% of the ones it should not |
| The model was validated on a hold-out test set with cross-validation | We tested it on data it had never seen, to be sure it works on new cases and not just familiar ones |
That fourth row is worth pausing on, because the obvious translation of it is wrong. A false negative rate under 5% means the system finds more than 95% of the cases it should catch. The 98% figure that people reach for comes from the false positive rate and describes performance on the other class entirely. Swapping those two overstates the system's ability to catch what matters, in front of the audience least equipped to notice, and it is the kind of error an oversight review does eventually find.
The governing principle is simple. Explain what the system does and why it matters. Do not explain how the math works. Executives do not need gradient descent or regularization. They need to know what the system does and whether it can be trusted, and every minute you spend on mechanism is a minute not spent on the decision you came to get.
Federal guidance points the same way. The Office of Management and Budget's 2024 memo on agency AI use (M-24-10) asks agencies to identify AI uses that affect the public and to track real outcomes and impacts, not just whether a model runs. The National Institute of Standards and Technology AI Risk Management Framework, the main U.S. playbook for trustworthy AI, frames everything in terms of measurable harms and benefits.
The fifteen-minute spoken briefing
When you have fifteen minutes with a busy executive, spend it in a fixed order. The structure below allocates roughly two minutes to each section with a longer stretch on risk, and its value is that the executive learns what to expect from you and can compare your project against others on identical terms.
Problem statement. Quantify what is broken and lead with numbers rather than technology. Weak: "We need AI to improve our system." Strong: "We process 10,000 benefit applications monthly. Average processing time is 40 days. 30% are delayed beyond our 30-day service standard. Managing the backlog costs 2 full-time equivalent staff daily, roughly $180,000 annually in staff time alone." The executive now understands the problem concretely enough to weigh it against everything else on her list.
Proposed solution. Plain English, no jargon. Weak: "Supervised learning classification algorithm with gradient boosting optimization." Strong: "An AI system that learns from past decisions to flag clear-cut cases for fast-track processing and routes complex cases to specialists, which makes decisions more consistent and frees specialist time from routine work."
Expected outcomes. Specific, measurable, tied to mission value. Weak: "Increase accuracy and improve fairness." Strong: "Reduce processing time from 40 days to 24 days, a 40% improvement. Process 50% more applications with existing staff. Maintain 99% accuracy on fast-tracked decisions. Keep demographic disparity under 5 percentage points across all groups."
Investment and return. Clear costs, timeline and payback. Weak: "We will need to spend some money on an AI system." Stronger: "Total investment $500,000, made up of $300K vendor and $200K staff and training, over a 12-month timeline." Then be careful with the payback line. A payback period has to be computable from the numbers you have just shown. If the only quantified cost of the problem is $180,000 a year and the investment is $500,000, then divide one by the other and say what you get, because the executive who is going to defend this number will do exactly that arithmetic and you want to have done it first.
Risks and mitigation. Give this section the most time. Weak: "Our system will be perfect." Strong: name each risk with its countermeasure. Accuracy may not reach 99% on all cases, mitigated by a phased rollout starting with low-stakes cases and the ability to roll back. Staff may resist the new system, mitigated by training, communication and gradual rollout. The system may show bias against some populations, mitigated by pre-deployment bias testing, continuous monitoring and fairness constraints in training. Executives trust the person who names the failure modes before being asked.
Next steps. One clear ask and a timeline. Weak: "We will keep working on this and let you know." Strong: "We need your approval to proceed with the 12-month project, with quarterly checkpoints, and preliminary results after 3 months that will inform the go or no-go decision for full deployment." End there. Do not end on "any questions," which hands the decision back to a room that was ready to give it to you.
The one-page artifact
The spoken briefing needs a written counterpart that survives the meeting. Use a fixed five-part structure on a single page or a single slide, with the detail in an appendix you bring but do not present. The outcome goes in one sentence with no tools and no acronyms. The cost and the crossover follow: what it costs to build, what it costs per year to run, and the month savings overtake spending. Then the confidence: what the system gets right, what it gets wrong, what happens to the wrong answers, and the name of the human backstop. Then the top risk with its owner and the rollback. Then the single decision you need today.
Copy the table below onto one page and fill every row in plain language. If you cannot fill a row, that is a conversation to have before the briefing rather than during it.
| Section | What goes here | Worked example (DMV chatbot) |
|---|---|---|
| Outcome | One mission sentence, no jargon | Cut renewal wait times by redirecting 3 FTE of staff time across 9 offices |
| Cost to build | One-time dollars and weeks | $95,000, 14 weeks |
| Cost to run | Annual recurring dollars | About $31,000 per year for hosting, monitoring and model usage |
| Crossover | When savings beat spending, with the dollar value of the freed time shown | Month 11, on current volumes |
| What it gets right and wrong | Plain accuracy plus the backstop | About 91% correct; uncertain answers route to a human within 30 seconds |
| Top risk and owner | Most likely failure, named owner, rollback | Wrong legal-deadline information; owned by Licensing Ops; bot disabled in under 1 hour if the error rate exceeds 5% |
| Drift plan | Who re-checks performance, how often | Monthly accuracy review by the data team, reported in this same forum |
| The ask | One specific decision today | Approve a 90-day pilot at the Riverside office |
One row on that page does more work than the rest combined, and it is the crossover. A crossover month is only as good as the dollar value you attach to the time being freed, and that value is what a resource owner will ask for first. Show the salary or loaded-cost assumption behind "3 FTE of staff time" on the page itself. A crossover month with no visible arithmetic behind it reads as a hope, and a resource owner has seen enough hopes.
A worked briefing, end to end
Here is the same structure applied to a different project, a federal hiring screening system, so you can see the shape independent of the example. The problem: the agency posts 200 positions annually, averaging 500 applicants per position, and manual screening takes 40 hours per position. That is 8,000 hours annually, equivalent to 4 FTE staff. Screening is also inconsistent, since some hiring managers are thorough and others are quick, and strong candidates get rejected because they were overlooked rather than assessed.
The solution, in plain language: an AI system that screens applications and highlights candidates meeting the position requirements, while managers still make every final hiring decision. The outcomes: reduce screening time from 40 hours to 8 hours per position, an 80% reduction; flag at least 95% of qualified candidates; and maintain hiring quality with no degradation in new hire performance. The investment: $250K in total. Notice what the freed time actually is here, since an 80% cut against a 4 FTE baseline does not free 4 FTE, and a payback claim built on the full baseline rather than the saved portion will not survive its first serious question.
The risks, named honestly: the system might miss qualified candidates, mitigated by setting a high recall threshold and verifying with sampling; hiring managers might trust the AI too much, mitigated by training that emphasizes human decision-making; and the system might show bias, mitigated by bias testing and monitoring. The ask: approval to run a 3-month pilot with 50 positions, with results in month 3 informing the go or no-go decision for full rollout. That is a complete briefing in six moves and it fits comfortably inside fifteen minutes.
The questions you will be asked
Executives ask a predictable set of hard questions. Having an answer ready is not scripting; it is the difference between a project that looks governed and one that looks improvised.
"What if it fails?" Describe the phased approach, low-stakes cases first, the ability to turn the system off, and the rollback plan. Make the rollback concrete rather than theoretical: name the trigger, the person who pulls it, and how long it takes, the way the DMV example specifies disabling the bot in under an hour if the error rate exceeds 5%. A rollback nobody has ever rehearsed is a paragraph, not a control.
"Will this eliminate jobs?" The honest answer is that this system is designed to change jobs rather than eliminate them, by automating routine decisions and freeing staff for complex cases, and that the plan is to use the efficiency gains to reduce backlog and improve service. Be careful not to convert that plan into a promise about headcount. Staffing levels are decided in budget cycles you do not control, and an assurance you cannot keep will be quoted back at you long after the project is forgotten. Say what the plan is, say who decides, and say what you will do if the decision changes.
"Is this going to be biased?" Bias testing is critical. Say that you test for fairness before deployment using demographic analysis, that you monitor continuously in production, and that if bias emerges you address it. Civil rights compliance is non-negotiable, and saying so plainly in the room is worth more than any metric you could put beside it.
"What about audit and compliance?" Say who was involved, meaning legal and compliance teams, and what exists: documented decisions, a maintained audit trail, and access for oversight bodies to examine the system at any time. Resist the blanket sentence "the system is compliant with relevant regulations," because that is a conclusion for a reviewer to reach rather than a claim for you to assert, and asserting it in front of an oversight audience invites the one question you cannot answer. Name the reviews completed, the reviews outstanding, and who signed each one.
"What is the real cost?" Explain the basis of the estimate, whether from comparable projects, that contingency is included, and that monthly cost reports will go to them with immediate notice of any overrun.
"How do we know it is working?" Describe the monitoring cadence: daily operational monitoring, a monthly dashboard covering accuracy, fairness, volume and user adoption, and immediate reporting when problems emerge rather than at the next scheduled update.
Managing expectations over time
The single most damaging thing you can do is let a leader believe AI is "done" at launch. Set the expectation early that an AI system is a managed service, not a finished product. Say it in the first briefing: "We will report accuracy in this same forum every month, because these systems can drift, and I would rather you hear it from me." That one sentence converts a future surprise into a routine update, and it makes you the trusted source instead of the person explaining a headline after the fact.
Be equally explicit about what will not happen. The system will improve efficiency, consistency and speed. It will not eliminate all manual review, because some decisions are genuinely too complex, and it will not solve every problem, because some challenges need process change rather than technology. Say that in the first meeting while it sounds like rigor. Said later, the same sentence sounds like an excuse.
Then commit to the machinery around the promise. Timelines with named phases, so a twelve-month project has a pilot with results in month 3 and full deployment in months 9 through 12, with support continuing indefinitely afterward. Decision gates, so that after the pilot you report and the executive decides: strong results proceed, weak results iterate or stop, and the decision rests on data rather than momentum. Governance, meaning quarterly reports, monthly dashboards and clear escalation paths. Under-promise and over-deliver, because saying 98% and delivering 99% builds trust that saying 99% and delivering 98% destroys, even though the second system is the better one.
Reporting progress between approvals
Keep a steady rhythm using the same rows every time. A leader who sees the crossover month slip from November to January, with a plain reason, stays calm. A leader who hears nothing for two quarters and then gets bad news does not. Predictable beats impressive, and quarterly is the minimum below which anxiety fills the silence.
A good update has five parts. What we promised: the original targets in their original words. What we achieved: the actual numbers with an honest label on each, such as a 42% time reduction against a 40% target, or 98% accuracy against a 99% target, marked as one point below and being investigated. Business impact translated into mission terms, such as processing time down from 40 days to 23, roughly 1.5 FTE of time freed and about $150K annually in cost avoidance, with 50% more applications handled at the same staffing. Monitoring, including what is watched and how often, plus operational facts such as uptime and adoption. And adjustments, meaning what you are doing about the gap and when you will report back.
Compare that against a six-month update on a benefits processing system, which promised a reduction from 40 to 30 days, 98% accuracy and fair treatment across demographics. It delivered 28 days, which exceeded the target, 98% accuracy on target, and demographic disparity under 2 percentage points. It reported 20% more applications handled with the same staff, a one-point accuracy decline between months 5 and 6 under investigation and probably caused by an incoming data quality issue, and legacy integration running slower than expected against constrained IT resources with no impact on schedule. Then it asked for nothing except to continue phase 2 and report again in 3 months. That is what a boring, trustworthy update looks like, and boring is the goal.
Never answer "how is it going" with "pretty good, users like it." Vagueness reads as either inattention or concealment, and executives have seen both. Bring the numbers, name the gaps before they are asked about, and frame the problems as learning rather than as failures being managed.
Anti-Patterns to Avoid
Every one of these is a habit that feels safe in the moment and costs the project later.
- The jargon shield. Explaining gradient descent and neural networks to an executive who wants to know whether the system works. They tune out and the message never lands. Reaching for "embeddings" and "fine-tuning" when nervous signals that you cannot explain the value plainly, and if you cannot say it without acronyms you do not yet understand it well enough to fund it.
- The demo trap. A flashy live demo invites the one question you cannot control, in front of the person you most need to impress. Show the outcome math instead and offer the demo as a follow-up for anyone who wants it.
- The certainty trap. Promising that it will work sets you up to be wrong on a single visible error. Promising 95% accuracy and delivering 85%, or promising a date and missing it, destroys trust faster than any technical failure. Promise a measured range and a backstop.
- Hiding a problem until it is solved. The instinct is to fix it quietly and report the resolution. Do not. Report early, frame it as learning, and let leadership see the problem while it is small. Every executive has been surprised before and remembers who did it.
- Going quiet after approval. Leadership approved six months ago and has heard nothing since, so they are now anxious and inventing explanations. Update quarterly at minimum, even when there is nothing dramatic to say. Especially then.
- Vague status reporting. "Pretty good, users like it" is not an answer. Have metrics, have data, and have the specific number for the thing that is not going well.
- Promising headcount outcomes you do not control. "This will not cost anyone their job" is the most tempting sentence available and the least defensible. Describe the plan and the decision-maker, not the guarantee.
- Asserting blanket compliance. "The system is compliant with relevant regulations" is a conclusion for a reviewer, not a claim for a project lead, and it is exactly the sentence an oversight body will test. Name the specific reviews done, by whom, and what remains open.
- Naming a backstop without capacity. "Uncertain answers route to a human" only protects anyone if that human exists, has time, and is authorized to overrule the system. Show the staffing behind the backstop or it is a diagram rather than a control.
- Quoting a metric as something it is not. An F1 score is not the share of correct answers, and a low false positive rate is not a high catch rate. Translating loosely in front of a non-technical audience is how an overstatement enters the record with nobody in the room able to catch it.
Practice Prompts
Use a real project you own or want to propose, not a hypothetical one.
- Build the fifteen-minute briefing. Quantify the problem in mission terms, explain the solution in plain language, specify the expected outcomes, work out the return with the arithmetic visible, describe the risks and their mitigations, and state exactly what decision you need.
- Translate a technical concept. Take one from your own work. Write the technical explanation, then the non-technical version, then why it matters to the mission, then the sentence a leader would actually use when repeating it to someone above them.
- Draft answers to the hard questions. Write your responses to "What if it makes biased decisions?", "How do we know it is actually working?", "Will this reduce our staff?" and "What happens if there is a failure?" Check each answer for a promise you cannot personally keep.
- Write a quarterly update. Cover what you promised, what you achieved, the business impact, any challenges, the metrics and the next steps, and include at least one thing that is not going well.
- Test it on a colleague. Deliver the pitch to someone outside your technical area and ask them two questions: do you understand the business case, and could you explain it to your own boss? If they cannot, revise until they can.
Reflection
Answer these about a briefing you have actually given or are about to give.
- Who chairs the meeting I am preparing for, and does my first sentence match what that person cares about?
- Which numbers in my deck could I derive on a whiteboard if someone asked me to, and which would I have to look up?
- Have I ever promised a leader something about staffing, timelines or compliance that was not mine to promise?
- What is the worst honest sentence in my project right now, and is it in the briefing?
- If an inspector general read my last progress report, what would they ask about first?
Glossary
- Executive briefing. A concise summary for senior leaders, typically fifteen to thirty minutes, focused on business outcomes, the decision needed and risk mitigation.
- Business case. A project justification covering problem quantification, solution description, expected benefits, costs, timeline and return.
- Return on investment. A measure of benefit gained relative to investment, expressed either as a percentage or as a payback period.
- Payback period. The time until cumulative benefits equal the investment, and the number an executive is most likely to check by hand.
- Crossover. The month at which accumulated savings overtake accumulated spending, which requires a stated dollar value for any time saved.
- Escalation. The process for raising issues up the organizational hierarchy when senior attention is needed.
- Risk mitigation. Actions taken to reduce the probability or the impact of an identified risk.
- Drift. Degradation in a deployed model's performance over time as the incoming data stops resembling the data it was built on.
- F1 score. A single index combining precision and recall. It is not the share of answers a system gets right, and should not be presented as one.
- Human backstop. The named person or team that reviews the cases a system routes to them, effective only to the extent that they have the capacity and authority to overrule it.
Related Lessons
Communication is the layer over everything else in an AI project, and these lessons cover what you are communicating about.
- How AI Projects Differ from Traditional IT explains the probabilistic, drifting, public-trust differences that make an AI briefing its own genre.
- Requirements Gathering for AI covers the upstream work that produces the quantified problem statement this lesson opens with.
- Working with AI Vendors and Contractors addresses the cost and vendor detail behind the build and run rows of the one-page template.
- Measuring AI Impact is where the outcome numbers in your progress reports come from, and how to know they mean what you say they mean.
- Quality Assurance for AI Work Products covers the evidence behind the accuracy and fairness figures you are asked to defend.
- AI Output Confidence Calibration explains why a confidence figure is not a correctness guarantee, which matters the moment you put one on a slide.
Closing
Effective communication is not a soft skill bolted onto the technical work. It is how you influence decisions, secure resources and keep support for an initiative through the months when nothing visible is happening. Leaders think in mission, budget and risk terms, and speaking those terms is not manipulation. It is professionalism, and it is the same courtesy you would expect from someone briefing you on a subject you did not study.
The best technical work in the world does not matter if leadership does not understand it, does not approve it and does not fund it. Marcus rebuilt his briefing to one page. He opened with the Riverside line, put the $95,000 build and the month 11 crossover in the next two sentences, named Licensing Ops as the risk owner, and ended with a single ask: a 90-day pilot at one office. The deputy director approved it before his eleven minutes were up. Nothing about the chatbot had changed.
Key Takeaways
- Brief the chair, not your comfort zone. Open with the outcome for mission owners, the money for resource owners and the guardrails for risk owners, and find out who is chairing before you build the deck.
- Write for the oversight audience that is not in the room. Executive sponsors, Congress, the IG, GAO and civil rights offices all review this work eventually. A briefing that would embarrass you in a review is one to fix now.
- Translate every model metric into a mission number. Accuracy is not the product. The shorter line, the freed staff hours or the avoided cost is, and every link in that chain should be independently checkable.
- Use the metric that answers the question. An F1 score is not the share of correct answers and a low false positive rate is not a high catch rate. Loose translation in front of a non-technical audience is how overstatements enter the record.
- Follow a fixed structure so projects can be compared. Problem, solution, outcomes, investment, risks, ask, spoken in fifteen minutes and mirrored on one page with the detail in an unpresented appendix.
- Make every financial claim derivable. A payback period or crossover month must follow from the numbers on the same page, including the dollar value of any time saved, because the executive will do that arithmetic.
- Name the human backstop and its capacity. Leaders fund probabilistic systems when they know what catches the wrong answers, and a backstop without staffing or authority to overrule is a diagram.
- Never promise what you do not control. Headcount, blanket regulatory compliance and guaranteed accuracy are three assurances a project lead cannot give. Describe the plan, the decision-maker and the evidence instead.
- Set the drift expectation at launch. Promise routine monitoring in the same forum so future bad news arrives as a managed update rather than a surprise, and report problems early rather than after solving them.
- End with a decision, never with "any questions." Ask for one concrete approval the leader can give in the room, and keep quarterly updates coming afterward whether or not anything dramatic happened.
Frequently Asked Questions
What if I genuinely do not know the dollar value of the time we would save? Say so, and say what you would need to find out. A range with a stated basis beats a single number with none, and your finance office can usually give you a loaded hourly cost for a position class in one email. What you must not do is state a crossover month or a payback period that depends on a value you never obtained, because that figure is the one a resource owner checks, and being unable to defend it costs you the rest of the page as well.
The executive interrupts and derails my structure. Do I fight for it? No. The structure exists so nothing important is left out, not so it gets delivered in order. Answer the question asked, then bridge back to the section that answers it more fully. What you protect is the ask at the end, not the sequence in the middle. If you are running out of time, skip the solution description before you skip the risk section, because a leader who has heard the risks and the ask can decide, and one who has heard the architecture cannot.
Should I show the model's failure cases to leadership? Show the category, not the gallery. Leadership needs to know what kind of thing goes wrong, how often, what happens to it next and who owns it. A parade of specific errors invites the room to reason from the most vivid one rather than from the rate. Name the worst plausible failure directly, since if you do not, someone else will produce it later in a context you do not control.
How honest should a progress report be when the news is bad? Completely, and early. A one-point accuracy drop reported in the month it happened, with an explanation and a date for the follow-up, is a routine update. The same drop discovered by someone else two quarters later is an incident, and it becomes a question about your reporting rather than about the model. Executives forgive problems far more readily than they forgive being surprised in front of their own leadership.
My project is genuinely small. Does it need all this? Scale the artifact, not the discipline. A small project may need three sentences rather than a page, but it still needs an outcome in mission terms, a named owner, an honest statement of what the system gets wrong and one specific ask. The habits are what make you credible when the project is large, and the first time you build a one-pager should not be the first time it matters.
Skill.re