Verification Workflows
Camille Desrosiers, an operations manager at a logistics software firm, sent her CFO a quarterly summary at 2:54 PM, six minutes before the deadline. She had asked an AI to pull the numbers together and skimmed the output for half a minute before hitting send. Two hours later the CFO replied with one line: "Revenue figure looks off, is this the renewals number or net new?" It was net new, mislabeled as total. The mistake was small. The cost was not. For the next month, every report Camille sent got a second look it had not gotten before. She had not been careless about the work, only about the last five minutes of it. This lesson is about those five minutes: how to build verification into your AI-assisted work so it catches the errors that matter without turning you into a bottleneck.
What This Lesson Covers
Verification is the habit of checking AI output before you put your name on it. You are accountable for every decision you make and every message you send, even when AI did the drafting. When you sign off on something, your judgment is on the line, not the algorithm's.
You will learn when to verify, what to check, and how to make it a reflex rather than an afterthought. The aim is a playbook for rapid verification that catches critical issues without tipping into perfectionism. Done well, verification is cheap insurance. Camille's missing thirty seconds of checking cost her weeks of rebuilt credibility, and a false claim to a customer or regulator can cost far more.
What Is Actually at Stake
Unverified output can damage things that are slow and expensive to repair. A strategic decision built on misread data sends the team in the wrong direction. A single false claim in an analysis makes people start questioning everything else you produce, exactly what happened to Camille. Wrong information handed to a customer erodes trust the moment they discover it. An inaccurate regulatory or compliance claim becomes a real exposure. And when your team acts on bad information you gave them, they lose faith in your judgment.
There is an upside that is easy to miss. When you verify, you understand the output better. You are not blindly trusting the algorithm; you are seeing what it did and validating it. That deeper understanding compounds, and your judgment improves over time. Verification is not just defense. It is how you actually learn from AI-assisted work.
Five minutes of verification before something reaches an executive is far cheaper than the hours of damage control after an error is found. The arithmetic is simple.
The Five Verification Layers
One verification method catches some errors and misses others. The reason is that a claim can fail in different ways: it can be sourced badly, or sourced well but reasoned badly, or true but outdated, or accurate but framed to mislead. Layered verification checks across those dimensions. The five layers Camille learned to run:
- Spot-check layer. Sample the critical facts, do not check everything. In a report with fifty claimed facts, verify the five to ten that matter most, deeply. You are using judgment to pick which facts carry the decision. If a proposal claims "AWS costs drop 30% after migration," that is load-bearing: check it against the current invoice and the vendor proposal. The other forty facts you skim for plausibility.
- Consistency layer. Does the output match what you know from direct experience? An AI report says "the team has been underutilized for six months," but your one-on-ones tell you people are overworked. That contradiction is a red flag. Camille hit exactly this once, and digging in revealed the AI had analyzed billable hours only, missing all the unbillable internal work.
- Source layer. Where did this fact come from, and how reliable is that source? Is it primary data you collected, or secondary data someone else reported? "Competitor X launched a product last month" from a press release is reliable. From a forum comment, it is not. Ask whether you can verify it at the source.
- Stakeholder layer. Have the relevant people confirmed this? An analysis says "the sales team hates the new CRM." You check with the head of sales, who says the team is actually split. That nuance only surfaces when a human who lives the situation looks at the claim.
- Logic layer. Do the conclusions follow from the data? Data shows productivity down 15% and the AI concludes the new tooling caused it. But is that the only explanation? Maybe the project got more complex, or people are ramping on something new. The data supports the conclusion without proving it. You verify the logic by asking what else could explain this.
You will not run all five layers on every output. But on a high-stakes decision, relying on one layer is how a fact that is accurate yet misleading slips through. A fact can be sourced correctly and still be two weeks out of date. An individual claim can be fine while the set of them together paints a picture that is wrong. Different layers catch different failures, which is the entire argument for using more than one.
What to Always Verify
Some things get checked before they enter any decision or message, no exceptions:
- Numbers and metrics. Revenue, customer counts, percentages, conversion rates, timelines, budgets. Numbers are the most error-prone thing AI produces, and one wrong digit changes everything. Check against source data.
- Names, dates, and specific factual claims. Titles, company names, when something happened, exact quotes. AI hallucinates these confidently.
- Causation claims. "X caused Y" needs more than X and Y both happening. "Churn rose after we changed pricing" is correlation. Before you call it causation, check the timing and rule out other variables.
- Anything that drives a major decision. If a fact is the linchpin of your argument, verify it thoroughly. If you are restructuring a team based on an analysis, the analysis gets checked.
- Claims about competitors or the market. Verify against recent announcements and earnings, not AI's memory of current events.
- Regulatory or compliance claims. "This needs a SOC 2 review" carries legal weight. Check with your compliance people, not just the AI.
What to Spot-Check
You do not verify everything. For these, sample the key supporting facts rather than every detail. General claims ("the SaaS market is growing"): verify a couple of the supporting facts, not every data point. Recommendations ("we should invest in X"): check the two or three biggest assumptions that would break the recommendation if they were wrong. Narrative framing ("this trend is accelerating"): confirm the data actually shows acceleration, not just steady growth, because "faster than ever" and "stable" are very different stories. Tone for important messages: read it aloud and ask whether it sounds like you and fits the audience.
The Data-Governance Layer
There is one more dimension that deserves its own attention, especially at this level: data governance. Verification is not only about whether the output is good. It is also about whether your handling of data along the way was appropriate. After an AI-assisted task, run a quick review of three questions. Did I share anything I should not have, any personally identifiable information or proprietary data, with a tool not approved for it? This matters most under time pressure, which is exactly when these mistakes happen. Is the output safe to share onward, or might it inadvertently make individuals identifiable even though I asked for no names? Did I use the right tool for this sensitivity level; a free consumer tool is wrong for proprietary financial data no matter how good the output looks?
Build this as a standard final step: after checking facts, logic, consistency, and tone, ask "is this data-governance clean?" At Level 2, make it explicit and write it into your checklist. Build the habit now, before you are working with AI more independently and at larger scale, so you never have to unlearn a bad practice later.
What You Can Usually Accept
Plenty does not need verification, though you should still read it. General phrasing and structure, if the memo is organized well, are probably fine. Minor details, an adjective here, a comma there, do not each need checking unless they affect the decision or will be public. Standard explanations of well-known concepts, if they sound right and you have general knowledge of the topic, usually are. Spend your verification budget on what is novel, disputed, or load-bearing, not on what is routine.
Verification Checklists by Output Type
Camille built three short checklists and a time budget for each, so verification became fast and routine instead of an anxious open-ended task.
Email (2 to 3 minutes). Read it aloud to catch tone and awkward phrasing. Verify any names, dates, and numbers. Confirm it actually says what you meant. Check that the voice sounds like you. Do a quick red-flag scan. Then send.
Report or analysis (5 to 10 minutes). Verify the three to five most important numbers against source data. Check all proper names and titles. Read the conclusions and ask whether they follow from the data. Ask whether anything is technically true but misleading. Check for missing context and caveats. Confirm the tone is executive-ready. Then send.
Plan or decision (15 to 20 minutes). Verify the key assumptions the plan rests on. Check feasibility with the people who would execute it. Look for missing considerations and risks. Validate with one trusted person before going wide. Adjust on their feedback. Then present.
Two notes on how to use these. The steps are ordered deliberately: reading aloud comes first on an email because tone problems are the ones you stop noticing after the third silent read. And the time budgets are targets, not minimums. If a two-minute email check keeps taking twelve, something in your process needs fixing, and if it consistently takes forty seconds, you are probably not checking the numbers.
A Worked Example: Verifying an Approval Email
The QBR is the big case. Most verification, though, happens on ordinary messages, so it is worth walking one all the way through. Camille needed her executive sponsor to approve a dashboard redesign project. She asked an AI for a draft, got one back in about thirty seconds, and then spent five minutes on it. Here is what those five minutes bought.
The draft was competent. It requested approval for a Q2 dashboard redesign, listed the timeline as April to September, a team of one engineer, one designer, and herself as PM, a budget of $45K for contractor support, and an expected outcome of "improved customer retention and market differentiation." The business case cited 22 customer requests for enhanced reporting, three enterprise customers blocked on deals worth $500K in ARR, and a competitor's recent dashboard redesign. It closed with an ROI summary totalling $1M in benefit against $45K in cost, a payback of half a month.
Reading it aloud surfaced the first problem, which was not factual at all. It sounded professional and flowed well, but "improve customer retention" was doing no work. Everything improves customer retention. A sponsor cannot approve a budget against a phrase that vague.
Checking names, dates, and numbers produced the most valuable catch of the exercise. The timeline said April to September, but Q2 is April to June. The draft had quietly committed her to a project running a full quarter longer than the one she was pitching. The 22 requests checked out against the feedback list. The $500K figure traced to sales, who confirmed three deals at roughly $180K each, so $540K was the accurate number. The competitor launch was two weeks old, current enough to cite. And the $45K budget reconstructed correctly from half a contractor at sixteen weeks, which came to about $48K before rounding.
Confirming it said what she meant caught a framing problem. The ask itself was right. But the $1M benefit figure treated three different outcomes with three very different likelihoods as one confident total, which is precisely the kind of claim a sponsor remembers when it does not materialize.
Checking tone came back acceptable. A bit corporate for her voice, right formality for an executive sponsor, confident without overreaching.
Her corrections followed directly from what she found. She changed the timeline to "April to June (Q2), with Q3 contingency if needed," which is both honest and easier to approve. She replaced "improve customer retention" with the specific version: reduce churn by closing the reported reporting gaps. She broke the ROI into its parts with a confidence level attached to each, roughly $540K in at-risk enterprise revenue at high confidence, about $300K in reduced SMB churn at medium confidence based on exit interviews, and about $200K in new enterprise deals at medium confidence pending sales. And she added two sections the draft had omitted entirely: what happens if the project is not approved, and what the project costs in her own time and the contractor's weekly rate.
Thirty seconds of drafting, five minutes of verification, five and a half minutes total. That is a normal ratio and a good one. The spot-check caught a timeline that misstated the commitment, the consistency check caught an ROI framed more confidently than the evidence supported, and the logic check produced the caveats that made the numbers believable. Had she wanted a fifth layer, a two-minute call to the sales lead would have confirmed the $540K directly rather than through a forwarded note. The email that went out was more credible than the draft and considerably more likely to get approved, which is the point worth remembering: verification is not only about avoiding errors. It usually makes the work better.
A Worked Example: Verifying a QBR Report
Two months after the CFO incident, Camille asked an AI to draft her Q3 business review from data across finance, sales, and the product roadmap. It produced a clean three-page report. This time she ran her report checklist, eight minutes, layer by layer. Here is what the draft claimed and what verification turned up.
The draft's headline numbers: revenue $2.3M against a $2.0M target (115%), four new enterprise customers, retention up to 92% from 84%, six major features shipped at 95% on-time, NPS 62. The challenges section noted an average sales cycle of 4.5 months and support tickets up 20%. The recommendation: hire two customer success managers at $180K annual cost, with an expected 5% revenue lift.
Running the layers against source data:
- Numbers (spot-check against source). The $2.3M revenue checked out against finance. The four enterprise customers checked out, with one pending final signature. But "six major features" was wrong: the roadmap showed five shipped and one still in progress. And "4.5 months" sales cycle was actually 4.2 in the CRM. Two number errors caught.
- Consistency. "Support tickets up 20%" was technically true, but the helpdesk data showed 60% of that increase was onboarding-related, not core support breaking down. The bare number implied a support-quality problem that did not exist. Context needed.
- Logic. The recommendation to hire two CSMs rested on the ticket increase. Did the logic hold? The main driver was onboarding, and CSMs do help with onboarding, so the chain held. But the "5% revenue lift" claim had no clear basis. Camille asked the sales lead, who called 5% optimistic and put it nearer 2 to 3%. She corrected the claim.
- Misleading framing. "115% of target" sounds great, but only if the target was realistic. Finance confirmed it was. And 92% retention is only good against a benchmark; she confirmed it sat above the roughly 90% norm, so the framing was fair.
She also checked for context the draft had left out, which is a step people skip because nothing on the page looks wrong. Was $2.3M monthly recurring revenue, annual recurring revenue, or gross revenue? The report did not say, and executives would each assume something different. Was "won three deals from top competitors" a meaningful business result or marketing language? And the lengthening sales cycle raised an obvious question the draft never answered: was that our execution or the market? Missing context is not a factual error, but it is what turns a report into a round of clarifying questions.
The corrections she made: changed "six features" to "five shipped, one shipping in Q4," changed 4.5 to 4.2 months, added "60% of the ticket increase is onboarding-related; core support quality held," softened the revenue-impact claim from 5% to "2 to 3%, pending validation," and added a one-line clarification that revenue was MRR and retention was monthly cohort. Eight minutes of verification caught two factual errors, one overstated claim, and one piece of missing context, the exact kinds of things an executive would have flagged in the room. The report went out credible, and this time no one replied asking what the numbers really meant.
A Worked Example: Stress-Testing a Plan
Plans are the hardest case, because there is often nothing to fact-check. The numbers may all be right and the plan still wrong. Camille found this out when she asked an AI to help her think through restructuring her team from eight individual contributors reporting to her into a structure with two promoted technical leads and five ICs.
The draft plan was well organized. Its stated goal was faster decision-making and less management overhead. It described the current state, eight ICs with no leadership layer and every decision requiring her sign-off, and the proposed state with two leads promoted from within. It listed benefits: faster decisions, better scaling, visible career paths, distributed domain expertise. It gave a five-week implementation timeline, identify two people in week one, announce in week two, transition over weeks three and four, complete in week five. It even listed risks and mitigations: promoted people might lack leadership skills, so provide training; remaining ICs might feel overlooked, so communicate clearly; leads might make inconsistent decisions, so hold a weekly sync.
Nothing in it was false. Verification still took it apart.
Verifying the assumptions (logic layer) hit the foundation immediately. The whole plan rested on "decision-making is slow." Was that true, or was it Camille's impression? She checked how long decisions actually took and asked the team whether they were frustrated. They were not. The premise was hers, unvalidated, and the AI had built a confident structure on top of it without ever questioning it, because questioning premises is not what it does. The second assumption was just as load-bearing: would promoting two leads fix the problem, or was the real issue something else entirely, such as unclear priorities or undefined domain ownership?
Checking feasibility (stakeholder layer) exposed what the plan carefully did not say. It never named who would be promoted. Two specific people would have to want the role and be good at it, and neither of those is knowable from a document. Nor did it address whether the five remaining ICs would read the change as opportunity or as rejection.
Checking for missing considerations found four gaps, all of them the kind that surface at the worst moment. Does a lead get paid more? What exactly does a lead do that an IC does not? What can a lead decide without her, and what still needs her sign-off? And what does career progression look like for an IC who has no interest in leading anyone?
Validating with one trusted person was the step that paid for the whole exercise. She took it to her own manager, who asked a question that appeared nowhere in the plan or in her head: what is your succession plan when those two newly promoted leads get recruited onto other teams? She also raised the idea informally with one likely candidate, framed as a hypothetical, which told her more about whether the plan was wanted than any amount of analysis would have.
Checking the timeline confirmed what the rest had implied. Identify people in week one and announce in week two leaves no room for the conversations that determine whether the restructure works at all, including the conversations with the people who are not promoted.
The revised plan looked different in kind, not just in detail. She started over from the actual problem rather than the assumed one. She identified specific candidates privately and sounded them out before proposing anything. She wrote in the missing definitions: the lead role, the compensation change, the decision authority. She extended the timeline from two weeks to four to six. She added the retention risk her manager had raised, with a salary adjustment as the mitigation. And she added an explicit growth path for ICs who do not want to lead.
Twenty minutes, on a decision affecting eight people's working lives. That is the right budget, and the lesson generalizes: for a plan, the AI gives you structure, and structure is genuinely useful. What it cannot give you is whether the plan solves the real problem, whether your people want it, or what a colleague with more scars will spot in thirty seconds. Use it as a thinking partner, never as a decision-maker.
Where This Goes Wrong
Verification becomes perfectionism. You spend thirty minutes on an email that was 95% fine, re-reading it five times and tweaking every word. The AI's speed advantage disappears and the cost-benefit goes negative. Set a time budget per output type, verify what matters rather than every adjective, and accept that a slightly awkward phrase is not the same as a false claim. If you are past your budget, the output is probably good enough. Stop.
No verification at all. This was Camille's original mistake: rushed, skimmed for thirty seconds, sent. The error went out, and damage control cost far more than five minutes would have. Build verification in as a non-negotiable step, like the act of hitting send. And remember the counterintuitive part: when you are rushed you need verification more, not less, because mistakes are likelier under pressure. If you are too rushed to verify, you are too rushed to send.
Trusting one channel. You spot-check the numbers, find them right, and assume everything else is sound. But a report with correct numbers can still draw a conclusion the data does not support, or frame an accurate fact to mislead. "Churn went 3%, 4%, 5%" is accelerating and the numbers are right, yet if 5% is a normal seasonal level the framing is still misleading. Use multiple layers on high-stakes work; if you are only fact-checking, you are missing the bigger picture.
Human Judgment Checkpoints
For each AI-assisted output, ask yourself five questions:
- What matters most here? What would be costly to get wrong, and what drives the decision?
- Did I verify those things? Not everything, just the critical few.
- Am I confident enough to put my name on this? Would I feel good defending it publicly?
- If someone challenged it, could I defend it? Can I show my sources, reasoning, and checks?
- Did I use more than one layer? Facts, logic, consistency, and where it matters, stakeholders.
Practice and Reflection
Verification becomes a reflex through repetition, and the fastest way to calibrate how much you actually need is to measure yourself. Six exercises, roughly in order of usefulness.
- Run the checklist three times and track it. For your next three pieces of AI-assisted work, explicitly walk the checklist for that output type and note how long it took and what you caught. Be honest about whether you found real errors and whether the output genuinely improved. Three data points tell you more about your own error rate than any general advice can.
- Calculate your risk before you verify. On your next high-stakes message or decision, list the facts where being wrong would actually damage something: credibility, a decision, a relationship, money. Verify those specifically. Spelling is not on the list. This is how you spend a fixed verification budget where it earns the most.
- Bring in one stakeholder. Next time you verify AI-assisted work, ask one relevant person a narrow question: does this match your understanding, would you do it differently, what am I missing? Do not ask them to check everything; ask them to validate the key assumptions. Notice how much of what they catch you could not have caught alone.
- Audit your sources. Take one AI-assisted report and mark every claim with where it came from. Which claims can you trace? Which could you defend to a skeptical executive? The ones you cannot trace are the ones that need verification, and doing this once teaches you a great deal about what sources the AI was actually drawing on, which is often less clear than it looks.
- Price the failure you avoided. Pick something you sent without real verification. Imagine an error had been in it. What would have happened, who would have found it, and how long would the cleanup have taken? That thought experiment is what makes five minutes feel cheap rather than annoying.
- Run a speed trial. Set a timer, take one AI-generated email or short report, and verify it as fast as feels safe. Did you hit your time budget? If not, find what is slowing you down: are you over-verifying, or is your checklist too aggressive for the kind of work you actually do? Adjust one of the two.
Then spend two minutes on this reflection. Think of one thing you sent this past week where a verification habit would have changed your approach. What would you have checked, what might you have found, and what would have been different about the outcome? Writing down that connection between the concept and your own week is what turns a checklist into a habit.
Related Lessons
Verification is the first of four skills that together make up human oversight of AI-assisted work.
- Data Interpretation Support comes immediately before this one and supplies much of what you will be verifying. Understanding how AI reads and summarizes data is what lets you tell a misreading from a legitimate conclusion.
- Knowing When to Override AI picks up where verification ends. Checking tells you whether the output is accurate; overriding is the separate judgment about when to set aside an accurate output because you know something it does not.
- Feedback Loops and Iteration turns each catch into an improvement. Every error you find is information about how to prompt better next time, and the managers who close that loop verify less over time because they need to.
- Documenting AI Assisted Work is what makes your verification defensible after the fact. Being able to show what you checked, and against which source, is the difference between "I verified it" and a record someone else can trust.
Key Takeaways
- Verification is non-negotiable. Build it in like hitting send. Camille's skipped thirty seconds cost her a month of rebuilt credibility; five minutes of checking is far cheaper than hours of damage control.
- Verify what matters, not everything. Focus on numbers, names, causation claims, and the facts a decision hinges on. Accept that minor imperfections are fine; you are catching errors that matter, not chasing perfection.
- Use layered verification. Facts, logic, consistency, sources, and stakeholders each catch different errors. A claim can be accurate and still misleading, and only multiple layers catch both.
- Keep a time budget. Email 2 to 3 minutes, report 5 to 10, plan 15 to 20. Longer means you are over-verifying; much shorter means you are under-verifying.
- Verification is not what slows you down. Skipping it is. Unverified errors cost far more time through cleanup, correction emails, and lost credibility than the checking would ever have cost. Treat it as an investment with a reliable return.
- Plans need a different kind of checking. When there are no facts to verify, verify the assumptions, the feasibility with the people who must execute, and the gaps, then run it past one trusted person before it goes wide.
- Add a data-governance check. After facts and logic, ask whether you shared anything you should not have, whether the output is safe to pass on, and whether the tool was approved for this data. Build the habit now.
- You own everything you send. Even when AI helped, your name is on it and the buck stops with you. Verification is how you earn the right to put your name on the work.
- Red flags deserve a deeper look. If a number seems wrong, the logic seems off, or the tone feels strange, that instinct is part of your toolkit. Verify deeper.
Skill.re