Evaluating AI Output
Marcus Delgado manages a 12-person operations team at a regional logistics company. One Thursday afternoon he asked an AI assistant to draft a customer update about a shipping delay, skimmed it, and sent it to a key account. Two hours later the client called: the email had cited a new 48-hour resolution guarantee that Marcus had never offered and the company did not provide. The AI had invented it, and Marcus had not caught it. The cleanup cost him a tense apology call and a small credit. After that, he built himself a five-minute verification habit. He has never sent an unchecked AI draft since, and over the next quarter his team caught dozens of similar errors before they reached anyone outside the room.
What This Lesson Covers
Evaluating AI output is the human judgment layer that sits between an AI response and the moment you actually use it. The AI gives you a draft. Your job is to decide: is this good enough to send, present, or act on, and if not, what needs fixing? This lesson teaches a repeatable habit for answering that question fast.
You will learn the five dimensions to check on every AI output: accuracy, tone, completeness, relevance to your context, and alignment with your values. You will see how to spot a hallucination (when AI confidently states something false), how to verify facts and numbers, and how to work through real cases: an email draft, a meeting summary, and a feedback analysis. By the end you will have a checklist you can run in two to five minutes.
Why Verification Is the Whole Game
The single biggest mistake managers make with AI is trusting output without checking it. AI writing reads fluently and sounds confident, which is exactly what makes unchecked errors dangerous. A fluent, confident, wrong answer is harder to catch than an obviously clumsy one.
Poor verification has real costs. It sends emails with tone-deaf sections. It presents information that is slightly wrong in a meeting where being wrong is expensive. It puts decisions on top of inaccurate summaries. And it quietly erodes your credibility, because the moment a colleague catches an AI mistake you missed, they start discounting everything you forward.
You are accountable for what you send, not the AI. Verification is how you hold up your end of that accountability.
The good news is that verification is a skill that compounds. Marcus timed himself. His first careful reviews took eight to ten minutes. After about ten drafts he was down to two or three minutes. After fifty, he could spot the usual problems in under a minute. The habit gets cheap fast; the cost of skipping it does not.
Dimension 1: Accuracy and Hallucinations
The question: is the factual content correct? This is where the most damaging errors live, because AI invents facts that sound plausible. A hallucination is when the model states something false with full confidence, often a specific number, a study, a date, or a policy that does not exist.
How Marcus checks accuracy now. For facts he knows well, he verifies them directly against his own knowledge. For facts he is unsure about, he treats them as suspect until confirmed. For any statistic or specific number, he applies extra scrutiny, because confidently invented numbers are the most common failure. And for anything tied to recent events, he verifies independently, since the model training data has a cutoff and may be out of date.
The red flags worth memorizing: specific numbers that sound plausible but you cannot source; references to studies or events you do not recognize; and oddly precise claims like the average manager works 47 hours per week. That kind of false precision is a classic tell. When Marcus sees a number he did not provide, his default assumption is now prove it rather than probably fine.
What to do when you hit an unverifiable claim. You have two clean options. Verify it against a reliable source, or rewrite to remove the specific claim. If an AI draft says research shows 73% of managers report feeling overwhelmed and you have never seen that figure, you either confirm it with a quick search or you rewrite to many managers report feeling overwhelmed, which makes a defensible point without staking your credibility on a number you cannot back up. Never forward a specific factual claim you have not checked, especially when others will read it.
Dimension 2: Tone and Appropriateness
The question: does this sound right for the context and the relationship? Content can be perfectly accurate and still land badly because the tone is off: too formal, too casual, too harsh, too soft, or simply too generic to sound like you.
The fastest check is to read the draft once as the person receiving it. Imagine it arriving in their inbox. Does it match your actual voice and your relationship with them? Watch for output that sounds corporate or machine-generated, that is missing warmth where warmth matters, or that is too blunt or too padded for the moment.
When the tone is wrong, you have a quick fix: ask the AI to adjust it. Make this warmer or this is too formal, make it conversational usually gets you most of the way. When an AI draft told Marcus team to expedite the completion of the aforementioned project, he laughed, because nobody on his team talks that way, and asked for a plain, friendly rewrite. The rule he keeps: if the tone feels wrong, do not send it, even when the content is right. Tone often matters more than polish.
Dimension 3: Completeness
The question: did the AI leave out anything important? Incompleteness is sneaky because what is on the page can look fine; the problem is what is missing. Compare the draft against what you actually meant to communicate. If you sent this as is, would the recipient have everything they need to act?
Watch for missing key information, instructions that skip a step, context that the reader needs but the draft assumes, and the absence of a clear call to action when one is required. Marcus once asked AI for a meeting agenda and got a clean list of topics with no time allocations. Without timing, attendees could not tell whether a topic needed five minutes or forty-five. He added the time blocks before sending. The list was not wrong; it was incomplete, and the gap would have shown up as a chaotic meeting.
Dimension 4: Relevance to Your Context
The question: does this account for your specific situation, your team, and your constraints? AI does not know your team dynamics, your organization culture, or the constraints you carry in your head unless you told it. So its output defaults to generic, and generic advice sometimes actively conflicts with your reality.
Check whether the draft reflects what you know: your team working style, your culture, the real constraints in play. The red flag is a suggestion that would never work in your organization, or one that ignores something obvious about your situation. When an AI proposed a competitive, leaderboard-style brainstorming format, Marcus immediately knew it would backfire: his team is collaborative and would have hated it. He reframed it around shared problem-solving. The fix is to add your context and ask for a revision, or rewrite the sections that miss your world.
Dimension 5: Alignment With Your Values
The question: does this communicate what you actually believe, in a way you would stand behind if questioned? Your name goes on it, so it has to sound like your principles, not a generic professional voice.
Read it and ask: does this reflect how I see my role? Would I be comfortable explaining this later? The warning sign is framing you disagree with or language that feels inauthentic. When an AI drafted performance feedback that opened with your performance has been inconsistent and below expectations in several areas, Marcus stopped. That harshness is not his coaching style. He rewrote it: I have noticed some inconsistency. Let us talk about what is going on and how I can support you better. Same underlying message, delivered in a way he could stand behind. If a draft feels inauthentic, do not send it; treat AI output as a first draft, never a final word.
A Worked Example: Running the Full Checklist
Here is the verification checklist applied end to end, the way Marcus runs it. Suppose he asked AI to draft a team email about a deadline change and got this:
Hello team, due to unexpected circumstances, we have had to modify the project deadline. The new completion date is Friday, March 20th. This may require some adjustment to your schedules. Please update your plans accordingly and reach out if you have concerns.
He scores it across the five dimensions:
- Accuracy: Is Friday, March 20th correct? He checks the calendar. Yes. Did the deadline actually move? Yes. Passes.
- Tone: For his team he is usually warmer and more direct. This reads corporate and distant. Fails.
- Completeness: It never says why the deadline moved, does not acknowledge the impact, and offers no real support. Incomplete.
- Context: Three people on his team are already overloaded this sprint. The draft ignores that completely. Misses his situation.
- Values: He would never hide behind unexpected circumstances. He names what is actually happening. Fails authenticity.
Verdict: a usable skeleton, but it needs real customization. Marcus rewrites it to explain the cause, acknowledge the load on the team, and offer to help reprioritize. Total verification and rewrite time: about four minutes. The result reflects his voice and would not have triggered a single confused reply.
The same checklist catches subtler problems in summaries and analysis. When he asked AI to summarize customer-meeting feedback, the draft mentioned invoicing as something one customer raised, but Marcus had been in those meetings and knew invoicing came up repeatedly and mattered. The summary was accurate on the surface and wrong on significance. He asked the AI to reorganize around what is working, what is frustrating customers, and what they want next, and to add sentiment. When he asked it to analyze team feedback themes, it bucketed 46% of responses as miscellaneous other concerns, which was a signal it had missed a real pattern. He pushed back, and the AI surfaced that requests for more feedback and connection time were actually a major theme, on par with career clarity. In both cases the checklist, not the AI, caught what mattered.
Anti-Patterns to Avoid
Skipping verification to save time. AI is usually good, I will just send it. One unchecked mistake can cost you more credibility than a month of verification saves you in minutes. Build two minutes of checking into your process and keep it there.
Trusting AI on factual claims. It said the statistic, so it must be true. AI confidently generates false numbers; that is a known, well-documented behavior. Always verify facts, especially numbers, before they leave your hands.
Not customizing when you should. Good enough, sending as is. Generic output that does not reflect you, your values, or your situation reads as inauthentic and erodes trust over time. Spend the two to three minutes to make it yours.
Assuming AI understands your context. It knows what I need. It does not, unless you told it. Always compare the output against what you actually wanted, not against whether it sounds polished.
Building the Verification Habit
Verification is a muscle. The first few reps are slow and feel like overhead. Within a couple of weeks they become automatic, and the speed gains are real: from ten minutes to under one for routine drafts. A few practices make the habit stick.
Run the five questions every time. Are the facts correct? Does this sound like me? Is anything missing? Does this fit my situation? Would I put my name on this? If you answer no to any of them, refine before using.
Know when to start over. Sometimes an AI draft is so far off that editing it takes longer than rewriting from scratch. Recognize that moment and abandon the prompt rather than polishing a bad foundation.
Keep a personal red-flags list. After a few weeks, you will notice the specific mistakes AI tends to make in your work: invented numbers, a too-formal default, missing calls to action. Write them down. Your personal list makes your verification faster and sharper than any generic checklist.
The Checkpoint Before You Send
The five dimensions are how you read an output. The checkpoint is what you do at the moment of decision, when the draft is finished and you are about to send it, present it, or act on it. Marcus treats that moment as a hard stop. Nobody downstream is going to catch what he misses, and once it leaves his hands it is his, not the model's.
Five questions, and a specific obligation attached to each "no."
- Are the facts correct, and should I verify anything before this goes out? If there is a number, date, name, or claim you cannot source, you either verify it or you rewrite it out. Those are the only two options.
- Does this sound like me, and does it fit this context? If it does not, ask for a tone adjustment or rewrite the offending passages yourself. Do not send a message that will read as if someone else wrote it.
- Is anything missing, and did the AI actually understand what I asked for? If the answer is no, name the gap explicitly and ask for an expansion, or add the missing piece by hand. The reader should not have to come back with questions.
- Does this account for my specific situation? If the draft ignores a constraint, a relationship, or a personality you know matters, supply that context and ask for a revision rather than sending something generic.
- Would I be comfortable putting my name on this? This is the last gate and the one that overrides the others. If the answer is no, it does not go out, however accurate and well-structured it is.
A "no" on any question means refine before using. There is no version of this checkpoint where a no becomes a shrug. The whole point is that you are the checkpoint, and skipping it moves the risk onto whoever reads what you sent.
Responsible Use: Accountability and Reputation
Verification is not just a quality practice. It is what makes using AI at work defensible in the first place.
The accountability does not transfer. When Marcus's client called about the 48-hour guarantee that did not exist, no one was interested in the fact that a model had written the sentence. It went out under his name, to his account, from his company. That is the permanent arrangement: AI produces drafts, and you remain the responsible party for everything you release. Verification is simply the mechanism by which you hold up your end of that. If you are not willing to check the output, you are not in a position to send it.
Your reputation is built slowly and spent quickly. A single uncaught AI error can undo a long stretch of good work, because the damage is not the error itself but what colleagues conclude from it. Once someone catches an invented statistic in something you forwarded, they start reading everything you send with suspicion, and you rarely get told that it happened. Two minutes of checking is remarkably cheap insurance against that.
Verification is a muscle, and it strengthens. The first few reviews feel like pure overhead and take the better part of ten minutes. That is normal and temporary. Within a few weeks you are running the same checks in under a minute because you know where the problems tend to hide in your particular kind of work.
Know when to abandon rather than repair. Sometimes an output is far enough off that editing it costs more than starting fresh, and the sunk time you have already spent makes it hard to see. Recognizing that moment is part of responsible use too. Discard the draft, rewrite the prompt with better context, and try again rather than polishing something built on the wrong foundation.
Practice and Reflection
These are best done on real work, with output you have actually received rather than an invented example.
- Run the full checklist once, slowly. Take an AI output from the past week and work through all five dimensions deliberately. Write down what you would change. Notice which dimension caught the most.
- Chase one factual claim. Find a specific fact, number, or reference in an AI output and try to verify it against a reliable source. Was it true? Was it verifiable at all? That exercise recalibrates how much benefit of the doubt you extend.
- Audit the tone from the other side. Read an AI draft as if it had landed in your inbox from a colleague. Does it sound like a person you know? What is missing that a human would have included?
- Customize and compare. Take a raw AI output and spend three minutes adding your context, voice, and the constraints only you know about. Put the two versions side by side and judge honestly how much the customization bought you.
- Time yourself. How long did verification actually take, and what did it catch? Track this for a week. The ratio of minutes spent to problems caught is usually more persuasive than any argument for the habit.
- Start your red-flags list. What problems keep recurring in AI output for your specific work? Write the top three down and keep the list where you can see it. That personal list will outperform any generic checklist within a month.
Key Takeaways
- Always verify before using. Even good AI output needs a human check. You are accountable for what you send, not the model.
- Use the five-dimension checklist. Accuracy, tone, completeness, context, and values cover the ways AI output goes wrong. Run all five.
- Verify facts and numbers with extra care. AI confidently invents statistics, studies, and dates. If you cannot source a specific claim, verify it or rewrite it out.
- Expect to customize. AI output is rarely a perfect fit for your team, your tone, or your values. Treat it as a first draft, never the final word.
- Tone and context matter as much as accuracy. A factually correct message can still fail if it does not sound like you or ignores your situation.
- Verification gets fast. What takes ten minutes today takes under one within a few weeks. The habit is cheap; skipping it is not.
- Keep your own red-flags list. The mistakes AI makes in your specific work are predictable. Track them and you will catch them faster every time.
Skill.re