Coaching Team Members on AI Quality Standards
Hiroshi Tanaka leads a six-person content team at a software company, producing help articles, release notes, and customer-facing guides. When the company rolled out an AI writing assistant, Hiroshi expected a productivity jump. Instead he got a quality problem he did not see coming. One of his writers, Lena, published a how-to article that confidently described a settings menu that did not exist - the AI had invented it, and nobody had checked. A customer flagged it. Hiroshi realized that giving his team a powerful tool was not the same as teaching them to use it well. This lesson follows how he turned that incident into a coaching system, including a quality scorecard he used to make "good" concrete enough to actually coach against.
Why people need coaching, not just access
Hiroshi's first mistake was assuming that capable people would automatically use AI well. They do not. Some over-rely on it and skip verification, like Lena. Some reach for it on tasks where it does not belong, such as anything touching confidential customer data. And some avoid it entirely because they misunderstand when it is appropriate, quietly falling behind. Left alone, a team drifts into all three patterns at once.
The stakes are real and concrete. Quality suffers when outputs go out unverified, as the phantom settings menu showed. Risk rises when people feed sensitive material into the wrong tool. People never develop judgment if no one coaches them. And the efficiency gain the company was counting on simply does not materialize. Hiroshi reframed his job: not to police every keystroke, which would make him a bottleneck, but to build his team's judgment about when to use AI, how to verify it, and when to stop and ask.
Step one: define what good actually looks like
Hiroshi could not coach toward "good" until he had defined it. Vague standards produce vague work. So for each common use case on his team he answered a few sharp questions. What accuracy level is required? What verification has to happen before output is used? Which tasks are safe for AI and which are off-limits? How much editing is expected? And how should people note that AI was involved?
For his content team, the article-drafting standard came out like this: AI drafts the first pass; the writer is responsible for the final article; every factual claim, menu path, and product name must be checked against the actual product or source documentation; screenshots and UI references must be verified live, never trusted from the model; and the writer notes in the doc that AI assisted the draft. He wrote it on a single page, posted it in the team wiki, and referenced it constantly. A standard nobody can see is a standard nobody follows.
Step two: coach by showing, not telling
The most persuasive coaching Hiroshi did was simply working in the open. In a team session he drafted a release note with the AI on the shared screen, then walked through his own process out loud: prompt, output, review, final. "I used AI here because it is a routine note. I spent two minutes checking the version numbers against the changelog, and I changed one sentence that overstated what the feature does." Then he did a harder one: "This customer apology needed real judgment, so I rewrote most of the draft - the AI's tone was too corporate for how we actually talk to people."
When he showed his reasoning, not just his result, the team learned the judgment underneath the rule. They saw that verification is not bureaucratic box-checking; it is the moment where a hallucinated menu gets caught before a customer sees it.
Worked example: a weighted quality scorecard, before and after coaching
To coach Lena after the phantom-menu incident, Hiroshi needed something more objective than "be more careful." He built a weighted quality scorecard so he could score AI-assisted output on consistent criteria, attach a real number to it, and show measurable improvement after coaching. Weighted scoring means each criterion gets a weight reflecting how much it matters, you score each on the same scale, and the weighted scores combine into one total.
He used five criteria, scored 1 to 5, with weights that sum to 100 percent:
- Factual accuracy - weight 35 percent. Are all claims, menu paths, and product names verified against the real product? This carries the most weight because it is where AI fails most dangerously.
- Completeness - weight 20 percent. Does the article actually cover what the user needs, with no missing steps?
- Tone and voice - weight 20 percent. Does it sound like the company, plain and human, not robotic?
- Clarity and structure - weight 15 percent. Is it easy to follow, well-ordered, scannable?
- Disclosure and sourcing - weight 10 percent. Is AI assistance noted and are sources cited where needed?
The total score is the sum of each criterion's score multiplied by its weight, then expressed out of 5. Here is Lena's phantom-menu article scored before any coaching:
- Factual accuracy: 2 of 5, multiplied by 0.35 = 0.70. The invented menu was the headline failure.
- Completeness: 4 of 5, times 0.20 = 0.80.
- Tone and voice: 4 of 5, times 0.20 = 0.80.
- Clarity and structure: 4 of 5, times 0.15 = 0.60.
- Disclosure and sourcing: 1 of 5, times 0.10 = 0.10. No verification notes, no sources.
That totals 3.00 out of 5. On the surface the article read well, which is exactly why the problem was dangerous: it scored fine on tone and structure while failing on the one criterion that mattered most. The weighting made the real risk visible. A flat average would have hidden it.
Hiroshi then coached Lena, privately and with a teaching tone rather than blame. He showed her the scorecard, pointed at the 0.70 and the 0.10, and said the quality bar was a verification habit, not raw writing skill, which she clearly had. Together they agreed on a simple routine: after every AI draft, open the actual product, click through every path the article mentions, and add a one-line source note. Two weeks later he scored her next published article the same way:
- Factual accuracy: 5 of 5, times 0.35 = 1.75. Every path checked live.
- Completeness: 4 of 5, times 0.20 = 0.80.
- Tone and voice: 5 of 5, times 0.20 = 1.00.
- Clarity and structure: 4 of 5, times 0.15 = 0.60.
- Disclosure and sourcing: 5 of 5, times 0.10 = 0.50.
That totals 4.65 out of 5, up from 3.00. The jump came almost entirely from the two heavily weighted criteria she had been failing. The scorecard turned a fuzzy "do better" into a precise, fair, and visibly improving target. Hiroshi later shared the rubric (without Lena's name on the low score) so the whole team knew exactly what good looked like.
The common issues, and how to coach each
Across his team, Hiroshi kept seeing the same handful of patterns, and he learned a coaching move for each. The thread through all of them is to lead with curiosity, not accusation.
Trusting output without verification. This was Lena's original issue. He named it gently ("did you get a chance to check this against the product?"), showed a concrete error, taught the habit ("draft, then verify against the source"), made it small enough to stick ("three minutes, read once for accuracy"), and followed up later. Habits form through repetition and follow-up, not a single talk.
Using AI for the wrong tasks. One writer pasted unreleased customer data into a consumer AI tool to summarize it. Hiroshi first understood the why ("what were you trying to solve?"), set a clear boundary ("check with me before putting any customer or unreleased data into an external tool"), and then wrote that rule into the team standard so the lesson outlived the conversation.
Over-reliance on AI. Another team member used AI to generate three-line Slack messages that took longer to prompt than to type. Hiroshi pointed out the inefficiency without judgment and offered a simple rule of thumb: if it would take more than about five minutes to write yourself, AI is probably worth it; if not, just write it. Judgment about when not to use AI is part of using it well.
Not documenting AI use. Disclosure matters for client transparency and for letting teammates calibrate how much to trust an output. He explained the why, modeled it himself, and reminded people early until it became automatic.
Stalled skill development. Some writers used AI competently but never improved. Hiroshi assessed their actual workflow ("walk me through how you used it here"), spotted gaps like never iterating on a weak draft, taught a better technique, and built on each success toward the next level.
Four conversation shapes for coaching
Hiroshi found that most of his coaching fit one of four conversation shapes, and naming them helped him pick the right one. The recognition conversation reinforces something done well: name the specific behavior, explain why it is right, encourage more of it ("you verified every fact against the docs before publishing - that is exactly right, keep doing that"). The redirecting conversation, always held privately, adjusts something off-track: acknowledge their thinking, explain the constraint, offer a clear alternative. The teaching conversation fills a missing skill: share the technique, demonstrate it, give them a chance to practice. And the judgment-development conversation builds the harder muscle of knowing when to use AI at all: invite their reasoning, validate that the uncertainty is normal, share how you think it through, and treat judgment as something that keeps developing.
Building a culture, not just correcting individuals
Individual coaching only scales so far. Hiroshi worked the team level too. In team meetings he normalized the topic by asking who had been using AI and what they had learned, and by spotlighting wins ("Sofia cut her summary time roughly in half with a clever prompt - Sofia, show us"). He paired experienced AI users with newer ones for peer coaching, which often landed better than anything from him because it carried no power dynamic. He built standards incrementally, starting loose and tightening as the team learned what worked, and he involved the team in the updates so the rules felt like theirs. And he documented learning in a shared knowledge base: techniques that worked, and mistakes discussed openly and without blame, because a team that hides its AI mistakes cannot learn from them.
Scaling the scorecard without becoming a bottleneck
Hiroshi did not want to personally score every article forever; that would turn him into the exact micromanaging gatekeeper he was trying not to be. So he used the scorecard as a teaching device first and a spot-check device second. For the first month after he introduced it, he scored two pieces per writer so each person saw concretely where they stood. After that, he shifted to scoring a random sample, roughly one in five published pieces, just enough to keep a pulse on quality without inspecting everything.
He also handed the scorecard to the writers themselves. Once people could self-score honestly, they caught their own weak spots before publishing. Lena, who had once scored a 3.00, started running her own drafts against the five criteria and flagged a borderline factual-accuracy issue to Hiroshi before it shipped, exactly the behavior he wanted. The rubric stopped being something done to the team and became something the team used. When a tool for judging quality becomes shared rather than imposed, it raises the floor for everyone instead of just catching individuals.
The team-level view was useful too. By tracking average scores per criterion across the team, Hiroshi could see patterns no single conversation would reveal. When the whole team's disclosure-and-sourcing scores sat low, that was not five separate coaching problems; it was a signal that the standard itself was unclear and needed a better example in the wiki. He fixed the standard, and the scores rose together.
Traps to avoid
Hiroshi watched other managers fall into predictable holes. Overly restrictive standards ("nobody uses AI without my sign-off") make the manager a bottleneck and kill adoption; better to set clear standards, trust people to follow them, and spot-check. No standards at all lets quality and risk run wild. Correcting someone in public, as one manager did in a team meeting, breeds embarrassment and defensiveness; corrections go private, recognition can go public. Assuming a single explanation equals understanding ignores that judgment forms through practice and feedback over time. And freezing standards in place ignores everything you learn after writing them; Hiroshi reviews his team's AI standards roughly every quarter and adjusts.
Key Takeaways
- Access is not the same as good use. People predictably over-rely on AI, misuse it on the wrong tasks, or avoid it entirely. Your job is to build their judgment, not to police every use.
- Define "good" concretely before coaching. For each use case, specify the required accuracy, the verification step, what is off-limits, and how to disclose AI use. A standard nobody can see is a standard nobody follows.
- A weighted scorecard turns vague feedback into a fair target. Weight the criteria that matter most (factual accuracy carried 35 percent for Hiroshi), score consistently, and you can show measurable before-and-after improvement - Lena went from 3.00 to 4.65 out of 5.
- Coach by showing your reasoning. Work in the open and narrate your verification and tone decisions. The team learns the judgment underneath the rule, not just the rule.
- Lead every correction with curiosity. Ask why someone did what they did before redirecting. Always correct privately and recognize publicly.
- Make verification a small, repeatable habit. "Three minutes, read once for accuracy" sticks where "be more careful" does not. Follow up later to cement it.
- Build culture, not just individual fixes. Normalize discussion, pair experienced users with newer ones, and discuss mistakes openly without blame so the whole team learns.
- Let standards evolve. Review them roughly quarterly and tighten as the team matures. Frozen standards ignore everything you learn after writing them.
Skill.re