Ethical Judgment in Practice
Naomi Hartwell leads a 10-person design team at a healthcare software company. One evening, racing to finish performance reviews, she pasted six months of a designer's meeting notes, Slack messages, and project tickets into an AI tool and asked it to "assess engagement and flight risk." In nine seconds it told her the designer, Theo, was "likely disengaged, predicted departure within six months, recommend reducing his project scope." Naomi's cursor hovered over the review draft. It would have been so easy to write Theo up as a flight risk and quietly shift his best project to someone else. Then she stopped. She had never actually talked to Theo. She was about to make a decision that could shape his career based on a pattern an algorithm found in data he never knew she was collecting. The question was not "can I do this?" Of course she could. The question was "should I?" That gap, between can and should, is where this lesson lives.
The Gap Between "I Can" and "I Should"
AI hands managers new powers almost weekly. You can draft sensitive feedback in seconds. You can analyze team sentiment from meeting transcripts. You can predict who might quit. The trouble is that "I can" is profoundly different from "I should," and AI is silent on the second question. It will do whatever you ask without ever telling you whether you should have asked.
This lesson does not hand you rules. Rules are too rigid for the situations you actually face, and they break the moment reality gets complicated. Instead it builds your judgment: the ability to feel the ethical tension in a moment, ask the right questions, and make a call you can stand behind. Think of it as a muscle. By the end you will have a repeatable way to work through the gray areas of bias, privacy, transparency, attribution, and knowing when to say no.
The gap between what AI lets you do and what you should do is exactly where your ethics live. No tool can close that gap for you. That is your job, and it is the part of management that cannot be automated.
Three Forces That Shape Every AI Ethics Call
Before the framework, three forces are worth naming, because they show up in every difficult decision.
The efficiency-versus-authenticity tension. AI is fast, and speed is seductive. But the fast path often costs the very thing that makes you effective. An AI-drafted feedback note that reads "I have observed that your communication could benefit from more structured organization of your key points" is efficient and sounds like an algorithm. What you would actually say is "your ideas are good, but in exec meetings try leading with your recommendation, it helps them follow you." One sounds like software. One sounds like a manager who cares. The tension is real; pretending it does not exist is how you drift into corporate, hollow leadership.
The responsibility amplifier. AI multiplies your reach, which means it multiplies your mistakes. If you handwrite one feedback email with a bad tone, one person sees it. If AI helps you draft feedback for your whole team and the tone is systematically too harsh, you have just damaged 10 relationships at once. Bias amplifies the same way: an AI screen trained on biased history scales that bias across 100 applications instead of 10. More power means more responsibility, always.
The consent question. The sharpest test is not "would they mind?" It is "do they know?" Does Theo know Naomi analyzed his messages? Do candidates know AI screened them and on what criteria? If the honest answer is no, and they would feel violated to find out later, that discomfort is a signal worth listening to.
A Four-Test Framework for the Hard Calls
When Naomi caught herself at that cursor, she ran a sequence she now uses every time AI touches something sensitive. None of the tests gives you the answer. Together they make the tradeoff visible so your values can decide.
1. The Purpose Clarity Test. What am I actually trying to do here, and does AI serve that purpose or contradict it? Naomi's real purpose was to help Theo grow. An algorithmic flight-risk label did not serve that purpose; it replaced understanding with a prediction. Be honest about what you want. If your purpose is genuine, caring feedback, an AI draft only serves it if you rewrite it to sound like genuine care from you.
2. The Transparency Test. Would I tell the person I did this? Naomi asked herself whether she would tell Theo she had run his six months of messages through an AI to assess his loyalty. The answer was no, and that no was the whole answer. If you would not tell them because it might sound creepy or because you are not sure it is your right, that hesitation is your integrity talking.
3. The Reversibility Test. If I get this wrong, can I undo it? Giving feedback is reversible; you can always have another conversation. Putting an AI-synthesized "flight risk" note in someone's permanent record is not; the record follows them. The less reversible the decision, the slower and more careful you should be. Naomi's planned move, quietly reducing Theo's scope, was hard to reverse and easy to misread as a demotion.
4. The Accountability Test. Can I explain and defend this decision to the person it affects? If you cannot look Theo in the eye and explain how you decided, you should not decide it that way. "An algorithm flagged you" is not a defense any human should have to accept.
Run through that scenario and the framework was unanimous. Naomi closed the AI draft and put a one-on-one with Theo on the calendar instead.
Walking the Worked Scenario to Its End
Here is what happened when Naomi chose conversation over prediction. She did not open with accusations or hidden suspicion. She said: "I want to check in on how you are feeling about your role. You have seemed a little less energized lately, and I want to make sure everything is okay."
Theo's answer demolished the algorithm's story. He was not disengaged or planning to leave. His mother had been seriously ill, which had pulled his focus for a few weeks, and separately he was hungry to grow into a lead role and unsure whether a path existed. The AI had seen reduced activity and predicted "departure." The reality was a stressed employee who wanted more responsibility, not less. Those are opposite problems with opposite solutions. Had Naomi acted on the prediction, she would have shrunk the role of the one person asking for a bigger one and likely pushed out someone who wanted to stay.
This is the core lesson in miniature. Algorithms predict trends; people are more complicated than trends. AI is excellent at generating a hypothesis ("maybe Theo is disengaged, worth a check-in"). It is dangerous when you treat that hypothesis as a verified fact and act on it without a single conversation.
Five Gray Areas You Will Actually Face
The framework generalizes. Here are the recurring gray areas where managers most often slip, each with the line that separates responsible use from misuse.
Drafting feedback. Responsible: use AI to organize your thinking or find a precise word, then rewrite in your own voice with specific examples you actually observed. Over the line: sending an AI draft as-is. It might be technically fair feedback, but it does not sound like it came from a manager who knows the person and cares about their growth.
Synthesizing team patterns. Responsible: use a pattern the AI surfaces (say, two people rarely speak to each other in meetings) as a hypothesis to explore through one-on-ones. Over the line: secretly analyzing months of transcripts to build a profile of "problem" dynamics and acting on it. That is surveillance, and it violates the implicit trust between you and your team.
Decisions about people. Responsible: AI generates a hypothesis you verify through conversation before acting. Over the line: treating an AI prediction as evidence and making life-affecting calls (reassignment, documentation, succession planning) on pattern-matching rather than dialogue. This is exactly the trap Naomi nearly fell into.
Efficiency versus relationship. Responsible: AI handles the admin (a draft agenda template) so you are more present and prepared in the actual conversation. Over the line: automating one-on-ones so heavily they become a checklist. The relationship is the core of your job; optimizing it away solves the wrong problem.
Bias at scale. Responsible: when you use AI for anything touching hiring, performance, or evaluation, you actively check whether it treats different groups differently and you never let it make the final call alone. Over the line: trusting an AI screen or score because it feels objective. AI trained on biased history launders that bias into a number that looks neutral and is not.
A Worked Fairness Check on Performance Reviews
Later that quarter Naomi used AI properly on all 10 reviews, and her process is worth copying. She used AI to synthesize multiple sources (her notes, peer feedback, self-assessments, project outcomes) and to organize each review into clear sections. Then she applied a deliberate fairness checklist before any review left her hands:
- Read every word against what I actually know. For the reviews where the AI draft felt off, she rewrote significantly rather than accepting tidy corporate phrasing.
- Add genuine specifics. "You showed up when the platform went down at 2am, that is leadership" beats "demonstrates leadership readiness," which sounds like it came from a template.
- Check for consistency. Is the feedback applied to the same standard across people? Does it quietly penalize the quieter designers for being quiet?
- Watch for amplified bias. If her peer-review inputs skewed toward people who communicate like the team's majority, the AI would faithfully amplify that. She corrected for it.
- Catch missing context. One designer's lower output traced to a family situation mentioned once in passing. The AI saw the data, not the reason. Naomi did.
The numbers made the stakes concrete. Across 10 reviews, the checklist surfaced two where the AI draft had understated a quiet contributor's impact and one where context completely changed the story. Three meaningful corrections out of 10 is not a rounding error; it is the difference between fair reviews and unfair ones. The AI saved her hours of structuring. Her judgment saved three people from being misjudged.
Knowing When to Say No
Part of ethical judgment is recognizing the situations where the right move is to not use AI at all, or to slow down. A few honest red flags tell you when you have drifted:
Using AI to avoid a hard conversation. "I will just synthesize their feedback instead of asking them" is the tell. Ask the person first; AI can help organize afterward.
Treating AI output as fact. The moment "the AI says X" becomes your reason for a decision about a human being, stop. AI generates hypotheses, not verdicts.
Hiding your AI use. If you are reaching for AI but planning never to mention it, ask why. People will eventually find out, and discovery damages trust far more than transparency ever would.
Choosing speed over authenticity. "This draft is good enough" when it does not sound like you is a quiet surrender. Some things, feedback, mentorship, difficult conversations, are worth doing slowly. Slow sometimes means better.
Your values are the deciding vote. Before any sensitive AI use, name what actually matters to you as a manager (authenticity, trust, fairness, development, psychological safety) and let those values, not the convenience of the tool, guide where you use AI and where you put it down.
Practice and Reflection
Judgment is a muscle, and muscles only grow under load. The exercises below are the load. Work through them with a notebook rather than in your head, because writing forces the honesty that thinking alone lets you dodge. Naomi did all five over a single quarter, and by the end she could feel the ethical tension in a decision before she had finished reading the prompt she was about to type.
- Run one live decision through the four tests. Pick something sensitive you are using or considering using AI for right now. Answer in writing: What is my purpose here, honestly? Speed, better thinking, fairness, or convenience? What am I risking? Authenticity, trust, fairness, or privacy? Is this transparent, meaning would I be comfortable if the people involved knew? Is the tradeoff actually worth it? And could I defend this decision to the person it affects? If the last answer is no, reconsider the decision, not the wording.
- Take a values inventory. Write down what genuinely matters to you as a manager. Not what should matter, what actually does. Choose your top three from authenticity, trust, fairness, development, psychological safety, efficiency, respect, and growth. Then hold your current and planned AI use up against them. Does the way you use AI amplify those values or quietly undermine them? Where do they conflict? If authenticity sits at the top, you will rewrite every draft. If efficiency does, you will accept more output as-is. Neither is wrong, but you need to know what you are trading.
- Test consent where AI touches people. List the places where AI affects people significantly: feedback, performance synthesis, pattern analysis, decisions about who gets what. For each, ask whether you would tell them. If the answer is "not really" or "probably not," stop and ask why. Is it because they might mind, because it would sound intrusive, because you are not certain it is your right? Sometimes discomfort is not squeamishness. It is your integrity talking, and it is worth listening to before rather than after.
- Map the gray areas of your own role. Every role has its own. Yours might be hiring decisions, performance management, team dynamics analysis, or career development conversations. For each one, write four things: how you currently use or plan to use AI there, what could go wrong, how you would feel if the person found out exactly what you did, and how you intend to navigate it. The gray areas you cannot name are the ones that will catch you.
- Set boundaries you will actually hold. This is the exercise that turns reflection into behavior. Do not write what you should do; write what you will do. Naomi's list began with three lines: she would not analyze team communication patterns without telling the team first, she would use AI to organize her thinking but always rewrite feedback in her own voice, and she would never let an AI screen make a hiring call without human validation. Write yours down, keep them somewhere you will see them, and revisit them each quarter as your tools and your understanding change.
Related Lessons
Two lessons extend this one directly. Bias Awareness and Mitigation goes deep on the specific ethical risk you met here in passing, showing you how to spot bias amplification and prevent it rather than discover it later. Maintaining Authenticity and Trust takes the efficiency-versus-authenticity tension seriously and works through how to keep AI from eroding your authenticity, and how to recognize the moments when the right move is to slow down.
Key Takeaways
- "I can" is not "I should." AI is silent on the second question, and that gap is where your ethics live. Build judgment, not a rule book, because rules break under real complexity.
- Run the four tests on anything sensitive. Purpose (does AI serve what I am actually trying to do?), Transparency (would I tell them?), Reversibility (can I undo a mistake?), and Accountability (can I defend this to the person affected?). They reveal the tradeoff so your values can decide.
- AI predicts trends; people are more than trends. Treat AI output as a hypothesis to verify through conversation, never as a verified fact you act on. A single one-on-one can demolish a confident wrong prediction.
- Transparency beats secrecy every time. People judge you less for using AI and more for whether you were honest about it. Hidden use discovered later erodes trust far worse than open use up front.
- AI amplifies bias and mistakes at scale. When AI touches hiring, performance, or evaluation, actively check for unequal treatment and never let it make the final call alone. Objective-looking numbers can launder old bias.
- Efficiency that costs authenticity solves the wrong problem. Some work (feedback, mentorship, hard conversations) is worth doing slowly. Use AI for the admin so you can be more present in the relationship, not less.
- Know when to say no. Avoiding a hard conversation, treating output as fact, hiding your use, or accepting "good enough" that does not sound like you are all signals to slow down or set the tool aside.
Skill.re