←
AI for Managers
Aware · M1 · lesson 1 of 26 · in progress
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

AI Capabilities and Limitations

17 min

Naomi Adeyemi runs a five-person marketing operations team at a healthcare-software company, and she nearly lost her credibility to an AI tool in her first month using one. She asked it to pull the latest quarterly benchmarks for email open rates in her industry and drop them into a board slide. The numbers looked authoritative, so she used them. In the meeting her VP asked for the source, Naomi could not produce one, and a quick check showed two of the figures were simply invented. The tool had not lied on purpose; it had done exactly what it does, which is generate plausible text, and plausible is not the same as true. That afternoon Naomi stopped asking whether AI is "good" and started asking a sharper question: good at what, specifically, and how do I tell before I trust it? This lesson is the map she built.

What This Lesson Covers

The single most useful thing a manager can know about today's AI is where it is reliable and where it is not, because that knowledge tells you which work to hand it and which to keep. AI here means a large language model, the kind of system behind chat assistants, which works by predicting likely text one piece at a time. It is not a database, not a calculator, and not a reasoning engine with guaranteed answers. Understanding that one fact explains almost every strength and every failure that follows.

You will learn the tasks AI does genuinely well, the tasks where it reliably fails and why, how to calibrate how much to trust any given output, and a simple scoring rubric you can apply to any team task to decide whether AI should touch it. We will follow Naomi as she scores real work from her own team.

The reason this matters is that the common mistake runs in both directions. Treat AI as a magical solution and you hand it work it cannot do; dismiss it as hype and you spend hours on tasks it would finish in minutes. Matching tasks to capability protects four things Naomi cares about: her credibility, because asking a tool to do something it plainly cannot do makes you look uninformed; her productivity, because the right tasks genuinely save hours; her risk exposure, because some delegations create real problems; and her team's trust, because they are watching whether their manager uses this stuff sensibly.

What AI Genuinely Does Well

AI is strongest when a task reshapes language you can immediately judge, has more than one acceptable answer, and gets reviewed by a human before it goes anywhere. That describes a large slice of a manager's week.

  • Drafting: first drafts of emails, memos, announcements, job descriptions, and proposals. You see the output and judge it instantly, and "good writing" has many valid forms, so there is no single right answer to miss.
  • Summarizing: condensing a long thread, document, or transcript into key points and action items. You can hold the summary against the original and verify it in seconds.
  • Classifying and extracting: sorting fifty customer comments into bug, feature request, or praise, or pulling every name and date out of a contract. The criteria are clear and you can spot-check a sample.
  • Brainstorming: generating ten angles on a problem when you only had two. Quantity is the point, and you filter, so a few weak ideas cost nothing.
  • Translating and explaining: rendering a message in another language, or explaining a concept at a chosen level for a specific audience.
  • Code help: drafting a formula, a script, or a query for a technical teammate to review and run.

Two quieter strengths belong on the same list because they have the same shape. AI is very good at reformatting and restructuring: turning a paragraph-form process document into a numbered step-by-step guide, converting a list into a table, reshaping an email into a memo. The original sits right beside the output, so any error is obvious at a glance. It is also good at reading patterns in text: the sentiment running through customer comments, the three complaints that recur across thirty testimonials, the assumptions buried in a draft. Naomi uses both constantly, and in each case she verifies by sampling rather than by faith.

The thread connecting all of these is the same: the output is visible, judgeable, and reviewed by you before it counts. When Naomi rebuilt her benchmark slide, she did the research herself, then handed AI the verified numbers and asked it to write the paragraph around them. That played to its strength, drafting, instead of its weakness, knowing facts.

Where AI Reliably Fails

The failures are not random; they follow from how the system works, which means you can predict them.

Hallucination. Because it generates plausible text rather than retrieving verified facts, AI will confidently produce wrong names, fake statistics, and citations to articles that do not exist. It does not know it is wrong, because it has no concept of true. This is what burned Naomi. Any output that asserts a specific fact must be verified against a real source before you rely on it.

No real-time or private knowledge. The model was trained on data up to a cutoff date and cannot see this morning's news, your internal documents, or last week's numbers unless you paste them in. Ask it for "the latest" anything and you will get either an old answer or an invented one.

Math and counting. It predicts the text of an answer rather than calculating it, so arithmetic, counting items, and multi-step numeric reasoning are genuinely unreliable. Use a spreadsheet for the math and let AI write the explanation around the result.

No guaranteed reasoning. It can produce text that looks like careful logic and still reach a wrong conclusion, especially on anything with many steps or hidden constraints. It is a pattern-matcher, not a thinker, however convincing the prose.

No context about your world. Your team's history, politics, relationships, and values are not in the model. It cannot know that a "reasonable" plan ignores the one colleague who will block it.

No accountability. If an AI-assisted decision goes wrong, the responsibility is entirely yours. The tool cannot be answerable for a livelihood, a legal outcome, or a person's career.

The Middle Ground: Useful, But Only With Real Review

Between the work AI does well and the work it fails at sits a wide band that Naomi found the most confusing at first. These are tasks AI can genuinely help with, as long as you treat its output as a skeleton you finish rather than an answer you accept.

Analytical summaries. AI can compare two approaches and lay out the pros and cons, evaluate options against criteria you define, describe trends in data you supply, and suggest what they might imply. When Naomi was choosing between two marketing-automation vendors, she gave the tool her three criteria, cost, integration capability, and support quality, and got a clean comparison in a minute. What she watched for is what it always misses: context she never supplied, analysis that stays at the surface, confidence stated more strongly than the evidence supports, and assumptions inherited from its training data. Her habit now is to push back on every analysis with two questions, "what are you basing this on?" and "what might I be missing?"

Planning and project management. Ask for a timeline, milestones, likely risks, a checklist for a complicated task, or the questions you should be asking stakeholders, and you will get a usable starting structure. The gap is everything specific to you. The model does not know your team's real availability, the politics around the project, the dependency nobody wrote down, or the budget cycle that constrains the whole plan, so the output tends to read as something that would work for any organization. Naomi treats it as a first skeleton and says so out loud when she shares it: this timeline assumes full team availability, which we do not have, so here is what I changed.

Decision support. AI can list the considerations in a decision, suggest a framework for thinking it through, play devil's advocate against your preferred option, ask clarifying questions, and generate scenarios you had not pictured. What it cannot do is decide, because it does not know what you actually value, what your organization stands for, or which relationships the choice will strain. When Naomi was weighing whether to promote from within or hire externally, she used AI to widen her thinking, not to pick. The framing that keeps this honest is simple: it is a thinking partner, and you are the decision maker.

Calibrating How Much to Trust an Output

Trust is not all-or-nothing; it scales with the task. Before relying on any AI output, Naomi runs five quick questions, and the more "yes" answers, the more human checking she adds.

  • Is it factual or creative? If it asserts facts, verify them. If it is a draft you will shape, lighter review is fine.
  • Does it need context the AI lacks? If your organization's specifics matter, you supply them and review carefully.
  • What happens if it is wrong? Low consequence means use freely; high consequence means verify or escalate.
  • Could it affect someone's job, pay, or livelihood? If yes, AI may organize information but never make the call.
  • Is it reversible? An email you can correct is forgiving; a public statement or a hiring decision is not.

This is how Naomi turned a vague unease into a routine. Drafting a team newsletter scores low across the board, so she uses AI freely. Choosing which of two people to promote scores high on nearly every question, so AI never touches the decision itself.

The Work That Stays Entirely Human

Some tasks are not a matter of careful review; they are simply off-limits for AI as the decision-maker, no matter how good the draft looks. The common thread is that a mistake lands on a person's life and the responsibility is legally and morally yours. Naomi keeps a short, firm list.

  • Performance evaluations with consequences. A review that affects pay, promotion, or someone's job needs your knowledge of the person's growth, context, and potential, none of which the model has. Naomi uses AI only to surface themes from feedback she has gathered; she writes every word of the judgment herself.
  • Sensitive HR and people conversations. Termination, accommodation, conflict, or discipline carry legal weight and depend on tone the model cannot read. Draft with AI if you like, but a human, often HR or legal, reviews before anything is said.
  • Anything with legal force. AI is not a lawyer and will advise wrongly with total confidence. Use it to summarize a concept so you understand it, then get a real answer from someone qualified.
  • Irreversible, high-stakes calls. Hiring, firing, public statements about a serious issue, and major strategy shifts cannot be undone. AI can lay out options and trade-offs; the decision is yours, because only you are accountable for it.

Two more belong on that list. Compensation decisions are off-limits because the model does not know real market rates, has no idea of your organization's compensation philosophy, and cannot see the equity and discrimination questions that sit underneath every pay number. Use it to help you understand general market concepts if that helps, then decide with your own judgment and your organization's approach. Ethical trade-offs are off-limits for a related reason: ethics requires values, and the model has none of its own, only patterns from what it read. Naomi will happily ask it to walk through what might happen if her team took a particular approach, because surfacing consequences is useful. The judgment about whether that approach is right stays with her, because she is the one who has to answer for it.

When a customer-data privacy question landed on Naomi's desk, she felt the pull to ask AI for a quick answer. Instead she sent it to her company's privacy contact, got the real ruling, and only then used AI to help her explain the rule to her team in plain language. The tool served her where it was strong, explaining, and stayed out of the place where it was dangerous, deciding.

A Worked Example: Scoring Real Team Tasks

To make the calibration concrete and shareable, Naomi built a small task-suitability rubric. She scores any task on three factors, each from 1 to 5, then reads the total. The factors are:

  • Verifiability (5 = I can instantly tell if the output is good; 1 = errors would slip past me unseen).
  • Stakes (5 = a mistake is cheap and reversible; 1 = a mistake is costly or irreversible). Note the scale is inverted on purpose, so high is always safe.
  • Context independence (5 = the task needs no private or organizational knowledge; 1 = it depends heavily on context only I hold).

She reads the total of 3 to 15 like this: 12 to 15, use AI confidently with a light review; 8 to 11, use with careful review; 3 to 7, do it yourself or use AI only to organize inputs you will fully redo. Here are four real tasks from her team scored against it.

  1. Draft the monthly client newsletter: Verifiability 5 (she reads it and knows), Stakes 4 (easy to edit, low risk), Context independence 4 (mostly generic). Total 13: use AI confidently. She drafts with AI and edits in ten minutes.
  2. Summarize 60 survey responses into themes: Verifiability 4 (she can spot-check), Stakes 4 (informs but does not decide), Context independence 4 (the text is self-contained). Total 12: use confidently, then sample a few responses to confirm the themes hold.
  3. Pull current competitor pricing for a strategy deck: Verifiability 2 (wrong numbers look right), Stakes 2 (a board deck is high-visibility), Context independence 3. Total 7: do it yourself. This is the exact trap she fell into; she now gathers the numbers manually and uses AI only to write the surrounding paragraph.
  4. Write a direct report's performance review: Verifiability 2, Stakes 1 (it affects a career and carries legal weight), Context independence 1 (it depends entirely on her knowledge of the person). Total 4: do it yourself. AI may surface themes from collected feedback, but the judgment and the words are hers.

The rubric took Naomi an afternoon to build and now takes thirty seconds per task. Its real value is that it gave her whole team a shared, defensible way to decide, so the question "should AI do this?" stopped being a gut call and became a quick, repeatable check. The exact thresholds are hers to tune; the discipline of scoring before delegating is what matters.

Common Traps to Avoid

Treating output as final. "The AI wrote it, so it is done" is how errors escape. For anything that counts, AI produces the first draft and your review is the last step.

Expecting it to know your organization. Your culture, history, and constraints are not in the model. If context matters, you have to supply it every time; the system does not remember yesterday's conversation.

Asking it for subjective judgment as if it were fact. "Is this email professional enough for our culture?" assumes the model knows your culture. It can describe how the text reads; you decide whether that fits.

Trusting its risk assessment blindly. "The AI said the risk is low" misses that it only sees patterns from training data, not the risks specific to your situation. Use it to list possible risks, then assess them with your own domain knowledge.

Expecting the same answer twice. Because generation is probabilistic, the same question can produce noticeably different wording and even different emphasis from one day to the next, and nothing carries over between sessions. Naomi learned to stop treating consistency as automatic: she saves the outputs that matter, restates the key context at the start of a new conversation, and writes decisions down so she can hand them back to the tool. "Yesterday we decided this, so given that, what should we do?" gets her continuity the model cannot provide on its own.

Using This Responsibly

Being honest about the limits of your own verification is part of using AI well. If a draft touches an area where you cannot personally check the facts, say so rather than passing it along as though you had. "AI drafted this and I am not an expert here, so treat it as a starting point, not a verified source" is a responsible sentence, and Naomi uses it without embarrassment.

The accountability point deserves repeating because it drives everything else. Anything you send that AI helped produce is yours. The tool carries no responsibility for a wrong number in a board deck or an unfair line in a review, which means the amount of review you do should scale with what you would have to answer for. That is the whole logic behind the rubric.

The same honesty applies when you introduce AI to your team. Overselling it sets them up to be burned exactly as Naomi was, and dismissing it wastes the hours it could genuinely save. What she tells her team now is plain: it is very good at summarizing and drafting, it is not good at knowing our specific situation or making judgment calls, so we use it to save time on the first draft and we keep the decisions.

Practice and Reflection

Work through these with your own tasks rather than in the abstract; the value is in the specifics.

  • Sort your real work. List five tasks you do regularly and score each one on Naomi's rubric. Which land in the confident band, which need careful review, and which should never leave your hands?
  • Design the verification. Pick one task you want to hand to AI and describe exactly what checking it would take. What could go wrong, and how would you catch it before anyone else did?
  • Name the missing context. Think of a decision in front of you right now. What does the model not know about your situation, and which part of the call requires judgment no tool can supply?
  • Explain it to your team. In your own words, how would you describe what AI is good and bad at in your organization specifically? Write the two or three sentences you would actually say.
  • Find the hidden risk. Identify one task that sounds like an obvious fit for AI but carries a risk that is not visible at first glance. What makes it riskier than it looks?

This map sits on top of two earlier ideas and is worth reading alongside them.

  • What AI is and Isn't establishes what these systems actually do, which is pattern recognition and generation rather than thinking. Everything in this lesson about strengths and failures is a consequence of that definition.
  • How Generative AI Works explains the token-by-token, probabilistic mechanism behind the output. If you want to understand why hallucination, weak arithmetic, and inconsistent answers all show up together, that is where the explanation lives.
  • The Managers Role in an AI World picks up where this lesson stops. Once you know which work the tool can carry, the natural next question is what that leaves for you, and that lesson answers it.

Key Takeaways

  • AI predicts plausible text; it does not know facts. Every strength and every failure flows from this one mechanism, and understanding it lets you predict where the tool will let you down.
  • Play to the strengths: drafting, summarizing, classifying, brainstorming, translating, code help. These share visible, judgeable output that you review before it counts.
  • Verify anything factual. Hallucinated statistics and fake citations look authoritative. Gather facts yourself, then let AI write the prose around them.
  • Calibrate trust to the task, not to the tool. Factual, high-stakes, context-heavy, or irreversible work demands more human review; cheap reversible drafting demands almost none.
  • Score tasks before you delegate them. A simple verifiability, stakes, and context rubric turns "should AI do this?" from a gut call into a quick, defensible, shareable check.
  • Analysis, planning, and decision support are skeletons, not answers. They save real time as long as you supply the missing context and make the call yourself.
  • Accountability never delegates. The tool cannot answer for a wrong decision, a harmed career, or a legal outcome. That responsibility stays with you, and it should shape how much you verify.

Frequently Asked Questions

If AI makes up facts, how can I trust anything it produces? By separating the two things it does. It is unreliable at knowing facts but reliable at shaping language. So you supply the verified facts and let it draft, summarize, or reformat them. The trust is not in the content it invents; it is in its ability to rework content you already trust. That division is the whole skill.

Why is AI bad at math when computers are great at math? A calculator computes; a language model predicts the text of an answer based on patterns it has seen. They are different kinds of programs. The model will often get simple arithmetic right because the answer appeared in its training, and then confidently fumble a multi-step calculation. For anything numeric, do the math in a spreadsheet and let AI explain the result.

How do I explain to my team what AI is and is not good at? Keep it honest and concrete: "It is great at first drafts and summaries, and we use it to save time there. It is bad at facts, our specific context, and judgment calls, so we always verify those and never let it make decisions about people." Pair that with the scoring rubric so the guidance is a shared practice, not just a slogan.