←
AI for Leader
Aware · M3 · lesson 3 of 28 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Building Your AI Intuition as a Leader

10 min

Opening

AI intuition is the ability to make better decisions about AI faster. It's not about understanding neural networks or knowing linear algebra. It's about developing a felt sense for what AI can and can't do, where the blind spots are, and what questions to ask. David, a CFO at a mid-market insurance company, sat across from a vendor demo. The vendor showed him a claims classification AI. "It's 99% accurate," the vendor said. Every executive in the room was impressed. David asked: "99% accurate on what data? Claims from the past three years? What if the next recession creates claim types we've never seen before?" The vendor deflated. The rest of the room realized: they had no idea what they were evaluating. They had no framework for spotting what the vendor wasn't saying. David had intuition. Most executives don't. Intuition is learnable. It comes from asking better questions, understanding what AI actually does (pattern matching from historical data), recognizing when that breaks (novel situations, data shifts, changes in the environment), and seeing the organizational obstacles that separate good technology from good business outcomes. This lesson teaches you to build that intuition, so you can evaluate AI proposals with clarity instead of hope.

Why This Matters

AI intuition directly affects your decision quality and your organization's ability to capture AI's value without wasting resources on the wrong initiatives. Without intuition, you rely on what vendors tell you, what the loudest person in the room believes, or what seems trendy. You fund the wrong pilots. You hire the wrong consultants. You miss the projects that would actually move the needle. With intuition, you ask better questions, spot red flags early, and make faster decisions with higher confidence. Consider the stakes: a single misguided AI initiative costs $500K to $5M (licensing, consulting, staff time, opportunity cost). Organizations that build strong AI intuition in their leadership reduce wasted spend by 40% and cut decision time in half. They also become better at spotting which vendors are bullshitting and which ones are real. Most importantly, they're no longer at the mercy of their technical teams to interpret what's possible. They can evaluate tradeoffs themselves.

Let's put numbers to the cost of getting this wrong. Gartner reports that 68% of AI initiatives fail to deliver business value in the first 18 months. The reasons? Mostly not technical. Mostly organizational. But it starts with misunderstanding what's real. Leaders allocate $2.1M to an AI initiative expecting a 30% efficiency gain. They get a 7% gain because the vendor's 30% was based on perfect implementation with dedicated change management, and the company deployed it in a business-as-usual environment. That's a $1.4M gap between expectation and reality.

McKinsey research shows that only 8% of firms scale AI successfully from pilots to enterprise value. The other 92% get stuck. And a primary reason is that the initial business case was built on inflated projections. Stakeholders funded the pilot based on a 'conservative' 20% uplift claim. The pilot delivered 8% uplift. Stakeholders feel betrayed. Funding for the next AI initiative becomes political. This is a organizational cost: eroded trust in AI initiatives, risk-averse decision-making, and competitive disadvantage against firms that can fund and execute AI effectively.

The other cost is opportunity. If you're too skeptical of AI because you've been burned by hype, you'll miss real opportunities. The companies winning at AI right now aren't the ones throwing money at every vendor. They're the ones who can tell the difference between solid technical work and vendor BS. They say 'yes' to the real opportunities. They say 'not now' to the premature ones. They allocate capital effectively. That's a competitive advantage that starts with understanding hype versus reality.

The Core Idea

AI intuition rests on understanding five core ideas: (1) AI is pattern matching. It finds patterns in historical data and applies them to new data. That's powerful. But the moment the new data is fundamentally different from the old data, the pattern breaks. Recession creates new claim types. Consumer preferences shift. Fraud tactics evolve. Your model was trained on historical data that no longer predicts the future. (2) Data quality matters more than model sophistication. A simple model on clean, representative data beats a complex model on garbage data. Most executives obsess about the algorithm. The real leverage is data. (3) The accuracy metric hides what matters. A model can be 95% accurate and still fail. If it's 95% accurate at predicting common cases but 60% accurate at predicting rare cases (the ones that matter), you have a problem. The metric is deceptive. Ask: accurate at what? On what data? What happens when you get it wrong? (4) Organizational readiness determines whether AI creates value. A brilliant model sitting in a lab is worthless. A mediocre model integrated into operations is valuable. Most failures aren't technical failures. They're organizational failures. (5) Every AI system has blind spots and failure modes. No system works in all contexts. Your job isn't finding a perfect system. It's understanding where it fails and building guardrails around those failure modes.

To understand this more deeply, let's build a framework. Mature AI technology (worked on real business problems for 5+ years):

  • Classification: Is this email spam? Is this image a cat? Is this transaction fraudulent? This works well.
  • Regression: Given these inputs, predict this number. Will this customer spend $X in the next quarter? This works well.
  • Anomaly detection: Is this data point unusual relative to the pattern? Has network behavior changed? This works well.
  • Recommendation: Given what users like, what should we recommend next? This works well in specific domains with good data.

    These technologies have real track records. They save money. They improve processes. They've been in production for years. When a vendor claims these capabilities, you can be reasonably confident the technology itself is solid. The question becomes: Will it work on your data? Will adoption succeed? Are the economics real?

    Emerging AI technology (2-5 years in production, rapidly improving):

    • Generative language models: Writing, coding, reasoning across domains, explaining, summarizing. Real capability. Real limitations. Hallucination is a real problem. These tools are genuinely useful but require human oversight.
    • Vision models: Specialized to specific domains. Very good at specific tasks. Don't generalize well to new domains. The headline accuracy is often deceptive.
    • Time-series forecasting with deep learning: Better than traditional methods in some cases. Not in others. Requires careful validation.

      When vendors claim these capabilities, you should probe more. The technology is newer. The failure modes are less well understood. Implementation requires more experimentation.

      Vaporware AI (claimed but not production-ready):

      • 'Our AI will replace your whole customer service team.' Nope. It's a tool that handles 35-45% of routine inquiries, requiring human review on complex cases.
      • 'This AI system doesn't need maintenance.' Nope. All systems drift. All systems require monitoring and retraining.
      • 'Our AI understands your business problems after reading your documentation.' Nope. Understanding comes from experimentation with your actual data and processes.

        Here's the key distinction: Real AI advantages in production come from bounded, well-defined problems with good data. Real disadvantages come from oversized change management, data quality issues, and integration complexity. The hype focuses on the capability. Reality includes the integration.

Think of It Like This

AI intuition is like developing taste in wine. A novice drinks wine and says "I like this or I don't." A sommelier tastes wine and says "This has notes of berry, the tannins are well-balanced, it pairs with this food but not that food, and here's why." Both are tasting the same wine. But the sommelier has frameworks for understanding what they're experiencing. They know what matters and what doesn't. They ask better questions and make better choices. AI intuition works the same way. Without it, you hear "This model is 95% accurate" and think that's good. With intuition, you ask: "95% on what? What fails? What data was it trained on? What happens when the environment changes?" You're asking better questions because you have frameworks for what matters.

Let's extend this analogy further. When you're evaluating a new manufacturing process, you'd ask:

  • Where was this tested? In a lab? In a pilot facility? In production for two years?
  • On what products? The ones we make? Similar products? Very different products?
  • What assumptions does it rely on? Specific labor skills? Specific equipment? Specific material quality?
  • How sensitive is the gain to those assumptions? If labor quality drops 10%, does the gain drop 5% or 50%?

    AI evaluation follows the same logic, but translated into data and model language. A vendor claims their system improves loan approval accuracy by 18%. Ask:

    • Tested on what data? Data from your bank? Data from similar banks? General lending data?
    • What types of loans? Mortgages? Personal loans? Small business loans? All types?
    • What's the baseline accuracy? Compared to what? Manual review? An older system?
    • How does accuracy vary by applicant demographic? (This is legally important.)
    • How often will the system recommend 'escalate to human'? (This is operationally important.)
    • What's the worst-case scenario? If the system is wrong, what happens? Is it reversible?

      A 18% improvement that's tested on your data, across your loan types, with demographic parity and clear escalation paths is different from an 18% improvement that's based on academic datasets and hasn't been tested on your applicants. Same accuracy number. Different reality.

What This Looks Like in Real Life

An operations director at a logistics company was evaluating an AI system for route optimization. The vendor said the system had been running at 34 major shipping companies and saved them an average of $2.3M per year. The director asked: "On what data was this trained? How much was the data similar to our operations?" The vendor said "The best way to know is to run a pilot." Translation: the vendor doesn't know if it will work for you. The director had intuition. She pushed back. "Show me the pilot methodology. How will we know if it's working? What metrics matter?" The vendor got defensive. That's a red flag. (Real vendors are confident in their pilots because they've run dozens.) The project was declined, saving the company from a $1.2M mistake. A healthcare CFO was pitching an AI system for patient readmission prediction. The vendor said it was 89% accurate. The CFO asked: "89% accurate at predicting what? Accurate for which patient populations? High-risk or low-risk patients? What's the false negative rate?" The vendor backtracked. They hadn't broken down accuracy by patient population. It turns out the model was 94% accurate at predicting low-risk readmissions (easy to predict) and 68% accurate at high-risk readmissions (the ones that matter). The CFO's intuition saved the organization from deploying a system that would fail exactly where it needed to work. A manufacturing VP evaluated a predictive maintenance system. The vendor said it would identify equipment failures 72 hours in advance. The VP asked: "Trained on what equipment? What if we upgrade a machine? What if we change maintenance schedules? How does the model adapt?" The vendor said these edge cases would require retraining. The VP had intuition: the model was brittle. It wouldn't adapt to change. She declined and built an internal system instead that could learn continuously. Intuition matters.

Let's walk through a fourth example in detail. A logistics company with 800 employees and $400M annual revenue evaluated an AI system to optimize their delivery routes. The vendor showed a case study where a similar company reduced delivery costs by 22%. Impressive claim. Before committing $3.2M to the implementation, the company did a detailed pilot.

The pilot revealed several reality gaps:

First, the 22% in the case study was for the vendor's 'standard' delivery environment: urban delivery, predictable traffic patterns, stable fleet size. The logistics company operated in three environments: urban (30% of volume), suburban (40%), and rural (30%). The vendor's system was highly optimized for urban. On suburban and rural routes, the system's recommendations often created longer drive times because they didn't account for the sparse pickup/delivery pattern. The 22% gain compressed to 6% across all routes.

Second, the case study assumed the system would run on historical data. But the company wanted the system to optimize routes in real-time. Real-time optimization requires the system to know traffic conditions, driver availability, and customer timing constraints as they evolve. The vendor's system was good at 'given these constraints, here's the best route.' It was poor at 'these constraints are changing; adjust now.' Retraining and redevelopment would cost another $800K and take 6 months.

Third, the case study didn't account for driver adoption. Drivers who had been optimizing routes themselves for years didn't trust an AI system's recommendations, especially when those recommendations contradicted their experience. The company needed 4 months of change management, driver training, and iterative adjustments before drivers actually followed the AI's recommendations.

The result: A 6% delivery cost reduction (instead of 22%) took 9 months to implement (instead of the projected 4 months) and required $4M in total investment (instead of $3.2M). The system is valuable. It's working. But the gap between vendor claim and delivered value was substantial. The company now has a realistic view of what the system does. And they know that next time they evaluate AI, they'll pilot on their actual data and conditions, not just trust the case study.

Where People Get This Wrong

Mistake 1: Assuming accuracy means the AI is ready to deploy. Accuracy is one metric. It hides failure modes. A model can be 95% accurate and still make terrible errors if it's inaccurate on the cases that matter most. Always ask: accurate on what subset of data? What does it miss? Mistake 2: Treating vendor claims as facts. Vendors have incentives to overstate. A 95% accuracy number might come from a dataset that doesn't match your situation. A $2M savings claim might come from a company 10x bigger with different operations. Always ask: under what conditions? On what data? On companies like ours? Mistake 3: Underestimating how much data quality and organizational readiness matter. Leaders obsess about model sophistication and ignore data pipelines and change management. The sophistication doesn't matter if the data is garbage or the organization isn't ready to use the system.

Let's add three more mistakes that leaders often make:

Mistake six: 'If we implement this AI system, it will fix our underlying data quality problems.' Wrong direction. AI amplifies bad data. If your data quality is poor, an AI system trained on poor data will make poor decisions confidently. You fix data quality first, then add AI. A customer analytics AI system trained on messy customer data will confidently categorize customers incorrectly. It won't suddenly become insightful. Fix the data. Then add AI.

Mistake seven: 'This AI system is a one-time investment. Build it and we're done.' No. AI systems require ongoing maintenance. Models drift over time. New data patterns emerge. New regulations require new constraints. The model you build in month six won't perform the same in month eighteen. Budget for continuous monitoring, retraining, and optimization. Most failed AI initiatives failed because the organization budgeted for implementation but not for operation.

Mistake eight: 'The vendor handles all the risk. If the AI doesn't work, it's their problem.' Legally and operationally, it becomes your problem. Your brand suffers if the AI makes bad recommendations in your name. Your risk exists. You need governance, monitoring, and the ability to turn the system off. Vendors can't take that responsibility away. They can share it. But they can't eliminate it.

Practical Takeaways

(1) When evaluating an AI proposal, start by asking: "What historical data was this trained on? How similar is that to our situation?" This reveals whether the vendor has any basis for claiming the system will work for you. (2) Ask for disaggregated accuracy. Don't accept a single accuracy number. Ask: Accurate at what? Accurate for which customer segments? What's the false positive rate? False negative rate? This reveals what the system gets wrong. (3) Understand the constraints. Ask: What happens when the environment changes? When we add new data? When customer behavior shifts? Real vendors have answers. Vendors selling hype will get defensive. (4) Look for organizational readiness red flags: "Do we have the data to train this?" "Do we have ownership and incentives aligned?" "Do we have change management planned?" Most AI failures aren't technical. They're organizational. (5) Ask for comparable case studies. Not the glossy marketing case study. Real references from companies like yours, of similar size, in similar industries, dealing with similar constraints. (6) Build your intuition by regularly asking hard questions and observing whether vendors can answer them. Each time you ask a hard question and get evasiveness instead of clarity, your intuition strengthens.

Sixth, establish an AI evaluation checklist for your organization. What information do you need before you fund an AI initiative? (Testing on your data? Reference customers? Failure mode analysis? Pilot costs? Change management plan?) Standardize the questions. Everyone uses the same framework. This prevents the situation where one leader asks tough questions and another leader approves the initiative without those answers.

Seventh, after an AI system launches, publish a 'reality report.' Compare vendor claims to actual results. 'Vendor claimed 40% efficiency gain. We achieved 12%. Here's why: [data quality, adoption friction, implementation scope].' This builds organizational learning. It teaches your team to hear vendor claims with appropriate skepticism. And it focuses attention on the real levers that determine success: adoption, data quality, and integration, not just the AI algorithm.

Key Insight

AI intuition is asking better questions about what the vendor isn't saying: What was it trained on? What does it get wrong? What happens when the world changes?

Before You Move On

Find one AI proposal in your organization right now. Ask it the hard questions: What historical data? Disaggregated accuracy? What are the failure modes? What organizational readiness is required? Notice which questions the team can answer with confidence and which ones trigger "we'll learn that in the pilot." Those gaps are where your intuition lives.

Reflect on a recent AI initiative in your organization (or your industry). What were the original projections? What has the actual impact been? What accounts for the gap, if any? Is the gap because of hype, or because of valid reasons like implementation complexity or change management friction?

Now do this: Find one claim you're tempted to believe about AI. It might be 'AI will replace 40% of white-collar jobs by 2027' or 'Our AI system will improve accuracy by 30% with no organizational change needed.' Write down why you believe it. What's your evidence? What could prove you wrong? Run it against this lesson's framework. Is it a bounded claim about a specific technology on specific data? Or is it an oversize claim that sounds good but lacks specifics? This is the habit that separates decision-makers from people who get burned by hype.