Red Flags in AI Proposals and Vendor Pitches
Opening
Some proposals are so broken that experienced leaders can spot it in seconds. Missing risk disclosure. Unrealistic timelines. Vague metrics. Building governance by sensing these flags early saves you from learning the hard way. This lesson teaches you the red flags and what they mean.
You're reviewing a proposal for an AI system to optimize supply chain costs. The team projects 'conservative estimate: 25% cost reduction.' Your instinct tingles. No one says 'conservative' when talking about brand new technology. You ask: 'How did you arrive at 25%?' Long pause. 'We benchmarked against industry reports.' You follow up: 'Have you modeled this with our specific data?' Silence. 'Not yet. We wanted approval first.' There's the red flag. This lesson teaches you to recognize these signals instantly. Your instincts are worth trusting once you know what to look for.
This problem appears everywhere. In boardrooms, vendors pitch AI systems that promise dramatic outcomes. In email, executives debate whether an AI initiative is worth funding. In planning meetings, teams argue about which AI projects are real opportunities versus hype. The language is unfamiliar. The claims are large. The stakes are real. And most leaders don't have a framework for cutting through to what's actually true.
Sarah, the Chief Risk Officer at a $1.2B insurance company, recently sat through a pitch for an AI system that would improve underwriting accuracy. The vendor claimed 22% better accuracy and lower claims loss. Impressive. But Sarah didn't know what questions to ask. Is 22% real? How was it measured? On what data? Against what baseline? The vendor's answer was well-rehearsed but didn't actually address what Sarah needed to know. She left the meeting uncertain, which is worse than skeptical. At least skepticism has a clear direction. Uncertainty leads to inaction or to defaulting to whoever speaks with the most confidence.
You're about to change that. You're going to learn to ask the right questions and understand the difference between real AI capability and vendor aspiration.
Why This Matters
Experienced investors can sense a bad deal in five minutes. They know the red flags: missing risk disclosure, unrealistic timelines, vague success metrics, 'we'll figure it out in production.' You need to develop the same sense for AI proposals. Red flags tell you the proposer hasn't thought carefully enough about risk.
Experienced venture investors reject 95% of pitches they see. Not because the entrepreneurs are bad, but because certain patterns—visible in 10 minutes—predict failure. A leader experienced with AI proposals can develop the same pattern recognition. Red flags like 'Machine learning will find patterns' (instead of a specific approach) or 'We'll handle edge cases in production' (instead of upfront validation) reveal deep problems with proposal thinking. Proposals containing three or more red flags have an 82% failure rate. Proposals with zero red flags have a 12% failure rate. Red flags aren't disqualifying by themselves. They're signals to dig deeper.
Let's put numbers to the cost of getting this wrong. Gartner reports that 68% of AI initiatives fail to deliver business value in the first 18 months. The reasons? Mostly not technical. Mostly organizational. But it starts with misunderstanding what's real. Leaders allocate $2.1M to an AI initiative expecting a 30% efficiency gain. They get a 7% gain because the vendor's 30% was based on perfect implementation with dedicated change management, and the company deployed it in a business-as-usual environment. That's a $1.4M gap between expectation and reality.
McKinsey research shows that only 8% of firms scale AI successfully from pilots to enterprise value. The other 92% get stuck. And a primary reason is that the initial business case was built on inflated projections. Stakeholders funded the pilot based on a 'conservative' 20% uplift claim. The pilot delivered 8% uplift. Stakeholders feel betrayed. Funding for the next AI initiative becomes political. This is a organizational cost: eroded trust in AI initiatives, risk-averse decision-making, and competitive disadvantage against firms that can fund and execute AI effectively.
The other cost is opportunity. If you're too skeptical of AI because you've been burned by hype, you'll miss real opportunities. The companies winning at AI right now aren't the ones throwing money at every vendor. They're the ones who can tell the difference between solid technical work and vendor BS. They say 'yes' to the real opportunities. They say 'not now' to the premature ones. They allocate capital effectively. That's a competitive advantage that starts with understanding hype versus reality.
The Core Idea
Ten red flags: (1) 'We'll get 40% revenue uplift' with no plan how. (2) 'Machine learning will find patterns' instead of specific approach. (3) No clear success metrics. (4) 'Data quality is fine' with no testing. (5) 'We'll handle edge cases in production.' (6) No plan for monitoring after deployment. (7) No clear ownership. (8) Timeline seems too fast or too slow. (9) 'Everyone is doing this' as justification. (10) Proposer can't articulate what could go wrong.
Twelve red flags to recognize:
- VAGUE VALUE: 'We'll get 40% revenue uplift.' How? 'Machine learning will find patterns.' That's not a plan—it's a prayer.
- NO BASELINE: 'We'll improve accuracy from X to Y.' But what's X? How was X measured? If they don't cite a baseline, you can't evaluate progress.
- HANDWAVING DATA: 'Our data quality is fine.' How do you know? 'The team thinks so.' Red flag. Data needs validation, not assumptions.
- EDGE CASE DEFERRAL: 'We'll handle edge cases in production.' That's when your customers get bad experiences. Edge cases should be identified upfront.
- NO MONITORING PLAN: 'We'll deploy and see how it goes.' Dangerous. You need alerts for performance degradation.
- UNCLEAR OWNERSHIP: 'The data science team will build it, and IT will deploy it.' But who owns performance post-launch? If no one's named, you have a problem.
- UNREALISTIC TIMELINE: '6 months from concept to production deployment at full scale.' That's 6-12 months faster than industry standard. Why? Red flag.
- NO SUCCESS DEFINITION: 'Success is that it works.' Works how? Works for whom? Be specific.
- PROPOSER CAN'T ARTICULATE FAILURE: You ask 'What could go wrong?' and get a blank stare or 'Uh, nothing obvious.' That's deeply concerning. Every system has failure modes.
- COMPETITIVE URGENCY: 'Our competitor is doing this, so we need to move fast.' FOMO is not a business case. Evaluate the project on its merits.
- NO RESOURCE PLAN: 'We'll staff this as needed.' How many people? From where? If they haven't thought through resources, they haven't thought this through.
- MAGIC THINKING: 'Once we deploy this, it will automatically improve over time.' Systems don't self-improve without someone actively retraining and validating them. The most common red flags are vague value (44% of failed proposals), edge case deferral (38%), and no monitoring plan (35%).
Think of It Like This
Red flags are like smoke signals: something's off. You might not know exactly what. But something's off. Trust that signal.
Red flags are like a physician's diagnostic signs. A patient with three symptoms—persistent cough, unexplained weight loss, fatigue—doesn't necessarily have cancer. But those are signs that warrant investigation. Your doctor doesn't ignore them. Red flags in AI proposals work the same way. They're signals to dig deeper, ask harder questions, or request more work. Experienced leaders can sense a weak proposal like experienced physicians can sense disease.
Let's extend this analogy further. When you're evaluating a new manufacturing process, you'd ask:
- Where was this tested? In a lab? In a pilot facility? In production for two years?
- On what products? The ones we make? Similar products? Very different products?
- What assumptions does it rely on? Specific labor skills? Specific equipment? Specific material quality?
- How sensitive is the gain to those assumptions? If labor quality drops 10%, does the gain drop 5% or 50%?
AI evaluation follows the same logic, but translated into data and model language. A vendor claims their system improves loan approval accuracy by 18%. Ask:
- Tested on what data? Data from your bank? Data from similar banks? General lending data?
- What types of loans? Mortgages? Personal loans? Small business loans? All types?
- What's the baseline accuracy? Compared to what? Manual review? An older system?
- How does accuracy vary by applicant demographic? (This is legally important.)
- How often will the system recommend 'escalate to human'? (This is operationally important.)
- What's the worst-case scenario? If the system is wrong, what happens? Is it reversible?
A 18% improvement that's tested on your data, across your loan types, with demographic parity and clear escalation paths is different from an 18% improvement that's based on academic datasets and hasn't been tested on your applicants. Same accuracy number. Different reality.
What This Looks Like in Real Life
A proposal came to your desk: 'AI system will increase revenue 40%.' How? 'Machine learning will find patterns.' That's not a plan. Red flag. Another proposal: 'We need $5M over 18 months.' What's the business case? 'We'll build capabilities first and figure out applications later.' That's a bet, not a plan. Red flag.
Three proposals arrive on your desk the same week:
Proposal A: 'AI system to optimize inventory. We'll reduce holding costs 15%. Baseline: $40M annual holding cost. Target: $34M. We'll pilot with three warehouses, measure results, expand if successful. Timeline: 6 months for pilot. Success metric: Actual reduction vs. projected reduction.'
Proposal B: 'AI-powered customer analytics. We'll gain insights into customer behavior and improve retention.' (No numbers. No baseline. Vague.)
Proposal C: 'Predictive maintenance system. We'll reduce equipment downtime 30%. Based on ML analysis of historical maintenance data.' (No mention of edge cases, no monitoring plan, unclear ownership.)
Your assessment:
- Proposal A: No major red flags. Clear value, specific metric, phased approach.
- Proposal B: Major red flag—vague value, no baseline, no success definition.
- Proposal C: Red flags—edge case deferral, no monitoring plan, competitive pressure implied.
You approve A for pilot. You send B back and ask for specific metrics. You reject C and ask them to address monitoring and ownership. You now know that spotting these patterns in 10 minutes prevents months of wasted effort.
Let's walk through a fourth example in detail. A logistics company with 800 employees and $400M annual revenue evaluated an AI system to optimize their delivery routes. The vendor showed a case study where a similar company reduced delivery costs by 22%. Impressive claim. Before committing $3.2M to the implementation, the company did a detailed pilot.
The pilot revealed several reality gaps:
First, the 22% in the case study was for the vendor's 'standard' delivery environment: urban delivery, predictable traffic patterns, stable fleet size. The logistics company operated in three environments: urban (30% of volume), suburban (40%), and rural (30%). The vendor's system was highly optimized for urban. On suburban and rural routes, the system's recommendations often created longer drive times because they didn't account for the sparse pickup/delivery pattern. The 22% gain compressed to 6% across all routes.
Second, the case study assumed the system would run on historical data. But the company wanted the system to optimize routes in real-time. Real-time optimization requires the system to know traffic conditions, driver availability, and customer timing constraints as they evolve. The vendor's system was good at 'given these constraints, here's the best route.' It was poor at 'these constraints are changing; adjust now.' Retraining and redevelopment would cost another $800K and take 6 months.
Third, the case study didn't account for driver adoption. Drivers who had been optimizing routes themselves for years didn't trust an AI system's recommendations, especially when those recommendations contradicted their experience. The company needed 4 months of change management, driver training, and iterative adjustments before drivers actually followed the AI's recommendations.
The result: A 6% delivery cost reduction (instead of 22%) took 9 months to implement (instead of the projected 4 months) and required $4M in total investment (instead of $3.2M). The system is valuable. It's working. But the gap between vendor claim and delivered value was substantial. The company now has a realistic view of what the system does. And they know that next time they evaluate AI, they'll pilot on their actual data and conditions, not just trust the case study.
Where People Get This Wrong
Mistake 1: Not knowing what red flags look like. Mistake 2: Ignoring them. 'They seem confident.' Confidence isn't governance. Mistake 3: Not having a fallback if red flags are present. Do you kill the project? Ask for more work?
80% of organizations see red flags but don't act on them systematically. Mistake 1: Seeing a red flag but ignoring it because the proposer 'seems confident.' Confidence isn't evidence. Red flags are evidence. Mistake 2: Treating red flags as disqualifying. They're not. They're signals to do more work. Sometimes proposals with red flags get sent back, reworked, and come back strong. Mistake 3: Not having a process for what happens when red flags are present. Do you kill the project? Send it back? Escalate? Have a clear decision rule.
Let's add three more mistakes that leaders often make:
Mistake six: 'If we implement this AI system, it will fix our underlying data quality problems.' Wrong direction. AI amplifies bad data. If your data quality is poor, an AI system trained on poor data will make poor decisions confidently. You fix data quality first, then add AI. A customer analytics AI system trained on messy customer data will confidently categorize customers incorrectly. It won't suddenly become insightful. Fix the data. Then add AI.
Mistake seven: 'This AI system is a one-time investment. Build it and we're done.' No. AI systems require ongoing maintenance. Models drift over time. New data patterns emerge. New regulations require new constraints. The model you build in month six won't perform the same in month eighteen. Budget for continuous monitoring, retraining, and optimization. Most failed AI initiatives failed because the organization budgeted for implementation but not for operation.
Mistake eight: 'The vendor handles all the risk. If the AI doesn't work, it's their problem.' Legally and operationally, it becomes your problem. Your brand suffers if the AI makes bad recommendations in your name. Your risk exists. You need governance, monitoring, and the ability to turn the system off. Vendors can't take that responsibility away. They can share it. But they can't eliminate it.
Practical Takeaways
When you see a red flag, don't approve the initiative. Either kill it or send it back for more work. What more work? Tell them: fix the red flag. 'We'll handle edge cases in production' isn't good enough. They need to identify and address edge cases upfront. Red flags aren't final verdicts. They're signals to dig deeper.
Red flag awareness will make you the most trusted decision-maker in your organization. Create a checklist of the 12 red flags. When a proposal arrives, spend 15 minutes checking for them. If you see three or more, the project needs major rework before approval. If you see one or two, dig deeper on those specific issues. If you see zero, the proposal is well-thought-out—proceed with normal due diligence. This takes minimal time but dramatically improves decision quality.
Sixth, establish an AI evaluation checklist for your organization. What information do you need before you fund an AI initiative? (Testing on your data? Reference customers? Failure mode analysis? Pilot costs? Change management plan?) Standardize the questions. Everyone uses the same framework. This prevents the situation where one leader asks tough questions and another leader approves the initiative without those answers.
Seventh, after an AI system launches, publish a 'reality report.' Compare vendor claims to actual results. 'Vendor claimed 40% efficiency gain. We achieved 12%. Here's why: [data quality, adoption friction, implementation scope].' This builds organizational learning. It teaches your team to hear vendor claims with appropriate skepticism. And it focuses attention on the real levers that determine success: adoption, data quality, and integration, not just the AI algorithm.
Key Insight
Some proposals have obvious problems visible in five minutes of scrutiny. Learn the red flags and spot them early.
Some proposals have obvious structural problems visible in 15 minutes of scrutiny. Learning to spot red flags early prevents the vast majority of AI initiative failures.
Before You Move On
Get copies of three AI proposals your organization received. Read them. Circle the red flags you now see. That pattern recognition is valuable.
Grab three AI proposals your organization received recently. Read each one and circle every red flag you now recognize. Compare your findings to the outcomes of those projects (if they've progressed). How predictive were the red flags?
Reflect on a recent AI initiative in your organization (or your industry). What were the original projections? What has the actual impact been? What accounts for the gap, if any? Is the gap because of hype, or because of valid reasons like implementation complexity or change management friction?
Now do this: Find one claim you're tempted to believe about AI. It might be 'AI will replace 40% of white-collar jobs by 2027' or 'Our AI system will improve accuracy by 30% with no organizational change needed.' Write down why you believe it. What's your evidence? What could prove you wrong? Run it against this lesson's framework. Is it a bounded claim about a specific technology on specific data? Or is it an oversize claim that sounds good but lacks specifics? This is the habit that separates decision-makers from people who get burned by hype.
Skill.re