Why Buying Beats Building (Usually)
The slide on the screen says "Phase 1: 16 weeks," and the room believes it. The VP of engineering is pitching an in-house AI tool to summarize claims files, and the pitch is genuinely stirring: our data is unique, our workflow is unique, no vendor understands us, and besides, we already built the data platform, so the hard part is done. Heads nod. Eighteen months later the tool handles about 60 percent of claim types, the two engineers who built it have been pulled onto a compliance deadline, the backlog of fixes has a name and a spreadsheet, and a competitor who signed a vendor contract the same quarter has been running summaries in production for a year. Nobody in that first meeting was foolish. They were just betting against a base rate nobody had put on the slide: in MIT's GenAI Divide research, externally purchased and partnered AI solutions succeeded about 67 percent of the time, roughly twice the success rate of internal builds. This lesson is about that number: why the gap exists, when it does not apply to you, and how to turn it into a decision you can make in one meeting instead of one fiscal year.
The Base Rate You Walk In With
In the first lesson of this chapter you met the headline: 95 percent of enterprise GenAI pilots produced no measurable profit-and-loss return, and only about 5 percent of custom-built tools ever crossed from pilot into production. Buried in the same MIT report was a quieter finding with sharper budget consequences. When organizations bought a solution from a vendor or built a partnership around one, those deployments succeeded about 67 percent of the time. Internal builds succeeded at roughly half that rate. Same companies, same models, same year. The difference was not the technology. It was who owned the burden of making the tool survive contact with real work.
Treat that finding the way an underwriter treats an actuarial table. A base rate is not a prophecy about your specific case; it is the price of entry for arguing that your case is different. When a smoker applies for life insurance, the underwriter does not refuse the policy, and does not pretend the mortality table does not exist. The premium moves. That is exactly how a readiness professional should treat "we are going to build our own": not as a forbidden sentence, but as a sentence that arrives carrying roughly a coin-flip penalty against the purchased alternative, and therefore owes the room evidence, not enthusiasm.
This is the burden-of-proof framing from lesson one, applied to the single most expensive decision in most AI programs. The default posture is buy or partner. Building is a position you argue into, with named evidence, against the base rate. Most organizations run the meeting backwards: building is the default because it is flattering, and buying must be argued for by whoever is willing to look unambitious. Reversing that polarity is worth more than most six-figure consulting engagements, and it costs one meeting. By the end of this lesson you will have the agenda for that meeting.
One caution before we go further, because this lesson's title has a load-bearing word in parentheses. "Usually" is doing real work. There are organizations for which building is correct, and we will name them precisely, because a readiness professional who chants "never build" is as useless as the one who chants "never buy." The point of a base rate is not obedience. It is honest pricing.
Why the Room Wants to Build
If external solutions succeed at twice the rate, why does the build instinct keep winning meetings? Because the forces pushing toward "build" are emotional, structural, and completely unrelated to the probability of success. You need to be able to name them out loud, politely, in real time, because every one of them will be in the room.
Pride: building feels like strategy, buying feels like shopping
An internal build produces a thing the organization made. It gets a code name, a demo at the town hall, a line in the CIO's yearly letter. A vendor contract produces a procurement file. No executive was ever profiled in an industry magazine for negotiating good license terms, and everyone in the meeting knows it. This asymmetry in glory is real, and it is worth exactly nothing in the success-rate column.
The "our data is special" reflex
This is the most respectable-sounding argument, so examine it hardest. In most cases the workflow is standard and the data is 90 percent identical in structure to every peer's: invoices are invoices, claims files are claims files, contracts are contracts. The genuinely special 10 percent usually needs configuration, prompt and retrieval tuning, or a vendor's professional-services engagement, not a bespoke product built from the foundation up. The question that deflates this reflex is gentle and factual: "Which specific fields, documents, or steps in this workflow could a configurable vendor product not handle, and have we asked one?" Silence after that question is diagnostic.
The control fantasy
"If we build it, we control it." True, in the way that you fully control a boat you built yourself, in a storm, alone. Control of software means control of its defect list, its model deprecations, its security patches, and its 2 a.m. failures. Vendor dependence is a real cost, and we will price it honestly later, but internal control is not free ownership; it is unshared ownership. Most teams pitching control are picturing the steering wheel, not the bilge pump.
Engineer enthusiasm, honestly priced
Your best developers want to build with AI. That is a genuine organizational asset and a genuine retention issue, and it deserves respect rather than mockery. But it should be priced as what it is, a talent-development and retention benefit, and weighed as that, instead of being laundered into the business case as strategy. There are cheaper ways to give engineers meaningful AI work (internal tooling, evaluation harnesses, integration work on a purchased product, which is harder than it sounds) than betting a production workflow at half the success rate.
The sunk platform
"We already spent two million on the data platform, so building the application layer is marginal cost." This is the sunk-cost fallacy wearing a hard hat. The platform spend is gone regardless of the decision in front of you; the only question that matters is the forward-looking cost and success probability of each path from today. A past investment that makes a new bet feel cheap is how organizations end up doubling down on the wrong side of a base rate.
Notice what all five forces have in common: not one of them is evidence about whether the build will succeed. They are reasons the build would be pleasant, flattering, or narratively satisfying. A steering committee that can tell the difference between an argument and a preference has already left the bottom half of the failure statistics.
The Invisible 80 Percent: What Building Actually Commits You To
Here is the structural reason the gap exists, and once you see it you will see it in every build proposal for the rest of your career. A working AI feature, the thing that appears in the demo, is the tip of the iceberg: visually impressive and maybe 20 percent of the total mass. Below the waterline sits everything that makes software survive in production, and an internal build signs you up for all of it, indefinitely:
- Product iteration. The tool that demos at 80 percent quality has to climb toward reliable, and the climb never ends. A vendor climbs it fueled by bug reports and edge cases from dozens or hundreds of customers. Your internal tool climbs only when your two engineers are free, and they will not be free, because the compliance deadline, the ERP upgrade, and the next executive priority all outrank "make the summarizer slightly better."
- Integration and its maintenance. Connecting to the document store, the case-management system, and the identity provider is a project; keeping those connections alive through every upstream upgrade is a job. Lesson one taught you that tools beside the workflow die; being inside the workflow means plumbing, forever.
- Model updates. The model your build depends on will be deprecated, repriced, or replaced, likely more than once in three years. Each swap means re-running evaluations, retuning prompts, and chasing regressions. Vendors amortize this churn across their whole customer base; your build absorbs it alone.
- Security patching and access control. An AI tool that reads your claims files or contracts is a high-value target with a permissions model somebody must design, audit, and patch. This is unglamorous, mandatory, and invisible in every kickoff deck.
- Support, documentation, and training. When the tool misbehaves at 4:50 p.m. on the last day of the quarter, somebody answers. A vendor has a support queue and an SLA (a service-level agreement, the contractual promise of response times). An internal build has whoever wrote it, if they still work there.
Now the uncomfortable sentence that explains the entire success gap: most enterprises are not software product organizations, and building AI tools is a software product activity. A product organization has standing teams funded year over year, a support function, on-call rotations, release discipline, and a roadmap that survives personnel changes. Most enterprises run project funding: a team assembles, builds, celebrates go-live, and disbands. The tool becomes an orphan on day one of production, which is precisely the day the real work begins. Recall the learning gap from lesson one: pilots died because tools never improved from correction. An internally built tool with no standing team does not merely risk the learning gap; it guarantees it, by construction. Nobody is staffed to close the loop.
BCG's 10-20-70 rule (10 percent of the effort is algorithms, 20 percent is technology and data, 70 percent is people and process) gives you the arithmetic. The build proposal on the table is a plan for the 10 percent, priced as if it were the whole. The invisible 80-plus percent below the waterline does not appear in the kickoff deck, but it appears, with compounding interest, in years two and three.
What a License Fee Actually Buys, and What It Really Costs
Flip the iceberg around and you can see what a purchase price actually purchases. It is not the software. It is a seat on the vendor's iteration loop. Every one of the vendor's customers is running edge cases through the product; every bug one of them files gets fixed for all of them; every workflow wrinkle in your industry that any customer hits becomes a feature you inherit at the next release. When you buy, you are renting a learning loop that is fed by a market instead of by your two overcommitted engineers. That, and not the code, is what the 67 percent success rate is made of.
Honesty requires the other column, because buying has real costs and a readiness professional prices them instead of discovering them at renewal:
- Lock-in. Switching costs grow with every workflow you wire into the product. Before signing, know how your data comes back out: format, completeness, cost, and timeline of export. A vendor who cannot answer the export question crisply is telling you something.
- Data terms. Where does your data go, who can see it, and does the vendor train models on it? These are contract clauses, not vibes. Many vendors now offer private deployments and no-training commitments; the clause you did not ask for is the one you do not get.
- Roadmap dependence. The feature your process desperately needs is the vendor's ticket number 4,317. You influence a shared roadmap; you do not command it. If a capability is existential for you and marginal for their other customers, that is a signal pointing toward partner or build.
- Renewal economics. Year-one pricing is a courtship. Model the renewal, cap escalation contractually where you can, and never let the contract term outrun your kill criteria.
And treat vendor claims with the same discipline this program applies to every number: a vendor's accuracy figure is a benchmark to verify on your documents, never a guarantee. The market you are buying from includes real products and costumes. Gartner, surveying the agentic AI vendor landscape in mid-2025, judged that of the thousands of vendors claiming agentic capabilities, only around 130 were real at the time, a phenomenon it politely called "agent washing." Buying beats building on the base rate, but buying badly is still buying a failure, just faster and with better catering at the sales dinner.
The wrapper trap
One buying failure deserves its own warning label. Some products are a thin wrapper: a prompt, a logo, and a subscription price layered over a general-purpose model your organization could access directly for a fraction of the cost. The test is one question in the demo: "What does this product do that a well-written prompt against a frontier model does not?" Legitimate answers exist and are specific: deep integration with your systems of record, a maintained evaluation suite for your document types, permissioning and audit trails, workflow routing, retrieval over your data done properly. If the answer is a chat window with your industry's vocabulary sprinkled in, you are not buying an iteration loop; you are buying markup. The wrapper trap matters because it is the failure mode that turns a sensible "buy" posture back into cynicism about buying, and hands the next meeting to the build faction for the wrong reason.
The Partner Path, and the Honest Case for Building
The middle door most meetings forget
The build-versus-buy framing hides a third door, and MIT's data suggests it is the best one: partnership. A partnered deployment means a vendor's product, co-developed and customized on your specific workflow: their engineers and yours redesigning the process together, their product absorbing what it learns from your edge cases, your team keeping ownership of the process change and the measurement. In the MIT findings, externally partnered solutions were the standout performers; the roughly two-to-one success advantage belongs to this externally sourced family of approaches, and the most deeply partnered versions of it performed best. The logic is simple once you see the iceberg: partnership buys the vendor's iteration loop and adds your workflow specificity on top, which is exactly the combination the 5 percent of successful pilots had. It costs more than a plain license, in money and in your team's time, and it earns that cost back in the only currency that matters: probability of surviving contact with production.
Partnering is also where the 70 in BCG's 10-20-70 gets done. A pure purchase can tempt an organization into believing the process will adapt itself to the tool. A real partnership forces the workflow conversation, because the co-development sessions are the workflow conversation.
When building genuinely is the right call
Now the promised honesty. There are four conditions under which the base rate bends toward building, and you should be able to recite them:
- The workflow is your competitive moat. If the process the AI touches is the thing customers actually pay you a premium for (a lender's underwriting judgment, a logistics firm's routing engine, a fund's research process), handing its brain to a vendor who sells to your competitors is strategically incoherent. Build the moat. Buy everything around the moat.
- No credible vendor category exists. If your genuine need has no market serving it, buying is not on the menu. Verify this honestly: "no vendor" usually means "we searched for an afternoon." A real scan of the category takes days, not hours, and includes asking peers and analysts.
- You are actually a software product organization. Some enterprises truly have standing product teams, support functions, on-call rotations, and multi-year funding: large banks, tech-forward retailers, software companies themselves. If the below-the-waterline commitments in the previous section describe capabilities you already run, the base rate was never really about you.
- Your data genuinely cannot leave. Rarer than claimed, now that private deployments and contractual no-training terms exist, but real in some regulated, classified, or treaty-bound contexts. Exhaust the deployment-model options before invoking this one.
One of these alone is a conversation. Two or more, documented, is a defensible build case. And even then, the decision inherits every evidence discipline from lesson one: a written baseline, pre-committed success and kill criteria, a named owner, and a funded plan for the invisible 80 percent. A justified build is still a bet against the base rate; the justification buys you the right to make the bet, not an exemption from measuring it.
Buy the commodity. Partner on the workflow. Build only the moat, and only if you can staff it like a product, not a project.
The Artifact: The One-Meeting Build/Buy/Partner Filter
Here is this lesson's deliverable for your readiness portfolio: six questions that convert the base rate into a decision, in one meeting, on one page. Ask them in order, write the answers down with one line of evidence each, and let the pattern of answers point at a door. No single question decides alone; the filter's output is a default plus a burden of proof.
- Is there a mature vendor category? Concretely: can you name at least three credible vendors with reference customers in your industry for this specific workflow? If yes, buying is viable and the meeting continues with buy as the default. If no, the realistic doors are partner (with the nearest-fit vendor) or build. Mini example: invoice data extraction has dozens of credible vendors; automated triage of reinsurance treaty disputes may have none. Do the scan properly before answering.
- Is this workflow a differentiator or a commodity? The test question: if we performed this process 30 percent better than every competitor, would customers notice, switch, or pay more? For expense processing, HR answers, and meeting summaries the answer is no: commodity, buy it. For the process your margin actually lives in, the answer may be yes, and the arrow swings toward partner or build. Be brutal here; every department believes its workflow is special, and lesson one taught you what beliefs without evidence are worth.
- Do we have a standing product team with support capacity? Not "do we employ developers," but: is there a permanently funded team that will own this tool in year three, with a support channel and someone on call? If the honest answer is "we have project developers who will move on after go-live," building is off the table no matter how the other questions land. This single question disqualifies most build proposals, which is exactly its job.
- Can our data leave the building under acceptable terms? Review the actual terms: processing location, training rights, retention, deletion, private-deployment options. Usually the answer, after real diligence, is yes with conditions, which keeps buy and partner open. A genuine, documented no (regulatory, contractual, or classification-based) is one of the few findings that legitimately forces the build door.
- What is the three-year total cost of each path? Build costs must include the iceberg: salaries, infrastructure, integration upkeep, model-update churn, support time. A useful rule of thumb: maintaining custom software over its life costs on the order of the original build again, so a build estimate that shows year two and three near zero is fiction, and you should say so. Buy costs include license, integration, administration, and a renewal escalation assumption. Partner costs sit between, plus your team's co-development hours. Compare at three years, never at year one, because building front-loads the pride and back-loads the bill.
- What is our exit if this choice fails? For buy: contract term, data-export clause, and a sketched switching plan. For build: written kill criteria with dates, and a stated plan for the team and the code if the criteria trigger. For partner: who owns the customizations and what survives a separation. If nobody can describe the exit, the organization is not ready to make the entrance; this is lesson one's kill-criteria discipline wearing procurement clothes.
Read the answers as a pattern. Mature category plus commodity workflow plus no product team: buy, and spend your energy on integration and workflow redesign. Differentiating workflow plus a near-fit vendor: partner, and negotiate co-development explicitly. Moat workflow plus real product organization plus data constraints: build, with the full evidence discipline attached. The filter does not think for you. It makes the room think in the right order, with the base rate on the table instead of under it.
A Claims Desk Runs the Filter: A Worked Example
Let us run the whole lesson through one concrete decision. The company and its numbers are hypothetical, built to be realistic rather than reported from a real firm; use them as a template, not a citation.
Harbor Mutual is a mid-market property insurer: 600 employees, 90 claims adjusters. On complex claims, an adjuster spends about 45 minutes assembling a summary of the file (police reports, photos, correspondence, policy terms) before any decision work begins. Roughly 550 complex claims arrive per week across the team. The VP of engineering proposes building a claims-summary tool: two engineers, 14 months to production, then ongoing maintenance. Fully loaded, that is about $700,000 over three years: two salaries for the build period, then a realistic maintenance tail of integration upkeep, model swaps, and fixes, which the first draft of the proposal had, of course, budgeted at approximately zero. The alternative: a claims-document AI vendor at $60,000 per year, $180,000 over three years, plus about $45,000 of integration work. A third option surfaces during diligence: the same vendor offers a co-development tier at $95,000 per year, where their team configures the product on Harbor's actual claim types and meets monthly to tune it, roughly $330,000 over three years including Harbor's staff time.
The filter, in one meeting. Question one: mature category? Yes, the scan finds four credible vendors with insurance reference customers. Buy is viable. Question two: differentiator or commodity? The room wants to say differentiator; the test question deflates it. Harbor wins business on underwriting appetite and claims decisions, not on how fast summaries get typed. The summary is an input to the moat, not the moat. Commodity. Question three: standing product team? Harbor has twelve developers, all project-funded, currently mid-flight on a policy-administration migration. Honest answer: no. This alone ends the build case, and the CFO visibly relaxes. Question four: can data leave? After legal review: yes, under a private-cloud deployment with a contractual no-training clause and EU data residency for the European book. Question five: three-year cost: roughly $700,000 to build versus $225,000 to buy versus $330,000 to partner. Question six: exit? The vendor agrees to a one-year initial term and a documented export format; kill criteria are written before signature: if adjuster time per complex claim has not fallen by at least 20 percent against the baseline within six months of go-live, Harbor exits at term.
Harbor chooses the partner tier, not the plain license, for a reason that should now be predictable: the co-development sessions are where the workflow gets redesigned, and lesson one taught what happens to tools bolted onto unchanged work. They capture a baseline before go-live (45 minutes average, measured across four weeks, not remembered), name the claims operations manager as the single owner of the number, and launch.
The three-year retrospective, still hypothetical, still realistic. Summary time on complex claims falls to about 20 minutes: 25 minutes saved, times roughly 550 complex claims a week, is about 230 adjuster hours weekly, close to six full-time equivalents of capacity redeployed into faster cycle times and an overtime line that stops growing. Against roughly $330,000 spent, the value is provable within the first two quarters because the baseline existed. The honest costs also arrive on schedule: the renewal price rises 11 percent in year two, one feature Harbor wants sits on the vendor roadmap for 14 months, and the export copy of claims summaries is maintained quarterly because the exit plan says so. Meanwhile the counterfactual build, per its own optimistic plan, would have reached production around month 14 at the earliest, trained on one company's edge cases, maintained by engineers who were reassigned twice. The vendor shipped 31 releases in the same period, fed by every insurer on the platform. Harbor got the iteration loop it could never have staffed. That, in one story, is the entire lesson.
What to Do Monday Morning
This lesson turns into practice the first time the filter meets a live decision. The sequence:
- Inventory the AI initiatives you can see and label each one build, buy, or partner. Most organizations have never looked at the portfolio through this single lens, and the distribution is often a surprise worth showing to a steering committee.
- Pick one live or pending build proposal and book the one meeting. Put the six filter questions on the agenda, in order, and require one line of written evidence per answer. Do not let the meeting start with the demo.
- Demand the three-year picture for every path, and challenge any build estimate whose maintenance line is near zero. Say the rule of thumb out loud: lifetime maintenance tends to rival the original build.
- Run the wrapper test on one purchased or shortlisted tool. Ask what it does beyond a well-written prompt against a frontier model, and require a specific answer: integration, evaluations, permissions, audit trail.
- Check one existing vendor contract for the exit. Find the data-export clause, the term, and the renewal escalation. If you cannot find them, that is this week's finding.
- File the completed filter as the second artifact in your readiness portfolio, beside lesson one's failure-mode checklist. The capstone will ask for both.
Key Takeaways
- Anchor every build-versus-buy conversation to the base rate: MIT found externally purchased and partnered AI solutions succeeded about 67 percent of the time, roughly twice the rate of internal builds, so the burden of proof sits with whoever proposes building.
- Name the five build-instinct forces in the room (pride, "our data is special," the control fantasy, engineer enthusiasm, sunk platform investments) and note that none of them is evidence about success probability.
- Price the invisible 80 percent before approving any build: product iteration, integration upkeep, model updates, security patching, and support, sustained by a standing team that most enterprises, which are not software product organizations, do not have.
- Understand what a license actually buys: a seat on the vendor's iteration loop across many customers, which is the structural reason bought solutions climb toward reliability while single-customer builds stall.
- Negotiate the real costs of buying up front: lock-in and data export, training rights and residency terms, roadmap dependence, and renewal escalation, and verify every vendor benchmark on your own documents.
- Prefer the partner path when the workflow matters: co-development on your process combines the vendor's learning loop with your specificity, and the most deeply partnered external deployments performed best in MIT's findings.
- Grant the build case only on evidence: a moat workflow, a genuinely empty vendor category, a real product engineering organization, or data that truly cannot leave, and even then attach baselines, kill criteria, and a named owner.
- Run the six-question one-meeting filter (vendor category, differentiator or commodity, standing product team, data terms, three-year cost, exit plan) before any door is chosen, and file the answers as a portfolio artifact.
Skill.re