←
AI Readiness & Process Transformation
Strategic · M5 · lesson 5 of 25 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
📖
in this lesson

Build, Buy, or Partner at Portfolio Scale

15 min

It is the third time this quarter, and everybody in the room recognises the shape of it before the sentence finishes. The wave-two planning session is forty minutes old, the invoice-exception item is on the screen with its baseline and its business case, and the head of enterprise architecture leans forward and says the words: "Before we go to market, I think we should look at building this ourselves." What follows is not a decision. What follows is three weeks: vendor demos scheduled to prove a point, a capability memo, a counter-memo, a skip-level conversation, a spreadsheet comparing a licence fee against four engineers' salaries as though those were the only two numbers involved, and a resolution that has less to do with evidence than with who had more political capital that month. You have now spent nine weeks of portfolio time this quarter on an argument that will recur, in identical form, on every remaining item. The problem is not that anyone in the room is wrong. The problem is that your organisation is deciding sourcing retail, one shouting match at a time, when sourcing is a policy question you were supposed to answer once.

The 2x Finding, and the Two Ways Everyone Misreads It

One finding in the enterprise AI failure record is more directly actionable than any other. MIT's GenAI Divide study, the August 2025 research behind the famous 95 percent (95 percent of enterprise generative AI pilots delivered no measurable profit-and-loss return) and the quieter companion figure (roughly 5 percent of custom-built AI tools ever crossed from pilot into production), also found this: externally partnered and purchased solutions succeeded roughly twice as often as internally built ones.

Twice. A doubling of the success rate, attributable not to the model or the use case but to the sourcing decision, which your organisation makes with a memo and a signature in under a month. Most levers this program teaches are slow: process redesign takes quarters, data remediation takes years, cultural change takes a leadership generation. Sourcing is fast, cheap, reversible before signature, and worth a doubling. It is the highest-leverage sentence in the failure record, which is exactly why it gets misread, in both directions, usually by two people in the same meeting.

Misreading one: "so we never build anything"

This is the strategist's misreading, seductive because it sounds like discipline. Armed with a 2x, the newly evidence-driven leader announces a blanket policy: we buy, we do not build, end of discussion. It fails expensively. It surrenders the handful of places where building is genuinely correct, which tend to be the places where competitive advantage lives, and it destroys the internal engineering capability that makes you a competent buyer and a credible partner. An organisation that cannot build anything cannot own anything, and blanket-buy organisations end up unable to evaluate what they are being sold, which is how the agent-washing market eats them.

Misreading two: "that statistic is about bad builders, not us"

This is the internal team's misreading, durable because it is unfalsifiable in advance: those failed builds were done by mediocre teams, our team is good, therefore the base rate does not apply. Sometimes the team really is good. That is the trap. The mechanisms below are structural rather than talent-driven, and they punish excellent teams almost as hard as mediocre ones, because they operate in the years after the build, when the excellent team has moved on. Every internal team in the failure record explained the statistic away with this argument before joining it.

The strategist's job is neither misreading. It is to convert a research finding into a sourcing policy: a decision rule set once, at portfolio level, category by category, with the build exceptions named in advance and defended by written criteria rather than by whoever is most enthusiastic in the room. Policy beats per-project debate for the same reason a selection scorecard beats lobbying: a rule written before you know who it will favour is a fair rule, and rules that exist end arguments that would otherwise recur eleven times in an eleven-item portfolio.

Decide sourcing once, by category, before you know whose proposal it kills. Anything you decide per project, you will decide by politics.

The artifact you finish this lesson holding is the Sourcing Policy Matrix: capability categories down the side, the default verdict across, the build-exception criteria written out, and a review cadence so the policy ages honestly. It fits on one page, it ends the three-week argument, and it does something subtler: it moves the conversation from "who wins" to "which category is this," a question an architect and a strategist can answer together without either losing face.

Why Internal Builds Underperform, Mechanically and Without Contempt

Before the policy, the mechanism. You need this not to win arguments but because you may be the internal builder's sponsor, and a policy you cannot explain is a policy you cannot enforce. Four mechanisms sit behind MIT's 2x, and none of them is "your engineers are not good enough."

Mechanism 1: the maintenance tail nobody budgets

Ask a team what the build costs and they tell you what the build costs: four engineers, nine months, a number. That number is not the cost; it is the deposit. The cost of an internal AI capability is the five years after launch: model upgrades when the provider deprecates what you built on, dependency upgrades, security patching, retraining as data drifts, the evaluation harness somebody keeps running, and the most expensive item of all, which appears in no business case: the departure of the two people who understood how it worked.

Operators know this shape: it is Total Cost of Ownership (TCO), counting an asset's whole life rather than its purchase price, applied rigorously to forklifts and carelessly to software. In AI it bites harder because the substrate moves underneath you, and that gap is an engineering budget nobody approved.

Now the asymmetry that produces the 2x: the vendor pays this tail too and amortises it across every customer. Their model-upgrade project costs them once and is funded by four hundred contracts. Yours costs you once and is funded by you. That is arithmetic, not a comment on competence, and arithmetic does not care how good your team is.

Mechanism 2: the feature-velocity gap

A category vendor's entire company exists to improve one product, shipping against thousands of customers who all complain about different things, which populates their roadmap with real-world failure modes at a rate no single enterprise can generate. They ship weekly. Your internal build ships when it wins a prioritisation meeting against a regulatory change, a system migration, and an outage: quarterly at best, annually in practice.

Now compound that. A good internal build often matches the market at launch, and that is the dangerous moment, because parity reads as vindication. Eighteen months later the vendor has shipped seventy improvements and you have shipped six. That is not a gap in quality of people. It is a gap in number of people multiplied by time, and it widens every quarter automatically.

Mechanism 3: the domain-versus-tooling confusion

This is the intellectual heart of the failure, and the argument that wins internally almost every time it is made: "our domain is unique, off-the-shelf tools are built for the generic case, therefore we must build."

The first clause is usually true: your claims-handling process, product taxonomy, regulatory position, and exception patterns genuinely are unlike anyone else's. The conclusion does not follow, because domain uniqueness almost never lives in the software layer. It lives in configuration (the rules, thresholds, taxonomies, and routing logic you set), in data (your history, your labels, your corrections), and in workflow (the sequence of steps, handoffs, and controls around the tool). Category software is built to be configured precisely because every buyer says their domain is unique and vendors learned, expensively, that they all mean configuration.

So put the question on the table in exactly these words: what specifically is unique here: the workflow, the data, or the software? Make the team answer for each separately. The workflow is unique perhaps four times in five, the data perhaps one in three, the software almost never. When the answer is "the software," ask what it does that a configured product cannot, and listen for whether the answer describes an algorithm or a preference.

Mechanism 4: the evidence disadvantage

The fourth mechanism is accumulated scar tissue. A category vendor arrives carrying the failure modes of every customer they have served. They already know that this document type breaks extraction, that this integration pattern deadlocks at volume, that this confidence threshold produces a silent error class surfacing in month four, that this user population abandons the tool unless a correction takes fewer than two clicks. None of that is in their marketing; all of it is in their product. Your internal build discovers each of those failure modes retail, one production incident at a time, paid for in your own operational disruption and your users' patience. Buying is, among other things, a way of buying other people's autopsies.

Put the four together and the 2x stops being surprising. It is not a judgment on internal engineers; it is a structural statement about amortised maintenance, shipping cadence, misdiagnosed uniqueness, and access to other people's scars. Say it that way when you present the policy, because the version that says "our people are not good enough" is both false and unusable.

The Sourcing Policy Matrix: Three Modes and Five Categories

Three modes, defined precisely (they are not a spectrum)

Most organisations treat build, buy, and partner as three points on a dial from least control to most. That framing is why partner gets treated as a compromise, which is why the highest-performing mode is the least used. They are three distinct operating models with different cost structures and different failure modes.

BUY means you license a product, accept the vendor's roadmap, and customise through configuration. You get speed (weeks to value, not quarters), the vendor's accumulated evidence, and their maintenance amortisation. You give up roadmap control and accept lock-in, which is not a dirty word but is a price: the correct response to lock-in is to price it and cap it, not to pretend to avoid it.

PARTNER means a vendor product plus a substantive co-development or heavy-configuration engagement: your data, your workflow, and usually your process redesign embedded deeply into their platform, with negotiated roadmap influence and often joint ownership of what gets built at the edges. This is MIT's advantaged category, and it is worth being precise about why, because "partner is best" as a slogan is useless.

Partnership works because it imports the vendor's failure-mode knowledge while preserving your process specificity. The build lacks the first. The naive buy sacrifices the second: the tool is generic, the workflow was never redesigned around it, and you land in the adoption-without-transformation pattern the entire failure record is built from. Partnership is the only mode that structurally addresses both, and it is the one that puts a workflow-redesign engagement inside the contract, which matters because workflow redesign is among the strongest measured drivers of impact. It costs more in coordination (governance meetings, joint roadmaps, a named relationship owner, slower decisions), and that coordination cost is the price of the 2x.

BUILD means internal ownership of the software itself: full control, full tail, full velocity gap. It is justified rarely, under conditions we will now write down.

The five categories and their default verdicts

Here is the spine of the artifact. Every AI capability your portfolio will contain falls into one of five categories, and each gets a default verdict set once, at portfolio level, before anyone knows which project it will hit.

CategoryWhat it looks likeDefault verdictWhy
Horizontal capabilityDocument extraction, transcription, summarisation, translation, generic chat assistance, meeting notesBUY. No exceptions worth havingCommodity categories with brutal vendor competition and per-unit pricing that falls yearly. Building one is a hobby with a budget code.
Workflow-embedded AIThe invoice-exception case: a generic AI capability sitting inside a process that is genuinely yoursBUY the capability, BUILD the workflow around itThe dominant enterprise pattern and the most frequently mis-sourced. The uniqueness is real but it is in the workflow, not the model.
Domain-specific with a real marketIndustry tools: claims triage, clinical coding, trade surveillance, underwriting supportBUY or PARTNER, with hard due diligenceReal vendors exist, and so do imitators: Gartner's agent-washing finding (of thousands of vendors claiming agentic capability, only around 130 judged real) says verify the product, not the deck.
Proprietary-advantage capabilityThe capability itself is the competitive moat and no vendor serves the categoryBUILD, if and only if all four conditions passThe narrow, real build case. Rare by construction: if a vendor could serve it profitably, one already would.
Glue and integrationConnectors, data pipelines, orchestration, the human-review interfaces, the audit loggingInternal, alwaysNot really a sourcing decision: connective tissue between your systems is yours by definition. Name it glue so it stops counting as "we built our AI."

Two notes the table cannot carry. On horizontal capability: a competent team gets a demo running in a fortnight, which is exactly the trap, because the production system with error handling, throughput, monitoring, and a five-year upgrade path is two years, at the end of which you own a slightly worse version of something whose price is falling. Every exception you grant in this row will be granted on the strength of a demo. On glue and integration: classify it explicitly, or it gets counted as "we built our own AI," inflating your apparent build footprint and legitimising the next build by precedent, and it gets left out of business cases, because glue is invisible until it breaks.

Workflow-embedded AI is the row your portfolio is mostly made of, and its split verdict is the whole lesson in one line. The invoice-exception item is the canonical shape: the exception-handling process (who reviews what, in what order, at what thresholds, escalating to whom, evidenced how) is genuinely specific to your finance function; the underlying capability (read a document, classify a variance, propose a resolution) is a commodity with fourteen credible vendors. Teams look at the composite, correctly observe that it is unique, and build all of it. Then they own a commodity extraction engine forever as the price of owning a workflow they could have owned for free. Split the verdict, and fund the build half properly: workflow design, integration, human-review interfaces, controls, and the learning loop.

The four-condition build test

The proprietary-advantage row needs a gate: four conditions that must all hold. Not three. Three-of-four is how every failed build passed its own review.

  1. The capability is the competitive advantage, not merely important to it. Ask: if a competitor bought this capability off the shelf tomorrow, what would we have lost? If the honest answer is "some efficiency," this is not condition 1. Condition 1 means customers choose you partly because of this thing.
  2. The data is genuinely unique and proprietary. Not large, not sensitive: unique. Nobody else has it, it cannot be reconstructed from public or purchasable sources, and it materially improves the capability. Sensitive-but-generic data is a hosting question, and hosting is negotiable in a contract.
  3. No vendor serves the category, and none plausibly will inside your horizon. This condition requires actual market scanning, not an afternoon of assumptions, and it must be re-tested annually because it is the condition that expires. Note the economics: if a vendor could serve this profitably, one usually already does.
  4. Sustained engineering capacity exists, funded and staffed beyond the build. Named team, budgeted maintenance for five years, a succession answer for the two people who will understand it, an evaluation harness somebody owns. "We have four great engineers available this year" is not condition 4; it is the opposite of it.

Run this gate honestly and it passes rarely. In a portfolio of eleven items, passing twice is normal and passing zero times is common and fine. If it passes five times the gate is being run generously, and the generous condition is usually number 3, followed by number 4.

Four Portfolio-Level Disciplines the Policy Must Carry

A policy that only answers "build or buy" is half a policy. The other half is what happens once you are buying at portfolio scale, which is a different game from buying once.

The concentration ledger

Sourcing has a portfolio dimension no single decision can see. Every decision to put one more item on the vendor you already trust is locally correct: integration is cheaper, the relationship exists, the team knows the platform, the discount improves. Five locally correct decisions produce one globally uncomfortable fact: a single vendor runs five of your processes, and the word for that at renewal time is not "efficiency."

The instrument is a concentration ledger: one row per vendor listing the portfolio items they carry, the annual spend, the processes affected, and the entry that does the actual work, an estimated exit cost in money and weeks (data migration, re-integration, reconfiguration, retraining, parallel running). You will not know it precisely. Estimate it anyway, to the nearest ten thousand and the nearest fortnight, and write down your assumptions, because a soft estimate still converts vague unease into a number a steering committee can hold a position on.

The discipline here is timing: price exit before signing the next wave's contracts, not during the renewal that surprises you. Concentration risk gets a full treatment later in the program, where you manage a vendor estate rather than a portfolio. At this level the move is simply: know the number, set a tolerance, review it at the wave boundary.

The exit-clause standard, set once as policy

Terms negotiated per deal get negotiated by whoever is closing that deal, under deadline pressure, against a vendor whose lawyers do this weekly. Terms set once as policy are negotiated by your institution instead. Four clauses belong in the standard.

  • Data export in usable formats. Not "data will be made available." Named formats, defined completeness (including derived fields, labels, and metadata), a maximum turnaround, and a right to test the export while the contract is live. An export you have never tested is a promise, not a control.
  • Model-change notification. Advance written notice before the underlying model or its behaviour changes materially, with a defined notice period and the right to re-run your acceptance tests. This clause exists because of one recurring incident: the vendor silently updates a model, output distribution shifts overnight, your accuracy monitoring shows a step change with no corresponding change on your side, and you spend nine days proving the problem is not yours. Anyone who has lived that fortnight writes this clause into every contract afterwards. That is how mature governance accumulates: one scar, one clause.
  • Notice periods and continuity. Minimum notice for termination, price change, or product sunset, plus a transition-assistance obligation. Vendors get acquired and product lines get discontinued; the clause is not about distrust, it is about calendar.
  • Ownership of your corrections and configuration. The question most contracts leave silent and the one that most determines exit cost. Who owns the labelled corrections your reviewers produced over eighteen months, your rules, thresholds, taxonomies, and routing logic, and can you extract them in a form that means anything elsewhere? A contract that quietly assigns them to the vendor has converted your operating investment into their moat.

These four elevate the control questions you already ask of any AI system into procurement standard: at project level a control specification, at portfolio level a clause library procurement applies without you in the room.

Pilot before portfolio

A vendor enters your portfolio by earning it through one gated wave item with real baselines and pre-committed success criteria, not by signing a master agreement on the strength of a demo and an enterprise discount. The pattern to refuse is familiar: an impressive demonstration, a three-year agreement covering six capabilities at an attractive rate, then the discovery that capability two was a roadmap item, capability four does not integrate, and capability six is a partner's product resold.

This is the enterprise-scale version of the skepticism you developed the first time a demo did something in eleven seconds that your process could not do in eleven days. As a rule: prove it on one item, with measurement, before it becomes a platform decision. Volume discounts are real and worth having, and the safe way to have them is a pre-negotiated expansion price contingent on the pilot item clearing its gate.

The annual review, and the graceful retirement

The policy has a shelf life, because the market moves faster than your governance calendar. Last year's legitimate build exception is this year's commodity category with nine vendors and falling prices. Set an annual review date, name its owner, and give it a fixed agenda: re-test condition 3 for every internal build you own, re-check the concentration ledger, refresh the exit-cost estimates, re-classify any category the market has changed.

The hard part is human, not analytical. Retiring an internal build a proud team spent a year on is politically expensive and financially obvious, and the gap between those two facts is where zombie systems live. Do not open with the migration; open with the market: "the category has four credible vendors now that did not exist when we started, and condition 3 no longer holds." Separate the judgment on the build from the judgment on the builders out loud, because in a well-run review the build was right when it was made and is wrong now, and those are different sentences. Then give the team the next thing: the migration, the integration layer, the evaluation harness, the configuration ownership. A retirement that redeploys the team reads as a portfolio decision; one that strands them reads as a verdict, and the next team to hear "we should build this" will fight you twice as hard.

The Policy in Practice: Eleven Items, and One Platform That Should Never Have Been Built

Take the enterprise from the running portfolio storyline: eleven items across three waves, an information technology (IT) function with genuine build capability and appetite for using it, and an institutional memory that includes a scrapped year and one abandoned internal build nobody discusses at length. All figures below are illustrative.

The matrix applied

Running the eleven items through the five categories takes about ninety minutes with the right four people in the room, which compares well with nine weeks of arguing.

  • Seven BUY. Four horizontal capabilities (document extraction, transcription for the service centre, summarisation in two functions) and three workflow-embedded items where the capability is bought and the workflow is designed internally, invoice exception handling among them.
  • Three PARTNER. The deep bet, with a discovery-and-redesign engagement because the redesign is the value and the vendor has done it forty times; the agentic item, where the control specification requires co-design and the agent-washing caution makes diligence heavy; and the retrieval-grounding infrastructure, where your content and taxonomy sit deep inside someone else's platform.
  • One BUILD. The integration and orchestration layer, reclassified honestly as glue. It is real engineering with a real budget and a named owner, and calling it what it is stops it being cited next year as proof that the organisation builds AI.

What happens when IT arrives with three build proposals

The IT function arrives at the sourcing review proposing to build three of the seven BUY items. This is the moment the policy either works or becomes another memo, and the difference is whether you use the four-condition test as an instrument or as a veto.

Proposal one is document extraction, argued on the grounds that the organisation's documents are unusual. It fails condition 1 in four minutes (nobody chooses this company because of how it reads documents) and fails condition 3 outright (eleven credible vendors exist). It is withdrawn by its own author, which is the outcome you want, because a proposal withdrawn by its author creates no grievance.

Proposal two, a summarisation and knowledge-assistant layer, fails condition 1 for the same reason and condition 4 on inspection: the team is available this year because a large program just ended, and no maintenance line is funded for years two through five. Naming condition 4 reframes the refusal from "we do not trust you to build it" to "we have not funded anyone to own it in 2031," a statement about the organisation rather than a judgment on the engineers.

Proposal three is the interesting one: a workflow-embedded item where the process really is distinctive and the team really does have relevant capability. It fails condition 3, because vendors exist, but the team's objection to those vendors is substantive rather than defensive: none handles the process shape without significant co-development. That is the definition of a PARTNER engagement, so it becomes one. A vendor is selected, the internal team co-develops the process-specific layer inside the vendor's platform, and IT keeps the integration layer and the evaluation harness permanently. The ambition is redirected rather than humiliated, and the head of enterprise architecture leaves with more scope than they came in with.

That is the policy's real job. It exists not to stop people building but to route internal ambition toward the places where internal ownership actually pays, using criteria rather than authority, because authority has to be spent again next quarter and criteria do not.

The concentration ledger, first pass

With sourcing decided, the ledger gets its first row. Vendor A, having won two wave-one items and then two more on the strength of integration convenience, now carries four of the eleven. Illustrative annual spend: 210,000 dollars. Processes touched: four, across two functions. Estimated exit cost: roughly 90,000 dollars and ten weeks, assuming the contracted data export works, plus re-integration of two systems, reconfiguration, and six weeks of parallel running.

The steering committee's tolerance: no vendor above five items or above an estimated 150,000 dollar exit cost without explicit board-level acceptance. Vendor A sits inside both and is flagged for the wave-three boundary, where two more items are candidates for the same platform. Nobody is alarmed, and everybody knows the number six months before it would have surfaced inside a renewal negotiation. That is the ledger's value: not preventing concentration, which is often right, but pricing it while you still have options.

The failure story: the platform that was genuinely good

A mid-size insurer, four years earlier. The data team is talented and correct about the thing they are correct about: the company's documents really are unusual, a tangle of legacy policy formats, broker submissions in a dozen house styles, and handwritten annotations no generic extraction product of the era handled well. They propose an internal document-processing platform: fourteen months, four engineers, roughly 1.6 million dollars fully loaded, against the licence costs it would displace. This is a good team making a defensible argument.

They build it, and here is what makes the story worth telling: it works. At launch it benchmarks at parity with the commodity vendors it was measured against, and on the insurer's gnarliest document classes it beats them. An internal case study is written. For about seven months this is a success story.

Then the mechanisms arrive, in order and on schedule. Vendors ship quarterly; the platform ships twice in year two, because the team was partly reassigned to a core-systems migration that could not wait. By month twenty it is behind on the document classes it used to win. In month twenty-six the lead engineer leaves for a company that will let her build things full time, and four months later the engineer who understood the retraining pipeline follows her. The maintainers who remain can keep it running but cannot meaningfully improve it, a distinction that never appears on a status report.

By year three the platform is owned by nobody in particular, and two business units quietly buy software-as-a-service (SaaS) products for their own document flows, because the internal backlog is eleven months long and their auditors are not patient. Learn to read that tell: when business units buy around your internal build, the market has passed it, whatever the roadmap says. The platform is formally retired in year four at a migration cost of about 300,000 dollars, on top of the 1.6 million to build it and roughly 900,000 in maintenance in between.

The postmortem contains one line worth the whole exercise: the documents were unique; the software never was. Condition 2 passed. Conditions 1, 3, and 4 failed, and condition 3 failed most decisively, because even in year one there were vendors who could have been pushed to handle those document classes as a partnered engagement for a fraction of 1.6 million. The four-condition test would have caught this in an afternoon, four years and 2.8 million dollars before the retirement meeting, not because the test is clever but because it forces four questions while the answers are still cheap.

Hold this next to the wider record: MIT's 95 percent of pilots producing no measurable return and roughly 5 percent of custom tools reaching production, S&P Global's 42 percent of companies scrapping most of their AI initiatives in a single year, McKinsey's 88 percent of organisations using AI against about 39 percent able to attribute any earnings impact. The insurer's platform is not an outlier there. It is a well-executed, entirely typical instance of the most common expensive mistake in the record, and what would have prevented it costs one page and one honest afternoon a year.

What to Do Monday Morning

The Sourcing Policy Matrix is a one-page artifact and a two-hour first draft. Here is the sequence that puts it into force before the next argument starts.

  1. Draft the five-category matrix with default verdicts. Categories down the side (horizontal, workflow-embedded, domain-specific with a market, proprietary-advantage, glue), verdict across, one sentence of reasoning per row. Do it before you look at your live portfolio, so the rules exist before you know whom they favour. Two hours, one page.
  2. Apply the four-condition build test to every live internal build proposal. All four conditions, in writing, with an evidenced yes or no on each, run with the proposing team and never at them. Expect conditions 3 and 4 to do most of the work, and expect at least one proposal to convert into a partner engagement rather than die.
  3. Start the concentration ledger. One row per AI vendor in your estate: items carried, annual spend, processes touched, estimated exit cost in dollars and weeks with assumptions written down. Set a tolerance with your steering committee and put the review at your next wave boundary, not at renewal.
  4. Send procurement the four-clause exit standard. Data export in tested usable formats, model-change notification with a notice period and re-test rights, notice and continuity terms, and explicit ownership of your corrections and configuration. Ask that it be attached to every AI contract from now on, including the one currently in redlines.
  5. Book the annual policy review and name its owner. Put a date in the calendar twelve months out with a standing agenda: re-test condition 3 for every internal build, refresh exit-cost estimates, re-classify any category the market has changed. An unbooked review is a policy with an undeclared expiry date.

Sourcing is now decided in advance, which removes one recurring argument from every item in your portfolio. What it does not remove is the harder honesty problem: the items correctly selected, correctly sourced, properly funded, and quietly not working. Deciding what to start is the easy half of portfolio management. The chapter closes with the other half: stage gates, pre-committed kill criteria, and the organisational trick of making a killed project a celebrated outcome rather than a career event.

Key Takeaways

  • Convert MIT's finding that externally partnered solutions succeed roughly twice as often as internal builds into policy rather than argument: it is the fastest, cheapest lever a portfolio owner has, decided with a signature rather than a culture change.
  • Reject both misreadings: "never build anything" surrenders the few places where ownership is the advantage and destroys your ability to evaluate vendors, while "that statistic is about bad builders" is what every internal team says immediately before joining it.
  • Explain the 2x with the four mechanisms rather than with contempt: the unbudgeted maintenance tail vendors amortise across customers, the feature-velocity gap that compounds after launch parity, the domain-versus-tooling confusion, and the evidence disadvantage of discovering every failure mode retail.
  • Ask the diagnostic question in exactly these words when a team claims uniqueness: what specifically is unique here, the workflow, the data, or the software? The workflow usually is, the data sometimes is, and the software almost never is.
  • Set default verdicts by category, not by project: horizontal capabilities BUY with no exceptions, workflow-embedded items BUY the capability and BUILD the workflow, domain-specific categories BUY or PARTNER with agent-washing diligence (only around 130 of thousands of claimed agentic vendors judged real), proprietary-advantage BUILD only on all four conditions, and glue named as glue.
  • Gate every build on all four conditions at once (the capability is the advantage, the data is genuinely unique, no vendor serves the category, sustained capacity is funded beyond the build), and expect eleven items to pass it about twice, or not at all.
  • Carry the portfolio disciplines alongside the verdicts: a concentration ledger pricing exit in money and weeks before the next wave signs, a four-clause exit standard including model-change notification, pilot-before-portfolio for new vendors, and an annual review that retires builds the market has overtaken.
  • Run retirement as a market judgment rather than a verdict on the team, and redeploy the builders onto the migration, the integration layer, and the evaluation harness, because a policy that humiliates internal ambition gets fought on every item that follows.