Managing AI Tool Proliferation Across Teams
Marcus Bellweather runs a 22-person product organization split across three teams: design, engineering, and product management. When he finally audited what AI tools his people were actually paying for, he counted 17 different subscriptions. Two teams were paying for nearly identical writing assistants. Three people had individually expensed the same transcription tool. One designer was pasting unreleased product specs into a free consumer chatbot with no data agreement. The total monthly spend was 38 percent higher than Marcus had assumed, and nobody owned the mess. This is tool proliferation, and it happens to every team that adopts AI faster than it organizes AI.
What This Lesson Covers
Tool proliferation is the sprawl that results when individuals and sub-teams each adopt AI tools independently, with no shared view of what is in use, what it costs, or what data it touches. As a manager across several teams, you sit at exactly the level where this can be fixed: high enough to see the overlap, close enough to know the work. This is operational governance, not enterprise IT policy.
This lesson covers why proliferation happens and why it is more than a budget problem, how to run a tool audit, how to rationalize a sprawling toolset using a simple scoring approach, how to set a light approval process that prevents the next round of sprawl, and how to do all of this without becoming the manager who kills useful experimentation. We follow Marcus from his 17-tool audit to a rationalized, governed set.
Why Proliferation Happens, and Why It Matters
Proliferation is not a sign of an undisciplined team. It is the natural result of a good thing: people trying tools to solve real problems. Each individual decision was reasonable. The designer wanted better copy; the analyst wanted faster transcripts. The problem is that nobody was looking at the sum.
The costs of sprawl are not only financial, though Marcus's 38 percent overspend was real. Duplicate tools mean duplicate spend and fractured workflows where two teams cannot share assets because they live in different systems. Worse is the invisible risk: when you do not know which tools touch company data, you cannot answer a security or compliance question, and the designer pasting specs into an unvetted consumer app is a data exposure waiting to surface. And sprawl fragments learning, because the prompt your engineer perfected in one tool is useless to the designer in another.
The problem is not that your team adopts AI tools. It is that they adopt them in the dark, where nobody can see the overlap, the cost, or the risk.
The Anatomy of Sprawl
Marcus's design team researched honestly and picked the best fit for their problem; engineering hit a similar problem and picked a marginally different tool. Both choices were sound locally, and the organization now carries two where one would have served. Each addition drags a tail nobody counts at adoption: people to train, a vendor relationship, a security review, an integration to keep working, a line of spend to track. Some tools do deserve to coexist, because forcing a content generation tool to do document processing is worse than paying for both. The question is never whether fewer tools is better, but whether this particular multiplication buys you anything. Five patterns say it does not.
- Redundant tools. Two tools solve the same problem with minor variations, like Marcus's writing assistants, so you pay twice and train people on two interfaces where one would cover both groups.
- Incompatible outputs. One team's tool emits highly structured output and another's emits loose text, so anything downstream of both needs a person to reformat by hand before work continues.
- Vendor risk. You depend on too many vendors for capabilities you now rely on. If one raises prices or changes the product, no fallback is ready, and each extra vendor is another party holding your data.
- Expertise fragmentation. You get thin familiarity with many tools instead of deep command of two or three. Nobody sees the whole picture, silos form, and good practice stops travelling between teams.
- Cost duplication. You pay repeatedly for capability consolidation would buy once, most visibly with per-seat pricing that multiplies across every team that signed up separately.
The Real Cost, Direct and Indirect
Licence fees are the visible cost and usually the smallest. Start instead with your own bandwidth: a manager coordinating across teams on different tools spends time learning each well enough to review the work, translating outputs between formats, and maintaining a long list of vendor relationships. That is attention consumed by plumbing rather than judgment. Onboarding multiplies too, because a new hire in Marcus's product team must learn not only their own team's tool but enough of design's and engineering's to collaborate. Then there is the infrastructure tax, since every tool needs integration, a security review, and monitoring; one tool's worth is unremarkable, seventeen tools' worth is a standing charge nobody approved.
The largest cost never appears on an invoice. Had Marcus's teams built deep expertise in three tools instead of shallow familiarity with 17, their prompting would be sharper and their processes more refined. Sprawl does not just cost money. It caps how good you can get.
Why Consolidation Makes You Better, Not Just Cheaper
Commit to one tool for a use case and you start investing in depth: someone builds prompts that reliably work, someone wires it into the systems where the work lives, and the team gets trained properly instead of left to guess. Depth produces quality, and depth is only affordable when it is concentrated. Spread thin, every team can afford surface customization only, so capability plateaus early. This is why organizations that consolidate so often find the tool they thought was weaker outperforms, once properly customized, the tool they thought was better when barely configured. Marcus saw a small version of it within a quarter, as shared prompt patterns pushed the standardized writing assistant past what either original tool had done.
Running the Tool Audit
You cannot rationalize what you cannot see, so Marcus started with an honest inventory. He asked each person, with explicit amnesty, to list every AI tool they used for work, what they used it for, what it cost, and whether it touched any company data. The amnesty mattered: people hide shadow tools when they fear punishment, and a hidden tool is the dangerous one.
He captured each tool in a simple table: name, owner, purpose, monthly cost, number of users, and data sensitivity. The picture that emerged sorted into four buckets. Redundant: multiple tools doing the same job, like the two writing assistants. Risky: tools touching sensitive data with no agreement, like the consumer chatbot. Underused: tools one person expensed and abandoned. Keep: tools delivering clear value to multiple people. That four-way sort turned a vague sense of mess into a concrete work-list.
Rationalizing the Toolset
With the audit done, Marcus had to decide what to consolidate, what to kill, and what to standardize on. For the overlapping categories he scored the competing tools rather than picking by who shouted loudest, using a simple weighted scoring approach with four criteria his teams agreed mattered: capability fit (weight 40 percent), data security and agreement status (30 percent), cost per user (20 percent), and ease of integration with existing work (10 percent).
Take the two writing assistants. Tool A scored 8 on capability, 9 on security, 6 on cost, and 7 on integration, for a weighted score of 7.8. Tool B scored 7, 4, 9, and 6, for 6.4. Tool B was cheaper, which is exactly why both teams' gut instinct had leaned toward it, but its weak security score and missing data agreement dragged it down once the criteria were weighted. Marcus standardized on Tool A, negotiated a team rate that came in below the combined cost of the two separate subscriptions, and migrated everyone within a month. The scoring made the decision defensible: when the team that lost its preferred tool pushed back, Marcus could show the numbers rather than argue taste.
The risky bucket got handled immediately and separately. The consumer chatbot was banned for company data the same week, and the designer was moved to the approved, agreement-backed alternative. Risk does not wait for a scoring exercise.
Turning the Survivors Into a Catalog
The output of rationalization should not be a decision Marcus remembers. It should be a document his teams can read. For each surviving tool he recorded what problems it solves, which teams use it, cost per user, security and compliance status, real capabilities alongside known limitations, where to find learning resources, and who to ask when something breaks. That last field does more work than it looks like it should, because naming a contact turns an anonymous subscription into something a person owns. The rest governs quietly: someone starting a project sees what is already approved, adopts it immediately, and skips both the request and the security review, because that work was done once on everyone's behalf. Over time new hires learn the tools in the catalog rather than discovering a private landscape of one-off subscriptions, so expertise accumulates instead of dispersing.
Preventing the Next Round of Sprawl
An audit is a one-time cleanup. Without a process, the 17 tools become 25 by next year. Marcus put in a deliberately light gate so that experimentation stayed alive while sprawl did not.
His rule: anyone can trial any tool freely with no company data and no approval, because killing experimentation is its own failure. But moving a tool into regular use, paying for it, or letting it touch company data required a one-page request and a yes from him, with a target turnaround of one week. He maintained a visible, shared list of approved tools per category so that most people's first instinct became "use the approved writing assistant" rather than "go find my own." The approved list does most of the governance work, because it makes the compliant path the easy path.
He also set a quarterly 30-minute review to re-audit lightly: what got added, what fell out of use, what should be retired. Proliferation is a recurring tide, not a one-time spill, so the prune is permanent.
Standard Problems and Novel Problems
What makes a light gate workable is the line between standard problems and novel ones. A standard problem is one your organization has already solved, or one your industry solves routinely. Document processing is the classic case: several functions need it and there is no reason each should shop separately, so the policy is to use the approved tool and the catalog says which one. A novel problem is one nobody here has solved, where the approved set has no answer, and that is where a pilot belongs. Marcus made pilots time-limited on purpose, with a defined end date, and required whoever ran one to write down what they learned whether the answer was yes or no. A pilot that ends in a decision is a cheap experiment. A pilot that quietly becomes permanent because everyone forgot to end it is how you get 17 tools.
Keep the review to a few questions. Does this solve a problem we actually care about? Does it meet our security and compliance requirements? What does it cost? Is there an approved alternative that would do the job, even imperfectly? Those fit in a short conversation, and asking them consistently holds the turnaround at a week. Make the policy visible so teams can predict the answer before they ask, because predictability is what stops people routing around you.
Working Groups Around Shared Tools
For any tool more than one team depends on, put its users in a room on a rhythm, monthly or quarterly depending on how fast the tool moves. Marcus ran one for the standardized writing assistant with members from all three teams: they compare what is working, walk through advanced use cases, and escalate to the vendor as one voice rather than three support tickets. The payoff is that discovery stops being repeated. When an engineer worked out a prompt structure that noticeably lifted output quality, she brought it to the group and within a week the other two teams were using it. Without the group, each team finds its own version months apart, or never. It also makes using a tool feel like belonging to a practice rather than sitting alone with a subscription.
The Deeper Portfolio Audit
The quarterly check keeps drift small. Once or twice a year, look at the portfolio whole. How many teams actually use each tool? What does each cost in total, seats included? Are there tools with exactly one user who could be served by something already approved? Are there tools nobody has opened in months that can be sunset outright? And, running the other way, are there tools many teams quietly rely on that were never approved and should now be formalized? That last question matters as much as the others, because the audit is not only pruning. It is your chance to consolidate overlaps, discontinue the dormant, and promote what has proven itself, since some of the best catalog entries arrive because a team was right before the policy caught up.
Sunsetting With a Plan
When a tool has to go, the decision is the easy half. Simply switching off support turns a sound portfolio call into a morale problem, because people built real working habits on top of it. Give every retirement an announced end date far enough out that migration is possible, a named approved alternative, and actual help getting there: teaching the replacement, sitting with people while they rebuild processes they had tuned, and accepting a dip in output for a few weeks. Handled that way, a retirement reads as respect for work already done. Handled abruptly, it reads as indifference, and the resentment surfaces the next time you ask a team to adopt anything.
Governing Without Killing Experimentation
The trap Marcus avoided is the overcorrection. A manager burned by sprawl often slams the door, requires heavy approval for everything, and watches the team route around the policy into deeper shadow IT than before. Too-tight governance produces the exact problem it was meant to solve.
The balance is to gate by risk, not by novelty. Free trials with no data carry no risk and need no gate. Spending money and touching data carry real risk and deserve a quick, lightweight check. By drawing the line at data and dollars rather than at curiosity, Marcus kept his team experimenting in the open, where he could see the good discoveries and adopt them, instead of driving the experimentation underground.
Where Judgment and Communication Still Matter
No policy resolves every case. When a team says the approved tool does not meet their needs, understand the problem underneath the request before restating the rule. Often you can close the gap by customizing the approved tool, which is cheaper than another vendor and deepens expertise you already hold. Sometimes the gap is real and their tool deserves proper evaluation. The harder case is a team that has built serious capability on an unapproved tool and would take a genuine hit by switching; that switching cost belongs in the decision, whether you grandfather them in during review or evaluate the tool on its merits for everyone. Balance does not mean never saying no. It means deciding contextually rather than categorically, since some tools deserve broad approval, some should stay restricted to the one team with the specific need, and some should be retired.
Whatever you decide, keep saying it out loud. Marcus repeated the policy to every new hire, to every team weighing a tool, and every time he approved or rejected a request, explaining what made it a yes or which criterion it failed. Inconsistency destroys the system: approve one team's tool and reject another's similar one without explaining the difference, and people conclude that favor decides outcomes, then stop asking and start hiding, which is where the 17 tools came from.
Three Ways Managers Get This Wrong
- The forbidden tool. Banning specific tools without a stated reason, and without understanding why people wanted them, does not remove the tools. It removes your visibility of them. Decide against published criteria and explain every no.
- Uncontrolled experimentation. Waving everything through in the name of innovation grows the portfolio past manageability while cost climbs and expertise scatters. Frame experimentation instead of stopping it: pilots with end dates, learning captured, a real decision about whether the tool graduates.
- Bureaucratic approval. When approval takes six weeks and several sign-offs, people use unapproved tools rather than wait. Marcus's one-page request and one-week decision exist precisely because friction is what creates shadow IT.
Terms Worth Knowing
- Approved tools catalog. The curated list of AI tools cleared for organizational use, documenting each tool's capabilities, cost, security status, and intended use cases.
- Pilot program. A time-limited trial of a new tool by one or more teams, run to decide whether it should be approved for wider use.
- Shadow IT. Technology in active use outside any formal governance, usually a symptom of policy that was too restrictive or too slow.
- Tool consolidation. Reducing the number of tools by migrating teams off similar tools onto a single shared standard.
Practice and Reflection
Work these as if you were in Marcus's chair, and write the answers down, because the reasoning is the point.
- Rationalize a portfolio. Picture 12 tools across 8 teams: five have one user each, three are used by multiple teams, two were recently piloted and now sit idle. Name what you would sunset, consolidate, and approve, justifying each.
- Handle an off-catalog request. A team wants an unapproved tool and insists it is perfect for their problem. Walk through the questions you would ask, the criteria you would apply, and how you would communicate the outcome either way.
- Design a working group. Pick a tool several teams use. What should the group accomplish, how often should it meet, what happens in a session, and how will you keep people participating rather than merely attending?
- Write the policy communication. Draft the message explaining why an approval process exists, how someone requests a new tool, and how a tool moves from pilot to approved.
Then turn the questions on your own situation. What is actually in your inventory, and where do you already see overlap? If you rationalized it, what would you consolidate, what would you discontinue, and on what criteria would you defend those calls? And what shape of governance fits your organization, light enough that people keep experimenting in the open, firm enough that the sprawl does not quietly rebuild?
Coordination, Not Control
The point of all this is not control. It is coordination, making sure the organization gets full value from what it already spends on AI by preventing chaos without preventing invention. Start where Marcus started, with your current state: map what you have, understand which tools solve which problems, then build the machinery of a catalog, a fast approval path, working groups, and a regular audit. From there you manage forward instead of cleaning up, deciding new proposals against clear criteria, promoting what proves itself, and retiring what has gone redundant. Done consistently this compounds, because expertise deepens when it is concentrated and knowledge crosses boundaries instead of pooling behind them. Capability with AI is not built by having the most tools. It is built by stewarding a focused portfolio well.
Key Takeaways
- Proliferation is reasonable decisions summing to a problem. Each tool adoption made sense individually; the harm is in the unseen total of cost, overlap, and data risk.
- Audit with amnesty first. You cannot govern tools you cannot see, and people only reveal shadow tools when revealing them is safe.
- Sort the inventory into redundant, risky, underused, and keep. The four-way sort turns a vague mess into a concrete work-list.
- Rationalize with weighted scoring. Score overlapping tools on capability, security, cost, and integration so consolidation decisions are defensible, not political.
- Handle risk immediately, separately. A tool exposing sensitive data with no agreement gets fixed the same week, outside the scoring exercise.
- Gate by data and dollars, not by curiosity. Let people trial freely without company data; require a quick one-page approval only to pay, integrate, or touch data, and keep a visible approved-tools list.
- The indirect costs dwarf the licence fees. Managerial bandwidth, multiplied onboarding, repeated security and integration work, and the capability you never built are where sprawl actually hurts.
- Standard problems use approved tools; novel problems get time-limited pilots. Every pilot ends on a date, with its learning captured and an explicit approval decision.
- Keep a real catalog and a working group per shared tool. Document what each tool solves, who uses it, cost, security status, limits, resources, and a named contact, then let the people using it share what they learn.
- Audit yearly and retire with a plan. Consolidate overlaps, sunset the dormant, formalize what teams already rely on, and give every retirement an early end date with migration support.
Skill.re