Multi-Cloud, Multi-Model Strategy
Daniel Reyes, chief technology officer at a 4,000-person federal grants agency, signed a single-vendor cloud deal in 2023 that looked like a bargain. One provider, one bill, one account team, a negotiated 18% discount. Eighteen months later he was trapped. The agency's new AI program needed a model the vendor did not offer, the data-egress fees to move 40 terabytes elsewhere came to $310,000, and the one contracting officer who understood the deal had retired. When a frontier model the agency wanted to pilot launched on a competitor's platform, Daniel could not touch it without re-architecting everything. "I optimized for the discount," he told his governance board, "and I bought myself a cage." This lesson is about building agency-wide AI platforms that keep their options open.
A multi-cloud, multi-model strategy means deliberately designing your AI platform so that no single cloud provider and no single AI model is load-bearing. You can run workloads on more than one cloud, and you can route a given task to whichever model fits best. Be precise about what that buys you, because this is where agencies fool themselves. Designing for portability makes a move possible; it does not make the move cheap, fast or proven. A contract that gives you the right to your data is not a working migration. Until somebody has actually moved a real workload and measured what it cost, your portability is a hypothesis.
Start with an inventory question that usually embarrasses people: how many cloud contracts does your agency actually have? Most large agencies have three or four, and often nobody can name them all. One division negotiated one agreement, another signed a purchase agreement with a different provider, and somewhere a team is still running models on a cluster they stood up years ago. The problem is not having several clouds. The problem is having several clouds and no strategy for them, which produces all of the cost and none of the optionality.
Why Multi-Cloud Matters, and What It Actually Delivers
Avoiding lock-in is the obvious reason. If every AI system depends on one vendor, that vendor holds the pen. They raise prices and you pay. They discontinue a service and you migrate on their timetable. You arrive at renewal with no credible alternative, which is the same as arriving with no negotiating position. But there are deeper reasons that matter more in government than they do commercially, and each one comes with a condition attached.
- Resilience. If one provider has an outage, the agency does not go dark. That holds only for workloads that genuinely run in more than one place with failover that has been tested under load. Multi-cloud on paper is single-cloud during an incident.
- Performance. Different workloads behave differently on different platforms, and the only way to know how yours behave is to run them.
- Compliance. Some data may only be processed in specific regions or under specific security models, which can force a placement decision regardless of price.
- Cost. You can place each class of work where it is cheapest, provided you constrain that routing by the compliance, latency and residency rules first.
- Positioning. Agencies have institutional reasons to avoid appearing dependent on any single supplier, and that concern is legitimate even when the technical case is neutral.
Why Multi-Model Matters
The same argument applies one layer up, at the models themselves. In an ideal world you would use one model family for everything, with one serving stack and one skill set on the team. Reality is messier, and it is messy for reasons that are mostly not your fault. You inherit older models built in different ecosystems, and each ecosystem brings its own tooling and its own operational habits. You adopt newer models from suppliers that run only on the supplier's own service. You run open-weight models inside your own perimeter because the data is too sensitive to leave it. And you keep specialized models, recommendation engines or time-series forecasters, that work best in a particular framework.
Without a strategy, this becomes a mess: several serving systems, inconsistent governance, and a team perpetually learning new tools. With a strategy, the same heterogeneity becomes a set of conscious trade-offs you can explain to an auditor. The point is not to reduce the number of frameworks to one. It is to know why each one is there and what it would take to retire it.
Comparing Clouds Without Writing a Scoreboard
Agencies keep asking which provider is best. That question has no durable answer, because the answer changes with each price update, each new authorized service and each change in your own workload mix. What does not change is the set of dimensions you should compare on. Build your own scorecard against these, measured on your workloads, and refresh it on a schedule rather than trusting a comparison someone published last year.
- Cost structure. Providers price differently: some bill granularly per unit of compute, some bundle into enterprise licensing that is predictable but harder to attribute, and running your own hardware trades high capital cost for lower per-use cost and fixed overhead. Granular billing is easier to optimize and easier to overspend.
- Geographic reach and residency. Region footprints differ widely. If you have data residency requirements, geography is a hard constraint rather than a preference, and authorized government regions are a smaller set than the commercial map suggests.
- Data gravity. Ask one question: if I put data here, what does it cost to take it out? Storage is generally cheap and transfer out generally is not, across providers. That is not malice, it is how the pricing works, and it means every terabyte you land increases the price of leaving.
- Managed machine learning services. Fully managed platforms reduce the work your team does and increase what is specific to that provider. Open tooling gives you more flexibility and hands you more operational burden. Both are defensible; only one of them is reversible cheaply.
- Compliance posture. For federal work, verify current authorization status at the right impact level for the specific service you intend to use, not for the provider in general. Authorization is granted service by service and it changes. State and local requirements vary, and some jurisdictions still require data to stay on premises.
Notice what this list refuses to do. It does not rank suppliers, because a ranking written into a lesson is stale before the lesson is taught, and because a procurement decision recorded as a preference for a named vendor is a procurement problem as well as an architecture one. Score the dimensions, document the evidence, and let the scores fall where your measurements put them.
The Lock-In Trap Daniel Fell Into
Single-vendor lock-in rarely arrives as one bad decision. It accretes. Daniel's discount was real. But each convenient native service he adopted, the proprietary database, the specific model-hosting tooling, the custom data format, added a strand to the rope. Lock-in shows up in four places: data, where egress fees and proprietary formats make moving expensive; services, where proprietary features have no equivalent elsewhere; skills, where the team only knows one platform; and contracts, where the terms punish leaving. Daniel had all four, and no single one of them would have been enough to trap him.
The structural lesson is that lock-in is not a contract clause you can negotiate away later. It is an architecture you either designed against from the start or did not. By the time the cost of leaving is visible on a spreadsheet, the decisions that created it are years old and were each individually reasonable.
Multi-Model: Routing the Task to the Right Model
Multi-cloud is about where you run. Multi-model is about which model you call. No single model is best at everything, and the leaderboard reshuffles constantly. A mature platform treats models as interchangeable parts behind a common doorway.
That doorway is an orchestration layer, sometimes called a model gateway or router: a thin piece of software your applications talk to instead of talking to any model directly. The router decides which model handles a given request, based on rules you set. Consider how Daniel's rebuilt platform routes work. A constituent-facing chatbot answering routine grant-status questions goes to a smaller, cheaper, fast model, roughly $0.50 per million words processed. A complex policy-analysis task goes to a larger frontier model, roughly $15 per million words, used sparingly because it costs thirty times more. Anything touching personally identifiable information routes only to models hosted inside the agency's authorized environment, never to a public commercial endpoint.
Because every application talks to the router rather than to a named model, Daniel can swap a model, add a new one, or re-route by cost or sensitivity by changing rules in one place. When that frontier model launched on a competitor's cloud, the rebuilt platform would have let him add it as one more option behind the router instead of re-architecting his agency. The router is also where the sensitivity rule becomes real: a routing policy that is enforced in code is a control, while the same rule written in a policy document is an aspiration.
The Abstraction Layer and Its Escape Hatches
The same idea applies to infrastructure. Rather than having every engineer learn every provider's interfaces, you build a unified layer that teams talk to, and the layer translates. A data scientist submits training data, specifies the model type, the hyperparameters and the compute they need; the layer decides where the job runs; the trained model is registered in a central registry. Kubernetes for compute abstraction, workflow orchestrators such as Airflow, distributed computing frameworks such as Ray, or custom middleware are all common building blocks for this.
Two warnings. First, the abstraction layer is now a system you own, staff and secure, and it is a single point of failure sitting in front of everything. The cognitive load you saved on provider interfaces reappears as operational load on the layer. Second, your abstraction will cover most workloads and not all of them. If you insist on generality, the workloads that do not fit will simply route around the layer, and you will end up operating two platforms instead of one. Accept escape hatches, document when they are allowed, and review who is using them.
Cost Routing, Bounded by Compliance
Routing work to the cheapest place sounds simple and is not, because cost per unit of useful work varies by workload type. Batch processing tends to be inexpensive everywhere, so the sensible rule is to run it where the data already lives and avoid paying to move it. Real-time inference is sensitive to latency, which varies by region and by model type. Training is dominated by accelerator pricing and availability, and both differ by provider and change often. The only defensible decision rule is to measure your own costs for your own workloads on each platform, and to re-measure on a schedule, because prices move.
Bound that routing with constraints before you let it optimize. Cost is one input, not the objective function. Moving a workload to a cheaper location without checking data residency, latency requirements or authorization status buys a saving and sells a control. Saving ten percent on infrastructure while opening a compliance hole is a bad trade in any agency, and it is the kind of trade that surfaces in an audit rather than in a budget review.
Model Availability Is a Portability Constraint
Before committing to any model, establish where it can actually run. Some models are available only through the supplier's own service. Some are offered through several clouds under different terms. Some are open-weight and can run in your own environment, with licensing that varies more than people expect and that deserves a legal read rather than a glance. Some exist on exactly one platform. The rule follows directly: if a model runs in only one place, then for that application you are not multi-cloud, no matter what the platform diagram says. Write that dependency down as a named risk with an owner, and revisit it whenever the application becomes load-bearing for a mission.
Where the Data Lives
If your data spans clouds, you have to choose a consistency posture, and there are only three. Replicate everything everywhere, which is expensive and leaves you managing stale copies and consistency conflicts. Keep data where it is and move computation to it, which is cheaper in storage and expensive in network transfer. Or run a hybrid: hot, recently used data is synchronized, and cold historical data stays where it was created. Most large organizations end up on the hybrid, and the reason is prosaic rather than architectural. It is the only one of the three whose cost stays roughly proportional to the data people actually use.
Model Taxonomy and Tiered Governance
Once several models run at once, governance gets harder faster than the model count grows. The fix is a taxonomy, so you can reason about classes rather than instances. Classify by domain, such as benefits, tax compliance or personnel; by type, such as classification, regression, ranking or generation; by freshness, real-time or batch; by trust level, experimental, production or sunsetting; and by framework. If you have a hundred models you cannot govern them individually, but you can govern the classes.
Then tier the governance to match criticality. A benefits eligibility model and a research prototype should not face the same process. Applying one heavy standard to both is either too strict, which stalls research, or too loose, which endangers the systems that decide things about people. Experimental models get lighter governance and tight limits on what they may touch; production models get the heavier treatment. The tiers should reflect a risk appetite your leadership has actually stated, not one your platform team inferred.
Model Progression, Routing Logic and Fallbacks
Routing also decides which version of a model serves a given request. Take an eligibility determination system that classifies applications. Model A is the incumbent production model, trained on 2023 data, achieving 96% accuracy on a recent test set. Model B is a candidate trained on 2024 data, achieving 97% on the same test set, but tested on only 10% of application types. A defensible routing rule sends applications that fall inside that tested 10% to Model B and everything else to Model A. As Model B accumulates evidence across more application types, more traffic moves to it, until it becomes production and Model A is archived. This is not an experiment for its own sake. It is conscious progression gated on evidence.
That is also the difference between two approaches you will hear compared. Classic split testing randomly divides traffic, compares outcomes and picks a winner, which means treating a worse model exactly like a better one for the whole test period. Progressive routing, often called a bandit approach, shifts traffic toward the better performing option as evidence accumulates. For government work the progressive approach is usually preferable because it limits how many people are affected by the weaker option while you learn. It is still live experimentation on real decisions about real people, so it belongs in your change approval process and your impact assessment, not in the platform team's backlog alone.
Finally, every production model needs a fallback, decided in advance. The options are to fall back to the previous version, to a simpler and more reliable model, to human decision-making, or to a temporary suspension of service, which is the least desirable. Government agencies should almost always have a human fallback. Telling someone that the system is unavailable and a person will decide their case is more honest, and usually more defensible, than serving a degraded automated decision and calling it service continuity.
What to Standardize, What to Leave Flexible
Multi-cloud done naively means running everything everywhere, which doubles cost and complexity. The discipline is choosing what stays portable and what is allowed to be provider-specific. Daniel's rule: standardize the things that are expensive to change, stay flexible on the things that are cheap to change.
He standardized on portable foundations, containers as a packaging format that runs the same way anywhere, open data formats, and an identity system that works across providers. He allowed teams to use a provider's native service only when it offered a real advantage and only after documenting an exit path. The test for any native service was one question: if we had to leave this provider in 90 days, what would this decision cost us? A team that cannot answer that question has not evaluated the service, only enjoyed it.
A Multi-Cloud, Multi-Model Platform Decision Framework
Daniel turned his hard-won lessons into a one-page framework his architecture board now applies to every platform decision. Score each dimension; the pattern reveals the risk.
- Exit cost. If we had to leave this provider or model in 90 days, what would it cost in dollars, time, and disruption? A number you cannot say out loud is a red flag.
- Data portability. Is our data in an open format, and what are the egress fees to move it? Estimate the full bill, not the per-gigabyte rate.
- Exit rehearsal. Has anyone ever actually moved a workload out, even a small one, and what did we learn? An untested exit path is a plan, not a capability.
- Model independence. Do our applications call models through a router, or are they wired to one named model?
- Sensitivity routing. Is there an enforced rule, in code rather than in a document, that sensitive data only reaches authorized, in-perimeter models?
- Skills concentration. Does more than one person understand each critical system, and do skills span providers?
- Cost transparency. Can we see what each workload and each model actually costs per month, by task?
- Service lock-in. For each proprietary native service we use, is there a documented equivalent elsewhere and a migration path?
- Contract leverage. Do our terms let us scale down, leave, or move workloads without punitive fees, and is renewal competitive?
- Governance owner. Who is the named, accountable official for this platform's resilience and exit readiness?
Should Your Agency Pursue Multi-Cloud At All?
Not every agency should. The answer is probably yes if you already run models at scale, if data residency requirements differ across divisions, if you want real leverage in a negotiation, or if single-supplier dependency is itself a security concern your leadership has raised. The answer is probably no if you run only a handful of models, if one provider genuinely meets your needs today, if you have limited engineering capacity to operate an orchestration layer, or if your team is already stretched thin. A badly run multi-cloud platform is worse than a well run single-cloud one.
If you go ahead, start small. Pick one new provider for one workload, learn what actually breaks, then expand. If you choose to stay on a single provider, that is a legitimate decision, but make it explicitly: write it down, record the criteria that would trigger an exception, and revisit the decision on a fixed cycle rather than discovering at renewal that nobody ever reconsidered it.
The Cost of Resilience, Honestly Stated
A multi-cloud, multi-model platform is not free. The orchestration layer, the portability discipline, and the second set of skills add overhead, often 10 to 20% above the cost of going all-in on one vendor. Daniel does not pretend otherwise. He frames it to his board as insurance rather than waste: the $310,000 egress bill and the frozen pilot were the uninsured loss. The premium buys negotiating leverage at renewal, the freedom to adopt a better model the week it ships, and continuity if a provider raises prices, suffers an outage, or loses an authorization. For an agency that must serve the public for decades, resilience is not a luxury feature. It is a core platform requirement, and the NIST AI Risk Management Framework's attention to third-party and supply-chain dependency is this same discipline applied to AI.
Anti-Patterns to Avoid
Calling designed portability proven portability. A platform built on containers, open formats and a router can move. That is not the same as having moved. Until a real workload has been relocated and the cost measured, the exit plan is untested, and untested exit plans fail at exactly the moment you need them. Rehearse a small migration on a fixed cycle and publish what it cost.
Multi-cloud without a strategy. Workloads land on three providers for three reasons nobody remembers, on three contracts with three security models. Migration is painful because no one can say why anything is where it is. Document which workloads run where and why, and review it on a fixed cycle.
Treating the abstraction layer as free. The layer that hides provider differences is a system your agency now builds, staffs, secures and can be taken down by. Budget for it as a product with an owner, not as glue.
An abstraction so rigid that teams route around it. If the layer serves most workloads and the remainder are refused, the remainder will go direct and you will operate two platforms. Design escape hatches deliberately and track their use.
Governing all models the same way. One standard applied to both a benefits eligibility model and a research prototype is simultaneously too slow for research and too light for the system that decides someone's benefits. Tier the governance to criticality.
Cost optimization without guardrails. Routing purely on price will eventually place regulated data somewhere it should not be, or put latency-sensitive inference somewhere too far away. Constrain the router by compliance, residency, security and performance first, then optimize inside those bounds.
Practice Prompts
1. Current state inventory. List every cloud your agency uses, including on-premises environments. For each, record what runs there, how it was originally selected, who owns the contract, and the annual spend. Note every item you could not answer without asking someone else.
2. Orchestration assessment. Describe how you deploy to more than one environment today. Do teams talk to a unified interface, or does each team learn each platform separately? How long does it currently take to move one model from one environment to another, and what is the single most painful step?
3. Model governance review. For your current production models, count the frameworks in use, write down how you decide which model is authoritative for a given decision, state the fallback for each, and describe how and how often old models are retired.
4. Exit rehearsal plan. Choose the smallest workload whose migration would still teach you something. Write the plan to move it, estimate the cost including egress, then compare your estimate with what it actually costs when you run it.
5. Sensitivity routing test. Take the rule that sensitive data may only reach authorized in-perimeter models. Find where that rule is enforced in code. If it exists only in a policy document, write the change request that moves it into the router.
Reflection
What is the single biggest supplier dependency your agency carries right now, and what would it take to reduce it? If one of your providers had a major outage tomorrow, which mission functions would stop, and how quickly would you know? For each of your top three AI models, could you move it somewhere else if you had to, and what specifically would prevent you? Finally, what does your current setup cost, and what would genuine, tested portability cost on top of that? If you cannot answer the last question, that gap is itself the finding.
Glossary
- Multi-cloud. An architecture that deliberately uses two or more cloud providers, typically to reduce lock-in, meet residency requirements, and preserve optionality.
- Data gravity. The friction created by the volume and location of data. Once large datasets sit with one provider, the cost of moving them pulls everything else toward staying.
- Orchestration layer. Software that abstracts differences between multiple deployment targets, so teams work against one interface instead of several.
- Model router. The component that decides which model or model version handles a given request, based on rules such as cost, task type, data sensitivity or accumulated evidence.
- Progressive routing. An approach that shifts traffic toward the better performing option as evidence accumulates, rather than splitting traffic evenly for a fixed test period.
- Fallback. The predetermined path when a production model is unavailable or untrustworthy, ideally ending in a human decision rather than a degraded automated one.
- Egress fees. Charges for moving data out of a provider's environment, which convert stored data into an exit cost.
- FedRAMP. The Federal Risk and Authorization Management Program, the government-wide process for authorizing cloud services for federal use, granted service by service at a defined impact level.
Related Lessons
This lesson sits between architecture and acquisition. Designing Agency-Wide AI Platforms covers the platform this strategy is a property of, and Data Mesh and Data Fabric for Government goes deeper on where data lives and how it is served across environments. Vendor Lock-In Prevention takes the lock-in analysis further into contracting practice, and AI Contract Negotiation covers the terms that turn an exit plan into an enforceable right. FedRAMP and AI Cloud Authorization is the authoritative treatment of the authorization gate described here, Third-Party AI Risk Management addresses supplier dependency as a risk discipline, and AI Infrastructure Cost Optimization pairs with the cost routing and measurement practices above.
Closing
Multi-cloud and multi-model strategies are not about technical elegance, and they are certainly not about fairness to suppliers. They are about organizational resilience: not waking up one day to discover that a decision made three years ago now constrains everything the agency wants to do. Daniel's mistake was not choosing a provider. It was choosing one without ever pricing the exit. Build the router, standardize what is expensive to change, keep the compliance constraints ahead of the cost optimizer, and rehearse a migration before you need one. The best time to design for optionality was before the first contract was signed. The second best time is at the next renewal.
Key Takeaways
- Discounts can buy a cage. Single-vendor savings are real, but lock-in across data, services, skills, and contracts can cost far more than the discount when you need to move.
- Designed portability is not proven portability. Containers, open formats and a router make a move possible; only an actual rehearsed migration tells you what it costs. Treat an untested exit path as a risk, not a control.
- Compare on dimensions, not on a supplier ranking. Cost structure, geographic reach and residency, data gravity, managed service depth and current authorization status, all measured on your own workloads and refreshed on a schedule.
- Route through an orchestration layer. Have applications call models through a router, not by name, so you can swap models, control cost, and enforce sensitivity rules in one enforced place.
- Match the model to the task. Send routine work to cheap, fast models and reserve costly frontier models for genuinely complex work; in Daniel's platform the price gap between the two tiers is thirty times.
- Constrain cost routing before you optimize it. Compliance, residency, security and latency are boundaries; price is what you optimize inside them.
- Tier governance to criticality, and always define the fallback. Experimental and production models need different processes, and every production model needs a predetermined fallback, which in government should almost always end with a human.
- Ask the 90-day exit question. For any provider, service or model, knowing the real cost of leaving in 90 days is the truest measure of your dependency, and the 10 to 20% portability premium is the price of keeping that number small.
Frequently Asked Questions
Is multi-cloud always the right answer for a government agency? No. If you run a handful of models, one provider meets your requirements, and your engineering capacity is thin, a well run single-cloud platform will serve you better than a badly run multi-cloud one. What is not optional is making the choice deliberately: document the decision, the criteria that would change it, and the date you will revisit it. The failure mode is not single-cloud. It is single-cloud by default, discovered at renewal.
Does multi-cloud guarantee we survive a provider outage? Only for the workloads that genuinely run in more than one place with failover you have tested. Having contracts with several providers does nothing during an incident. Identify the mission functions that must survive an outage, confirm they actually run in more than one environment, and test the failover under realistic load. Everything else is a single-cloud workload with extra invoices.
How do we handle a model that only runs on one platform? Use it if it is the right model, but record the dependency explicitly as a named risk with an owner, and note which applications inherit it. For those applications you are single-cloud regardless of the rest of your architecture. Revisit the risk whenever the application becomes load-bearing for a mission outcome, and check whether an equivalent has become available elsewhere.
What belongs in the orchestration layer and what does not? Put the decisions you want to make once and enforce everywhere: which model serves which task, what sensitivity rules apply, how requests are logged, and what happens when a model fails. Keep out anything so workload-specific that generalizing it would make the layer brittle. Expect escape hatches, define when they are permitted, and monitor their use, because an abstraction that teams route around has become overhead with no benefit.
How do we justify the portability premium to a budget office? Frame it against the cost of the alternative, in your own numbers. Daniel's board understood the argument only when he showed them the egress quote and the pilot he could not run. Present the premium as a percentage of platform cost, present the exit cost of your current largest dependency next to it, and state plainly what the agency loses at the next renewal if it arrives without an alternative. That is a comparison a budget officer can evaluate, which an architectural argument is not.
Skill.re