←
AI for Government
Strategic · M5 · lesson 5 of 47 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
AI Infrastructure Cost Optimization
📖
now learning

AI Infrastructure Cost Optimization

15 min

Eleanor Whitfield is the chief information officer of a cabinet-level state agency with 14,000 employees. Eighteen months ago her teams launched eleven separate AI pilots: a chatbot here, a document-summarizer there, a fraud-detection model in the revenue division. Each was approved on its own small budget. Then the consolidated cloud invoice landed on her desk. It read $2.3 million for the quarter, up from $410,000 the year before, and climbing 18 percent month over month. The appropriations staff had questions. Eleanor had a harder problem than the number itself: she could not explain it. Nobody could say what each pilot cost, which ones delivered value, or why the bill kept growing.

Why AI Costs Spiral in Government

AI infrastructure costs behave differently from traditional IT spending, and the differences are exactly what catches public-sector budgets off guard. Traditional systems have fixed, predictable costs: you buy a server, you run it, and the number on the budget line is the number you defend next year. AI systems have variable, usage-driven costs. Every query to a large language model, every image processed, every hour a graphics processing unit sits powered on adds to the bill. When usage grows, cost grows, often faster than anyone modeled.

Eleanor's chatbot was a victim of its own success. More residents used it, so it cost more, and nobody had budgeted for popularity. The controlling analogy for this lesson is the utility bill: AI infrastructure is metered like water or electricity, and you pay for what you consume. Just as a building's water bill explodes when a pipe leaks unnoticed, an AI bill explodes when an idle accelerator runs all weekend, an oversized model answers a question a smaller one could have handled, or a test loop calls a premium service ten thousand times.

This is not a state-agency problem or a Whitfield problem. Federal AI spending tripled between fiscal 2021 and fiscal 2024, and the Government Accountability Office has repeatedly flagged that agencies cannot explain what they are getting for the money. GAO has found that agencies often cannot separate AI spending from general IT spending, cannot attribute cloud bills to specific AI use cases, and cannot compare the unit economics of competing approaches. The line that follows from all three findings is the operating principle of this lesson: you cannot govern what you cannot cost, and you cannot cost what you have not tagged.

The failure mode has a shape. An engagement by 18F, the federal digital services consultancy that operated inside the General Services Administration, found that a single generative AI pilot at a cabinet department had spent $4.2 million on GPU capacity before anyone in the chief financial officer's shop realized it was AI spending at all. It had been invoiced as cloud compute under the department's enterprise cloud contract. The pilot delivered three demos and no production workload, and the program was cancelled at OMB passback. Nothing was hidden. Nothing was tagged either.

Who Owns the Number

AI infrastructure cost is one of the few technology problems where four executives have a genuine stake and no single one can solve it. The chief information officer owns the platform. The chief AI officer owns the use-case inventory and the risk categorization. The chief data officer owns the storage and movement that quietly dominates some workloads. The chief financial officer owns the appropriation and the answer to the appropriations committee. Without a shared vocabulary these four blame each other after the invoice arrives, which is a reliable way to lose a program.

Three artifacts give them something to argue about productively rather than personally: a total cost of ownership model, a cost dashboard that attributes spend to programs, and a build-buy-rent decision memo for each significant capability. Build those three and the conversation changes from whose fault the bill is to which lever we pull. The rest of this lesson is how to construct each one, starting with the cost model, because the other two depend on it and because almost every agency's version of it is wrong in the same direction.

The policy perimeter around these decisions is wider than most technologists expect. OMB Circular A-11 governs budget formulation, and A-130 governs information resource management. OMB M-24-10, Advancing Governance, Innovation, and Risk Management for Agencies' Use of Artificial Intelligence, issued in 2024, requires agencies to maintain a use-case inventory of safety-impacting and rights-impacting AI. OMB M-22-18 covers software supply chain security. FISMA sets the security baseline, the Clinger-Cohen Act and FITARA establish CIO authority, and the Antideficiency Act sets the outer limit on obligating money you do not have.

Build a Real Total Cost of Ownership Model

Total cost of ownership means the full cost of running an AI capability over its useful life, not just the headline subscription. Most agencies dramatically understate it because they count only the obvious line. When Eleanor's team rebuilt the model for the chatbot, the subscription they had been tracking turned out to be 31 percent of the real cost. Compute, egress, and the two engineers maintaining it made up the rest. An AI pilot priced on its subscription is like a car priced on its sticker and not its fuel, insurance, and mechanic: the cheap part is the part you noticed.

The shorthand most agencies actually track has five buckets. Compute covers the processing time for training and, more importantly for most deployed government AI, daily inference, because you train once and you run forever. Data covers storage, movement, and the egress fees cloud vendors charge to take your own data out. Model access covers per-query or per-token fees for commercial models used through an interface. People covers the engineers, data scientists, and security staff who keep it running, usually the biggest line and the one pilot budgets ignore. Governance covers accreditation, monitoring, and audit.

The fuller federal model, the one you need when the workload is large enough to have its own infrastructure, has seven buckets. Hardware covers servers, accelerators, storage, networking, and the redundancy required to meet high availability targets; accelerators depreciate faster than general-purpose compute because generational leaps shift the price-per-unit-of-work curve every 18 to 24 months. Facilities covers power, cooling, space, and physical security. Personnel is the hidden giant. Software covers operating systems, container orchestration, pipeline tooling, model-serving frameworks, and commercial monitoring licenses, which stack quickly.

Network covers egress, inter-region transfer, inter-cloud transfer, and the traffic-inspection overhead that Trusted Internet Connections 3.0 architectures impose. Integration covers connecting the AI platform to mission systems, identity and credential management, PIV and CAC authentication, continuous diagnostics logging, and financial systems; it runs routinely 30 to 50 percent of first-year spend and is almost always underestimated. Compliance and decommissioning covers FedRAMP authorization and continuous monitoring, sustaining the authority to operate, evidence collection, media sanitization under NIST SP 800-88, and records disposition under National Archives schedules.

Two of those buckets deserve concrete numbers because they surprise people. On facilities: a top-end training accelerator draws roughly 700 watts at peak, so a 200-unit cluster needs about 140 kilowatts of power and roughly 40 tons of cooling. For a federal data center that usually means new power distribution, new cooling units, and a facilities project that adds six to twelve months and $1 million to $3 million before you train a single model. On personnel: at a blended fully loaded federal rate of about $220,000, a ten-person platform team costs $2.2 million a year whether or not a single workload runs.

On network, the arithmetic is easy to check and easy to ignore. A federal workload moving 100 terabytes per month out of a government cloud region to an on-premises data lake pays roughly $9,000 per month in egress alone, before whatever the inspection provider charges. Nothing about that architecture looks expensive on a diagram. It looks like a sensible separation of concerns, and it bills every month forever. Egress is the silent budget killer, and it is the bucket most often missing entirely from a pilot's cost estimate.

Accelerator Economics

Accelerator pricing dominates AI infrastructure cost today, and the on-demand rate is the worst rate available. A single on-demand top-end training instance in a US government cloud region listed at roughly $98 per hour in 2026, which works out to about $860,000 per year if it is left running. A three-year reservation paid fully up front cuts that to roughly $480,000 per year. Capacity reservations that guarantee availability for a scheduled training run reduce the effective rate further when you can plan the run in advance rather than starting it the afternoon someone asks.

For inference the calculus is entirely different, and this is where most government money is wasted. A top-end training accelerator is wasted on most inference workloads. Inference-class accelerators handle most government retrieval-augmented generation workloads at roughly one quarter of the cost, and vendor-specific inference and training chips are cheaper still for workloads that fit their compiler constraints. The right question is not which accelerator is fastest. It is which is the cheapest accelerator that still meets the latency target the service is actually held to.

Reserved capacity is a commitment and should be treated as one. If mission needs shift, a reservation becomes sunk cost that you will be explaining rather than using. Federal cloud acquisition guidance and the software supply chain memorandum both push agencies toward shorter commitments and portable workloads for this reason. A pragmatic portfolio is roughly 40 percent reserved, 40 percent on a savings-plan style commitment, and 20 percent on demand, rebalanced quarterly as actual usage data accumulates rather than annually as a budget exercise.

Buy this capacity through the governmentwide vehicles your agency already uses, and document the cost basis as you go. The GSA Multiple Award Schedule and the governmentwide acquisition contracts run by NASA and by the National Institutes of Health's information technology acquisition center are the usual routes for commitment discounts and reserved capacity. Confirm which vehicle generation is currently open before you plan around it, because these vehicles are re-competed and consolidated on their own schedule. Write down why you chose the commitment mix you chose, because your CFO will need it in a GAO audit.

Turn the Right Knobs

Once you can see the bill you can cut it, often by 40 to 70 percent without losing the capability the program depends on. Read that claim carefully, because it is a claim about the workloads people have measured, not a law. Each lever below has a quality risk attached, and the discipline that makes the savings safe is measuring output quality on the same cases before and after, on a sample that includes the hard ones. A cost reduction with no quality measurement attached is not an optimization. It is an untested change to a production system.

Right-size the model

The most expensive habit in government AI is using a giant, premium model for a task a small one handles fine. Summarizing a one-page memo does not need the most capable model on the market. Eleanor's team moved 60 percent of chatbot queries, the simple and repetitive ones, to a smaller and cheaper model, and reserved the premium model for complex cases. That single change cut chatbot compute cost by 48 percent with no measurable drop in answer quality on the categories they evaluated. Note the qualifier: quality was measured on those categories, over that period, and it has to keep being measured as the query mix drifts.

Cache and reuse

Residents ask the same questions thousands of times. Caching common answers means you pay the model once rather than ten thousand times. A well-tuned cache served 35 percent of the agency's chatbot traffic without a single new model call. Caching carries its own correctness obligation: a cached answer to a policy question is wrong the moment the policy changes, so pair the cache with an invalidation trigger tied to the source of truth, not to a fixed expiry someone picked because it sounded reasonable.

Turn off idle compute and terminate orphans

Accelerators left running overnight and at weekends are the unnoticed leaking pipe. Automatic shutdown of non-production environments outside business hours is the closest thing to free money in the AI budget. Alongside it, run a monthly sweep for orphans: endpoints nobody calls, notebooks nobody opens, storage volumes detached from anything, and development clusters belonging to a project that ended. Right-size instances monthly against actual utilization data, and move cold data to archival storage tiers rather than leaving it on the tier you first uploaded it to.

Batch, schedule, and control egress

Document processing that does not need to be instant can run in scheduled batches at lower rates instead of on demand at premium rates. And keep data and compute in the same region and the same cloud wherever the mission permits, because every cross-region or cross-cloud move can carry a fee. Architecture choices made for convenience become recurring charges that nobody revisits, which is why the network bucket in your cost model should name the specific flows rather than showing a single monthly total.

Build, Buy, or Rent

For each capability, Eleanor faces a structural choice that drives long-term cost more than any tuning decision. The rule of thumb is rent to learn, build to scale. Renting is cheapest while demand is uncertain, because you pay only for what you use. Once a capability has high, steady, predictable volume, the per-query rental fee can exceed the cost of hosting your own model, and that crossover point is when building starts to pay. Eleanor kept her low-volume pilots on rented interfaces and moved only the high-volume fraud-detection model to owned infrastructure.

OptionWhat it meansBest whenWatch out for
Rent an interfacePay per use for someone else's modelVariable or uncertain demand; fast pilots; no custom training needed; data handling policy permits vendor processing under a documented agreementCosts scale with success; data residency limits; vendor lock-in
Buy a managed serviceVendor-operated platform you configure rather than runYou do not want to run infrastructure; the service meets your required authorization level; your ATO strategy can inherit the provider's controlsLicensing escalators; customization limits; the lock-in trade-off you accepted
Build on authorized cloudRun open or in-house models on cloud infrastructure you controlYou need elasticity; workloads are bursty; engineering is modest but capable; you want portability across cloudsYou own the platform work; portability costs architectural discipline
Build on premisesRun models on hardware your agency owns and housesData cannot leave the facility per statute or mission policy; sustained utilization above 70 percent; mature infrastructure engineering; a five-year authorized budgetLarge up-front and facilities cost; you own maintenance forever

The decision is rarely uniform across an agency portfolio, and it should not be. Most mature agencies run a mixed model: rent for generative assistants, buy for production vision and tabular machine learning, build on authorized cloud for custom fine-tuning, and build on premises for the workloads whose data genuinely cannot leave. The job of the architect is not to pick one answer. It is to make each choice explicit, tie it to a written criterion, and record the reasoning so that the decision can be revisited when the volume assumption behind it changes.

The FinOps Operating Model

FinOps is the operating model that keeps costs under control after the architecture decision is made, and it maps onto federal environments with some adaptation. It runs in three phases that repeat rather than complete. Inform makes spending visible and attributable. Optimize removes waste and buys down rates. Operate turns both into a governed monthly rhythm with named owners. Agencies that skip straight to optimize get a burst of savings and a slow return to the previous trajectory, because nothing they built keeps the attribution current.

Inform starts with tagging, and tagging is the whole game. Tag every resource with the program code, the appropriation, the capital planning investment identifier, the AI use-case identifier from your M-24-10 inventory, the environment, and the data sensitivity. Untagged resources are presumed waste and subject to monthly termination review, a policy that sounds harsh until you have watched an untagged cluster bill for a year. Ingest the cost and usage report into a data warehouse and build showback dashboards per program so that each mission owner sees their own number.

Optimize is the recurring hygiene described earlier, run on a schedule rather than in response to a bad invoice: monthly right-sizing against utilization data, cold data moved to archive, orphaned resources terminated, and commitment discounts applied to workloads with at least six months of stable history. The six-month rule matters. Committing on three months of a rising curve is how agencies end up with the over-reservation problem, and no amount of later optimization recovers money already obligated on a three-year term.

Operate is governance. Publish a monthly AI cost review co-chaired by the CIO and CFO, so the two people who can actually change something are in the same meeting. Set spending guardrails and budget alerts that fire before an overrun rather than after it. And route cost anomaly alerts into the security operations runbook, because a cost anomaly is frequently a security anomaly. Cryptocurrency mining on stolen credentials is the canonical example: it shows up first as an inexplicable compute bill, days before it shows up as anything else.

Instrument the workload so that each mission program can see the cost per inference, per training run, and per user. That is what makes an explainable trade-off possible: a program office that knows the unit cost of the premium model against the smaller one can decide for itself whether the accuracy difference is worth it on its own caseload. A program office that only sees an allocated share of a departmental total has no lever and no incentive, and will treat the whole subject as somebody else's budget problem.

Cost Failures You Will Recognize

Shadow accelerator clusters. A research group buys its own hardware under a single purchase, runs it in a closet, and the agency has no visibility into its cost, security posture, or utilization. GAO has flagged this pattern in multiple engagements. The mitigation is not a prohibition memo, which simply pushes the practice further underground. It is an agency-wide platform with a chargeback model good enough that research groups do not need to shadow-buy in the first place.

Egress sprawl. A mission program trains in a government cloud region and serves from an on-premises data center, so every inference pulls weights across the inspected boundary. At one civilian agency the first-month egress bill was $180,000. The mitigation is to co-locate training and serving, cache at the edge, or move the data lake to the cloud, and the time to choose is during architecture review, not after the first invoice makes the choice for you.

Unbounded vector databases. A retrieval pilot indexes every document in the agency's collaboration platform. The managed vector database bill grows linearly with the corpus while the marginal value of each additional document falls. The mitigation is tiered retrieval, decay by document age, and a corpus governance board that decides what gets indexed, because the alternative to curation is not neutrality. It is paying to search a decade of duplicated drafts.

Over-reserved capacity. One agency reserved 256 top-generation accelerators on a three-year term against a forecast assuming demand would double every year. Actual growth came in at 1.2 times per year. The lesson does not carry a figure for how much of that reservation sat idle, and you should not need one: the gap between doubling and 1.2 times compounds every year of the term, and the money was obligated up front. The mitigations are shorter terms, laddered commitments, and savings plans that apply across instance families rather than pinning you to one.

What Other Agencies Learned

Identity verification offers a unit-economics benchmark worth borrowing. After a 2022 controversy over a commercial identity-verification arrangement, the IRS pivoted to Login.gov. The cost model changed substantially, because Login.gov is a shared service operated by GSA with per-authentication pricing rather than a capacity-based private contract. Any agency planning to build identity AI in house should benchmark against that per-authentication economics before committing capital, if only to know what number it has to beat.

Bursty workloads reward elasticity, and the effect is larger than most cost models predict. The Environmental Protection Agency's ECHO platform handles very bursty analytical demand during quarterly enforcement cycles. A naive always-on reservation sized for the peak wastes 70 percent of its capacity between cycles. Moving to elastic cloud with interruptible and committed-discount capacity saved an estimated 45 percent while improving peak performance, which is the unusual case where the cheaper architecture is also the better one.

Storage-dominant workloads need a different playbook entirely. The Securities and Exchange Commission learned expensively on its consolidated audit trail that storage and query costs on a multi-petabyte event store dominate total cost of ownership, and that no amount of compute tuning addresses that. Columnar file formats, aggressive partitioning, and tiered storage are not optimizations at that scale. They are the difference between a system you can afford to query and one you can only afford to write to.

The Cost Governance Scorecard

Review every AI system on this scorecard quarterly. Its purpose is to turn the statement that the bill is too high into a specific, assignable action with an owner and a date. Run it as a standing agenda item in the monthly cost review rather than as an annual exercise, and record the answers, because the trend across quarters tells you more than any single row.

QuestionHealthy answerAction if not
Do we know the full cost of ownership across every bucket?Yes, updated quarterlyRebuild the model before the next budget cycle
Is every resource tagged to a program and a use case?Untagged spend near zeroRun a tag audit and a remediation plan
Is the model right-sized to the task?Smallest model that meets the quality barPilot a smaller model on simple queries, with quality measured
Is common output cached?Repeat queries served from cache, with invalidation tied to policy changesAdd caching for top queries
Is idle compute shut down off-hours?Auto-shutdown enabled on non-productionEnable scheduling this month
Are egress and region costs controlled?Data and compute co-locatedRe-architect the named data flows
Is the commitment portfolio reviewed?Reserved, committed, and on-demand mix rebalanced quarterlyList expiring commitments and decide renew, resize, or lapse
Does cost map to measured value?Each system has a value metricDefine a value metric or sunset the system
Is there a spend alert or budget cap?Alerts fire before overruns and route to the SOC as well as the CFOSet budget and anomaly alerts now

Two quarters after Eleanor rolled out the scorecard, the consolidated AI bill dropped from $2.3 million to $1.1 million while serving 40 percent more traffic. Just as importantly, she could explain every dollar to the appropriations committee. The value-metric row mattered most: two pilots that cost real money but had no value metric at all were quietly retired. Cost optimization is not only about spending less. It is about spending only on what works, and being able to prove which is which.

Anti-Patterns

  • Pricing a pilot on its subscription. The subscription is the part you noticed. Eleanor's chatbot subscription was 31 percent of the real cost; compute, egress, and two engineers were the rest. Cost every bucket, or the pilot's approved budget is a fiction from the day it is signed.
  • Cutting cost without measuring quality. Moving traffic to a smaller model, adding a cache, or batching work are all changes to a production system. Measure output quality on the same cases before and after, including the hard ones, and keep measuring as the query mix drifts. A saving with no quality evidence is an untested change.
  • Leaving resources untagged. You cannot govern what you cannot cost, and you cannot cost what you have not tagged. Untagged spend is not a reporting inconvenience; it is the mechanism by which a $4.2 million pilot invoices as ordinary cloud compute for a year.
  • Committing on a short, rising curve. Reserving multi-year capacity against an optimistic forecast obligates money that no later optimization recovers. Require at least six months of stable history before a commitment, ladder the terms, and prefer discounts that apply across instance families.
  • Splitting training and serving across the boundary. An architecture that trains in cloud and serves on premises looks like a clean separation of concerns and bills as egress every month forever. One agency's first month came to $180,000.
  • Indexing everything. A retrieval corpus that grows without curation costs linearly and returns less per document. Decide what gets indexed and what ages out, because the default is paying to search duplicated drafts.
  • Banning shadow purchases instead of serving them. A prohibition memo moves the closet cluster further out of view. A platform with credible chargeback removes the reason to buy one.
  • Treating a cost anomaly as only a cost problem. An inexplicable compute spike is frequently the first visible symptom of compromised credentials. Route anomaly alerts into the security runbook as well as the finance review.
  • Optimizing a system with no value metric. Making a worthless system cheaper is still spending. If no one can state what the system is for and how you would know it is working, the decision is retirement, not tuning.

Practice Prompts

  • Build the model. Construct a five-year cost of ownership model for your most expensive AI workload covering hardware, facilities, personnel, software, network, integration, compliance, and decommissioning. Compare on-premises, authorized cloud, managed service, and rented interface. Identify the utilization at which the ranking flips.
  • Run a tag audit. Export last month's cloud bill and classify every line item by program, environment, and AI use case against your inventory. Report the percentage of untagged spend and draft a remediation plan with a date.
  • Review the commitment portfolio. List every reservation and commitment expiring in the next twelve months. For each, decide renew, resize, or lapse, and document the rationale in language your CFO could defend in a GAO audit.
  • Write the decision memo. Pick one upcoming AI initiative and write two pages covering cost, risk, and schedule across build on premises, build on authorized cloud, buy managed, and rent. Route it through your CIO, chief AI officer, CFO, and general counsel.
  • Price the quality trade-off. For one production workload, measure the cost per inference and the output quality of the premium and the smaller model on the same evaluation set, including the hardest cases. State the trade-off in one sentence a program manager could act on.
  • Produce the plan. Write a one-page AI cost optimization plan identifying three actions you will take in the next 90 days to reduce AI infrastructure cost by at least 20 percent without degrading mission outcomes, with an owner and a measurement for each.

Reflection

Sit down with your agency's largest AI bill and try to do what Eleanor could not. Break it into the buckets, attribute each bucket to a program, and name the person who could reduce it. Where you cannot attribute a line, note whether the obstacle is tagging, contract structure, or the fact that nobody has ever been asked. That distinction determines whether the fix is technical, acquisition work, or a governance meeting, and agencies routinely apply the wrong one of the three.

Then ask the harder question about value. For each system on the list, write the sentence that says what it is for and how you would know it is working. The systems where that sentence does not come easily are not optimization targets. They are retirement candidates, and the reason agencies find that conclusion so uncomfortable is that somebody's approval memo is attached to each one. Eleanor retired two, and the appropriations conversation went better for it.

Glossary

  • Total cost of ownership. The sum of all direct and indirect costs of running a capability over its useful life, including the buckets that never appear on the subscription invoice.
  • Inference. Running a trained model to answer a query. For most deployed government AI it is the larger ongoing cost, because you train once and run forever.
  • Egress. The fee a cloud provider charges to move your own data out of its environment, billed per volume and recurring for as long as the architecture requires the movement.
  • FedRAMP. The Federal Risk and Authorization Management Program, which sets the security baseline for cloud services used by federal agencies. Compliant environments cost more than the cheapest commercial tier.
  • FinOps. An operating model for cloud spending with three repeating phases: inform, which makes spend visible and attributable; optimize, which removes waste and buys down rates; and operate, which governs both on a monthly rhythm.
  • Showback and chargeback. Showback reports each program its own consumption; chargeback bills it. Showback changes behavior through visibility, chargeback through consequence.
  • Tagging taxonomy. The set of labels applied to every resource, typically program, appropriation, investment identifier, use-case identifier, environment, and data sensitivity, which makes attribution possible at all.
  • Reserved and committed capacity. Discounted rates purchased against a commitment to a term or a spend level. The discount is real and so is the obligation, which becomes sunk cost if the mission shifts.
  • Right-sizing. Matching the model or the instance to the difficulty of the task rather than to the most capable option available, measured against a stated quality bar.
  • Unit economics. Cost expressed per inference, per training run, or per user, which is what lets a program office make an explainable trade-off between accuracy and price.

Closing

The technical levers in this lesson are real and they work, but they are not what separates agencies that control AI cost from agencies that do not. The separator is attribution. An agency that can say what each system costs, which program consumes it, and what value it returns can make every other decision on evidence: which model to run, which commitment to sign, which pilot to retire, and which number to take to an appropriations committee. An agency that cannot is managing a total, and a total offers no lever at all.

Eleanor's real problem was never the $2.3 million. Large numbers are defensible when you can explain them, and a growing bill for a service residents are using more is a success story if you can show the unit cost falling. Her problem was that eleven pilots had been approved separately, billed jointly, and tagged not at all, so the only fact available was the total. Build the attribution first. The savings follow, and the ones that matter most are the two pilots you discover nobody can justify.

Key Takeaways

  • AI costs are metered, not fixed. Treat the bill like a utility: it scales with usage, success itself drives spend, and popularity is a budget event nobody plans for.
  • Count every bucket. The five-line shorthand is compute, data, model access, people, and governance; the fuller federal model adds hardware, facilities, software, network, integration, and compliance with decommissioning. Eleanor's tracked subscription was 31 percent of the real cost.
  • Tagging is the precondition for everything else. You cannot govern what you cannot cost, and you cannot cost what you have not tagged. Untagged spend is how a $4.2 million pilot invoices as ordinary cloud compute.
  • The on-demand rate is the worst rate. A top-end training instance listed around $98 per hour in 2026, roughly $860,000 a year if left running, against roughly $480,000 on a three-year fully prepaid reservation. A 40 percent reserved, 40 percent committed, 20 percent on-demand mix, rebalanced quarterly, keeps the discount without the lock-in.
  • Right-sizing is the biggest lever, and it needs a quality measurement. Moving 60 percent of simple queries to a smaller model cut Eleanor's chatbot compute cost by 48 percent with no measurable quality drop on the categories evaluated. Keep evaluating as the query mix drifts.
  • Cache, shut down idle compute, and control egress. Repeated queries, weekend accelerators, and cross-region data moves quietly inflate every AI bill. One agency's split training and serving architecture billed $180,000 of egress in its first month.
  • Rent to learn, build to scale. Renting wins while demand is uncertain; owning pays off past the crossover point, which for on-premises builds means sustained utilization above 70 percent, data that cannot leave, mature engineering, and a five-year authorized budget.
  • Tie cost to measured value. Every system needs a value metric. A system nobody can justify should be retired rather than optimized, and a quarterly scorecard with budget and anomaly alerts is what surfaces which is which.

Frequently Asked Questions

Our pilot is small. Do we really need a cost model for it? The pilots are exactly where this goes wrong, because each one is approved on a number that counts only the subscription and none of the people, egress, or compliance work. Eleanor had eleven of them, each individually defensible and collectively unexplainable. Build the model at pilot scale, when it takes an afternoon, rather than at portfolio scale after an appropriations committee has asked a question you cannot answer.

Is on-premises cheaper than cloud for AI? Only under conditions you should check rather than assume: sustained utilization above 70 percent, mature infrastructure engineering, a five-year authorized budget, and usually a mission or statutory reason the data cannot leave the facility. The costs that sink on-premises models are facilities and people, not hardware. A cooling and power project can add six to twelve months and $1 million to $3 million before a single model trains, and a ten-person platform team runs about $2.2 million a year whether or not anything is running on it.

How do we know a smaller model is safe to switch to? By measuring, on your own cases, before you switch and continuously after. Build an evaluation set from real traffic that deliberately includes the hard, unusual, and consequential queries rather than a random sample dominated by easy ones. Then set a quality bar in advance and hold the switch to it. The 48 percent saving in this lesson came with an evaluation attached; without one, the same change is a cost reduction and an unmeasured risk transfer onto whoever receives the output.

Who should chair the monthly cost review? The CIO and the CFO together. Either one alone produces a meeting with no authority over half the problem: the CFO cannot re-architect a data flow and the CIO cannot move an appropriation. Bring the chief AI officer for the use-case inventory and the chief data officer for storage and movement. And put the security operations lead on the distribution for anomaly alerts, because the first sign of compromised credentials is often an unexplained compute bill.

What do we do about a three-year commitment we now regret? Take the lesson rather than the loss. The obligated money is gone and no optimization recovers it, so the useful work is preventing the next one: require six months of stable history before committing, ladder terms so they do not all expire together, prefer discounts that apply across instance families, and review the commitment portfolio quarterly. Document the reasoning each time, because a reviewer will eventually ask why the forecast and the outcome diverged.