It makes considerably less sense once the workload is running 24 hours a day and the finance team starts asking why the AI line item keeps outgrowing the AI budget.
The rent-first default is starting to show its age
Cloud-first was the correct call for a pilot, that math changes once usage settles into steady, round-the-clock demand. Recent industry surveys suggest a significant share of large organizations are now spending well into eight figures annually on public cloud with little elasticity benefit to show for it.
The utilization number that flips the decision
Multiple independent TCO analyses converge on a similar threshold: once a GPU cluster runs above roughly 60 to 70% utilization, owned infrastructure starts beating on-demand cloud pricing, and the gap widens fast from there.
On-demand H100 rates have ranged from about $2.25 an hour on specialized GPU providers to $6.88 or more on major hyperscalers through mid-2026. Lease pricing rose almost 40% between October 2025 and March 2026 as inference demand outpaced supply.
Idle capacity is expensive either way, and the difference is simply who eats the cost.

What the TCO analysis found
The analysis found cumulative cloud costs running four to six times higher than the owned equivalent for steady, high-utilization workloads.
The exact multiple will shift with your own hardware mix, contract terms, and workload profile; the report's assumptions are worth checking against your own numbers before treating the multiple as a given.
The direction still lines up with independent analysis, which puts on-premise infrastructure at 30 to 40% cheaper than cloud over five years at high utilization.
While methodologies differ across studies, the broader conclusion remains consistent: utilization is one of the primary drivers of AI infrastructure economics.
Repatriation stopped being a fringe opinion
Gartner separately projects that by 2028, over 40% of leading enterprises will have folded hybrid computing architectures into critical business workflows, up from just 8% today, a broader shift that touches the cloud-versus-on-prem question among several others.
IDC research from around the same period puts that in context: only 8 to 9% of companies planned to repatriate workloads in full. Most of this movement is happening workload by workload, not as a wholesale exit from the cloud.
The pattern that shows up across nearly every serious analysis:
- Training runs stay elastic. Bursty, days-to-weeks compute at massive scale benefits from cloud capacity that few enterprises could justify owning outright, so training tends to stay put or run hybrid with an on-prem baseline.
- Steady inference moves home. Production inference on a stable model, running continuously against predictable traffic, is precisely the workload where owned hardware wins on cost and where regulated industries want the data control anyway.

The costs a spreadsheet misses on the first pass
Egress fees are the quiet line item that turns a clean cloud estimate into a surprise. Moving training data and inference outputs across the cloud-to-on-prem boundary repeatedly can compound into a real percentage of total spend.
Add fast-storage premiums for high-throughput inference and the on-demand price on the vendor's homepage stops describing what actually lands on the invoice.
Owning has its own hidden costs, and pretending otherwise sets the CIO up for a rough budget conversation later:
Utilization risk sits entirely on the owner. A large-scale GPU deployment operating well below expected utilization may struggle to deliver the anticipated return on investment.
AI infrastructure moves on a fast cycle. NVIDIA has publicly committed to an annual cadence for its data center platforms, with each generation bringing meaningful gains in performance and efficiency. For owners, keeping pace means budgeting for refresh cycles from day one.
Renting can offer more flexibility here, but only if the contract allows it: on-demand users can move to newer instances as they launch, while reserved and committed-spend agreements often tie you to specific hardware for the term.
The framework that actually holds up
Sort workloads by two variables before anything else: how steady the utilization is, and how sensitive the data is. Steady and sensitive points toward owned infrastructure or colocation. Bursty and low-sensitivity points toward staying in the cloud.
Everything else lives in the hybrid middle, which is where most enterprises will spend the next several years regardless of which vendor's slide deck they read first.
Where to get the actual numbers
The utilization thresholds and cost multiples above are directional. Building an actual business case means running the same modeling against your own cluster size, workload mix, and contract terms, which is precisely what analyst Dan Olds set out to do in detail.
- Full 3- and 5-year cost breakdowns, covering hardware, power, cooling, staffing, and support across on-premises deployment and AWS, Google Cloud, and Oracle Cloud.
- A 248-GPU B200 case study, modeled at matched performance so the comparison holds up against real procurement decisions rather than back-of-envelope math.
- Transparent assumptions, built on real pricing rather than spot instances or best-case scenarios propping up the numbers.
Get the full modeling in Dan Olds's report, The Real Cost of AI Infrastructure.

