Buy or Rent GPUs? The 2026 Break-Even Math

Buy or Rent GPUs? The 2026 Break-Even Math

A single NVIDIA B200 lists somewhere between $30,000 and $40,000 as of mid-2026 — and the same GPU rents for roughly $4 to $7 per hour depending on where you shop. Divide one number by the other and you get the question every AI budget meeting now turns on: at what point does renting silicon stop making sense? I have sat through enough of these CFO-meets-CTO sessions to know the answer is not “it depends.” It is a utilization number, and you can calculate it this quarter.

This brief lays out the 2026 break-even math, where AWS and Google Cloud actually price Blackwell capacity, what the buy path really costs once you count lead times, and a decision rule you can defend in front of finance. One warning before we start: most of what ranks in search for this question is written by GPU clouds selling one side of the answer. Treat their calculators accordingly.

What changed

Blackwell went from paper launch to rentable reality. AWS now sells EC2 P6-B200 instances — eight B200s per node with 1,440 GB of GPU memory — in multiple US regions, plus P6e-GB200 UltraServers for rack-scale Grace Blackwell. Google Cloud’s A4 VMs put the same HGX B200 board on-demand at roughly $4.30 per GPU-hour at list. Meanwhile H100-class capacity, the workhorse of 2024–25, has become a buyer’s market: specialty clouds advertise it under $2/hour while hyperscaler list pricing still sits in the $4–8 range.

The other change is on the buy side. Blackwell supply improved through 2026, but enterprise orders still quote 6–12 month lead times for full systems once you include networking, power provisioning, and integration. That lag is a cost, and most buy-vs-rent models quietly leave it out.

The break-even math

The verdict first: at sustained utilization of roughly 60–70%, owning beats renting on a three-year horizon. Below that, rent. The table shows the shape of the market as of mid-2026 — list pricing varies, so run your own numbers with your negotiated rates.

Rent — specialty GPU cloudRent — hyperscalerBuy — on-prem or colo
H100-class, per GPU-hour~$1.50–3.00~$4–8 (list)~$1.30–2.00 effective at 70%+ utilization
B200-class, per GPU-hour~$4–7 on-demand~$4.30 (Google Cloud A4 list) and up~$2.20–2.80 effective at full utilization
Upfront capitalNoneNone (commits optional)$30–40K per B200 GPU + facility costs
Time to capacityHours to daysHours (quota permitting)6–12 months for full systems
Best fitBursty training, experimentsData-gravity and compliance-adjacent workSteady-state inference, 24/7 pipelines

Here is the arithmetic behind the crossover. A B200 at roughly $35K amortized over three years is about $1.33 per hour of raw depreciation. Add power, cooling, networking, colo space, and operations staff and the fully loaded figure lands near $2.20–2.80 per hour — if the card is busy every hour. At 60% utilization that effective cost climbs to roughly $3.70–4.60, which is exactly where Blackwell rental pricing sits. That is the break-even. Our on-prem GPU cluster cost guide walks through the facility-side line items most spreadsheets miss.

Why it matters

Because the default is drifting. In 2024 the safe answer was “rent everything — the hardware cycle is too fast to own.” In 2026 that reflex is costing real money for anyone running production inference around the clock. An inference fleet at 80% utilization on rented hyperscaler capacity can pay for equivalent owned hardware in under 18 months. Boards have noticed, and finance teams are asking why AI compute is 100% opex when the workload profile looks like a steady utility.

The reverse mistake is just as expensive. Teams that bought H100 clusters in 2024 for “future training needs” and ran them at 25% utilization effectively paid $6–8 per GPU-hour for capacity they could have rented for $3. Utilization, not unit price, decides this argument. The takeaway a VP can repeat: rent your spikes, own your baseline.

Renting: AWS, Google Cloud, and the neoclouds

AWS is the depth play. P6-B200 instances slot into the same VPC, IAM, and data-platform surface your teams already run, and Capacity Blocks let you reserve GPU windows for defined training runs instead of paying on-demand rates around the clock. The catch is price: AWS GPU list pricing runs well above specialty clouds on a per-GPU-hour basis, and the gap only closes with Savings Plans or negotiated commits. If your training data already lives in S3 and your security team has blessed the account structure, that premium buys real friction reduction. If not, you are paying hyperscaler rates for undifferentiated silicon.

Google Cloud is currently the sharpest hyperscaler pencil on Blackwell. A4 VMs at roughly $4.30 per GPU-hour at list undercut comparable AWS on-demand pricing meaningfully, and Dynamic Workload Scheduler plus spot capacity can push effective rates lower for interruptible training. Google also gives you an escape hatch NVIDIA-only shops lack: TPU capacity for workloads that fit it. The weakness is the familiar one — smaller enterprise footprint, and quota negotiations that favor large committed spenders. Google Cloud fits teams that treat compute as a market and shop it; it fits less well where the org is contractually welded to another cloud.

Below both sit the neoclouds — CoreWeave, Lambda, and a long tail — renting H100-class capacity at $1.50–3.00 per hour. Real savings, real trade-offs: thinner enterprise support, variable data-center quality, and contract terms that reward diligence.

Buying: the NVIDIA hardware path

NVIDIA sets the terms on the buy side, and its interests are not subtle: it sells to you, to AWS, to Google, and to every neocloud simultaneously. A B200 at $30–40K list is only the opening line item. HGX baseboards, NVLink switching, InfiniBand or Ethernet fabric, and 10kW+ per-node power budgets typically push a deployed cluster to 1.6–2x the GPU line. NVIDIA’s strength for buyers is the software moat — CUDA, NIM, and enterprise support make owned hardware genuinely productive on day one. The weakness is that you are buying at the top of a fast product cadence: Blackwell Ultra and the Rubin generation are already on NVIDIA’s public roadmap, which compresses the resale value of whatever you rack today.

Model the 6–12 month lead time as a cost, not a footnote. If you must rent B200 capacity at $5/hour while your purchased cluster clears procurement, facilities, and bring-up, a 64-GPU stopgap can add seven figures to the “buy” column before your hardware serves a single token. For inference-heavy shops, the buy decision also interacts with the build-vs-API question — our self-hosted LLM vs API cost analysis covers when owning the model layer pays.

Honest downsides on both sides

  • Renting: price volatility (Blackwell on-demand rates have moved double-digit percentages within a year), quota ceilings at exactly the moment everyone wants capacity, egress charges on training data, and the quiet ratchet of commits that turn “flexible opex” into a three-year contract anyway.
  • Buying: capital tied up in a depreciating asset on a roughly 18-month product cadence, lead-time exposure, the staffing reality that a GPU cluster needs datacenter and MLOps skills you may not have, and utilization risk — the whole case collapses if the workload you bought for gets cancelled.

What to do about it

Instrument first. Pull 90 days of actual GPU utilization before any procurement conversation. Then apply the rule: workloads sustained above ~70% utilization go on owned or colo hardware; anything below ~50% stays rented; the 50–70% band is your negotiation zone for reserved capacity and committed-use discounts. Rent Blackwell for training bursts and evaluation runs. Buy — or lease through colo — for the inference baseline you can forecast two years out. Re-run the math every two quarters, because rental pricing in this market does not sit still. And when a vendor’s calculator tells you their side wins, check whose logo is on the spreadsheet.

Frequently asked questions

Is it cheaper to buy or rent GPUs for AI?

It depends on utilization, not preference. As of mid-2026, owning wins on a three-year horizon once sustained utilization passes roughly 60–70%. Below that, rental pricing — especially H100-class capacity under $3/hour — is hard to beat.

How much does it cost to rent a B200 GPU per hour?

On-demand B200 pricing spans roughly $4–7 per GPU-hour as of mid-2026. Google Cloud’s A4 VMs list near $4.30 per GPU-hour; specialty clouds and spot capacity can run lower, hyperscaler on-demand can run higher.

How much does an NVIDIA B200 cost to buy?

List pricing runs roughly $30–40K per GPU, but a deployed cluster — baseboards, fabric, power, cooling, integration — typically lands at 1.6–2x the GPU line item, with 6–12 month lead times for full systems.

Should I use AWS or Google Cloud for GPU workloads?

Google Cloud currently prices Blackwell more aggressively at list; AWS offers deeper enterprise integration and reservation mechanics like Capacity Blocks. If your data and security posture already live on one of them, that gravity usually outweighs the list-price gap.

Enterprise Techie publishes vendor-honest analysis like this daily — get the brief by email, free.