Cloud GPU Pricing Comparison 2026: AWS vs Azure vs GCP
Key Takeaways What each provider offers, what their GPU instances actually cost, and how far their native tools go in helping you manage the bill. Ask what an H100 costs per hour on the major clouds and cloud GPU pricing gives you numbers between $5.19 and $12.29. Both are real, both come from the providers’ […]
The newest GPU is the cheapest. GCP’s B200 node lists at $8.06 per GPU-hour, 27% below its own H100 node. Defaulting to H100 out of habit is paying a premium for older silicon.
The headline rates aren’t comparable. AWS’s $5.19 is prepaid reserved capacity; GCP’s $11.06 and Azure’s $12.29 are metered on-demand. AWS publishes no directly comparable on-demand rate for these families.
Utilization beats provider choice. A node bills at 100% whether it’s working or idle, so at 40% utilization an $11.06 GPU-hour really costs $27.65. No rate negotiation comes close to that.
Two SKU traps on Azure. The non-InfiniBand H100 variant costs 10% less than the InfiniBand one, and the same H200 SKU varies 30% between East US 2 and West US.
Spot and Trainium sit well below everything else. GCP spot H200 at $4.65 per GPU-hour undercuts even AWS’s prepaid H100 reservation, and Trainium2 at $2.24 is 6.3x cheaper than a B300.
What each provider offers, what their GPU instances actually cost, and how far their native tools go in helping you manage the bill.
Ask what an H100 costs per hour on the major clouds and cloud GPU pricing gives you numbers between $5.19 and $12.29. Both are real, both come from the providers’ own pricing pages, and comparing them directly is close to meaningless.
The gap is not a discount. It is a different product, and it is the single biggest trap in cloud GPU pricing.
This cloud GPU pricing guide goes provider by provider: which GPU instance families each one offers, what they cost, how you can buy them, and what each platform gives you natively for viewing and controlling that spend. Then it covers the fixed-cost alternative, and where all three native toolsets stop.
If you want a number against your own usage first, our GPU pricing page tracks live rates across providers and the GPU pricing calculator turns them into a figure for your own hours.
The short answer
An H100 GPU-hour costs roughly $5.19 on AWS through prepaid Capacity Blocks, $11.06 on GCP on demand, and $12.29 on Azure on demand. AWS looks far cheaper, but Capacity Blocks are reserved and paid up front rather than metered.
Every one of those nodes bills at 100% whether the silicon is working or idle, so your effective rate is the list rate divided by your utilization. At 40% utilization, an $11.06 GPU-hour actually costs you $27.65.
AWS cloud GPU pricing
AWS has the widest accelerator range of the three and the only first-party silicon, which makes its cloud GPU pricing the hardest to read at a glance.
Two things the table below does not show. AWS lists P6e UltraServer configurations reaching 72 accelerators for the largest training jobs, and the Trn1 and Trn2 families run on Trainium, AWS’s own silicon rather than NVIDIA’s.
| Instance | GPUs | Node price/hr | Per accelerator-hr |
| p4d.24xlarge | 8x A100 | $11.80 | $1.48 |
| trn2.48xlarge | 16x Trainium2 | $35.76 | $2.24 |
| p5.48xlarge | 8x H100 | $41.53 | $5.19 |
| p5e.48xlarge | 8x H200 | $47.76 | $5.97 |
| p5en.48xlarge | 8x H200 | $54.92 | $6.87 |
| p6-b200.48xlarge | 8x B200 | $98.84 | $12.36 |
| p6-b300.48xlarge | 8x B300 | $112.32 | $12.36 |
Those are Capacity Blocks rates, which is the detail most comparisons drop. Capacity Blocks are a prepaid capacity product: you reserve a window of instance-hours and pay for it up front. AWS’s own example makes the consequence plain, reserving 192 instance-hours and using only 188 still bills all 192. You are buying a window, not consumption.
Two things sit on top. Operating system charges are separate, so RHEL adds $1.84 per instance-hour. And AWS states that reservation prices update regularly with supply and demand, with the next revision scheduled for October 2026.
Trainium2 at $2.24 per accelerator-hour is the cheapest thing on this board, under half an H100 and 5.5x cheaper than a B300. It runs AWS’s own toolchain, so check framework support before planning a migration.
How AWS helps you manage it. The native stack runs through Cost Explorer for spend by service, account, and tag, AWS Budgets for thresholds and alerts, Cost Anomaly Detection for unusual spend, and the Cost and Usage Report for hourly resource-level detail. Compute Savings Plans and Spot cover the discount side. It is a capable toolset, and it stops at the AWS boundary. Check the current feature set in AWS’s own billing documentation, since these services change frequently.
Azure cloud GPU pricing
Azure splits GPU compute into training-class and inference-class families, which matters more for budgeting here than on the other two.
Two things the table below does not show. ND is the training line, connected over InfiniBand for distributed work per Microsoft’s own ND H100 v5 documentation, while NC is the inference line. And Microsoft updates the ND family regularly, so check the current VM size list for newer generations alongside the rates.
| Instance | GPUs | Node price/hr | Per GPU-hr |
| ND96isr_H100_v5 | 8x H100 | $98.32 | $12.29 |
| NC-series H100 | 1x H100 | approx $6.98 | approx $6.98 |
Those figures are consistent across multiple third-party trackers, and the ND monthly equivalent matches an independent listing to the cent, but I could not confirm either on a Microsoft first-party pricing page. Verify in the Azure pricing calculator for your region before committing.
The practical trap on Azure is the node minimum. There is no single-GPU ND instance, so teams price from a per-GPU figure, provision an eight-GPU node, and receive an invoice eight times larger. If you only need one card, NC-series is the right family, though it drops NVLink and InfiniBand and is unsuitable for distributed training.
How Azure helps you manage it. Cost Management and Billing covers cost analysis, budgets, and alerts, with Azure Advisor surfacing rightsizing and idle-resource recommendations, and reservations and savings plans handling commitment discounts. As with AWS, the view is complete within Azure and blank outside it.
GCP cloud GPU pricing
Google’s naming is the most readable of the three, and its spot discounts are the steepest.
Two things the table below does not show. A2 is available in single-GPU shapes, which neither of the other two offer at the top of their training range, and Google adds newer accelerator families over time, so check the current list alongside the rates. We break the Google side down further in our GCP GPU pricing comparison.
| Instance | GPUs | On-demand/hr | Per GPU-hr | Spot per GPU-hr |
| a2-ultragpu-1g | 1x A100 80GB | $5.07 | $5.07 | n/a |
| a3-ultragpu-8g | 8x H200 | $84.81 | $10.60 | $5.30 |
| a3-highgpu-8g | 8x H200 | $88.49 | $11.06 | $4.79 |
| a3-megagpu-8g | 8x H100 | $93.40 | $11.68 | $5.04 |
Two findings from Google’s own accelerator pricing. First, the newer GPU is cheaper: the H200 node lists below the H100 node, so H200 comes out 4.2% cheaper per GPU with 141GB of memory instead of 80GB. Second, spot H100 capacity at $4.79 undercuts AWS’s prepaid Capacity Blocks rate, which reorders the headline entirely. The catch is eviction at very short notice, fine for checkpointed training, useless for live traffic. Google also publishes one and three year committed use rates on the same page.
How GCP helps you manage it. Cloud Billing reports break spend down by project, service, and label, with budgets and alerts on top, the Recommender surfaces idle and rightsizing opportunities, and committed use discounts handle commitments. Our GCP pricing calculator models Compute Engine configurations including commitment plans and Spot before you deploy.
OpenMetal: the fixed-cost alternative
All three hyperscalers meter by the hour, which means you pay a premium for elasticity whether or not you use it. If your GPUs run continuously, that premium buys you nothing.
OpenMetal takes the other approach: single-tenant bare metal H200 and NVIDIA RTX PRO 6000 servers on fixed monthly pricing with no metered hours, included egress, full root access, and multiple servers interconnected over a private 20 Gbps network for clusters. For sustained training or always-on inference, fixed monthly removes both the idle-hour premium and the egress line, and single tenancy removes the noisy-neighbour variance that makes multi-tenant inference latency unpredictable.
The trade runs both ways. Fixed cost wins on sustained load and loses on spiky, scale-to-zero work, where metered hourly is exactly the right instrument. OpenMetal quotes GPU pricing rather than publishing it, though they honour written quotes for 30 days, so model it against your own hours. Our GPU pricing calculator works that crossover point out from the rates tracked across providers. Here is a comparison guide against AWS if you want OpenMetal’s own framing of where that break-even sits.
The idle silicon tax, and where native tools stop
A GPU node bills at 100% from the moment it boots. It does not know whether your job is running or your data loader is starving the card. So the number that matters in cloud GPU pricing is the list rate divided by utilization:
| Utilization | Effective cost per useful H100-hour (GCP list $11.06) |
| 100% | $11.06 |
| 70% | $15.80 |
| 50% | $22.12 |
| 40% | $27.65 |
| 25% | $44.24 |
At 40% utilization you are paying more per useful GPU-hour than the most expensive list rate in this guide. No provider switch fixes that.
Each native toolset described above is genuinely capable inside its own cloud. Three things none of them do. They cannot show you AWS, Azure, and GCP GPU spend in one view, so a multi-cloud estate needs three consoles and a spreadsheet. They report by account, project, and tag rather than by model, training run, or customer. And they do not reconcile against the AI provider spend sitting in the same budget, which our guides to Anthropic cost monitoring and the LLM cost calculator cover separately.
How to cut cloud GPU pricing and spend
Six levers move cloud GPU pricing more than provider choice does.
- Measure utilization before you negotiate rates. A 10% discount on a node running at 40% is worth less than moving that node to 70%.
- Match the purchase model to the workload. Capacity Blocks for scheduled runs, on-demand for bursty work, Spot for interruptible jobs, fixed-cost bare metal for sustained load.
- Right-size the accelerator. Inference rarely needs a training-class GPU, and Trainium2 costs a fraction of an H100.
- Kill idle clusters automatically. Forgotten resources are the largest source of GPU waste and are invisible in any rate comparison.
- Account for egress early. All three hyperscalers meter data leaving the network per GB, billed separately from compute.
- Track GPU spend against output. Cost per training run is the only number that tells you whether the spend worked.
Our cloud cost optimization guide covers the broader practice, and top cloud cost management tools compares the platforms.
How Economize tracks GPU spend
GPU instances are among the hardest lines on a cloud bill to follow. They arrive under generic compute SKUs, a p5.48xlarge looks like any other EC2 row until you know the instance family, and the utilization that decides your effective rate lives in a monitoring tool rather than a billing one. Across three providers that becomes three formats and three consoles.
Economize connects to all three agentlessly and resolves that into one view. Concretely, it tracks GPU spend like this:
- It finds the GPU lines. Explorer surfaces individual instances at resource level across AWS, Azure, and GCP, so a p5, an ND96isr, and an a3-highgpu sit in the same list rather than three separate billing consoles.
- It attributes them to something meaningful. Virtual tagging applies allocation rules across your existing tags, including shared and untaggable costs, so a cluster resolves to a team, model, or training run instead of an account number.
- It identifies what is idle. Recommendations flag idle and oversized instances with an estimated saving attached, which is how the utilization gap above becomes a number someone can act on.
- It catches a spike in the week it happens. Anomaly detection with root cause analysis surfaces the cluster nobody spun down, rather than leaving it for month end.
- It tells whoever owns the workload. Alerts and scheduled reports route to Slack, Teams, Google Chat, or Discord.
- It sits beside your AI spend. OpenAI and Anthropic usage costs land in the same dashboards, so GPU infrastructure and model API spend are compared in one place rather than reconciled by hand.
Setup takes about five minutes with no agents to deploy. Pricing is published and flat rather than a share of spend: free up to $100,000 a month in combined cloud and AI spend, $249 up to $250,000, and enterprise from $2,499 with SSO, RBAC, and audit logs. Economize is SOC 2 certified, and customers including Hasura, DeepSource, Capchase, and Apna report savings up to 30%.
Frequently Asked Questions
On published rates AWS is lowest for H100 at $5.19 per GPU-hour, but that is prepaid Capacity Blocks rather than on-demand. Among on-demand rates GCP’s a3-highgpu-8g at $11.06 undercuts Azure’s ND96isr_H100_v5 at $12.29, and GCP spot at $4.79 is cheaper than all of them. Our GCP GPU pricing comparison goes deeper on the Google side.
Roughly $5.19 on AWS through Capacity Blocks, $11.06 on GCP on demand, and $12.29 on Azure on demand. All three are eight-GPU nodes priced at $41.53, $88.49, and $98.32 respectively.
Usually utilization. A node bills at 100% whether the GPUs are working or idle, so at 40% utilization your effective cost per useful GPU-hour more than doubles. Egress, separate OS charges, persistent storage, and node minimums add the rest.
It depends on utilization. Fixed monthly bare metal from providers such as OpenMetal removes the metered-hour premium and bundles egress, which wins on sustained training or always-on inference. Metered hourly wins on spiky, scale-to-zero workloads.
Trainium2 at $2.24 per accelerator-hour, under half the H100 rate and 5.5x cheaper than a B300. It uses AWS’s own toolchain, so check framework support first.
The cheapest GPU-hour is the one you actually use. Connect your accounts to Economize and see GPU spend across AWS, Azure, and GCP in one place.
Simran Sardar
FinOps enthusiastProduct Manager at Economize with over 3 years of experience, focused on FinOps strategies and cloud cost optimization. Dedicated to helping organizations streamline cloud expenses and drive financial efficiency.
Maximize Cloud Efficiency and Optimize Costs
Get started free in our sandbox or book a personalized call with our experts






















