# Cloud AI GPU Pricing Comparison

> Source: https://aiwiki.ai/wiki/cloud_gpu_pricing_comparison
> Updated: 2026-08-01
> Fact-checked: 2026-08-01
> Categories: AI Hardware, AI Infrastructure, Data Centers
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "Cloud AI GPU Pricing Comparison." aiwiki.ai, 1 Aug 2026. https://aiwiki.ai/wiki/cloud_gpu_pricing_comparison
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

Cloud AI GPU pricing is difficult to compare from a headline hourly rate alone. Providers sell different GPU variants, node sizes, CPU and memory bundles, regions, network configurations, and purchase models. The tables below therefore present a dated evidence snapshot, not a ranking. The main comparison uses provider pages captured between July 11 and July 27, 2026, before the research cutoff of July 28, 2026 at 23:59:59 UTC+7.[1][2][3][4]

The snapshot covers published US-dollar rates for NVIDIA [H100](https://aiwiki.ai/wiki/nvidia_h100), [H200](https://aiwiki.ai/wiki/nvidia_h200), and [B200](https://aiwiki.ai/wiki/nvidia_b200) offers from [RunPod](https://aiwiki.ai/wiki/runpod), [Together AI](https://aiwiki.ai/wiki/together_ai), [Nebius](https://aiwiki.ai/wiki/nebius), and [CoreWeave](https://aiwiki.ai/wiki/coreweave). Negotiated quotes, credits, taxes, managed inference, serverless endpoints, and non-NVIDIA accelerators are outside the numerical comparison.

## Method and units

Every numerical row is tied to a specific provider, GPU label, purchase model, billing object, and evidence date. The normalized figure is calculated as:

`normalized GPU-hour rate = published node or instance hourly rate / GPUs in that node or instance`

No division is needed when a provider already states that its rate is per GPU per hour. When division is used, the published system price and GPU count remain visible so that the arithmetic can be checked. A normalized GPU-hour is only a price unit. It does not make an eight-GPU HGX node equivalent to a one-GPU virtual machine.

This separation follows the broader cost-accounting distinction between a list unit price, the pricing category, the pricing unit, and effective cost. The FinOps Open Cost and Usage Specification, or FOCUS, treats these as separate fields and also separates list, contracted, billed, and effective cost.[5] NIST's definition of cloud computing includes measured service as an essential characteristic, but a provider's meter does not by itself describe workload performance or total cost.[6]

For this article, "not stated" means the cited pricing evidence did not specify that property. It does not mean that the provider lacks the feature. Prices do not prove that capacity was available for immediate deployment on the evidence date.

## Cutoff-captured price snapshot

All rates in this table are provider-published dollar-denominated rates. The source date is the date of the archived provider page. Values rounded to two decimals are marked as calculations.

| Provider and offer | GPU label | GPU count used for normalization | Purchase model | Published rate | Normalized USD/GPU-hour | Source date |
|---|---|---:|---|---:|---:|---|
| RunPod Secure Cloud Pods | H100 SXM, 80 GB | 1 | Secure Cloud Pod hourly rate | $2.99/GPU-hour | $2.99 | 2026-07-11 [1] |
| RunPod Secure Cloud Pods | H200, 141 GB | 1 | Secure Cloud Pod hourly rate | $4.39/GPU-hour | $4.39 | 2026-07-11 [1] |
| RunPod Secure Cloud Pods | B200, 180 GB | 1 | Secure Cloud Pod hourly rate | $5.89/GPU-hour | $5.89 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | H100 SXM, 80 GB | 1 | Community Cloud Pod hourly rate | $2.69/GPU-hour | $2.69 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | H200, 141 GB | 1 | Community Cloud Pod hourly rate | $3.59/GPU-hour | $3.59 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | B200, 180 GB | 1 | Community Cloud Pod hourly rate | $5.89/GPU-hour | $5.89 | 2026-07-11 [1] |
| Together GPU Clusters | NVIDIA HGX H100 | Not divided; provider states per GPU | On-demand | $3.99/GPU-hour | $3.99 | 2026-07-18 [2] |
| Together GPU Clusters | NVIDIA HGX H200 | Not divided; provider states per GPU | On-demand | $5.99/GPU-hour | $5.99 | 2026-07-18 [2] |
| Together GPU Clusters | NVIDIA HGX B200 | Not divided; provider states per GPU | On-demand | $8.19/GPU-hour | $8.19 | 2026-07-18 [2] |
| Nebius GPU instance | NVIDIA HGX H100 | Not divided; provider states per GPU | On-demand | $3.85/GPU-hour | $3.85 | 2026-07-12 [3] |
| Nebius GPU instance | NVIDIA HGX H200 | Not divided; provider states per GPU | On-demand | $4.50/GPU-hour | $4.50 | 2026-07-12 [3] |
| Nebius GPU instance | NVIDIA HGX B200 | Not divided; provider states per GPU | On-demand | $7.15/GPU-hour | $7.15 | 2026-07-12 [3] |
| CoreWeave North America | NVIDIA HGX H100, 80 GB | 8 | On-demand node | $49.24/node-hour | $6.16 calculated | 2026-07-27 [4] |
| CoreWeave North America | NVIDIA HGX H200, 141 GB | 8 | On-demand node | $50.44/node-hour | $6.31 calculated | 2026-07-27 [4] |
| CoreWeave North America | NVIDIA HGX B200, 180 GB | 8 | On-demand node | $68.80/node-hour | $8.60 calculated | 2026-07-27 [4] |

The CoreWeave calculations are $49.24 / 8 = $6.155, $50.44 / 8 = $6.305, and $68.80 / 8 = $8.60 before display rounding.[4]

### Bundle and billing context

The same rate can buy materially different surrounding resources. This table records the dimensions needed to interpret the rows above.

| Provider offer | Region | CPU and system memory shown with offer | Tenancy and network context | Billing minimum | Storage, egress, tax, and fee context |
|---|---|---|---|---|---|
| RunPod Secure Cloud and Community Cloud Pods | Pricing page does not bind the rows to one region | H100: 20 vCPUs and 125 GB RAM; H200: 24 vCPUs and 276 GB RAM; B200: 28 vCPUs and 283 GB RAM | Pods are described as dedicated GPU instances; the price table does not explain the operational difference between its Secure Cloud and Community Cloud labels, and network topology is not stated in the price row | Page presents per-hour and per-second views; no minimum is stated | Storage is priced separately, starting at $0.05/GB/month; egress and tax treatment are not stated in the GPU row [1] |
| Together GPU Clusters | Not stated in price table | Not stated | Cluster product; tenant isolation and fabric details are not stated in the price table | Hourly rate; no minimum is stated | Shared filesystem is separately listed at $0.16/GiB/month; egress and tax treatment are not stated [2] |
| Nebius GPU instances | Not stated in price table | Values displayed beside H100 and H200: 16 vCPUs and 200 GB RAM; beside B200: 20 vCPUs and 224 GB RAM. The allocation scope of these CPU and RAM values is not stated | Tenant isolation and network fabric are not stated in the GPU row | Per GPU-hour; no usage-duration minimum is stated in the GPU pricing row. The page separately states a $25 minimum for the first payment | Storage is separate; the page lists network ingress and egress as free and states that prices exclude applicable taxes, including VAT [3] |
| CoreWeave North America HGX nodes | North America | Each listed H100, H200, or B200 node: 128 vCPUs and 2,048 GB system RAM | Eight-GPU HGX node; tenant isolation and inter-node fabric are not stated in the rate table | Per node-hour; no minimum is stated | The node includes 61.44 TB local storage; other storage, egress, and tax treatment are not specified by these rows [4] |

## Purchase models are not interchangeable

On-demand, spot or preemptible, reserved, and scheduled-capacity prices answer different questions. They should not occupy one column labeled "cheapest" without qualification.

- **On-demand or pay as you go.** AWS describes its EC2 On-Demand model as compute billed by the hour or second, with a 60-second minimum and no long-term commitment.[7] Other providers can use the same label with different billing details.
- **Spot or preemptible.** These are provider labels that must be evaluated under the provider's exact operational terms. In the Nebius capture, $2.15 for H100, $2.45 for H200, and $3.95 for B200 are explicitly in the `Preemptible` column, not a commitment column.[3]
- **Reserved or committed.** Together publishes fixed-duration GPU Cluster rates. For 91-180 days, the captured rates are $3.09 for H100, $3.99 for H200, and $6.79 for B200 per GPU-hour. Its shorter 7-30 and 31-90 day columns carry different prices.[2]
- **Scheduled capacity blocks.** [Amazon Web Services](https://aiwiki.ai/wiki/amazon_web_services) states that EC2 Capacity Blocks reserve a scheduled amount of accelerator capacity. The reservation fee is charged up front, prices can change with supply and demand, and an operating-system fee is added when instances run.[8] That is not an EC2 On-Demand rate.
- **Marketplace offers.** [Vast.ai](https://aiwiki.ai/wiki/vast_ai) documents a real-time marketplace in which hosts set prices and supply and demand affect rates. It bills actual usage by the second.[9] A live marketplace quote therefore needs its own timestamp and configuration record.

The CoreWeave archive also lists North America spot node prices of $19.71 for eight H100 GPUs, $20.93 for eight H200 GPUs, and $34.11 for eight B200 GPUs. Those normalize to approximately $2.46, $2.62, and $4.26 per GPU-hour, respectively.[4] They are not substituted for the on-demand rows because the source labels them as a different purchase model.

## Providers omitted from the numerical snapshot

A missing row is not evidence that a provider lacks the GPU or charges more than a listed provider.

| Provider | Why no strict cutoff-captured on-demand number is shown |
|---|---|
| AWS | The captured official sources establish the On-Demand terms and Capacity Block model, but the evidence set does not contain an exact archived H100, H200, or B200 EC2 On-Demand rate row [7][8] |
| [Azure](https://aiwiki.ai/wiki/azure) | The retail records available to this review were retrieved after the cutoff, so they are not presented as archived July 28 observations |
| [Google Cloud](https://aiwiki.ai/wiki/google_cloud) | No clean official accelerator-pricing capture dated by the cutoff was available for this snapshot |
| [Lambda Labs](https://aiwiki.ai/wiki/lambda_labs) | The exact pricing-page archive lookup did not yield a near-cutoff capture; an older page was not treated as a July 2026 rate |
| Vast.ai | Its documented real-time marketplace model requires a timestamped offer-level capture, which is not part of this snapshot [9] |

This conservative exclusion prevents a current price, an old price, a scheduled reservation price, or a third-party estimate from being presented as a July 28 on-demand observation.

## How to compare offers for a workload

The provider-published hourly rate is an input to cost analysis, not the result. A practical comparison should keep the following dimensions fixed or model their effect explicitly:

1. **GPU identity and memory.** H100 PCIe, H100 SXM, H100 NVL, H200, and B200 are distinct labels on provider pages. Record accelerator memory and the exact provider SKU.[1][4]
2. **Node composition.** Record GPU count, vCPUs, system RAM, local storage, and whether the bill applies to one GPU or a whole node. The captured RunPod and CoreWeave offers demonstrate how those bundles differ.[1][4]
3. **Region, availability, and tenancy.** Record the region and whether the source states capacity or tenant isolation. Vast.ai's documentation says regional supply and demand affect its marketplace pricing.[9] Do not infer immediate availability or isolation from a pricing row that does not state them.
4. **Network topology.** Record any intra-node or inter-node fabric stated for the offer, and mark it unknown when the pricing evidence is silent. Research on transient cloud GPU training evaluates performance across different cluster configurations, which is one reason to preserve configuration context.[10]
5. **Purchase and interruption terms.** Read the provider's exact terms. When an offer is revocable, model checkpoint recovery, restart delay, and the chance of losing capacity. Research on transient cloud GPU training describes the trade-off among training time, cost, configuration, and revocation risk.[10]
6. **Billing granularity and idle time.** Multiply the applicable rate by billed time, not only active kernel time. For inference, a 2026 preprint reports that utilization and request concurrency can materially change effective cost per token on identical hardware.[11]
7. **Ancillary charges.** Add persistent storage, snapshots, data transfer, public IP addresses, support, software licenses, taxes, and any unused commitment or reserved capacity. FOCUS distinguishes list, contracted, billed, and effective cost for this reason.[5]
8. **Measured workload performance.** Compare cost per completed training run, successful experiment, or served output at a stated service level. Habitat presents a runtime-based method intended to help users make informed, cost-efficient GPU selections by predicting iteration execution time across GPUs.[12] A later research proposal also argues that time-based GPU pricing can misrepresent resource use for bandwidth-bound applications.[13]

As an illustrative accounting identity, rather than a universal provider billing formula, a multi-node training estimate can be written as:

`estimated job cost = billed node-hours * node price + storage + data transfer + software and support fees + taxes`

For an interruptible job, the model should also include expected recomputation and checkpoint overhead. For steady inference, measure throughput and latency under the expected request pattern, then calculate cost per output unit. These workload-level measures are more informative than sorting a mixed set of GPU-hour rates.

## Limitations and update policy

This article is an evidence-bounded historical snapshot. Provider pages are mutable, marketplace offers can move quickly, and negotiated enterprise prices are outside this snapshot. The four main provider captures were not simultaneous. The comparison does not infer missing prices, treat contact-sales offers as numbers, or use third-party price aggregators to fill gaps.

Before a purchasing decision, verify the provider's current price, region, quota or capacity, billing minimum, cancellation terms, interruption policy, network, storage and egress charges, taxes, and support level. Preserve the quote or API response with its timestamp and SKU so a later invoice can be compared against the same assumptions.

## References

1. RunPod, "GPU Cloud Pricing," official pricing page archived July 11, 2026. https://web.archive.org/web/20260711111931id_/https://www.runpod.io/pricing
2. Together AI, "Pricing," official pricing page archived July 18, 2026. https://web.archive.org/web/20260718003135id_/https://www.together.ai/pricing
3. Nebius, "AI Cloud pricing," official pricing page archived July 12, 2026. https://web.archive.org/web/20260712235400id_/https://nebius.com/prices
4. CoreWeave, "CoreWeave Cloud Pricing," official pricing page archived July 27, 2026. https://web.archive.org/web/20260727103436id_/https://www.coreweave.com/pricing
5. FinOps Foundation, "FinOps Open Cost and Usage Specification 1.4," June 2026. https://focus.finops.org/wp-content/uploads/2026/06/FOCUS_spec-v1_4.pdf
6. Peter Mell and Timothy Grance, "The NIST Definition of Cloud Computing," NIST Special Publication 800-145, September 2011. https://doi.org/10.6028/NIST.SP.800-145
7. Amazon Web Services, "Amazon EC2 On-Demand Pricing," official page archived July 28, 2026. https://web.archive.org/web/20260728160604id_/https://aws.amazon.com/ec2/pricing/on-demand/
8. Amazon Web Services, "Amazon EC2 Capacity Blocks for ML pricing," official page archived July 4, 2026. https://web.archive.org/web/20260704092824id_/https://aws.amazon.com/ec2/capacityblocks/pricing/
9. Vast.ai Documentation, "Pricing," official documentation archived June 28, 2026. https://web.archive.org/web/20260628131718id_/https://docs.vast.ai/guides/instances/pricing
10. Shijian Li, Robert J. Walls, and Tian Guo, "Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers," IEEE ICDCS 2020, arXiv:2004.03072. https://arxiv.org/abs/2004.03072v1
11. Chitral Patil, "Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation," arXiv:2606.11690, submitted June 10, 2026. https://arxiv.org/abs/2606.11690v1
12. Geoffrey X. Yu, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko, "Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training," USENIX ATC 2021, pages 503-521. https://www.usenix.org/conference/atc21/presentation/yu
13. Ian McDougall, Noah Scott, Joon Huh, Kirthevasan Kandasamy, and Karthikeyan Sankaralingam, "Agora: Bridging the GPU Cloud Resource-Price Disconnect," arXiv:2510.05111, submitted September 26, 2025. https://arxiv.org/abs/2510.05111v1

