Citation and evidence

Cloud AI GPU Pricing Comparison

28 min full readUpdated 36 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI HardwareAI InfrastructureData Centers

Cite this article

Cloud AI GPU pricing is difficult to compare from a headline hourly rate alone. Providers sell different GPU variants, node sizes, CPU and memory bundles, regions, network configurations, and purchase models. The tables below therefore present a dated evidence snapshot, not a ranking. The main comparison uses provider pages captured between July 11 and July 27, 2026, before the research cutoff of July 28, 2026 at 23:59:59 UTC+7.[1][2][3][4] A separate September 2026 section records a second dated check of the same providers, hyperscaler and Lambda list prices, B300 rates, and market-wide rental price indices published by Ornn, Silicon Data, and SemiAnalysis.[14][25][29][30]

The snapshot covers published US-dollar rates for NVIDIA H100, H200, and B200 offers from RunPod, Together AI, Nebius, and CoreWeave. Negotiated quotes, credits, taxes, managed inference, serverless endpoints, and non-NVIDIA accelerators are outside the numerical comparison.

Method and units

Every numerical row is tied to a specific provider, GPU label, purchase model, billing object, and evidence date. The normalized figure is calculated as:

normalized GPU-hour rate = published node or instance hourly rate / GPUs in that node or instance

No division is needed when a provider already states that its rate is per GPU per hour. When division is used, the published system price and GPU count remain visible so that the arithmetic can be checked. A normalized GPU-hour is only a price unit. It does not make an eight-GPU HGX node equivalent to a one-GPU virtual machine.

This separation follows the broader cost-accounting distinction between a list unit price, the pricing category, the pricing unit, and effective cost. The FinOps Open Cost and Usage Specification, or FOCUS, treats these as separate fields and also separates list, contracted, billed, and effective cost.[5] NIST's definition of cloud computing includes measured service as an essential characteristic, but a provider's meter does not by itself describe workload performance or total cost.[6]

For this article, "not stated" means the cited pricing evidence did not specify that property. It does not mean that the provider lacks the feature. Prices do not prove that capacity was available for immediate deployment on the evidence date.

Cutoff-captured price snapshot

All rates in this table are provider-published dollar-denominated rates. The source date is the date of the archived provider page. Values rounded to two decimals are marked as calculations.

Provider and offerGPU labelGPU count used for normalizationPurchase modelPublished rateNormalized USD/GPU-hourSource date
RunPod Secure Cloud PodsH100 SXM, 80 GB1Secure Cloud Pod hourly rate$2.99/GPU-hour$2.992026-07-11 [1]
RunPod Secure Cloud PodsH200, 141 GB1Secure Cloud Pod hourly rate$4.39/GPU-hour$4.392026-07-11 [1]
RunPod Secure Cloud PodsB200, 180 GB1Secure Cloud Pod hourly rate$5.89/GPU-hour$5.892026-07-11 [1]
RunPod Community Cloud PodsH100 SXM, 80 GB1Community Cloud Pod hourly rate$2.69/GPU-hour$2.692026-07-11 [1]
RunPod Community Cloud PodsH200, 141 GB1Community Cloud Pod hourly rate$3.59/GPU-hour$3.592026-07-11 [1]
RunPod Community Cloud PodsB200, 180 GB1Community Cloud Pod hourly rate$5.89/GPU-hour$5.892026-07-11 [1]
Together GPU ClustersNVIDIA HGX H100Not divided; provider states per GPUOn-demand$3.99/GPU-hour$3.992026-07-18 [2]
Together GPU ClustersNVIDIA HGX H200Not divided; provider states per GPUOn-demand$5.99/GPU-hour$5.992026-07-18 [2]
Together GPU ClustersNVIDIA HGX B200Not divided; provider states per GPUOn-demand$8.19/GPU-hour$8.192026-07-18 [2]
Nebius GPU instanceNVIDIA HGX H100Not divided; provider states per GPUOn-demand$3.85/GPU-hour$3.852026-07-12 [3]
Nebius GPU instanceNVIDIA HGX H200Not divided; provider states per GPUOn-demand$4.50/GPU-hour$4.502026-07-12 [3]
Nebius GPU instanceNVIDIA HGX B200Not divided; provider states per GPUOn-demand$7.15/GPU-hour$7.152026-07-12 [3]
CoreWeave North AmericaNVIDIA HGX H100, 80 GB8On-demand node$49.24/node-hour$6.16 calculated2026-07-27 [4]
CoreWeave North AmericaNVIDIA HGX H200, 141 GB8On-demand node$50.44/node-hour$6.31 calculated2026-07-27 [4]
CoreWeave North AmericaNVIDIA HGX B200, 180 GB8On-demand node$68.80/node-hour$8.60 calculated2026-07-27 [4]

Expanded article table

The CoreWeave calculations are $49.24 / 8 = $6.155, $50.44 / 8 = $6.305, and $68.80 / 8 = $8.60 before display rounding.[4]

Bundle and billing context

The same rate can buy materially different surrounding resources. This table records the dimensions needed to interpret the rows above.

Provider offerRegionCPU and system memory shown with offerTenancy and network contextBilling minimumStorage, egress, tax, and fee context
RunPod Secure Cloud and Community Cloud PodsPricing page does not bind the rows to one regionH100: 20 vCPUs and 125 GB RAM; H200: 24 vCPUs and 276 GB RAM; B200: 28 vCPUs and 283 GB RAMPods are described as dedicated GPU instances; the price table does not explain the operational difference between its Secure Cloud and Community Cloud labels, and network topology is not stated in the price rowPage presents per-hour and per-second views; no minimum is statedStorage is priced separately, starting at $0.05/GB/month; egress and tax treatment are not stated in the GPU row [1]
Together GPU ClustersNot stated in price tableNot statedCluster product; tenant isolation and fabric details are not stated in the price tableHourly rate; no minimum is statedShared filesystem is separately listed at $0.16/GiB/month; egress and tax treatment are not stated [2]
Nebius GPU instancesNot stated in price tableValues displayed beside H100 and H200: 16 vCPUs and 200 GB RAM; beside B200: 20 vCPUs and 224 GB RAM. The allocation scope of these CPU and RAM values is not statedTenant isolation and network fabric are not stated in the GPU rowPer GPU-hour; no usage-duration minimum is stated in the GPU pricing row. The page separately states a $25 minimum for the first paymentStorage is separate; the page lists network ingress and egress as free and states that prices exclude applicable taxes, including VAT [3]
CoreWeave North America HGX nodesNorth AmericaEach listed H100, H200, or B200 node: 128 vCPUs and 2,048 GB system RAMEight-GPU HGX node; tenant isolation and inter-node fabric are not stated in the rate tablePer node-hour; no minimum is statedThe node includes 61.44 TB local storage; other storage, egress, and tax treatment are not specified by these rows [4]

Expanded article table

Purchase models are not interchangeable

On-demand, spot or preemptible, reserved, and scheduled-capacity prices answer different questions. They should not occupy one column labeled "cheapest" without qualification.

  • On-demand or pay as you go. AWS describes its EC2 On-Demand model as compute billed by the hour or second, with a 60-second minimum and no long-term commitment.[7] Other providers can use the same label with different billing details.
  • Spot or preemptible. These are provider labels that must be evaluated under the provider's exact operational terms. In the Nebius capture, $2.15 for H100, $2.45 for H200, and $3.95 for B200 are explicitly in the Preemptible column, not a commitment column.[3]
  • Reserved or committed. Together publishes fixed-duration GPU Cluster rates. For 91-180 days, the captured rates are $3.09 for H100, $3.99 for H200, and $6.79 for B200 per GPU-hour. Its shorter 7-30 and 31-90 day columns carry different prices.[2]
  • Scheduled capacity blocks. Amazon Web Services states that EC2 Capacity Blocks reserve a scheduled amount of accelerator capacity. The reservation fee is charged up front, prices can change with supply and demand, and an operating-system fee is added when instances run.[8] That is not an EC2 On-Demand rate.
  • Marketplace offers. Vast.ai documents a real-time marketplace in which hosts set prices and supply and demand affect rates. It bills actual usage by the second.[9] A live marketplace quote therefore needs its own timestamp and configuration record.

The CoreWeave archive also lists North America spot node prices of $19.71 for eight H100 GPUs, $20.93 for eight H200 GPUs, and $34.11 for eight B200 GPUs. Those normalize to approximately $2.46, $2.62, and $4.26 per GPU-hour, respectively.[4] They are not substituted for the on-demand rows because the source labels them as a different purchase model.

Providers omitted from the numerical snapshot

A missing row is not evidence that a provider lacks the GPU or charges more than a listed provider.

ProviderWhy no strict cutoff-captured on-demand number is shown
AWSThe captured official sources establish the On-Demand terms and Capacity Block model, but the evidence set does not contain an exact archived H100, H200, or B200 EC2 On-Demand rate row [7][8]
AzureThe retail records available to this review were retrieved after the cutoff, so they are not presented as archived July 28 observations
Google CloudNo clean official accelerator-pricing capture dated by the cutoff was available for this snapshot
Lambda LabsThe exact pricing-page archive lookup did not yield a near-cutoff capture; an older page was not treated as a July 2026 rate
Vast.aiIts documented real-time marketplace model requires a timestamped offer-level capture, which is not part of this snapshot [9]

Expanded article table

This conservative exclusion prevents a current price, an old price, a scheduled reservation price, or a third-party estimate from being presented as a July 28 on-demand observation. The September 2026 price check below adds separately dated rows for AWS, Azure, Google Cloud, and Lambda; those rows are not backfilled into the July table.

September 2026 price check

The July snapshot above is kept unchanged as a historical record. The tables in this section were compiled from the same providers' public pages, plus hyperscaler and Lambda price sources, retrieved on September 27-28, 2026. They are also dated observations, not current quotes.[14][15][16][17]

Change since the July snapshot

Provider and offerGPU labelJuly 2026 rateSeptember 2026 rateChangeSource
RunPod Secure Cloud PodsH100 SXM, 80 GB$2.99/GPU-hour$3.49/GPU-hour+16.7% calculated[1][14]
RunPod Secure Cloud PodsH200, 141 GB$4.39/GPU-hour$4.59/GPU-hour+4.6% calculated[1][14]
RunPod Secure Cloud PodsB200, 180 GB$5.89/GPU-hour$6.79/GPU-hour+15.3% calculated[1][14]
RunPod Community Cloud PodsH100 SXM, 80 GB$2.69/GPU-hour$2.69/GPU-hourunchanged[1][14]
RunPod Community Cloud PodsH200, 141 GB$3.59/GPU-hour$3.59/GPU-hourunchanged[1][14]
RunPod Community Cloud PodsB200, 180 GB$5.89/GPU-hour$5.98/GPU-hour+1.5% calculated[1][14]
Together GPU Clusters, on-demandNVIDIA HGX H100 / H200 / B200$3.99 / $5.99 / $8.19 per GPU-hour$3.99 / $5.99 / $8.19 per GPU-hourunchanged[2][15]
Together GPU Clusters, H100 reservedNVIDIA HGX H100, 7-30 / 31-90 / 91-180 days$3.59 / $3.29 / $3.09 per GPU-hour$3.69 / $3.45 / $3.19 per GPU-hour+2.8% / +4.9% / +3.2% calculated[2][15]
Nebius GPU instance, on-demandNVIDIA HGX H100$3.85/GPU-hour$3.85/GPU-hour; $4.50 effective October 1, 2026+16.9% calculated on October 1[3][16]
Nebius GPU instance, on-demandNVIDIA HGX H200$4.50/GPU-hour$4.50/GPU-hour; $5.40 effective October 1, 2026+20.0% calculated on October 1[3][16]
Nebius GPU instance, on-demandNVIDIA HGX B200$7.15/GPU-hour$7.15/GPU-hour; $8.50 effective October 1, 2026+18.9% calculated on October 1[3][16]
CoreWeave North America, on-demand nodeNVIDIA HGX H100 / H200 / B200, 8 GPUs$49.24 / $50.44 / $68.80 per node-hour$49.24 / $50.44 / $68.80 per node-hourunchanged[4][17]

Expanded article table

RunPod's page was labeled "Updated September 27, 2026" when retrieved.[14] Together's H200 and B200 reserved rates were unchanged, and its September table added a "Preemptible Compute" column at $1.99 for H100, $2.99 for H200, $4.09 for B200, and $4.99 for HGX B300 per GPU-hour.[15] Nebius's September page replaced the fixed July preemptible prices with dynamic spot prices "from" $0.79 per GPU-hour for H100 and H200 and "from" $0.99 for B200 and B300. It says spot prices can change as often as every 15 minutes and that the highest rate is one cent below the on-demand rate.[16] CoreWeave's North America spot node prices were also unchanged from July at $19.71 (H100), $20.93 (H200), and $34.11 (B200) per eight-GPU node-hour.[17]

On September 17, 2026, Investing.com, citing a Bloomberg News report of a Nebius emailed statement, reported that Nebius would raise prices for its on-demand GPU services on October 1 and that the increases would apply to H100, H200, B200, and B300 resources.[31] The per-GPU figures in the table come from Nebius's own price page, which lists a separate "Effective October 1, 2026" column.[16]

Hyperscaler and Lambda list prices, September 2026

The hyperscaler rows below are pay-as-you-go Linux list prices for whole instances, divided by the provider-stated GPU count. They include CPU, memory, and local storage bundles that differ from the neocloud rows above, so the normalized figure is only a price unit.

Provider and instanceGPU label and countRegion and price modelPublished rateNormalized USD/GPU-hourSource
AWS p5.48xlarge8 x H100US East (N. Virginia), Linux On-Demand$55.04/instance-hour$6.88 calculated[18][19][20]
AWS p5.4xlarge1 x H100US East (N. Virginia), Linux On-Demand$6.88/instance-hour$6.88[18][20]
AWS p5en.48xlarge8 x H200US East (N. Virginia), Linux On-Demand$63.296/instance-hour$7.91 calculated[18][19]
AWS p6-b200.48xlarge8 x B200US East (N. Virginia), Linux On-Demand$113.9328/instance-hour$14.24 calculated[18][19]
AWS p6-b300.48xlarge8 x B300US East (N. Virginia), Linux On-Demand$142.416/instance-hour$17.80 calculated[18][19]
Azure Standard_ND96isr_H100_v58 x H100 80 GBEast US, Linux pay-as-you-go$98.32/instance-hour$12.29 calculated[21][22]
Google Cloud a3-highgpu-8g8 x H100On-demand, page default region Iowa (us-central1)$88.49/instance-hour$11.06 calculated[23]
Google Cloud a3-megagpu-8g8 x H100On-demand, page default region Iowa (us-central1)$93.40/instance-hour$11.68 calculated[23]
Google Cloud a3-ultragpu-8g8 x H200On-demand, page default region Iowa (us-central1)$84.81/instance-hour$10.60 calculated[23]
Google Cloud a4-highgpu-8g8 x B200On-demand price shown as "N/A"; DWS Flex-start $64.44/instance-hour$64.44/instance-hour (Flex-start)$8.06 calculated (Flex-start)[23]
Lambda instance, 8x planNVIDIA H100 SXM, 80 GBSelf-serve on-demand, per GPU$3.99/GPU-hour$3.99[24]
Lambda instance, 1x planNVIDIA H100 SXM, 80 GBSelf-serve on-demand, per GPU$4.29/GPU-hour$4.29[24]
Lambda instance, 8x planNVIDIA B200 SXM6, 180 GBSelf-serve on-demand, per GPU$6.69/GPU-hour$6.69[24]

Expanded article table

The AWS figures come from the price data file behind the EC2 On-Demand pricing page, which carried a publication date of September 25, 2026; the GPU counts come from AWS's Capacity Blocks and P5 instance pages.[18][19][20] Azure's figure comes from Microsoft's public Retail Prices API; the API lists December 1, 2023 as the effective start date of that price record.[21] Microsoft documents the ND96isr H100 v5 size as having eight H100 80 GB GPUs.[22] Google's accelerator-optimized pricing page lists GPU counts in each machine type, notes that these machine types are not eligible for sustained use discounts, and showed B200 on-demand pricing as "N/A" when retrieved.[23] Lambda's per-GPU prices rise as the instance gets smaller (for example, H100 SXM at $3.99 on an eight-GPU instance and $4.29 on a one-GPU instance) and exclude sales tax, VAT, or GST. Its 1-Click Clusters, sold for two weeks to one year, were listed at $5.54 to $6.16 per H100 GPU-hour and $8.87 to $9.86 per B200 GPU-hour, depending on GPU count.[24]

AWS's EC2 Capacity Block rates, a reservation product rather than On-Demand, were $5.191 per H100 accelerator-hour for P5 in US Regions, $5.97 for P5e (H200), $6.865 for P5en (H200) in US Regions, $12.355 for P6-B200, and $14.04 for P6-B300 when retrieved in September 2026. The same rates were announced on the July 2026 archived page as effective July 7, 2026, and the September page said prices were scheduled to be updated next in October 2026.[8][19]

B300 and GB300 rental prices

NVIDIA's Blackwell Ultra B300 appeared on several public price lists by September 2026. Rack-scale GB300 NVL72 systems were "Contact sales" or "Contact us" at CoreWeave, Nebius, and Together.[15][16][17]

Provider and offerPurchase modelPublished rateNormalized USD/GPU-hourSource date
RunPod Secure Cloud Pod, B300 288 GBHourly Pod$7.89/GPU-hour (July 11: $7.39)$7.892026-09-27 [14]
RunPod Community Cloud Pod, B300 288 GBHourly Pod$6.94/GPU-hour (July 11: $6.94)$6.942026-09-27 [14]
Together GPU Clusters, NVIDIA HGX B300On-demand$9.99/GPU-hour$9.992026-09-27 [15]
Together GPU Clusters, NVIDIA HGX B300Preemptible$4.99/GPU-hour$4.992026-09-27 [15]
Nebius, NVIDIA HGX B300On-demand$7.85/GPU-hour; $9.50 effective October 1, 2026$7.85; $9.502026-09-27 [16]
CoreWeave North America, NVIDIA HGX B300, 8 GPUsSpot$35.84/node-hour (on-demand: contact sales)$4.48 calculated2026-09-27 [17]
AWS p6-b300.48xlarge, 8 x B300On-Demand, US East (N. Virginia)$142.416/instance-hour$17.80 calculated2026-09-25 publication [18]
AWS p6-b300.48xlarge, 8 x B300Capacity Block, US Regions$112.32/instance-hour$14.042026-09-27 [19]

Expanded article table

Market-wide rental price indices

Provider rate cards show what a seller asks. Several firms also publish benchmark series that try to summarize what the market pays. Their methods differ, so their values for the same GPU on the same week can differ by more than the spread between providers.

Publisher and seriesGPUValueAs ofStated basis
Ornn Compute Price Index, OCPI-H100H100 SXM$2.54/GPU-hourSeptember 26, 2026Volume-weighted, winsorized average of transacted neocloud rental prices, settled daily [25][27]
Ornn, OCPI-H200H200$4.78/GPU-hourSeptember 26, 2026Same method [27]
Ornn, OCPI-B200B200$8.04/GPU-hourSeptember 26, 2026Same method [26]
Ornn, OCPI-A100A100 SXM4$0.95/GPU-hourSeptember 26, 2026Same method [27]
Ornn B300 index, as stated by OrnnB300$11.32/GPU-hourSeptember 24, 2026 postNot in Ornn's free public tier; figure from Ornn's own post [28]
Silicon Data H100 Rental Price Index, neo-cloud reading (SDH100RT)H100$2.70/GPU-hourSeptember 27, 2026Observations standardized for specs, rental terms, platform performance, and geography, with outliers removed [29]
Silicon Data B200 index (SDB200RT)B200$5.87/GPU-hourDisplayed September 27, 2026Silicon Data method [29]
Silicon Data B300 indexB300$6.80/GPU-hourDisplayed September 27, 2026Silicon Data method [29]
Silicon Data H200 indexH200$3.33/GPU-hourDisplayed September 27, 2026Silicon Data method [29]
SemiAnalysis H100 1-year contract indexH100, 1-year term$2.35/GPU-hourMarch 2026Monthly survey of 100+ market participants (25th to 75th percentile range), checked against transactions [30]

Expanded article table

Ornn. Ornn publishes GPU price indices and, through Ornn Compute, also rents reserved capacity priced against its own index; its site says "Ornn both prices compute and rents it."[25] Its OCPI pages say the indices are built only from executed trades in the neocloud market, not listings or provider rate cards, and that the values are not a provider list price, bid, offer, or inventory quote.[27] The public H100 series was volatile in September 2026: the 14 daily settlements from September 13 to 26 ranged from $2.53 (September 15) to $3.05 (September 23), and Ornn reported an 8.0% decline over 30 days from $2.76 and a three-month range of $2.45 to $3.17. The B200 series, by contrast, was up 33.3% over 30 days from $6.03, with a three-month range of $4.38 to $8.04.[25][26] Intercontinental Exchange and Ornn have announced plans for cash-settled GPU compute futures based on OCPI, subject to regulatory approval.[32] See GPU compute futures.

In a promotional post on September 24, 2026, Ornn described NVIDIA GPUs as "a high-yield asset." It stated that its H100 index was $2.91 per hour, up 49% from a year earlier; that its new B300 index was $11.32 per hour, up 66% from June; that H100 marketplace utilization was 87% and B300 utilization 90%; and that resale values had held within -5% to +7%.[28] Ornn's published OCPI-H100 settlement for that date was $2.92.[25] The post's worked examples (an H100 bought in September 2025 for about $19,000, and a B300 bought in July 2026 for about $54,000) are Ornn's own calculations; its charts state that they assume 80% utilization and operating costs of $0.30 per hour for H100 and $0.45 per hour for B300. The post also gave two-year projections to September 2028 from Ornn's forward curves. Those are Ornn's forecasts of owner returns, not observed prices, and the year-on-year and since-June changes could not be checked against Ornn's free public data, which covers only three months and does not include B300.[25][27][28]

Silicon Data. Silicon Data publishes the H100 index as a neo-cloud reading (ticker SDH100RT) with a separate hyperscaler reading. It says observations come from neo-cloud providers, hyperscalers, colocation markets, brokered cluster sales, and private rental platforms, and are standardized before publication.[29] CME Group's product page states that NYMEX will launch Silicon Data H100 Rental Index futures, 730 GPU-hours per contract and financially settled against the average of the index's on-demand settlement prices, pending regulatory review periods.[33] On September 21, 2026, the CFTC extended its review of the H100 and B200 contracts to the end of November 9, 2026, and a revised CME Clearing advisory of September 25 listed the effective date as to be announced.[35][36] The settlement series is defined in CME's contract terms and should not be assumed to be identical to the public SDH100RT reading in the table above (see CME Compute Futures).

Divergence between indices. On nearly the same dates in late September 2026, Ornn's B200 index ($8.04) was 37% above Silicon Data's B200 index ($5.87), and the B300 figure in Ornn's post ($11.32) was 66% above Silicon Data's B300 index ($6.80). The two H100 series were closer ($2.54 and $2.70). These percentages are calculated from the published values.[25][26][28][29] Transaction-weighted, listing-based, and survey-based series answer different questions, so a buyer should check which one a contract or report uses before comparing it with a provider quote.

The 2026 price trend

Several independent sources describe rising GPU rental prices through 2026. SemiAnalysis noted that before late 2025 the prevailing expectation had been that Hopper rental prices would drop considerably as Blackwell deployments ramped.[30]

  • Contract market. SemiAnalysis wrote on April 2, 2026 that H100 one-year rental contract pricing had risen "almost 40%" from a low of $1.70 per GPU-hour in October 2025 to $2.35 by March 2026. It said on-demand capacity was sold out across all GPU types, that providers had shifted to longer terms and higher prepayments, and that capacity coming online through August to September 2026 had already been booked.[30]
  • How posted prices move. The same report said providers usually hold on-demand prices fixed and change them only occasionally, so posted on-demand prices stay flat for long periods and then jump, and utilization is a better high-frequency signal than price.[30] The provider pages checked for this article fit that pattern: between July and late September 2026, CoreWeave's posted on-demand and spot node prices and Together's on-demand rates did not change, while RunPod raised several Secure Cloud rates and Nebius scheduled increases of about 17% to 21%.[14][15][16][17]
  • Outside the United States. The Seoul Economic Daily reported on September 22, 2026 that the Korean neocloud Vessl AI had raised the hourly H100 rate on its VESSL Cloud service from $2.39 to $2.98 on September 1, a 24.7% increase, and described the move as following rate increases by US neoclouds.[34]
  • Divergent readings. Not every series rose over the same window. Ornn's public H100 index fell 8.0% in the 30 days to September 26, 2026, while its B200 index rose 33.3%.[25][26]

How to compare offers for a workload

The provider-published hourly rate is an input to cost analysis, not the result. A practical comparison should keep the following dimensions fixed or model their effect explicitly:

  1. GPU identity and memory. H100 PCIe, H100 SXM, H100 NVL, H200, and B200 are distinct labels on provider pages. Record accelerator memory and the exact provider SKU.[1][4]
  2. Node composition. Record GPU count, vCPUs, system RAM, local storage, and whether the bill applies to one GPU or a whole node. The captured RunPod and CoreWeave offers demonstrate how those bundles differ.[1][4]
  3. Region, availability, and tenancy. Record the region and whether the source states capacity or tenant isolation. Vast.ai's documentation says regional supply and demand affect its marketplace pricing.[9] Do not infer immediate availability or isolation from a pricing row that does not state them.
  4. Network topology. Record any intra-node or inter-node fabric stated for the offer, and mark it unknown when the pricing evidence is silent. Research on transient cloud GPU training evaluates performance across different cluster configurations, which is one reason to preserve configuration context.[10]
  5. Purchase and interruption terms. Read the provider's exact terms. When an offer is revocable, model checkpoint recovery, restart delay, and the chance of losing capacity. Research on transient cloud GPU training describes the trade-off among training time, cost, configuration, and revocation risk.[10]
  6. Billing granularity and idle time. Multiply the applicable rate by billed time, not only active kernel time. For inference, a 2026 preprint reports that utilization and request concurrency can materially change effective cost per token on identical hardware.[11]
  7. Ancillary charges. Add persistent storage, snapshots, data transfer, public IP addresses, support, software licenses, taxes, and any unused commitment or reserved capacity. FOCUS distinguishes list, contracted, billed, and effective cost for this reason.[5]
  8. Measured workload performance. Compare cost per completed training run, successful experiment, or served output at a stated service level. Habitat presents a runtime-based method intended to help users make informed, cost-efficient GPU selections by predicting iteration execution time across GPUs.[12] A later research proposal also argues that time-based GPU pricing can misrepresent resource use for bandwidth-bound applications.[13]

As an illustrative accounting identity, rather than a universal provider billing formula, a multi-node training estimate can be written as:

estimated job cost = billed node-hours * node price + storage + data transfer + software and support fees + taxes

For an interruptible job, the model should also include expected recomputation and checkpoint overhead. For steady inference, measure throughput and latency under the expected request pattern, then calculate cost per output unit. These workload-level measures are more informative than sorting a mixed set of GPU-hour rates.

Limitations and update policy

This article is an evidence-bounded historical snapshot. Provider pages are mutable, marketplace offers can move quickly, and negotiated enterprise prices are outside this snapshot. The four main provider captures were not simultaneous. The comparison does not infer missing prices, treat contact-sales offers as numbers, or use third-party price aggregators to fill gaps. The September 2026 check was a single pass on September 27-28, 2026; Nebius's October 1 prices were scheduled, not yet in effect, on those dates. Index values are publishers' benchmarks, not quotes from any one provider.

Before a purchasing decision, verify the provider's current price, region, quota or capacity, billing minimum, cancellation terms, interruption policy, network, storage and egress charges, taxes, and support level. Preserve the quote or API response with its timestamp and SKU so a later invoice can be compared against the same assumptions.

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16RunPod, "GPU Cloud Pricing," official pricing page archived July 11, 2026. web.archive.org/...pricing
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Together AI, "Pricing," official pricing page archived July 18, 2026. web.archive.org/...pricing
  3. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Nebius, "AI Cloud pricing," official pricing page archived July 12, 2026. web.archive.org/...prices
  4. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10CoreWeave, "CoreWeave Cloud Pricing," official pricing page archived July 27, 2026. web.archive.org/...pricing
  5. ^1 ^2FinOps Foundation, "FinOps Open Cost and Usage Specification 1.4," June 2026. focus.finops.org/...FOCUS_spec-v1_4.pdf
  6. ^Peter Mell and Timothy Grance, "The NIST Definition of Cloud Computing," NIST Special Publication 800-145, September 2011. doi.org/...NIST.SP.800-145
  7. ^1 ^2Amazon Web Services, "Amazon EC2 On-Demand Pricing," official page archived July 28, 2026. web.archive.org/...on-demand
  8. ^1 ^2 ^3Amazon Web Services, "Amazon EC2 Capacity Blocks for ML pricing," official page archived July 4, 2026. web.archive.org/...pricing
  9. ^1 ^2 ^3Vast.ai Documentation, "Pricing," official documentation archived June 28, 2026. web.archive.org/...pricing
  10. ^1 ^2Shijian Li, Robert J. Walls, and Tian Guo, "Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers," IEEE ICDCS 2020, arXiv:2004.03072. arxiv.org/...2004.03072v1
  11. ^Chitral Patil, "Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation," arXiv:2606.11690, submitted June 10, 2026. arxiv.org/...2606.11690v1
  12. ^Geoffrey X. Yu, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko, "Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training," USENIX ATC 2021, pages 503-521. usenix.org/...yu
  13. ^Ian McDougall, Noah Scott, Joon Huh, Kirthevasan Kandasamy, and Karthikeyan Sankaralingam, "Agora: Bridging the GPU Cloud Resource-Price Disconnect," arXiv:2510.05111, submitted September 26, 2025. arxiv.org/...2510.05111v1
  14. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Runpod, "GPU Cloud Pricing | Per-Second H100, A100, RTX | Runpod," official pricing page (labeled "Updated September 27, 2026"), retrieved September 27, 2026. runpod.io/pricing
  15. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Together AI, "Pricing | Together AI," official pricing page, retrieved September 27, 2026. together.ai/pricing
  16. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Nebius, "NVIDIA GPU Pricing | Nebius AI Cloud," official pricing page, retrieved September 27, 2026. nebius.com/prices
  17. ^1 ^2 ^3 ^4 ^5 ^6CoreWeave, "CoreWeave Cloud Pricing," official pricing page, retrieved September 27, 2026. coreweave.com/pricing
  18. ^1 ^2 ^3 ^4 ^5 ^6 ^7Amazon Web Services, EC2 On-Demand price data for US East (N. Virginia), Linux (file publication date September 25, 2026), used by the Amazon EC2 On-Demand Pricing page, retrieved September 27, 2026. b0.p.awsstatic.com/...Linux
  19. ^1 ^2 ^3 ^4 ^5 ^6 ^7Amazon Web Services, "Amazon EC2 Capacity Blocks for ML Pricing," official page, retrieved September 27, 2026. aws.amazon.com/...pricing
  20. ^1 ^2 ^3Amazon Web Services, "Amazon EC2 P5 Instances," official product page, retrieved September 27, 2026. aws.amazon.com/...p5
  21. ^1 ^2Microsoft, Azure Retail Prices API, query for Virtual Machines H100 SKUs in East US, retrieved September 27, 2026. prices.azure.com/...prices
  22. ^1 ^2Microsoft Learn, "ND-H100-v5 size series - Azure Virtual Machines," retrieved September 27, 2026. learn.microsoft.com/...ndh100v5-series
  23. ^1 ^2 ^3 ^4 ^5Google Cloud, "Accelerator-optimized VM Pricing," official pricing page, retrieved September 27, 2026. cloud.google.com/...accelerator-optimized
  24. ^1 ^2 ^3 ^4Lambda, "GPU cloud pricing: rent NVIDIA H100, H200, and B200," official pricing page, retrieved September 27, 2026. lambda.ai/pricing
  25. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Ornn Data, "H100 SXM Price Index (OCPI-H100)," settled September 26, 2026, retrieved September 27, 2026. data.ornn.com/...h100-sxm
  26. ^1 ^2 ^3 ^4Ornn Data, "B200 Price Index (OCPI-B200)," settled September 26, 2026, retrieved September 27, 2026. data.ornn.com/...b200
  27. ^1 ^2 ^3 ^4 ^5Ornn Data, "GPU Market Prices: Compute Price Index (OCPI)," retrieved September 27, 2026. data.ornn.com/markets
  28. ^1 ^2 ^3 ^4Ornn (@OrnnExchange), post on X, September 24, 2026 (text and chart images retrieved via the fxtwitter API). x.com/...2103188722757910769
  29. ^1 ^2 ^3 ^4 ^5 ^6 ^7Silicon Data, "H100 Rental Price Index," retrieved September 27, 2026. silicondata.com/...h100
  30. ^1 ^2 ^3 ^4 ^5SemiAnalysis, "The Great GPU Shortage - Rental Capacity - Launching our H100 1 Year Rental Price Index," April 2, 2026. newsletter.semianalysis.com/...age-rental-capacity
  31. ^Investing.com, "Nebius to increase prices for Nvidia GPU resources from October," September 17, 2026. ca.investing.com/...rces-from-october-93CH-4843693
  32. ^Intercontinental Exchange, "ICE and Ornn to Launch GPU Compute Futures Contracts," press release, May 19, 2026. ir.theice.com/...default
  33. ^CME Group, "Silicon Data H100 Rental Index Futures - Contract Specs," retrieved September 27, 2026. cmegroup.com/...silicon-data-h100-rental-index
  34. ^Seoul Economic Daily, "H100 Rental Prices Jump 22% in a Month as GPU Shortage Hits Firms and Labs," September 22, 2026. en.sedaily.com/...ump-22-percent-in-a-month-as-gpu
  35. ^US Commodity Futures Trading Commission, Division of Market Oversight, "Extension of Initial Review Period of NYMEX Submission No. 26-370 and Submission No. 26-370S for an Additional 45-Days," letter to CME Group, September 21, 2026. cftc.gov/...orgdcmnymexcompcontr260921.pdf
  36. ^CME Clearing, "Revised New Product Summary: Initial Listing of Two (2) Compute Futures Contracts - Silicon Data H100 Rental Index Futures and Silicon Data B200 Rental Index Futures, Effective Date to be Announced," Advisory 26-274, September 25, 2026. cmegroup.com/...26-274

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

5 revisions · v6 · 5,504 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Re-stamp (xg10, 28 Sep 2026): added CFTC 21 Sep extension sentence checked against CFTC letter and CME advisory 26-274

Cite this page: AI Wiki. "Cloud AI GPU Pricing Comparison." aiwiki.ai, updated 27 Sept 2026, fact-checked 27 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/cloud_gpu_pricing_comparison

Suggest edit