Cloud AI GPU Pricing Comparison
Cloud AI GPU pricing is difficult to compare from a headline hourly rate alone. Providers sell different GPU variants, node sizes, CPU and memory bundles, regions, network configurations, and purchase models. The tables below therefore present a dated evidence snapshot, not a ranking. The main comparison uses provider pages captured between July 11 and July 27, 2026, before the research cutoff of July 28, 2026 at 23:59:59 UTC+7.[1][2][3][4] A separate September 2026 section records a second dated check of the same providers, hyperscaler and Lambda list prices, B300 rates, and market-wide rental price indices published by Ornn, Silicon Data, and SemiAnalysis.[14][25][29][30]
The snapshot covers published US-dollar rates for NVIDIA H100, H200, and B200 offers from RunPod, Together AI, Nebius, and CoreWeave. Negotiated quotes, credits, taxes, managed inference, serverless endpoints, and non-NVIDIA accelerators are outside the numerical comparison.
Method and units
Every numerical row is tied to a specific provider, GPU label, purchase model, billing object, and evidence date. The normalized figure is calculated as:
normalized GPU-hour rate = published node or instance hourly rate / GPUs in that node or instance
No division is needed when a provider already states that its rate is per GPU per hour. When division is used, the published system price and GPU count remain visible so that the arithmetic can be checked. A normalized GPU-hour is only a price unit. It does not make an eight-GPU HGX node equivalent to a one-GPU virtual machine.
This separation follows the broader cost-accounting distinction between a list unit price, the pricing category, the pricing unit, and effective cost. The FinOps Open Cost and Usage Specification, or FOCUS, treats these as separate fields and also separates list, contracted, billed, and effective cost.[5] NIST's definition of cloud computing includes measured service as an essential characteristic, but a provider's meter does not by itself describe workload performance or total cost.[6]
For this article, "not stated" means the cited pricing evidence did not specify that property. It does not mean that the provider lacks the feature. Prices do not prove that capacity was available for immediate deployment on the evidence date.
Cutoff-captured price snapshot
All rates in this table are provider-published dollar-denominated rates. The source date is the date of the archived provider page. Values rounded to two decimals are marked as calculations.
| Provider and offer | GPU label | GPU count used for normalization | Purchase model | Published rate | Normalized USD/GPU-hour | Source date |
|---|---|---|---|---|---|---|
| RunPod Secure Cloud Pods | H100 SXM, 80 GB | 1 | Secure Cloud Pod hourly rate | $2.99/GPU-hour | $2.99 | 2026-07-11 [1] |
| RunPod Secure Cloud Pods | H200, 141 GB | 1 | Secure Cloud Pod hourly rate | $4.39/GPU-hour | $4.39 | 2026-07-11 [1] |
| RunPod Secure Cloud Pods | B200, 180 GB | 1 | Secure Cloud Pod hourly rate | $5.89/GPU-hour | $5.89 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | H100 SXM, 80 GB | 1 | Community Cloud Pod hourly rate | $2.69/GPU-hour | $2.69 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | H200, 141 GB | 1 | Community Cloud Pod hourly rate | $3.59/GPU-hour | $3.59 | 2026-07-11 [1] |
| RunPod Community Cloud Pods | B200, 180 GB | 1 | Community Cloud Pod hourly rate | $5.89/GPU-hour | $5.89 | 2026-07-11 [1] |
| Together GPU Clusters | NVIDIA HGX H100 | Not divided; provider states per GPU | On-demand | $3.99/GPU-hour | $3.99 | 2026-07-18 [2] |
| Together GPU Clusters | NVIDIA HGX H200 | Not divided; provider states per GPU | On-demand | $5.99/GPU-hour | $5.99 | 2026-07-18 [2] |
| Together GPU Clusters | NVIDIA HGX B200 | Not divided; provider states per GPU | On-demand | $8.19/GPU-hour | $8.19 | 2026-07-18 [2] |
| Nebius GPU instance | NVIDIA HGX H100 | Not divided; provider states per GPU | On-demand | $3.85/GPU-hour | $3.85 | 2026-07-12 [3] |
| Nebius GPU instance | NVIDIA HGX H200 | Not divided; provider states per GPU | On-demand | $4.50/GPU-hour | $4.50 | 2026-07-12 [3] |
| Nebius GPU instance | NVIDIA HGX B200 | Not divided; provider states per GPU | On-demand | $7.15/GPU-hour | $7.15 | 2026-07-12 [3] |
| CoreWeave North America | NVIDIA HGX H100, 80 GB | 8 | On-demand node | $49.24/node-hour | $6.16 calculated | 2026-07-27 [4] |
| CoreWeave North America | NVIDIA HGX H200, 141 GB | 8 | On-demand node | $50.44/node-hour | $6.31 calculated | 2026-07-27 [4] |
| CoreWeave North America | NVIDIA HGX B200, 180 GB | 8 | On-demand node | $68.80/node-hour | $8.60 calculated | 2026-07-27 [4] |
The CoreWeave calculations are $49.24 / 8 = $6.155, $50.44 / 8 = $6.305, and $68.80 / 8 = $8.60 before display rounding.[4]
Bundle and billing context
The same rate can buy materially different surrounding resources. This table records the dimensions needed to interpret the rows above.
| Provider offer | Region | CPU and system memory shown with offer | Tenancy and network context | Billing minimum | Storage, egress, tax, and fee context |
|---|---|---|---|---|---|
| RunPod Secure Cloud and Community Cloud Pods | Pricing page does not bind the rows to one region | H100: 20 vCPUs and 125 GB RAM; H200: 24 vCPUs and 276 GB RAM; B200: 28 vCPUs and 283 GB RAM | Pods are described as dedicated GPU instances; the price table does not explain the operational difference between its Secure Cloud and Community Cloud labels, and network topology is not stated in the price row | Page presents per-hour and per-second views; no minimum is stated | Storage is priced separately, starting at $0.05/GB/month; egress and tax treatment are not stated in the GPU row [1] |
| Together GPU Clusters | Not stated in price table | Not stated | Cluster product; tenant isolation and fabric details are not stated in the price table | Hourly rate; no minimum is stated | Shared filesystem is separately listed at $0.16/GiB/month; egress and tax treatment are not stated [2] |
| Nebius GPU instances | Not stated in price table | Values displayed beside H100 and H200: 16 vCPUs and 200 GB RAM; beside B200: 20 vCPUs and 224 GB RAM. The allocation scope of these CPU and RAM values is not stated | Tenant isolation and network fabric are not stated in the GPU row | Per GPU-hour; no usage-duration minimum is stated in the GPU pricing row. The page separately states a $25 minimum for the first payment | Storage is separate; the page lists network ingress and egress as free and states that prices exclude applicable taxes, including VAT [3] |
| CoreWeave North America HGX nodes | North America | Each listed H100, H200, or B200 node: 128 vCPUs and 2,048 GB system RAM | Eight-GPU HGX node; tenant isolation and inter-node fabric are not stated in the rate table | Per node-hour; no minimum is stated | The node includes 61.44 TB local storage; other storage, egress, and tax treatment are not specified by these rows [4] |
Purchase models are not interchangeable
On-demand, spot or preemptible, reserved, and scheduled-capacity prices answer different questions. They should not occupy one column labeled "cheapest" without qualification.
- On-demand or pay as you go. AWS describes its EC2 On-Demand model as compute billed by the hour or second, with a 60-second minimum and no long-term commitment.[7] Other providers can use the same label with different billing details.
- Spot or preemptible. These are provider labels that must be evaluated under the provider's exact operational terms. In the Nebius capture, $2.15 for H100, $2.45 for H200, and $3.95 for B200 are explicitly in the
Preemptiblecolumn, not a commitment column.[3] - Reserved or committed. Together publishes fixed-duration GPU Cluster rates. For 91-180 days, the captured rates are $3.09 for H100, $3.99 for H200, and $6.79 for B200 per GPU-hour. Its shorter 7-30 and 31-90 day columns carry different prices.[2]
- Scheduled capacity blocks. Amazon Web Services states that EC2 Capacity Blocks reserve a scheduled amount of accelerator capacity. The reservation fee is charged up front, prices can change with supply and demand, and an operating-system fee is added when instances run.[8] That is not an EC2 On-Demand rate.
- Marketplace offers. Vast.ai documents a real-time marketplace in which hosts set prices and supply and demand affect rates. It bills actual usage by the second.[9] A live marketplace quote therefore needs its own timestamp and configuration record.
The CoreWeave archive also lists North America spot node prices of $19.71 for eight H100 GPUs, $20.93 for eight H200 GPUs, and $34.11 for eight B200 GPUs. Those normalize to approximately $2.46, $2.62, and $4.26 per GPU-hour, respectively.[4] They are not substituted for the on-demand rows because the source labels them as a different purchase model.
Providers omitted from the numerical snapshot
A missing row is not evidence that a provider lacks the GPU or charges more than a listed provider.
| Provider | Why no strict cutoff-captured on-demand number is shown |
|---|---|
| AWS | The captured official sources establish the On-Demand terms and Capacity Block model, but the evidence set does not contain an exact archived H100, H200, or B200 EC2 On-Demand rate row [7][8] |
| Azure | The retail records available to this review were retrieved after the cutoff, so they are not presented as archived July 28 observations |
| Google Cloud | No clean official accelerator-pricing capture dated by the cutoff was available for this snapshot |
| Lambda Labs | The exact pricing-page archive lookup did not yield a near-cutoff capture; an older page was not treated as a July 2026 rate |
| Vast.ai | Its documented real-time marketplace model requires a timestamped offer-level capture, which is not part of this snapshot [9] |
This conservative exclusion prevents a current price, an old price, a scheduled reservation price, or a third-party estimate from being presented as a July 28 on-demand observation. The September 2026 price check below adds separately dated rows for AWS, Azure, Google Cloud, and Lambda; those rows are not backfilled into the July table.
September 2026 price check
The July snapshot above is kept unchanged as a historical record. The tables in this section were compiled from the same providers' public pages, plus hyperscaler and Lambda price sources, retrieved on September 27-28, 2026. They are also dated observations, not current quotes.[14][15][16][17]
Change since the July snapshot
| Provider and offer | GPU label | July 2026 rate | September 2026 rate | Change | Source |
|---|---|---|---|---|---|
| RunPod Secure Cloud Pods | H100 SXM, 80 GB | $2.99/GPU-hour | $3.49/GPU-hour | +16.7% calculated | [1][14] |
| RunPod Secure Cloud Pods | H200, 141 GB | $4.39/GPU-hour | $4.59/GPU-hour | +4.6% calculated | [1][14] |
| RunPod Secure Cloud Pods | B200, 180 GB | $5.89/GPU-hour | $6.79/GPU-hour | +15.3% calculated | [1][14] |
| RunPod Community Cloud Pods | H100 SXM, 80 GB | $2.69/GPU-hour | $2.69/GPU-hour | unchanged | [1][14] |
| RunPod Community Cloud Pods | H200, 141 GB | $3.59/GPU-hour | $3.59/GPU-hour | unchanged | [1][14] |
| RunPod Community Cloud Pods | B200, 180 GB | $5.89/GPU-hour | $5.98/GPU-hour | +1.5% calculated | [1][14] |
| Together GPU Clusters, on-demand | NVIDIA HGX H100 / H200 / B200 | $3.99 / $5.99 / $8.19 per GPU-hour | $3.99 / $5.99 / $8.19 per GPU-hour | unchanged | [2][15] |
| Together GPU Clusters, H100 reserved | NVIDIA HGX H100, 7-30 / 31-90 / 91-180 days | $3.59 / $3.29 / $3.09 per GPU-hour | $3.69 / $3.45 / $3.19 per GPU-hour | +2.8% / +4.9% / +3.2% calculated | [2][15] |
| Nebius GPU instance, on-demand | NVIDIA HGX H100 | $3.85/GPU-hour | $3.85/GPU-hour; $4.50 effective October 1, 2026 | +16.9% calculated on October 1 | [3][16] |
| Nebius GPU instance, on-demand | NVIDIA HGX H200 | $4.50/GPU-hour | $4.50/GPU-hour; $5.40 effective October 1, 2026 | +20.0% calculated on October 1 | [3][16] |
| Nebius GPU instance, on-demand | NVIDIA HGX B200 | $7.15/GPU-hour | $7.15/GPU-hour; $8.50 effective October 1, 2026 | +18.9% calculated on October 1 | [3][16] |
| CoreWeave North America, on-demand node | NVIDIA HGX H100 / H200 / B200, 8 GPUs | $49.24 / $50.44 / $68.80 per node-hour | $49.24 / $50.44 / $68.80 per node-hour | unchanged | [4][17] |
RunPod's page was labeled "Updated September 27, 2026" when retrieved.[14] Together's H200 and B200 reserved rates were unchanged, and its September table added a "Preemptible Compute" column at $1.99 for H100, $2.99 for H200, $4.09 for B200, and $4.99 for HGX B300 per GPU-hour.[15] Nebius's September page replaced the fixed July preemptible prices with dynamic spot prices "from" $0.79 per GPU-hour for H100 and H200 and "from" $0.99 for B200 and B300. It says spot prices can change as often as every 15 minutes and that the highest rate is one cent below the on-demand rate.[16] CoreWeave's North America spot node prices were also unchanged from July at $19.71 (H100), $20.93 (H200), and $34.11 (B200) per eight-GPU node-hour.[17]
On September 17, 2026, Investing.com, citing a Bloomberg News report of a Nebius emailed statement, reported that Nebius would raise prices for its on-demand GPU services on October 1 and that the increases would apply to H100, H200, B200, and B300 resources.[31] The per-GPU figures in the table come from Nebius's own price page, which lists a separate "Effective October 1, 2026" column.[16]
Hyperscaler and Lambda list prices, September 2026
The hyperscaler rows below are pay-as-you-go Linux list prices for whole instances, divided by the provider-stated GPU count. They include CPU, memory, and local storage bundles that differ from the neocloud rows above, so the normalized figure is only a price unit.
| Provider and instance | GPU label and count | Region and price model | Published rate | Normalized USD/GPU-hour | Source |
|---|---|---|---|---|---|
| AWS p5.48xlarge | 8 x H100 | US East (N. Virginia), Linux On-Demand | $55.04/instance-hour | $6.88 calculated | [18][19][20] |
| AWS p5.4xlarge | 1 x H100 | US East (N. Virginia), Linux On-Demand | $6.88/instance-hour | $6.88 | [18][20] |
| AWS p5en.48xlarge | 8 x H200 | US East (N. Virginia), Linux On-Demand | $63.296/instance-hour | $7.91 calculated | [18][19] |
| AWS p6-b200.48xlarge | 8 x B200 | US East (N. Virginia), Linux On-Demand | $113.9328/instance-hour | $14.24 calculated | [18][19] |
| AWS p6-b300.48xlarge | 8 x B300 | US East (N. Virginia), Linux On-Demand | $142.416/instance-hour | $17.80 calculated | [18][19] |
| Azure Standard_ND96isr_H100_v5 | 8 x H100 80 GB | East US, Linux pay-as-you-go | $98.32/instance-hour | $12.29 calculated | [21][22] |
| Google Cloud a3-highgpu-8g | 8 x H100 | On-demand, page default region Iowa (us-central1) | $88.49/instance-hour | $11.06 calculated | [23] |
| Google Cloud a3-megagpu-8g | 8 x H100 | On-demand, page default region Iowa (us-central1) | $93.40/instance-hour | $11.68 calculated | [23] |
| Google Cloud a3-ultragpu-8g | 8 x H200 | On-demand, page default region Iowa (us-central1) | $84.81/instance-hour | $10.60 calculated | [23] |
| Google Cloud a4-highgpu-8g | 8 x B200 | On-demand price shown as "N/A"; DWS Flex-start $64.44/instance-hour | $64.44/instance-hour (Flex-start) | $8.06 calculated (Flex-start) | [23] |
| Lambda instance, 8x plan | NVIDIA H100 SXM, 80 GB | Self-serve on-demand, per GPU | $3.99/GPU-hour | $3.99 | [24] |
| Lambda instance, 1x plan | NVIDIA H100 SXM, 80 GB | Self-serve on-demand, per GPU | $4.29/GPU-hour | $4.29 | [24] |
| Lambda instance, 8x plan | NVIDIA B200 SXM6, 180 GB | Self-serve on-demand, per GPU | $6.69/GPU-hour | $6.69 | [24] |
The AWS figures come from the price data file behind the EC2 On-Demand pricing page, which carried a publication date of September 25, 2026; the GPU counts come from AWS's Capacity Blocks and P5 instance pages.[18][19][20] Azure's figure comes from Microsoft's public Retail Prices API; the API lists December 1, 2023 as the effective start date of that price record.[21] Microsoft documents the ND96isr H100 v5 size as having eight H100 80 GB GPUs.[22] Google's accelerator-optimized pricing page lists GPU counts in each machine type, notes that these machine types are not eligible for sustained use discounts, and showed B200 on-demand pricing as "N/A" when retrieved.[23] Lambda's per-GPU prices rise as the instance gets smaller (for example, H100 SXM at $3.99 on an eight-GPU instance and $4.29 on a one-GPU instance) and exclude sales tax, VAT, or GST. Its 1-Click Clusters, sold for two weeks to one year, were listed at $5.54 to $6.16 per H100 GPU-hour and $8.87 to $9.86 per B200 GPU-hour, depending on GPU count.[24]
AWS's EC2 Capacity Block rates, a reservation product rather than On-Demand, were $5.191 per H100 accelerator-hour for P5 in US Regions, $5.97 for P5e (H200), $6.865 for P5en (H200) in US Regions, $12.355 for P6-B200, and $14.04 for P6-B300 when retrieved in September 2026. The same rates were announced on the July 2026 archived page as effective July 7, 2026, and the September page said prices were scheduled to be updated next in October 2026.[8][19]
B300 and GB300 rental prices
NVIDIA's Blackwell Ultra B300 appeared on several public price lists by September 2026. Rack-scale GB300 NVL72 systems were "Contact sales" or "Contact us" at CoreWeave, Nebius, and Together.[15][16][17]
| Provider and offer | Purchase model | Published rate | Normalized USD/GPU-hour | Source date |
|---|---|---|---|---|
| RunPod Secure Cloud Pod, B300 288 GB | Hourly Pod | $7.89/GPU-hour (July 11: $7.39) | $7.89 | 2026-09-27 [14] |
| RunPod Community Cloud Pod, B300 288 GB | Hourly Pod | $6.94/GPU-hour (July 11: $6.94) | $6.94 | 2026-09-27 [14] |
| Together GPU Clusters, NVIDIA HGX B300 | On-demand | $9.99/GPU-hour | $9.99 | 2026-09-27 [15] |
| Together GPU Clusters, NVIDIA HGX B300 | Preemptible | $4.99/GPU-hour | $4.99 | 2026-09-27 [15] |
| Nebius, NVIDIA HGX B300 | On-demand | $7.85/GPU-hour; $9.50 effective October 1, 2026 | $7.85; $9.50 | 2026-09-27 [16] |
| CoreWeave North America, NVIDIA HGX B300, 8 GPUs | Spot | $35.84/node-hour (on-demand: contact sales) | $4.48 calculated | 2026-09-27 [17] |
| AWS p6-b300.48xlarge, 8 x B300 | On-Demand, US East (N. Virginia) | $142.416/instance-hour | $17.80 calculated | 2026-09-25 publication [18] |
| AWS p6-b300.48xlarge, 8 x B300 | Capacity Block, US Regions | $112.32/instance-hour | $14.04 | 2026-09-27 [19] |
Market-wide rental price indices
Provider rate cards show what a seller asks. Several firms also publish benchmark series that try to summarize what the market pays. Their methods differ, so their values for the same GPU on the same week can differ by more than the spread between providers.
| Publisher and series | GPU | Value | As of | Stated basis |
|---|---|---|---|---|
| Ornn Compute Price Index, OCPI-H100 | H100 SXM | $2.54/GPU-hour | September 26, 2026 | Volume-weighted, winsorized average of transacted neocloud rental prices, settled daily [25][27] |
| Ornn, OCPI-H200 | H200 | $4.78/GPU-hour | September 26, 2026 | Same method [27] |
| Ornn, OCPI-B200 | B200 | $8.04/GPU-hour | September 26, 2026 | Same method [26] |
| Ornn, OCPI-A100 | A100 SXM4 | $0.95/GPU-hour | September 26, 2026 | Same method [27] |
| Ornn B300 index, as stated by Ornn | B300 | $11.32/GPU-hour | September 24, 2026 post | Not in Ornn's free public tier; figure from Ornn's own post [28] |
| Silicon Data H100 Rental Price Index, neo-cloud reading (SDH100RT) | H100 | $2.70/GPU-hour | September 27, 2026 | Observations standardized for specs, rental terms, platform performance, and geography, with outliers removed [29] |
| Silicon Data B200 index (SDB200RT) | B200 | $5.87/GPU-hour | Displayed September 27, 2026 | Silicon Data method [29] |
| Silicon Data B300 index | B300 | $6.80/GPU-hour | Displayed September 27, 2026 | Silicon Data method [29] |
| Silicon Data H200 index | H200 | $3.33/GPU-hour | Displayed September 27, 2026 | Silicon Data method [29] |
| SemiAnalysis H100 1-year contract index | H100, 1-year term | $2.35/GPU-hour | March 2026 | Monthly survey of 100+ market participants (25th to 75th percentile range), checked against transactions [30] |
Ornn. Ornn publishes GPU price indices and, through Ornn Compute, also rents reserved capacity priced against its own index; its site says "Ornn both prices compute and rents it."[25] Its OCPI pages say the indices are built only from executed trades in the neocloud market, not listings or provider rate cards, and that the values are not a provider list price, bid, offer, or inventory quote.[27] The public H100 series was volatile in September 2026: the 14 daily settlements from September 13 to 26 ranged from $2.53 (September 15) to $3.05 (September 23), and Ornn reported an 8.0% decline over 30 days from $2.76 and a three-month range of $2.45 to $3.17. The B200 series, by contrast, was up 33.3% over 30 days from $6.03, with a three-month range of $4.38 to $8.04.[25][26] Intercontinental Exchange and Ornn have announced plans for cash-settled GPU compute futures based on OCPI, subject to regulatory approval.[32] See GPU compute futures.
In a promotional post on September 24, 2026, Ornn described NVIDIA GPUs as "a high-yield asset." It stated that its H100 index was $2.91 per hour, up 49% from a year earlier; that its new B300 index was $11.32 per hour, up 66% from June; that H100 marketplace utilization was 87% and B300 utilization 90%; and that resale values had held within -5% to +7%.[28] Ornn's published OCPI-H100 settlement for that date was $2.92.[25] The post's worked examples (an H100 bought in September 2025 for about $19,000, and a B300 bought in July 2026 for about $54,000) are Ornn's own calculations; its charts state that they assume 80% utilization and operating costs of $0.30 per hour for H100 and $0.45 per hour for B300. The post also gave two-year projections to September 2028 from Ornn's forward curves. Those are Ornn's forecasts of owner returns, not observed prices, and the year-on-year and since-June changes could not be checked against Ornn's free public data, which covers only three months and does not include B300.[25][27][28]
Silicon Data. Silicon Data publishes the H100 index as a neo-cloud reading (ticker SDH100RT) with a separate hyperscaler reading. It says observations come from neo-cloud providers, hyperscalers, colocation markets, brokered cluster sales, and private rental platforms, and are standardized before publication.[29] CME Group's product page states that NYMEX will launch Silicon Data H100 Rental Index futures, 730 GPU-hours per contract and financially settled against the average of the index's on-demand settlement prices, pending regulatory review periods.[33] On September 21, 2026, the CFTC extended its review of the H100 and B200 contracts to the end of November 9, 2026, and a revised CME Clearing advisory of September 25 listed the effective date as to be announced.[35][36] The settlement series is defined in CME's contract terms and should not be assumed to be identical to the public SDH100RT reading in the table above (see CME Compute Futures).
Divergence between indices. On nearly the same dates in late September 2026, Ornn's B200 index ($8.04) was 37% above Silicon Data's B200 index ($5.87), and the B300 figure in Ornn's post ($11.32) was 66% above Silicon Data's B300 index ($6.80). The two H100 series were closer ($2.54 and $2.70). These percentages are calculated from the published values.[25][26][28][29] Transaction-weighted, listing-based, and survey-based series answer different questions, so a buyer should check which one a contract or report uses before comparing it with a provider quote.
The 2026 price trend
Several independent sources describe rising GPU rental prices through 2026. SemiAnalysis noted that before late 2025 the prevailing expectation had been that Hopper rental prices would drop considerably as Blackwell deployments ramped.[30]
- Contract market. SemiAnalysis wrote on April 2, 2026 that H100 one-year rental contract pricing had risen "almost 40%" from a low of $1.70 per GPU-hour in October 2025 to $2.35 by March 2026. It said on-demand capacity was sold out across all GPU types, that providers had shifted to longer terms and higher prepayments, and that capacity coming online through August to September 2026 had already been booked.[30]
- How posted prices move. The same report said providers usually hold on-demand prices fixed and change them only occasionally, so posted on-demand prices stay flat for long periods and then jump, and utilization is a better high-frequency signal than price.[30] The provider pages checked for this article fit that pattern: between July and late September 2026, CoreWeave's posted on-demand and spot node prices and Together's on-demand rates did not change, while RunPod raised several Secure Cloud rates and Nebius scheduled increases of about 17% to 21%.[14][15][16][17]
- Outside the United States. The Seoul Economic Daily reported on September 22, 2026 that the Korean neocloud Vessl AI had raised the hourly H100 rate on its VESSL Cloud service from $2.39 to $2.98 on September 1, a 24.7% increase, and described the move as following rate increases by US neoclouds.[34]
- Divergent readings. Not every series rose over the same window. Ornn's public H100 index fell 8.0% in the 30 days to September 26, 2026, while its B200 index rose 33.3%.[25][26]
How to compare offers for a workload
The provider-published hourly rate is an input to cost analysis, not the result. A practical comparison should keep the following dimensions fixed or model their effect explicitly:
- GPU identity and memory. H100 PCIe, H100 SXM, H100 NVL, H200, and B200 are distinct labels on provider pages. Record accelerator memory and the exact provider SKU.[1][4]
- Node composition. Record GPU count, vCPUs, system RAM, local storage, and whether the bill applies to one GPU or a whole node. The captured RunPod and CoreWeave offers demonstrate how those bundles differ.[1][4]
- Region, availability, and tenancy. Record the region and whether the source states capacity or tenant isolation. Vast.ai's documentation says regional supply and demand affect its marketplace pricing.[9] Do not infer immediate availability or isolation from a pricing row that does not state them.
- Network topology. Record any intra-node or inter-node fabric stated for the offer, and mark it unknown when the pricing evidence is silent. Research on transient cloud GPU training evaluates performance across different cluster configurations, which is one reason to preserve configuration context.[10]
- Purchase and interruption terms. Read the provider's exact terms. When an offer is revocable, model checkpoint recovery, restart delay, and the chance of losing capacity. Research on transient cloud GPU training describes the trade-off among training time, cost, configuration, and revocation risk.[10]
- Billing granularity and idle time. Multiply the applicable rate by billed time, not only active kernel time. For inference, a 2026 preprint reports that utilization and request concurrency can materially change effective cost per token on identical hardware.[11]
- Ancillary charges. Add persistent storage, snapshots, data transfer, public IP addresses, support, software licenses, taxes, and any unused commitment or reserved capacity. FOCUS distinguishes list, contracted, billed, and effective cost for this reason.[5]
- Measured workload performance. Compare cost per completed training run, successful experiment, or served output at a stated service level. Habitat presents a runtime-based method intended to help users make informed, cost-efficient GPU selections by predicting iteration execution time across GPUs.[12] A later research proposal also argues that time-based GPU pricing can misrepresent resource use for bandwidth-bound applications.[13]
As an illustrative accounting identity, rather than a universal provider billing formula, a multi-node training estimate can be written as:
estimated job cost = billed node-hours * node price + storage + data transfer + software and support fees + taxes
For an interruptible job, the model should also include expected recomputation and checkpoint overhead. For steady inference, measure throughput and latency under the expected request pattern, then calculate cost per output unit. These workload-level measures are more informative than sorting a mixed set of GPU-hour rates.
Limitations and update policy
This article is an evidence-bounded historical snapshot. Provider pages are mutable, marketplace offers can move quickly, and negotiated enterprise prices are outside this snapshot. The four main provider captures were not simultaneous. The comparison does not infer missing prices, treat contact-sales offers as numbers, or use third-party price aggregators to fill gaps. The September 2026 check was a single pass on September 27-28, 2026; Nebius's October 1 prices were scheduled, not yet in effect, on those dates. Index values are publishers' benchmarks, not quotes from any one provider.
Before a purchasing decision, verify the provider's current price, region, quota or capacity, billing minimum, cancellation terms, interruption policy, network, storage and egress charges, taxes, and support level. Preserve the quote or API response with its timestamp and SKU so a later invoice can be compared against the same assumptions.
References
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16RunPod, "GPU Cloud Pricing," official pricing page archived July 11, 2026. web.archive.org/...pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Together AI, "Pricing," official pricing page archived July 18, 2026. web.archive.org/...pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Nebius, "AI Cloud pricing," official pricing page archived July 12, 2026. web.archive.org/...prices
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10CoreWeave, "CoreWeave Cloud Pricing," official pricing page archived July 27, 2026. web.archive.org/...pricing
- ^1 ^2FinOps Foundation, "FinOps Open Cost and Usage Specification 1.4," June 2026. focus.finops.org/...FOCUS_spec-v1_4.pdf
- ^Peter Mell and Timothy Grance, "The NIST Definition of Cloud Computing," NIST Special Publication 800-145, September 2011. doi.org/...NIST.SP.800-145
- ^1 ^2Amazon Web Services, "Amazon EC2 On-Demand Pricing," official page archived July 28, 2026. web.archive.org/...on-demand
- ^1 ^2 ^3Amazon Web Services, "Amazon EC2 Capacity Blocks for ML pricing," official page archived July 4, 2026. web.archive.org/...pricing
- ^1 ^2 ^3Vast.ai Documentation, "Pricing," official documentation archived June 28, 2026. web.archive.org/...pricing
- ^1 ^2Shijian Li, Robert J. Walls, and Tian Guo, "Characterizing and Modeling Distributed Training with Transient Cloud GPU Servers," IEEE ICDCS 2020, arXiv:2004.03072. arxiv.org/...2004.03072v1
- ^Chitral Patil, "Beyond Per-Token Pricing: A Concurrency-Aware Methodology for LLM Infrastructure Cost Estimation," arXiv:2606.11690, submitted June 10, 2026. arxiv.org/...2606.11690v1
- ^Geoffrey X. Yu, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko, "Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training," USENIX ATC 2021, pages 503-521. usenix.org/...yu
- ^Ian McDougall, Noah Scott, Joon Huh, Kirthevasan Kandasamy, and Karthikeyan Sankaralingam, "Agora: Bridging the GPU Cloud Resource-Price Disconnect," arXiv:2510.05111, submitted September 26, 2025. arxiv.org/...2510.05111v1
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12Runpod, "GPU Cloud Pricing | Per-Second H100, A100, RTX | Runpod," official pricing page (labeled "Updated September 27, 2026"), retrieved September 27, 2026. runpod.io/pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Together AI, "Pricing | Together AI," official pricing page, retrieved September 27, 2026. together.ai/pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9Nebius, "NVIDIA GPU Pricing | Nebius AI Cloud," official pricing page, retrieved September 27, 2026. nebius.com/prices
- ^1 ^2 ^3 ^4 ^5 ^6CoreWeave, "CoreWeave Cloud Pricing," official pricing page, retrieved September 27, 2026. coreweave.com/pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Amazon Web Services, EC2 On-Demand price data for US East (N. Virginia), Linux (file publication date September 25, 2026), used by the Amazon EC2 On-Demand Pricing page, retrieved September 27, 2026. b0.p.awsstatic.com/...Linux
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Amazon Web Services, "Amazon EC2 Capacity Blocks for ML Pricing," official page, retrieved September 27, 2026. aws.amazon.com/...pricing
- ^1 ^2 ^3Amazon Web Services, "Amazon EC2 P5 Instances," official product page, retrieved September 27, 2026. aws.amazon.com/...p5
- ^1 ^2Microsoft, Azure Retail Prices API, query for Virtual Machines H100 SKUs in East US, retrieved September 27, 2026. prices.azure.com/...prices
- ^1 ^2Microsoft Learn, "ND-H100-v5 size series - Azure Virtual Machines," retrieved September 27, 2026. learn.microsoft.com/...ndh100v5-series
- ^1 ^2 ^3 ^4 ^5Google Cloud, "Accelerator-optimized VM Pricing," official pricing page, retrieved September 27, 2026. cloud.google.com/...accelerator-optimized
- ^1 ^2 ^3 ^4Lambda, "GPU cloud pricing: rent NVIDIA H100, H200, and B200," official pricing page, retrieved September 27, 2026. lambda.ai/pricing
- ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Ornn Data, "H100 SXM Price Index (OCPI-H100)," settled September 26, 2026, retrieved September 27, 2026. data.ornn.com/...h100-sxm
- ^1 ^2 ^3 ^4Ornn Data, "B200 Price Index (OCPI-B200)," settled September 26, 2026, retrieved September 27, 2026. data.ornn.com/...b200
- ^1 ^2 ^3 ^4 ^5Ornn Data, "GPU Market Prices: Compute Price Index (OCPI)," retrieved September 27, 2026. data.ornn.com/markets
- ^1 ^2 ^3 ^4Ornn (@OrnnExchange), post on X, September 24, 2026 (text and chart images retrieved via the fxtwitter API). x.com/...2103188722757910769
- ^1 ^2 ^3 ^4 ^5 ^6 ^7Silicon Data, "H100 Rental Price Index," retrieved September 27, 2026. silicondata.com/...h100
- ^1 ^2 ^3 ^4 ^5SemiAnalysis, "The Great GPU Shortage - Rental Capacity - Launching our H100 1 Year Rental Price Index," April 2, 2026. newsletter.semianalysis.com/...age-rental-capacity
- ^Investing.com, "Nebius to increase prices for Nvidia GPU resources from October," September 17, 2026. ca.investing.com/...rces-from-october-93CH-4843693
- ^Intercontinental Exchange, "ICE and Ornn to Launch GPU Compute Futures Contracts," press release, May 19, 2026. ir.theice.com/...default
- ^CME Group, "Silicon Data H100 Rental Index Futures - Contract Specs," retrieved September 27, 2026. cmegroup.com/...silicon-data-h100-rental-index
- ^Seoul Economic Daily, "H100 Rental Prices Jump 22% in a Month as GPU Shortage Hits Firms and Labs," September 22, 2026. en.sedaily.com/...ump-22-percent-in-a-month-as-gpu
- ^US Commodity Futures Trading Commission, Division of Market Oversight, "Extension of Initial Review Period of NYMEX Submission No. 26-370 and Submission No. 26-370S for an Additional 45-Days," letter to CME Group, September 21, 2026. cftc.gov/...orgdcmnymexcompcontr260921.pdf
- ^CME Clearing, "Revised New Product Summary: Initial Listing of Two (2) Compute Futures Contracts - Silicon Data H100 Rental Index Futures and Silicon Data B200 Rental Index Futures, Effective Date to be Announced," Advisory 26-274, September 25, 2026. cmegroup.com/...26-274
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
5 revisions · v6 · 5,504 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Re-stamp (xg10, 28 Sep 2026): added CFTC 21 Sep extension sentence checked against CFTC letter and CME advisory 26-274
Cite this page: AI Wiki. "Cloud AI GPU Pricing Comparison." aiwiki.ai, updated 27 Sept 2026, fact-checked 27 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/cloud_gpu_pricing_comparison