GPU Cluster

RawGraph

A GPU cluster is a group of servers containing graphics processing units that are connected by high-bandwidth, low-latency networks and operated as one computational pool. Clusters used for distributed training combine a fast scale-up fabric inside a server or rack with a scale-out fabric between servers. They also depend on schedulers, collective-communication software, storage, checkpoint recovery, power distribution, and cooling. Recent systems use accelerators such as the NVIDIA H100, H200, Blackwell, and Rubin GPUs, while Google and Amazon operate large clusters built from their own accelerators.

Public descriptions of cluster size require care. A physical GPU count, an "H100-equivalent" performance estimate, a data-center power commitment, and a multi-site infrastructure program are different measures and cannot be compared as if they were the same. For example, SpaceXAI said in May 2026 that Colossus 1 contained more than 220,000 physical NVIDIA GPUs, while OpenAI said in April 2026 that its multi-site Stargate program had secured more than 10 gigawatts of United States capacity without publishing a corresponding physical GPU count.[11][21] An independent 2024 engineering estimate put the critical IT requirement of a hypothetical 100,000-H100 campus above 150 megawatts.[26]

How did GPU clusters originate?

NVIDIA announced DGX-1 on 5 April 2016. The 3U appliance contained eight Tesla P100 GPUs, connected them with an NVLink Hybrid Cube Mesh, and included local solid-state storage and InfiniBand networking.[1] DGX-1 made a multi-GPU deep-learning system available as an integrated product rather than a one-off assembly. Later DGX generations added faster GPUs and NVSwitch fabrics, and operators joined many such servers into larger systems.

Model growth and generative AI increased the need for multi-node training. The 2017 paper "Attention Is All You Need" demonstrated an architecture based on attention rather than recurrence or convolution.[2] GPT-3, described in 2020, had 175 billion parameters and was trained with several forms of model parallelism.[3] That year Microsoft disclosed an Azure supercomputer built for OpenAI with more than 285,000 CPU cores, 10,000 GPUs, and 400 Gbit/s of network connectivity for each GPU server.[4] OpenAI did not disclose GPT-4's hardware configuration or training compute in its 2023 technical report, so public GPU-count estimates for that training run are not established facts.[38]

The scale disclosed by operators continued to grow. Azure Eagle used 14,400 H100 GPUs in 2023.[5] Meta described two 24,576-H100 clusters in 2024 and a 129,000-H100 cluster in 2025.[7][8] xAI said Colossus was expanded to 200,000 GPUs in 92 days, and SpaceXAI later described Colossus 1 as having more than 220,000 H100, H200, and GB200 GPUs.[9][11] These disclosures do not imply that every accelerator was assigned to one synchronous training job.

What is the physical layout of a GPU cluster?

A cluster is assembled in layers, but terms such as pod, island, and superpod are not standardized across operators. A common node has eight GPUs connected locally. In NVIDIA's H100 SuperPOD data-center guide, a scalable unit contains up to 32 DGX H100 systems, or 256 GPUs, plus its InfiniBand leaf infrastructure. Up to four scalable units form the documented 127-system SuperPOD configuration because management appliances occupy one system position.[12] This corrects the sometimes repeated claim that an H100 scalable unit contains only eight servers.

Rack-scale systems move part of the node boundary into the rack. A GB200 NVL72 rack has 18 compute trays, each with two Grace CPUs and four Blackwell GPUs, for a total of 36 CPUs and 72 GPUs. Nine NVLink switch trays connect those GPUs into one NVLink domain.[30] The wider cluster then connects many racks through InfiniBand or RDMA-capable Ethernet.

NVIDIA's H100 reference design and the large-campus example modeled by SemiAnalysis illustrate a bandwidth hierarchy across these layers. Inside an H100 server, fourth-generation NVLink supplies up to 900 GB/s of aggregate GPU-to-GPU bandwidth per GPU.[14] In those examples, the scale-out network has lower per-device bandwidth and must also manage congestion, routing, and cable distance. A campus may divide compute into islands with high bandwidth inside each island and less bandwidth between islands. The exact taper and topology depend on the workload, switch radix, optics budget, building layout, and fault-domain design.[12][26]

Rail-optimized fabrics

In a rail-optimized design, corresponding network interfaces on different servers connect through the same leaf-switch plane. Traffic between GPUs on a matching rail can avoid additional leaf hops, which can help collective operations. NVIDIA's H100 SuperPOD reference architecture uses a rail-optimized InfiniBand leaf-and-spine fabric, but this is a reference design rather than proof that every large cluster uses the same topology.[12][13]

Rail optimization has tradeoffs. It can require longer cables and more optical links than placing a switch near the servers it connects. Operators may instead use middle-of-row or middle-of-rack arrangements, partial oversubscription, or separate fabrics for compute, storage, and management. Public deployment details are often incomplete, so topology should be stated only when an operator or vendor has documented it.

Fat-tree and dragonfly variants

Clos or fat-tree networks provide multiple paths between endpoints and can be built with full or oversubscribed bisection bandwidth. The number of tiers is a design result, not a universal property of a GPU cluster. Dragonfly and dragonfly-derived networks reduce the number of long global links by grouping routers. The Frontier supercomputer at Oak Ridge National Laboratory, for example, uses an HPE Slingshot interconnect rather than NVIDIA InfiniBand or Spectrum-X.[15]

AI clusters and scientific high-performance computers can use related topologies, but their traffic is not identical. Dense AI training often generates repeated all-reduce, all-gather, reduce-scatter, and all-to-all collectives. Network design must therefore match the model's tensor, pipeline, data, and expert-parallel communication patterns.

How are the GPUs interconnected?

NVLink is NVIDIA's scale-up interconnect. NVSwitch chips route traffic between connected GPUs so that software can treat a server or rack as a high-bandwidth domain. H100 provides up to 900 GB/s of aggregate NVLink bandwidth per GPU.[14] In a GB200 NVL72 rack, fifth-generation NVLink connects all 72 GPUs, while the rack's compute trays use separate network adapters for communication beyond the rack.[30]

Scale-up and scale-out bandwidth should not be conflated. The NVLink figure is aggregate GPU-to-GPU bandwidth inside the NVLink domain. A 400 or 800 Gbit/s InfiniBand or Ethernet port describes a network link outside that domain and uses different units and traffic assumptions.

InfiniBand

InfiniBand provides remote direct memory access, lossless transport features, adaptive routing, and support for collective offload. NVIDIA's H100 SuperPOD architecture uses Quantum-2 NDR InfiniBand at 400 Gbit/s for compute and storage fabrics.[13] Azure Eagle also used NVIDIA InfiniBand NDR.[6] These examples establish documented deployments, not a market-wide claim that InfiniBand is the dominant interconnect.

GPU Direct RDMA lets a network adapter transfer data to or from GPU memory without staging it through host memory. This reduces CPU involvement and supports collective libraries such as NCCL. Operators still need congestion control, topology-aware placement, and isolation between compute, storage, management, and tenant traffic.

RDMA over converged Ethernet (RoCE)

RDMA over Converged Ethernet supplies RDMA semantics on Ethernet networks. Meta's two 24,576-H100 clusters provide a direct public comparison: one used a RoCE fabric built with Arista 7800, Wedge400, and Minipack2 switches, while the other used Quantum-2 InfiniBand. Both connected 400 Gbit/s endpoints. Meta said it trained Llama 3 on the RoCE system without a network bottleneck after co-designing the network, software, and model architecture.[7]

xAI's initial 100,000-GPU Colossus deployment is another documented Ethernet cluster. NVIDIA said it used Spectrum-X Ethernet with SN5600 switches and BlueField-3 SuperNICs. NVIDIA attributed 95 percent data throughput and no application latency degradation or packet loss from flow collisions to that configuration.[10] Those are vendor-reported results for the cited deployment, not generic guarantees for all Ethernet clusters.

SHARP in-network reduction

NVIDIA's Scalable Hierarchical Aggregation and Reduction Protocol, or SHARP, performs parts of collective operations in the InfiniBand network. Switch-based aggregation can reduce the data that endpoints exchange and free CPU or GPU resources for other work. SHARP support depends on the network hardware, software stack, collective, message size, and topology, so no single speedup applies to every workload.[16]

What software runs an AI GPU cluster?

NCCL and HCCL

The NVIDIA Collective Communications Library, or NCCL, implements all-reduce, all-gather, reduce-scatter, broadcast, and point-to-point operations for NVIDIA GPUs. It supports PCIe, NVLink, InfiniBand, and Ethernet-based communication and selects algorithms according to the detected topology.[17] NCCL 2.27 added features including symmetric-memory support and scalable connection establishment aimed at large training and inference deployments.[18]

AMD's RCCL and Huawei's HCCL provide related collective interfaces for their accelerator platforms. Application performance still depends on framework integration, process placement, network configuration, and the match between the collective algorithm and the physical topology.

Orchestration: Slurm, Kubernetes, and custom schedulers

Schedulers assign nodes to jobs and try to place tightly communicating ranks within appropriate network and failure domains. Slurm is widely used in high-performance computing and supports gang-scheduled multi-node jobs. Kubernetes can manage GPU workloads with device plugins and workload-queue extensions. Large operators also use internal systems. Meta, for example, has described topology-aware scheduling changes that reduced upper-layer network traffic in its 24,576-GPU clusters, and later named Twine and MAST as systems it was evolving for long-distance training.[7][8]

Public sources do not establish a single scheduler as the most common across private frontier-model clusters. Likewise, an open-source example does not by itself establish the framework or scheduler used for a company's production training.

Frameworks: PyTorch, JAX, and parallelism libraries

PyTorch and JAX both support distributed accelerator computation. PyTorch exposes NCCL through torch.distributed and includes DistributedDataParallel and Fully Sharded Data Parallel. Megatron-LM, DeepSpeed, and other libraries add tensor, pipeline, sequence, expert, and optimizer-state parallelism. Ray and Kubernetes-native systems can coordinate workloads that mix training, evaluation, and reinforcement learning.

Framework choice is separate from hardware ownership. A company may use several frameworks for different workloads, and public repositories are insufficient evidence for an organization-wide production choice.

What publicly documented clusters illustrate current scale?

The table compares selected, officially documented systems and programs as of 17 August 2026. It is not a ranking. Physical accelerators, equivalent-compute figures, and power commitments are kept in separate columns.

DeploymentPublicly disclosed acceleratorsNetwork or system detailStatus described by source
Meta large H100 cluster129,000 H100 GPUsSpans multiple data-center buildings and adjacent facilitiesMeta described it as built in 2025 [8]
xAI Colossus 1More than 220,000 NVIDIA H100, H200, and GB200 GPUsSpectrum-X documented for the original 100,000-GPU deploymentPhysical count reported in May 2026 [10][11]
Microsoft Azure Eagle14,400 H100 GPUsNVIDIA InfiniBand NDR; 561.20 PFLOP/s LinpackRanked 3rd in November 2023 and 7th in June 2026 [5][6]
StargatePhysical GPU count not disclosed in the cited sourcesAbilene uses NVIDIA GB200 systems on Oracle Cloud InfrastructureOpenAI said the United States program had secured more than 10 GW by April 2026 [21]
Google TPU 8t superpodUp to 9,600 TPU 8t chips2 PB of shared high-bandwidth memoryTraining-optimized system announced in April 2026 [25]
Anthropic on AWSMore than 1 million Trainium2 chips in useProject Rainier and additional AWS capacityAnthropic reported the deployed chip count in April 2026 [22]

Meta's H100 clusters

In March 2024 Meta described two clusters with 24,576 H100 GPUs each. One used RoCE and the other InfiniBand, and both had 400 Gbit/s endpoints. Meta said the design supported Llama 3. Its performance graph measured normalized AllGather bandwidth: the optimized large clusters returned to the 90-percent-plus range achieved by its smaller clusters. That number should not be described as general NCCL utilization or model FLOP utilization.[7]

Meta reported a 129,000-H100 cluster in September 2025. It also said that construction of Prometheus, a planned 1-gigawatt cluster, was underway and that Hyperion was expected to start coming online in 2028 and could reach 5 gigawatts when complete.[8] Those future power figures are site or cluster capacity plans, not physical GPU counts.

xAI Colossus

xAI says Colossus was built in 122 days and then doubled to 200,000 GPUs in 92 days.[9] NVIDIA's October 2024 account of the original 100,000-GPU deployment says that 19 days elapsed between the first rack arriving and training beginning.[10] SpaceXAI updated the physical description in May 2026, saying Colossus 1 had more than 220,000 NVIDIA GPUs across H100, H200, and GB200 systems.[11]

SpaceXAI separately said at the end of 2025 that Colossus I and Colossus II together exceeded one million H100 GPU equivalents. That is a performance-equivalent measure across more than one site, not evidence that Colossus 1 contained one million physical GPUs. The more than 220,000 figure is therefore the appropriate first-party physical count for Colossus 1 in this article.[11][39]

DDN said it supplied EXAScaler and Infinia storage for the original 100,000-GPU Colossus deployment.[37] The source establishes the vendor and products, but not a complete public specification of the current storage system after later expansions.

Microsoft Azure Eagle and the OpenAI partnership

Microsoft's 2020 system for OpenAI contained 10,000 GPUs and exposed 400 Gbit/s of network connectivity to each GPU server.[4] Microsoft did not identify it in that announcement as the machine that trained GPT-3, so the two facts should not be joined without a separate source.

Microsoft later described Azure Eagle as having 14,400 H100 GPUs and Intel Xeon Sapphire Rapids processors.[5] TOP500 records an Rmax of 561.20 PFLOP/s and NVIDIA InfiniBand NDR. Eagle entered the list at number 3 in November 2023 and was number 7 in the June 2026 list.[6] Linpack performance is useful for comparison with high-performance computers but is not a direct measure of large-model training throughput.

Stargate

The Stargate Project was announced in January 2025 by OpenAI, SoftBank, Oracle, and MGX. The announcement said the companies intended to invest $500 billion in United States AI infrastructure over four years, with $100 billion to begin immediately. It named Arm, Microsoft, NVIDIA, Oracle, and OpenAI as initial technology partners.[19] It did not disclose 40-percent equity stakes for OpenAI and SoftBank, so those percentages are omitted.

In September 2025 OpenAI announced five additional United States sites and said the planned portfolio, together with Abilene, had reached nearly 7 gigawatts and more than $400 billion of investment over the following three years.[20] By April 2026 OpenAI said it had surpassed the initial goal of securing 10 gigawatts by 2029. It also said the Abilene data center operated on Oracle Cloud Infrastructure, used NVIDIA GB200 systems, and had trained GPT-5.5.[21] OpenAI did not publish a physical accelerator count for the full program in those sources.

On 17 August 2026 OpenAI announced a separate agreement to secure approximately 8 IT-gigawatts at the planned PORTS-Pike Technology Campus in Ohio. OpenAI said the first 800 megawatts were expected in 2028, with a six-year buildout through 2032, and that the site would exclusively host NVIDIA AI compute under a 20-year lease from SB Energy.[41] This is a future capacity agreement for frontier-model training and product demand, not an operating 8-gigawatt cluster.

Anthropic's compute partnerships

Anthropic trains and serves Claude through several infrastructure partnerships. In April 2026 it said it used more than one million Trainium2 chips, was committing more than $100 billion over ten years to Amazon Web Services technologies, and had secured up to 5 gigawatts of new AWS capacity. The company expected nearly 1 gigawatt of Trainium2 and Trainium3 capacity to be online by the end of 2026.[22] The statement described AWS as Anthropic's primary training and cloud provider for mission-critical workloads.

The same month, SpaceXAI announced that Anthropic would receive access to Colossus 1 to improve capacity for Claude Pro and Claude Max subscribers.[11] These are distinct arrangements, and the 5-gigawatt AWS agreement should not be presented as an NVIDIA GPU count.

Google's TPU pods

Google's TPU is a non-GPU accelerator platform used for training and inference. A TPU v5p pod contains up to 8,960 chips connected in a three-dimensional torus, with 4,800 Gbit/s of inter-chip interconnect bandwidth per chip.[23] Trillium, also called TPU v6e, followed in 2024.

Ironwood, announced in April 2025, was Google's first TPU designed specifically for inference. Google said a pod could contain 9,216 chips and deliver 42.5 exaflops of FP8 compute.[24] At Google Cloud Next in April 2026, Google announced TPU 8t for training and TPU 8i for inference. A TPU 8t superpod scales to 9,600 chips and 2 PB of shared high-bandwidth memory, while a TPU 8i pod directly connects 1,152 chips.[25]

How much power and cooling does a GPU cluster need?

Power scale and grid impact

Power consumption depends on accelerator type, server design, network, storage, utilization, and facility overhead. SemiAnalysis estimated in June 2024 that a hypothetical 100,000-H100 deployment would require more than 150 megawatts of critical IT capacity and consume 1.59 terawatt-hours in a year. Its component model used 700 W for each H100, about 575 W per GPU for the rest of each server, and roughly another 10 percent for non-server IT equipment.[26] Those figures are an attributed estimate, not a measurement from a named operating cluster, and critical IT power does not include every facility loss captured by total site power.

At larger scales, companies increasingly describe infrastructure in megawatts or gigawatts rather than accelerator counts. This can include future buildings and utility capacity that is not yet energized. Capacity announcements should therefore retain their date and status language, such as "secured," "under construction," or "expected," rather than being reported as live cluster draw.

Nuclear and dedicated generation deals

In September 2024 Constellation and Microsoft announced a 20-year power-purchase agreement tied to restarting Three Mile Island Unit 1 as the Crane Clean Energy Center. Constellation said the unit would return 835 megawatts to the grid and require about $1.6 billion of investment, subject to regulatory approvals. Its current project page expects service in 2027.[27]

Talen Energy and Amazon expanded their agreement in June 2025. Talen said a power-purchase agreement would supply up to 1,920 megawatts from the Susquehanna nuclear plant through 2042, with the full quantity expected no later than 2032 and possible acceleration.[28] The agreement supports AWS Trainium and other cloud infrastructure indirectly through Amazon's data-center operations; it is not a dedicated count of GPUs or Trainium2 chips.

Air versus liquid cooling

Cooling requirements must be stated for a documented system configuration, not inferred from a GPU's thermal design power alone. NVIDIA specifies an air-cooled DGX H100 server at 10.2 kW maximum. Four such servers occupy a reference rack footprint of 40.8 kW.[12] NVIDIA's DGX B300 is also an air-cooled server: its user guide specifies eight B300 GPUs and 14.5 kW system power.[29]

Rack-scale Grace Blackwell systems are different. A GB200 NVL72 rack contains 72 GPUs and 36 Grace CPUs, uses liquid-cooled cold plates for the CPUs and GPUs, leaves networking and storage components air-cooled, and consumes approximately 120 kW.[30] NVIDIA's GB300 NVL72 reference architecture describes a liquid-cooled rack requiring up to 142 kW.[31] It is therefore incorrect to use a DGX B300 server specification and a GB300 NVL72 rack specification interchangeably.

Documented systemAcceleratorsDocumented system powerCooling description
DGX H1008 H100 GPUs10.2 kW maximum; 40.8 kW for four systems per reference rackAir [12]
DGX B3008 B300 GPUs14.5 kWAir [29]
DGX GB200 NVL7272 Blackwell GPUsApproximately 120 kW per rackLiquid for CPUs and GPUs; air for other components [30]
GB300 NVL7272 Blackwell Ultra GPUsUp to 142 kW per rackLiquid-cooled rack [31]
Vera Rubin NVL7272 Rubin GPUs and 36 Vera CPUsNot stated on the cited product pageDirect liquid cooling; published specifications remain preliminary [32][40]

What comes after Blackwell?

NVIDIA's standard Vera Rubin rack-scale product is NVL72, not NVL144. NVIDIA lists 72 Rubin GPUs, 36 Vera CPUs, 20.7 TB of GPU memory, and NVLink 6, while marking the specifications as preliminary.[32] A separate Rubin CPX configuration has been described as NVL144 CPX, so the two names should not be merged.

By 21 July 2026, NVIDIA said Vera Rubin NVL72 production was ramping and that racks were running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius.[33] This is a more current status than a statement that Rubin remained wholly unavailable. NVIDIA did not state rack power on the cited product page, so this article does not infer it.

How often do large GPU clusters fail?

Failure rates at scale

Synchronous training exposes the job to failures across every participating accelerator, host, network link, and storage path. Meta's Llama 3 paper provides a measured example. During a 54-day snapshot of pre-training on 16,384 H100 GPUs, Meta recorded 466 interruptions: 47 planned and 419 unexpected. It attributed 58.7 percent of unexpected interruptions to GPU issues, including 148 faulty-GPU events and 72 HBM3-memory events. Network switch or cable failures accounted for 35 events.[34]

Meta reported more than 90 percent effective training time during that period. That metric is the fraction of elapsed time spent on useful training, not the fraction of failures handled automatically. The paper says significant manual intervention was required only three times; automation handled the remaining issues.[34]

SemiAnalysis separately modeled a 100,000-H100 network and estimated a 26.28-minute mean time to the first job-affecting link failure under its stated per-link assumptions and without fault recovery.[26] That result is a model, not an observed failure interval for every 100,000-GPU cluster.

Checkpointing and recovery

Checkpointing periodically saves model weights, optimizer state, and other state needed to resume a job. Large synchronous jobs must balance checkpoint frequency against write bandwidth and the amount of computation lost after a failure. Recovery can include restarting from persistent storage, reconstructing state from other ranks, replacing unhealthy workers, or rescheduling onto a different set of nodes.

The Llama 3 case demonstrates that automation can handle most interruptions while still leaving measurable downtime.[34] NCCL 2.27 also added connection-management and reliability features intended for large training and inference jobs.[18] Neither source supports a universal claim that a failed node can always be swapped within minutes.

Hot spares and silent data corruption

Some operators keep powered spare nodes and replacement components on site, but repair and reintegration time depends on the fault.[26] A separate risk is silent data corruption, in which hardware produces an incorrect result without an immediate error signal. Meta has described production mechanisms including FleetScanner, Ripple, and Hardware Sentinel. It says periodic tests, co-located tests, telemetry, and analysis are used to identify faults across its fleet.[35]

Meta's Llama 3 record included six unexpected interruptions classified as GPU silent data corruption.[34] This provides a documented example, but it does not establish a universal per-GPU incidence rate for other hardware or operators.

How do GPU clusters store and feed training data?

Training storage serves several different traffic patterns: sequential or shuffled dataset reads, metadata operations, checkpoint writes, checkpoint reloads, and movement of data between regions. Required bandwidth depends on model, batch size, data format, caching, and checkpoint cadence. Dataset token count alone is not enough to infer a terabytes-per-second requirement.

Meta's 2024 clusters used a FUSE interface backed by a flash-optimized version of Tectonic and a parallel NFS deployment co-developed with Hammerspace. Meta said the system enabled synchronized checkpoint saves and loads across thousands of GPUs.[7] In July 2026 it described hundreds of exabyte-scale storage clusters built on Tectonic, with object, file, and block APIs. For AI workloads, Meta reported an 80 percent average hit rate for a distributed host-memory data cache and described dynamic concurrency control for checkpoint egress spikes.[36]

xAI's Colossus used DDN EXAScaler and Infinia in the documented 100,000-GPU configuration.[37] These examples show that a cluster can combine regional object or block storage, parallel file access, local NVMe, and distributed caching rather than relying on one storage tier.

How does an AI cluster differ from traditional HPC?

AI clusters and traditional high-performance computing systems share accelerators, low-latency networks, parallel storage, and batch scheduling. The distinction is mainly workload and design emphasis. Large-model training repeatedly executes dense or sparse matrix operations at low precision and exchanges tensors through collective communication. Scientific HPC workloads span a wider range of numerical methods, precision requirements, and communication patterns.

Frontier at Oak Ridge illustrates the overlap. The system has 9,408 HPE Cray EX nodes and 37,888 AMD Instinct MI250X GPUs, uses HPE Slingshot networking, and delivered more than 1.1 exaflops on the High Performance Linpack benchmark at about 21 megawatts.[15] GPU training systems may use similar components but are commonly evaluated with model throughput, model FLOP utilization, collective bandwidth, training availability, and cost per useful token rather than Linpack alone.

See also

References

  1. ^NVIDIA, "NVIDIA Launches World's First Deep Learning Supercomputer", NVIDIA Newsroom, 2016-04-05. nvidianews.nvidia.com/...ep-learning-supercomputer. Accessed 2026-08-17.
  2. ^Ashish Vaswani et al., "Attention Is All You Need", arXiv:1706.03762, 2017-06-12. arxiv.org/...1706.03762. Accessed 2026-08-17.
  3. ^Tom B. Brown et al., "Language Models are Few-Shot Learners", arXiv:2005.14165, 2020-05-28. arxiv.org/...2005.14165. Accessed 2026-08-17.
  4. ^Jennifer Langston, "Microsoft announces new supercomputer, lays out vision for future AI work", Microsoft Source, 2020-05-19. news.microsoft.com/...openai-azure-supercomputer. Accessed 2026-08-17.
  5. ^Microsoft Azure, "New infrastructure for the era of AI: Emerging technology and trends in 2024", Microsoft Azure Blog, 2024-02-26. azure.microsoft.com/...chnology-and-trends-in-2024. Accessed 2026-08-17.
  6. ^TOP500, "Eagle - Microsoft NDv5, Xeon Platinum 8480C 48C 2GHz, NVIDIA H100, NVIDIA Infiniband NDR", TOP500. top500.org/...180236. Accessed 2026-08-17.
  7. ^Kevin Lee, Adi Gangidi, and Mathew Oldham, "Building Meta's GenAI Infrastructure", Engineering at Meta, 2024-03-12. engineering.fb.com/...g-metas-genai-infrastructure. Accessed 2026-08-17.
  8. ^Meta Engineering, "Meta's Infrastructure Evolution and the Advent of AI", Engineering at Meta, 2025-09-29. engineering.fb.com/...olution-and-the-advent-of-ai. Accessed 2026-08-17.
  9. ^SpaceXAI, "Colossus: Our gigafactory of compute", SpaceXAI. x.ai/colossus. Accessed 2026-08-17.
  10. ^NVIDIA, "NVIDIA Ethernet Networking Accelerates World's Largest AI Supercomputer, Built by xAI", NVIDIA Newsroom, 2024-10-28. nvidianews.nvidia.com/...t-networking-xai-colossus. Accessed 2026-08-17.
  11. ^SpaceXAI, "New Compute Partnership with Anthropic", SpaceXAI, 2026-05-06. x.ai/...anthropic-compute-partnership. Accessed 2026-08-17.
  12. ^NVIDIA, "Planning a Data Center Deployment", NVIDIA DGX SuperPOD: Data Center Design Featuring NVIDIA DGX H100 Systems. docs.nvidia.com/...planning. Accessed 2026-08-17.
  13. ^NVIDIA, "DGX SuperPOD Components", NVIDIA DGX SuperPOD Reference Architecture Featuring NVIDIA DGX H100 Systems. docs.nvidia.com/...dgx-superpod-components. Accessed 2026-08-17.
  14. ^NVIDIA, "NVLink and NVLink Switch", NVIDIA. nvidia.com/...nvlink. Accessed 2026-08-17.
  15. ^Oak Ridge Leadership Computing Facility, "Frontier", Oak Ridge National Laboratory. olcf.ornl.gov/frontier. Accessed 2026-08-17.
  16. ^NVIDIA, "Advancing Performance with NVIDIA SHARP In-Network Computing", NVIDIA Technical Blog, 2024-02-15. developer.nvidia.com/...sharp-in-network-computing. Accessed 2026-08-17.
  17. ^NVIDIA, "NVIDIA Collective Communications Library (NCCL)", NVIDIA Developer. developer.nvidia.com/nccl. Accessed 2026-08-17.
  18. ^John Bachan et al., "Enabling Fast Inference and Resilient Training with NCCL 2.27", NVIDIA Technical Blog, 2025-07-14. developer.nvidia.com/...nt-training-with-nccl-2-27. Accessed 2026-08-17.
  19. ^OpenAI, "Announcing The Stargate Project", OpenAI, 2025-01-21. openai.com/...announcing-the-stargate-project. Accessed 2026-08-17.
  20. ^OpenAI, "Five new Stargate sites advance $500 billion, 10-gigawatt commitment", OpenAI, 2025-09-23. openai.com/...five-new-stargate-sites. Accessed 2026-08-17.
  21. ^OpenAI, "Building the compute infrastructure for the Intelligence Age", OpenAI, 2026-04-29. openai.com/...rastructure-for-the-intelligence-age. Accessed 2026-08-17.
  22. ^Anthropic, "Anthropic and Amazon expand collaboration for up to 5 gigawatts of new compute", Anthropic, 2026-04-20. anthropic.com/...anthropic-amazon-compute. Accessed 2026-08-17.
  23. ^Google Cloud, "Cloud TPU v5p", Google Cloud Documentation. cloud.google.com/...v5p. Accessed 2026-08-17.
  24. ^Mark Lohmeyer and George Elissaios, "Introducing Ironwood TPUs and new innovations in AI Hypercomputer", Google Cloud Blog, 2025-04-09. cloud.google.com/...whats-new-with-ai-hypercomputer. Accessed 2026-08-17.
  25. ^Thomas Kurian, "Welcome to Google Cloud Next 26", Google Cloud Blog, 2026-04-22. cloud.google.com/...welcome-to-google-cloud-next26. Accessed 2026-08-17.
  26. ^Dylan Patel and Daniel Nishball, "100,000 H100 Clusters: Power, Network Topology, Ethernet vs InfiniBand, Reliability, Failures, Checkpointing", SemiAnalysis, 2024-06-17. newsletter.semianalysis.com/...sters-power-network. Accessed 2026-08-17.
  27. ^Constellation Energy, "Crane Clean Energy Center", Constellation Energy. constellationenergy.com/...crane-clean-energy-center. Accessed 2026-08-17.
  28. ^Talen Energy, "Talen Energy Expands Nuclear Energy Relationship with Amazon", Talen Energy, 2025-06-11. ir.talenenergy.com/...r-energy-relationship-amazon. Accessed 2026-08-17.
  29. ^NVIDIA, "NVIDIA DGX B300 User Guide", NVIDIA Documentation. docs.nvidia.com/...dgxb300-user-guide.pdf. Accessed 2026-08-17.
  30. ^NVIDIA, "Hardware", NVIDIA DGX GB Rack Scale Systems User Guide. docs.nvidia.com/...hardware. Accessed 2026-08-17.
  31. ^NVIDIA, "System Hardware and Components", NVIDIA NVL72 AI Factory Reference Architecture. docs.nvidia.com/...components. Accessed 2026-08-17.
  32. ^NVIDIA, "NVIDIA DGX Vera Rubin NVL72", NVIDIA. nvidia.com/...dgx-vera-rubin-nvl72. Accessed 2026-08-17.
  33. ^NVIDIA, "NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide", NVIDIA Blog, 2026-07-21. blogs.nvidia.com/...vera-rubin. Accessed 2026-08-17.
  34. ^Aaron Grattafiori et al., "The Llama 3 Herd of Models", arXiv:2407.21783, 2024-07-31. arxiv.org/...2407.21783. Accessed 2026-08-17.
  35. ^Meta Engineering, "How Meta keeps its AI hardware reliable", Engineering at Meta, 2025-07-22. engineering.fb.com/...eps-its-ai-hardware-reliable. Accessed 2026-08-17.
  36. ^Meta Engineering, "Meta's AI Storage Blueprint at Scale", Engineering at Meta, 2026-07-01. engineering.fb.com/...i-storage-blueprint-at-scale. Accessed 2026-08-17.
  37. ^DDN, "DDN's Data Platform Propels xAI's Colossus to World-Class Performance", DDN, 2024-11-18. ddn.com/...ais-colossus-to-world-class-performance. Accessed 2026-08-17.
  38. ^OpenAI, "GPT-4 Technical Report", OpenAI, 2023-03-27. cdn.openai.com/...gpt-4.pdf. Accessed 2026-08-17.
  39. ^SpaceXAI, "xAI Raises $20B Series E", SpaceXAI, 2026-01-06. x.ai/...series-e. Accessed 2026-08-17.
  40. ^Kyle Aubrey, "Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer", NVIDIA Technical Blog, 2026-01-05, updated 2026-03-16. developer.nvidia.com/...chips-one-ai-supercomputer. Accessed 2026-08-17.
  41. ^OpenAI, "OpenAI joins PORTS-Pike project", OpenAI, 2026-08-17. openai.com/...openai-joins-ports-pike-project. Accessed 2026-08-17.

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

6 revisions · v7 · 4,854 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independently rechecked corrected scale, power, cooling, networking, reliability, storage, and deployment claims against primary sources through 2026-08-17.

Cite this page: AI Wiki. "GPU Cluster." aiwiki.ai, updated 17 Aug 2026, fact-checked 17 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/gpu_cluster

Suggest edit