NVIDIA H100
NVIDIA H100 (also called the H100 Tensor Core GPU) is a data-center graphics processing unit built by NVIDIA on the Hopper microarchitecture, fabricated with over 80 billion transistors on a custom TSMC 4N (4…
Explore AI Infrastructure through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Infrastructure.
Showing 61-98 of 98 articles
NVIDIA H100 (also called the H100 Tensor Core GPU) is a data-center graphics processing unit built by NVIDIA on the Hopper microarchitecture, fabricated with over 80 billion transistors on a custom TSMC 4N (4…
NVIDIA HGX is a family of accelerated-server platform designs built around tightly connected data-center GPUs. It is not one immutable board specification.
NVIDIA MGX is a modular reference architecture that NVIDIA publishes so that server makers and contract manufacturers can build accelerated systems around NVIDIA GPUs, CPUs, DPUs and networking without…
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
NVIDIA Quantum-X Photonics is a family of co-packaged-optics network switches that NVIDIA announced at its GTC conference on March 18, 2025.
NVIDIA Spectrum-6 is an Ethernet switch ASIC and system architecture for large AI infrastructure.
NVIDIA Spectrum-X Photonics is a co-packaged optics (CPO) Ethernet switch platform from Nvidia, announced at the company's GPU Technology Conference (GTC) on March 18, 2025.
NVIDIA Spectrum-XGS, marketed in full as NVIDIA Spectrum-XGS Ethernet, is a networking technology from NVIDIA designed to link several geographically separated data centers into a single
NVLink Fusion is a program and silicon technology from NVIDIA that opens its NVLink high speed interconnect to third party chips.
Neuromorphic computing is brain-inspired computer hardware that processes information with spiking neural networks (SNNs) and event-driven, in-memory computation, co-locating memory and processing the way…
The OpenAI-Broadcom collaboration is a strategic partnership, announced on October 13, 2025, under which OpenAI designs its own custom AI inference accelerators and Broadcom co-develops, manufactures, and…
PCI Express (PCIe) is a high-speed serial interconnect standard used to attach processors to graphics cards, network adapters, storage devices, and accelerators inside nearly every modern computer and server.
Project Rainier is a distributed artificial intelligence supercomputer built by Amazon Web Services around its in-house AWS Trainium accelerators, created principally to train and serve the Claude models of…
RISC-V (pronounced "risk-five") is an open standard instruction set architecture (ISA), the contract that defines which instructions a processor executes and which registers software can see.
ROCm is AMD's open software stack for GPU computing and the main alternative to NVIDIA's CUDA platform.
A robot is a programmable machine that senses its environment, computes a decision, and acts on the physical world through a closed loop of perception, planning, and actuation.
SRAM, or static random-access memory, is a semiconductor memory that stores each bit in a latch built from cross-coupled inverters.
The SambaNova SN40L is a reconfigurable dataflow AI accelerator designed by SambaNova Systems and unveiled on September 19, 2023.
The SambaNova SN50 is a reconfigurable dataflow accelerator that SambaNova Systems unveiled in late February 2026 as the successor to its SN40L chip.
SemiAnalysis is an independent research and analysis firm covering the semiconductor and artificial intelligence industries.
Space-based data centers are data centers launched into Earth orbit (or, in some proposals, placed on the Moon), where satellites carrying AI accelerators draw power from solar arrays and reject heat by…
SpaceX Starmind is SpaceX's planned constellation of solar-powered artificial intelligence compute satellites, intended to function as orbital data centers that run AI workloads in space and beam results back…
Starcloud, Inc. is an American space infrastructure startup that builds space-based data centers: satellites carrying data-center-class GPUs that are powered by solar arrays and cooled by radiating waste heat…
TPU Ironwood (officially TPU v7 or TPU7x) is Google's seventh-generation Tensor Processing Unit and the first TPU designed specifically for inference, unveiled at Google Cloud Next 2025 in Las Vegas on April…
A TPU node is the legacy Google Cloud architecture for accessing Tensor Processing Unit (TPU) hardware, in which a user's virtual machine (VM) runs application code and communicates with a separate
A TPU Pod is a single Google supercomputer built from many Tensor Processing Unit (TPU) chips wired directly to each other by a high-speed Inter-Chip Interconnect (ICI) fabric arranged as a 2D or 3D torus, so…
A TPU worker is a virtual machine (VM) running Linux that has direct access to one or more Tensor Processing Unit (TPU) chips and executes the actual TPU computation on that attached hardware.
A Tensor Core is a specialized execution unit inside NVIDIA GPUs that computes a small matrix multiplication and accumulation, D = A x B + C
A Tensor Processing Unit (TPU) is a family of custom application-specific integrated circuits developed by Google to accelerate machine-learning computation.
Tenstorrent is a North American artificial intelligence hardware and intellectual property company that designs processors for AI training and inference on the basis of the open-standard RISC-V instruction set…
Terafab (styled "Terafab" on the project's official site, "TERAFAB" in Elon Musk's announcement post, and frequently written "TeraFab" in press coverage) is a planned semiconductor fabrication venture between…
Trusted Execution Environments for machine learning (TEEs for ML, sometimes marketed as "Confidential AI" or "confidential inference") are deployments of hardware-isolated execution environments to run…
UALink (Ultra Accelerator Link) is an open industry standard for scale-up interconnect between AI accelerators that lets up to 1,024 accelerators inside a single pod read and write each other's memory directly…
Ultra Ethernet is an open networking specification that reworks Ethernet into a high performance fabric for large AI training clusters and high performance computing, giving operators an interoperable
Wormhole is the second-generation AI accelerator application-specific integrated circuit (ASIC) designed by Tenstorrent, a Toronto-based hardware startup led by chip architect Jim Keller.
bfloat16 (short for Brain Floating Point Format, sometimes written BF16) is a 16-bit floating-point number format that uses one sign bit, eight exponent bits, and seven mantissa bits
d-Matrix is a privately held American semiconductor company headquartered in Santa Clara, California that builds accelerators, I/O cards and software for AI inference in data centers.
Raptor is the second-generation AI inference accelerator from d-Matrix, a Santa Clara semiconductor startup, and the first commercial chip built on the company's 3D stacked digital in-memory compute technology