AI Infrastructure

Explore AI Infrastructure through related topics and the articles other pages reference most.

Most referenced in this topic

Ranked by links from other AI Wiki pages.

Explore articles

Browse subtopics (52)

Articles that also belong to these categories. Counts cover all of AI Infrastructure.

Showing 1-60 of 281 articles

AI Accelerator

An AI accelerator is hardware designed or configured to execute artificial intelligence and machine learning workloads more efficiently than a general-purpose processor executing the same workload without…

AI HardwareAI Inference

AI Chip

An AI chip is an integrated circuit, or a tightly integrated multi-die semiconductor package, designed or selected to execute artificial intelligence workloads efficiently.

AI Hardware

AI Infrastructure

AI infrastructure is the hardware, facilities, networking, storage, and software used to develop, train, evaluate, deploy, and operate artificial intelligence systems. It includes more than accelerator chips.

AI energy consumption

AI energy consumption is the electricity, and the associated water, land, and emissions, required to train and operate artificial intelligence systems, principally generative AI and large language models…

AI Energy

AMD Helios

AMD Helios is a rack-scale artificial intelligence system from AMD that packages 72 Instinct MI455X GPUs and 18 sixth-generation EPYC "Venice" CPUs into a single double-wide cabinet

AI Hardware

AMD Instinct MI300X

The AMD Instinct MI300X is a data center GPU accelerator that Advanced Micro Devices released on December 6, 2023, built on the CDNA 3 architecture and paired with 192 GB of HBM3 memory at 5.3 TB/s of bandwidth

AI HardwareData Centers

AMD Instinct MI325X

The AMD Instinct MI325X is a data center GPU accelerator from AMD for AI training and inference, built on the CDNA 3 architecture with 256 GB of HBM3E memory and 6 TB/s of memory bandwidth

AI HardwareData Centers

AMD Instinct MI355X

The AMD Instinct MI355X is a data center GPU accelerator built on AMD's CDNA 4 architecture, announced at the AMD Advancing AI 2025 event on June 12, 2025, and reaching general availability in October 2025.

AI HardwareData Centers

AMD Instinct MI430X

The AMD Instinct MI430X is a data center GPU accelerator built for scientific computing and sovereign AI, and the high-precision member of the AMD Instinct MI400 series.

AI Hardware

AMD Instinct MI455X

The AMD Instinct MI455X is a data center GPU accelerator announced by AMD on 23 July 2026 at Advancing AI 2026 in San Francisco.

AI Hardware

AMD Pensando

AMD Pensando is the data processing unit (DPU) and AI networking product line of AMD, built around Pensando Systems, a startup AMD acquired in 2022 in a transaction valued at approximately $1.9 billion .

AI HardwareData Centers

AWS Trainium

AWS Trainium is a family of custom machine learning accelerator chips designed by Annapurna Labs for Amazon Web Services, purpose-built for training and, increasingly, for serving large neural networks.

AI Hardware

AWS Trainium 2

AWS Trainium 2 (also written as Trainium2 and abbreviated Trn2) is the second generation of Amazon Web Services' custom machine learning training accelerator

AI CompaniesAI Hardware

AWS Trainium 3

AWS Trainium 3 (also written as Trainium3 and abbreviated Trn3) is the third-generation custom AI training and inference accelerator from Amazon Web Services, designed by Amazon's in-house chip team Annapurna…

AI Hardware

Abilene data center (Stargate)

The Abilene data center, also called the Crusoe Abilene Stargate Campus or Stargate I, is the flagship artificial intelligence data center of the Stargate Project, built by Crusoe Energy on a roughly 1,000…

OpenAI

Actor model

The actor model is a mathematical model of concurrent computation whose universal primitive is the actor: an autonomous, isolated entity that owns private state and communicates with other actors only by…

Computer ScienceProgramming Languages

Alexandr Wang

Alexandr Wang (born January 1997) is an American entrepreneur who co-founded Scale AI in 2016 and, in June 2025, became the first Chief AI Officer of Meta

AI CompaniesPeople

Alibaba AI

Alibaba AI refers to the artificial-intelligence work of Alibaba Group Holding Limited, a Chinese multinational technology conglomerate headquartered in Hangzhou.

AI CompaniesChinese AI

Amazon Q

Amazon Q is a family of generative AI-powered assistants from Amazon Web Services (AWS), announced on November 28, 2023, at the AWS re:Invent conference and made generally available on April 30, 2024.

AI Code GenerationArtificial Intelligence

Amazon SageMaker

Amazon SageMaker is Amazon Web Services' fully managed machine learning platform for building, training, and deploying models at scale, first launched at AWS re:Invent on November 29, 2017 and rebranded in…

AI Tools & ProductsMachine Learning

Amazon Web Services

Amazon Web Services (AWS) is the cloud-computing business of Amazon. It supplies computing, storage, database, networking, analytics, security, and application services from infrastructure operated by Amazon.

Andrew Tulloch

Andrew Tulloch is an Australian machine-learning researcher and engineer known for building large-scale machine-learning infrastructure at Meta (formerly Facebook), for training frontier models at OpenAI, and…

People

Anyscale

Anyscale is an American technology company, founded in 2019 by the creators of Ray at the University of California, Berkeley, that develops the Anyscale Platform: a fully managed compute service built on Ray…

AI Companies

Apache MXNet

Apache MXNet (pronounced "mix-net") was an open-source deep learning framework that combined imperative and symbolic execution in one runtime, created around 2015 by the DMLC (Distributed Machine Learning…

Artificial IntelligenceDeveloper Tools

Applied Digital

Applied Digital Corporation (Nasdaq: APLD) is a United States designer, builder, and operator of large data centers for artificial intelligence and high-performance computing, headquartered in Dallas, Texas.

AI Companies

Automatic Differentiation

Automatic differentiation (abbreviated AD, also called algorithmic differentiation, autodiff, or autograd) is a family of techniques for computing exact derivatives of a function specified by a computer program

Machine LearningMathematics

Baseten

Baseten is an inference platform for deploying, serving, and scaling machine learning models in production.

AI CompaniesMLOps

Bittensor

Bittensor is a decentralized machine learning network that uses blockchain-based incentives to pay independent contributors for producing digital commodities such as model inference, training, data, and raw…

AI CompaniesMachine Learning

Blackhole (Tenstorrent)

Blackhole is the third-generation AI accelerator architecture from Tenstorrent, the Toronto and Santa Clara based fabless semiconductor company led by CEO Jim Keller.

AI Hardware

Bloom Energy

Bloom Energy Corporation (NYSE: BE) is an American power-generation company based in San Jose, California, that designs and manufactures solid oxide fuel cells, sold as the Bloom Energy Server, which produce…

AI Energy

Broadcom Tomahawk 6

The Broadcom Tomahawk 6 is an Ethernet switch chip built for the networks that connect large clusters of AI accelerators.

AI Hardware

CANN (Huawei)

CANN, short for Compute Architecture for Neural Networks, is the heterogeneous computing architecture and software stack that Huawei provides for its Ascend line of AI processors.

Chinese AIDeveloper Tools

CME Compute Futures

CME Compute Futures are a planned family of financially settled futures contracts tied to Silicon Data benchmarks for hourly rentals of Nvidia H100 and B200 graphics processors.

AI HardwareData Centers

CUDA

CUDA (Compute Unified Device Architecture) is NVIDIA's platform and programming model for general-purpose computation on its graphics processing units.

Developer ToolsNVIDIA

CUTLASS

CUTLASS is an open-source library of reusable building blocks for writing high-performance matrix kernels on NVIDIA GPUs.

Developer ToolsNVIDIA

Cerebras WSE-3

The Cerebras WSE-3 (Wafer-Scale Engine 3) is the third-generation wafer-scale AI chip developed by Cerebras Systems, announced on March 13, 2024, and is the largest semiconductor ever built.

AI HardwareData Centers

Chroma

Chroma is an open-source embedding database designed for artificial intelligence applications.

Open Source AI

Chunked prefill

Chunked prefill is a scheduling technique for large language model serving that splits the processing of a long input prompt (the prefill) into smaller, fixed-size token chunks and combines each chunk with the…

Machine Learning

Cloud TPU

Cloud TPU is Google Cloud's offering of Tensor Processing Units (TPUs), the family of custom application-specific integrated circuits (ASICs) that Google builds to accelerate machine learning training and…

AI HardwareMachine Learning

Cloud computing

Cloud computing is the on-demand availability of computer system resources, especially data storage and computing power, delivered over the internet without active management by the end user.

AI Hardware

Cloudflare

Cloudflare, Inc. is an American internet infrastructure company that has become one of the most consequential players in the AI economy by running serverless AI inference at the network edge and by setting the…

AI Companies

Co-Packaged Optics

Co-packaged optics (CPO) is a hardware architecture that places active optical engines on the same first-level package substrate as a host application-specific integrated circuit (ASIC).

AI HardwareData Centers

Constellation Energy

Constellation Energy (Nasdaq: CEG) is the largest operator of nuclear power plants in the United States and the central electricity supplier to the artificial intelligence buildout

AI Energy

Context Parallelism

Context Parallelism (CP) is a distributed training strategy that partitions the input sequence dimension of a transformer across multiple accelerators and uses ring-style point-to-point communication to…

Training & Optimization

Continuous Batching

Continuous batching is a scheduling technique for large language model (LLM) inference servers that inserts new requests into a running batch at the granularity of individual model iterations rather than…

AI Inference