Category

AI Infrastructure

262 AI Wiki articles on AI Infrastructure. The most referenced are Data Center, CUDA and NVIDIA H100.

262 articlesRSS

Showing 1-60 of 262 articles

3D NAND

3D NAND is the vertical architecture that NAND flash memory adopted when shrinking cells sideways stopped working. Instead of laying memory cells out on the...

AI HardwareComputer Science

AI Accelerator

An AI accelerator is hardware designed or configured to execute artificial intelligence and machine learning workloads more efficiently than a general-purpose...

AI HardwareAI Inference

AI Accelerator Comparison (H100 vs B200 vs MI300 vs TPU)

As of July 2026, the best AI accelerator depends on the metric you care about, so the honest answer to "H100 vs B200 vs MI300 vs TPU" is not a single winner....

AI HardwareData Centers

AI Chip

An AI chip is an integrated circuit, or a tightly integrated multi-die semiconductor package, designed or selected to execute artificial intelligence workloads...

AI Hardware

AI Infrastructure

AI infrastructure is the hardware, facilities, networking, storage, and software used to develop, train, evaluate, deploy, and operate artificial intelligence...

AI energy consumption

AI energy consumption is the electricity, and the associated water, land, and emissions, required to train and operate artificial intelligence systems,...

AI Energy

AMD Advancing AI 2026

Advancing AI 2026 (styled AAI 2026 in the company's own materials) was an annual product and strategy conference held by AMD in San Francisco on July 22 and...

AI CompaniesAI Hardware

AMD Helios

AMD Helios is a rack-scale artificial intelligence system from AMD that packages 72 Instinct MI455X GPUs and 18 sixth-generation EPYC "Venice" CPUs into a...

AI Hardware

AMD Instinct MI300X

The AMD Instinct MI300X is a data center GPU accelerator that Advanced Micro Devices released on December 6, 2023, built on the CDNA 3 architecture and paired...

AI HardwareData Centers

AMD Instinct MI325X

The AMD Instinct MI325X is a data center GPU accelerator from AMD for AI training and inference, built on the CDNA 3 architecture with 256 GB of HBM3E memory...

AI HardwareData Centers

AMD Instinct MI355X

The AMD Instinct MI355X is a data center GPU accelerator built on AMD's CDNA 4 architecture, announced at the AMD Advancing AI 2025 event on June 12, 2025,[1]...

AI HardwareData Centers

AMD Instinct MI430X

The AMD Instinct MI430X is a data center GPU accelerator built for scientific computing and sovereign AI, and the high-precision member of the AMD Instinct...

AI Hardware

AMD Instinct MI455X

The AMD Instinct MI455X is a data center GPU accelerator announced by AMD on 23 July 2026 at Advancing AI 2026 in San Francisco. It is the flagship, AI-focused...

AI Hardware

AMD Pensando

AMD Pensando is the data processing unit (DPU) and AI networking product line of AMD, built around Pensando Systems, a startup AMD acquired in 2022 in a...

AI HardwareData Centers

AQLM (Additive Quantization of Language Models)

AQLM, short for Additive Quantization of Language Models, is a weight-only post-training quantization method that compresses the weights of a large language...

Machine Learning

AWS Graviton

AWS Graviton is a family of Arm-based server processors designed by Amazon Web Services for use in its own cloud computing fleet. The first generation launched...

AI HardwareAI Inference

AWS Trainium

AWS Trainium is a family of custom machine learning accelerator chips designed by Annapurna Labs for Amazon Web Services, purpose-built for training and,...

AI Hardware

AWS Trainium 2

AWS Trainium 2 (also written as Trainium2 and abbreviated Trn2) is the second generation of Amazon Web Services' custom machine learning training accelerator,...

AI CompaniesAI Hardware

AWS Trainium 3

AWS Trainium 3 (also written as Trainium3 and abbreviated Trn3) is the third-generation custom AI training and inference accelerator from Amazon Web Services,...

AI Hardware

Abilene data center (Stargate)

The Abilene data center, also called the Crusoe Abilene Stargate Campus or Stargate I, is the flagship artificial intelligence data center of the Stargate...

OpenAI

Actor model

The actor model is a mathematical model of concurrent computation whose universal primitive is the actor: an autonomous, isolated entity that owns private...

Computer ScienceProgramming Languages

Agent Payments Protocol (AP2)

The Agent Payments Protocol (AP2) is an open specification that lets autonomous AI agents initiate, authorize, and settle payments on behalf of human users. It...

AI AgentsAgentic Commerce

Alexandr Wang

Alexandr Wang (born January 1997) is an American entrepreneur who co-founded Scale AI in 2016 and, in June 2025, became the first Chief AI Officer of Meta,...

AI CompaniesPeople

Alibaba AI

Alibaba AI refers to the artificial-intelligence work of Alibaba Group Holding Limited, a Chinese multinational technology conglomerate headquartered in...

AI CompaniesChinese AI

Amazon Nova

Amazon Nova is a family of foundation models developed by Amazon and offered through Amazon Bedrock, announced on December 3, 2024, at the AWS re:Invent...

Artificial IntelligenceGenerative AI

Amazon Q

Amazon Q is a family of generative AI-powered assistants from Amazon Web Services (AWS), announced on November 28, 2023, at the AWS re:Invent conference and...

AI Code GenerationArtificial Intelligence

Amazon SageMaker

Amazon SageMaker is Amazon Web Services' fully managed machine learning platform for building, training, and deploying models at scale, first launched at AWS...

AI Tools & ProductsMachine Learning

Amazon Web Services

Amazon Web Services (AWS) is the cloud-computing business of Amazon. It supplies computing, storage, database, networking, analytics, security, and application...

Andrew Tulloch

Andrew Tulloch is an Australian machine-learning researcher and engineer known for building large-scale machine-learning infrastructure at Meta (formerly...

People

Anthropic-Amazon Trainium expansion

The Anthropic-Amazon Trainium expansion is an enlarged compute and investment agreement between Anthropic and Amazon, announced on 20 April 2026, under which...

AnthropicData Centers

Anthropic-Google TPU deal

The Anthropic-Google TPU deal is a set of escalating compute and investment agreements between the AI lab Anthropic and Google Cloud, built around Anthropic...

AnthropicGoogle

Anyscale

Private Industry 2019; Berkeley, California Founders San Francisco, California, United States CEO Ray (open-source distributed computing framework) ...

AI Companies

Apache MXNet

Apache MXNet (pronounced "mix-net") was an open-source deep learning framework that combined imperative and symbolic execution in one runtime, created around...

Artificial IntelligenceDeveloper Tools

Applied Digital

Applied Digital Corporation (Nasdaq: APLD) is a United States designer, builder, and operator of large data centers for artificial intelligence and...

AI Companies

Automatic Differentiation

Automatic differentiation (abbreviated AD, also called algorithmic differentiation, autodiff, or autograd) is a family of techniques for computing exact...

Machine LearningMathematics

Baseten

Baseten is an inference platform for deploying, serving, and scaling machine learning models in production. The company converts ML models into...

AI CompaniesMLOps

Blackhole (Tenstorrent)

Blackhole is the third-generation AI accelerator architecture from Tenstorrent, the Toronto and Santa Clara based fabless semiconductor company led by CEO Jim...

AI Hardware

Bloom Energy

Bloom Energy Corporation (NYSE: BE) is an American power-generation company based in San Jose, California, that designs and manufactures solid oxide fuel...

AI Energy

Broadcom Tomahawk 6

The Broadcom Tomahawk 6 is an Ethernet switch chip built for the networks that connect large clusters of AI accelerators. Broadcom announced it on June 3,...

AI Hardware

CANN (Huawei)

CANN, short for Compute Architecture for Neural Networks, is the heterogeneous computing architecture and software stack that Huawei provides for its Ascend...

Chinese AIDeveloper Tools

CHIPS and Science Act

The CHIPS and Science Act is a United States federal law, enacted as Public Law 117-167 on August 9, 2022, that appropriated $52.7 billion for domestic...

AI HardwareAI Policy & Regulation

CUDA

CUDA (Compute Unified Device Architecture) is NVIDIA's platform and programming model for general-purpose computation on its graphics processing units. NVIDIA...

Developer ToolsNVIDIA

CUTLASS

CUTLASS is an open-source library of reusable building blocks for writing high-performance matrix kernels on NVIDIA GPUs. The name expands to CUDA Templates...

Developer ToolsNVIDIA

Central processing unit

A central processing unit (CPU) is the general-purpose processor that executes a computer's instruction stream. It fetches instructions from memory, decodes...

AI HardwareComputer Science

Cerebras WSE-3

The Cerebras WSE-3 (Wafer-Scale Engine 3) is the third-generation wafer-scale AI chip developed by Cerebras Systems, announced on March 13, 2024, and is the...

AI HardwareData Centers

China's semiconductor industry

China's semiconductor industry is the network of wafer fabrication plants, equipment and materials suppliers, chip design houses, packaging and test...

AI HardwareAI Policy & Regulation

Chroma

Chroma is an open-source embedding database designed for artificial intelligence applications. It stores vector embeddings alongside their associated documents...

Open Source AI

Chunked prefill

Chunked prefill is a scheduling technique for large language model serving that splits the processing of a long input prompt (the prefill) into smaller,...

Machine Learning

Cloud AI GPU Pricing Comparison

Cloud AI GPU pricing is difficult to compare from a headline hourly rate alone. Providers sell different GPU variants, node sizes, CPU and memory bundles,...

AI HardwareData Centers

Cloud TPU

Cloud TPU is Google Cloud's offering of Tensor Processing Units (TPUs), the family of custom application-specific integrated circuits (ASICs) that Google...

AI HardwareMachine Learning

Cloud computing

Cloud computing is the on-demand availability of computer system resources, especially data storage and computing power, delivered over the internet without...

AI Hardware

Cloudflare

Cloudflare, Inc. is an American internet infrastructure company that has become one of the most consequential players in the AI economy by running serverless...

AI Companies

ClusterMAX

ClusterMAX is a rating and ranking system for GPU cloud providers published by the research firm SemiAnalysis. It scores providers against ten criteria and...

Data CentersModel Evaluation

Constellation Energy

Constellation Energy (Nasdaq: CEG) is the largest operator of nuclear power plants in the United States and the central electricity supplier to the artificial...

AI Energy

Context Parallelism

Context Parallelism (CP) is a distributed training strategy that partitions the input sequence dimension of a transformer across multiple accelerators and uses...

Training & Optimization

Continuous Batching

Continuous batching is a scheduling technique for large language model (LLM) inference servers that inserts new requests into a running batch at the...

AI Inference

Contrastive decoding

Contrastive decoding (CD) is a decoding strategy for text generation from a large language model that selects tokens by contrasting two models of different...

Machine Learning

CoreWeave

CoreWeave, Inc. is an American artificial intelligence cloud computing company headquartered in Livingston, New Jersey, that rents large clusters of NVIDIA...

AI Companies

Crusoe

Crusoe is an American AI infrastructure company headquartered in Denver, Colorado, that builds and operates data centers for artificial intelligence workloads....

AI CompaniesData Centers

DRAM (Dynamic Random-Access Memory)

Dynamic random-access memory (DRAM) is the working memory of nearly every computer built since the late 1970s, and the physical substrate on which modern AI...

AI HardwareComputer Science