AI Infrastructure

Explore AI Infrastructure through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (52)

Articles that also belong to these categories. Counts cover all of AI Infrastructure.

Showing 61-120 of 281 articles

Contrastive decoding

Contrastive decoding (CD) is a decoding strategy for text generation from a large language model that selects tokens by contrasting two models of different sizes.

Machine Learning

CoreWeave

CoreWeave, Inc. is an American artificial intelligence cloud computing company headquartered in Livingston, New Jersey, that rents large clusters of NVIDIA GPUs to AI developers and enterprises and builds the…

AI Companies

Crusoe

Crusoe is an American AI infrastructure company headquartered in Denver, Colorado, that builds and operates data centers for artificial intelligence workloads.

AI CompaniesData Centers

Dask

Dask is an open-source Python library for parallel and distributed computing that scales the familiar APIs of libraries such as NumPy, pandas, and scikit-learn to process larger-than-memory datasets.

Data ScienceMachine Learning

Data Center

A data center is a purpose-built facility, or a dedicated part of a facility, that houses and interconnects information technology and telecommunications equipment together with the power…

AI Hardware

Data Parallelism

Data parallelism is a distributed training technique in which the same neural network model is replicated across multiple processing units (typically GPUs), each device trains on a different shard of the input…

Deep LearningMachine Learning

DeepInfra

DeepInfra is a serverless AI inference cloud that hosts open-source and open-weight AI models and serves them to developers through a single pay-per-token API.

AI CompaniesAI Inference

DeepSpeed

DeepSpeed is an open-source deep learning optimization library, originally developed by Microsoft, that makes distributed training and inference of large models efficient, easy to use, and cost-effective.

Deep LearningMachine Learning

DriveNets

DriveNets is an Israeli networking-software company that builds large telecom and AI networks out of standard "white box" hardware controlled by cloud-native software instead of traditional proprietary routers.

AI Companies

DualPipe

DualPipe is a bidirectional pipeline parallelism scheduling algorithm created by DeepSeek-AI to almost fully overlap computation with communication and shrink the idle time, called the pipeline "bubble," that…

AI Agents

EAGLE (speculative decoding)

EAGLE (Extrapolation Algorithm for Greater Language-model Efficiency) is a lossless speculative decoding method that speeds up large language model (LLM) inference by 2x to 6.5x by doing autoregression at the…

AI Inference

Edge computing

Edge computing is a distributed computing paradigm that runs computation and data storage close to where data is generated, at the "edge" of the network

AI Hardware

FPGA

A field-programmable gate array (FPGA) is an integrated circuit whose logic functions and internal wiring are set by the customer after the chip has been manufactured, and can be reset later.

AI HardwareAI Inference

Feature store

A feature store is a centralised data system that stores, serves, discovers, shares, monitors and reuses machine-learning features, separating feature computation from model training and inference so the same…

MLOps

Fermi America

Fermi America is a United States power and data center developer attempting to build what it calls the world's largest energy and computing complex

AI Energy

Firebase

Firebase is a backend-as-a-service (BaaS) platform developed by Google. It began in 2012 as a real-time database product launched by Andrew Lee and James Tamplin and was acquired by Google on October 21, 2014.

Developer ToolsGoogle

Frozen v2 (Google AI chip)

Frozen v2 is the informal internal codename for a specialized artificial-intelligence inference chip that Google is reportedly developing to run its Gemini models more cheaply and with far less energy.

AI HardwareGoogle

GE Vernova

GE Vernova (NYSE: GEV) is an American energy and electric-power equipment company, headquartered in Cambridge, Massachusetts, that builds and services much of the hardware used to generate and move…

AI Energy

GEM (Meta)

GEM (Generative Ads Recommendation Model) is a proprietary foundation model for advertising recommendation developed by Meta Platforms.

AI ModelsEnterprise AI

GPU Cluster

A GPU cluster is a group of servers containing graphics processing units that are connected by high-bandwidth, low-latency networks and operated as one computational pool.

AI HardwareData Centers

GPU computing

GPU computing is the use of a graphics processing unit (GPU) to perform general-purpose computation that was traditionally handled by the central processing unit (CPU).

AI HardwareDeep Learning

Genesis (simulator)

Genesis is an open-source, generative physics simulation platform for robotics and embodied AI, released on December 19, 2024 after a roughly two-year (24-month) collaboration involving more than 20 academic…

Open Source AIRobotics

Google Cloud

Google Cloud is the enterprise cloud business of Google and the name of a reportable segment in Alphabet's financial statements.

AI Companies

Groq Hardware

Groq hardware is a family of artificial-intelligence accelerators and multi-chip systems built around a statically scheduled streaming architecture.

AI HardwareAI Inference

HUMAIN

HUMAIN is a Saudi Arabian artificial intelligence company launched on 12 May 2025 and owned by the Public Investment Fund (PIF), the kingdom's roughly trillion-dollar sovereign wealth fund.

AI Companies

Horovod

Horovod is an open-source distributed training framework for deep learning that lets a single-GPU training script scale across many GPUs and many machines by adding only a few lines of code.

Developer ToolsOpen Source AI

Huawei Ascend 910B

The Huawei Ascend 910B is a data-center AI accelerator designed by Huawei's HiSilicon unit that became China's most widely deployed domestic alternative to restricted NVIDIA data-center GPUs across roughly…

AI HardwareChinese AI

Huawei Ascend 910C

The Huawei Ascend 910C is a data-center artificial intelligence accelerator developed by Huawei and positioned as China's leading domestic alternative to high-end NVIDIA GPUs that are barred from sale to…

AI HardwareChinese AI

Hyperscaler

A hyperscaler is a company that builds and operates computing infrastructure at a scale far beyond a conventional enterprise IT estate: globally distributed fleets of data centers holding millions of servers…

AI EnergyAI Hardware

IREN

IREN Limited is an Australian-founded data center company listed on the Nasdaq Global Select Market under the ticker IREN.

AI Companies

Inferact

Inferact is an artificial intelligence infrastructure company founded by creators and core maintainers of vLLM, the open-source inference engine that emerged from the University of California, Berkeley in 2023.

AI CompaniesAI Inference

InfiniBand

InfiniBand is a high throughput, low latency networking interconnect standard used to connect servers, storage, and accelerators inside high performance computing systems and large artificial intelligence…

AI Hardware

Internet of Things

The Internet of Things (IoT) is the network of physical objects ("things") embedded with sensors, software, and connectivity that lets them collect data, exchange it with other devices and systems over the…

AI HardwareEnterprise AI

Invisible Technologies

Invisible Technologies is an American artificial intelligence company that combines a global network of vetted human experts with its own orchestration software to provide AI training data, reinforcement…

AI Companies

Ion Stoica

Ion Stoica is a Romanian-American computer scientist, professor of electrical engineering and computer sciences at the University of California, Berkeley, and a serial entrepreneur whose academic and…

Computer SciencePeople

KAI Scheduler

KAI Scheduler is an open-source Kubernetes scheduler that optimizes the allocation of GPU resources for artificial intelligence and machine learning workloads.

MLOpsNVIDIA

KV-cache quantization

KV cache quantization is a family of large language model inference optimizations that store the attention key and value (KV) cache in low-bit numeric formats, typically 2 to 4 bits per value

Machine Learning

Kairos Power

Kairos Power is a United States advanced nuclear reactor developer, founded in 2016 and headquartered in Alameda, California

AI Energy

Kioxia

Kioxia is a Japanese flash memory manufacturer, the direct corporate descendant of the Toshiba division where flash memory was invented in the 1980s.

AI CompaniesAI Hardware

LMDeploy

LMDeploy is an open-source toolkit for compressing, deploying, and serving large language models, developed by the MMRazor and MMDeploy teams associated with the InternLM project at the Shanghai AI Laboratory.

Developer ToolsOpen Source AI