AI Infrastructure

Explore AI Infrastructure through related topics and the articles other pages reference most.

Explore articles

Browse subtopics (52)

Articles that also belong to these categories. Counts cover all of AI Infrastructure.

Showing 121-180 of 281 articles

Lambda Labs

Lambda Labs (operating as Lambda, Inc.) is an American AI infrastructure company that provides GPU cloud computing, on-premises GPU hardware, and deep learning software for artificial intelligence research and…

AI Companies

LanceDB

LanceDB is an open-source, developer-friendly vector database and multimodal lakehouse built on the Lance columnar storage format, designed to store vector embeddings, images, video, audio, and structured…

AI CompaniesDeveloper Tools

Land, Power, and Shell

Land, power, and shell (LPS) is an emerging commercial label for the physical site, deliverable electricity, and building structure needed before computing equipment can be installed in a large data center.

Data Centers

LangSmith

LangSmith is a commercial observability, evaluation, and deployment platform for large language model (LLM) applications and AI agents, developed and operated by LangChain Inc. It provides developers and…

AI CompaniesDeveloper Tools

Larry Ellison

Larry Ellison (born August 17, 1944) is an American businessman who co-founded the database company that became Oracle Corporation in 1977.

Enterprise AIPeople

Lenovo

Lenovo Group Limited is a Chinese multinational technology company that is the world's largest personal computer vendor by unit shipments and one of the largest builders of AI-optimized data center hardware…

AI CompaniesAI Hardware

Lepton AI

Lepton AI was an American AI cloud company, founded in 2023, that built a cloud-native inference platform for serving large language models, generative image models, and other AI workloads on NVIDIA GPUs.

AI CompaniesNVIDIA

MCP server

An MCP server is a program that implements the Model Context Protocol (MCP) to expose tools, resources, and prompts to MCP clients running inside an AI application or AI agent

AI AgentsAnthropic

MTIA

MTIA (Meta Training and Inference Accelerator) is a family of custom silicon chips that Meta designs for use in its own data centers rather than for sale.

AI HardwareAI Inference

Mark Zuckerberg

Mark Zuckerberg is the co-founder, chairman, and chief executive officer of Meta Platforms, and since the early 2010s he has directed one of the largest corporate artificial-intelligence programs in the world

AI CompaniesPeople

Medusa

Medusa is a large language model inference acceleration framework that speeds up text generation by adding multiple lightweight decoding heads on top of an existing model to predict several future tokens in…

AI Inference

Meta Compute

Meta Compute is a top-level organization that Meta created in January 2026 to plan, build, and run the gigawatt-scale data center capacity behind its push toward artificial general intelligence and what the…

Data CentersMeta AI

Meta MTIA

MTIA (Meta Training and Inference Accelerator) is a family of custom AI chips that Meta designs in-house to run its largest artificial intelligence workloads, beginning with the deep learning recommendation…

AI Hardware

Microscaling formats

Microscaling (MX) formats are a family of low-precision number formats for machine learning in which a small block of values, normally 32 of them, shares one common scale factor while each value is stored in a…

AI HardwareMachine Learning

Microsoft Azure

Microsoft Azure is the public cloud computing platform operated by Microsoft, and the world's second-largest cloud provider, behind Amazon Web Services (AWS) and ahead of Google Cloud.

Microsoft

Milvus

Milvus is an open-source vector database built for billion-scale similarity search, developed by Zilliz and governed under the Linux Foundation AI & Data Foundation.

Open Source AI

MoE load-balancing loss

The MoE load balancing loss is an auxiliary training objective used in sparse mixture of experts (MoE) neural networks to keep work spread evenly across the expert sub-networks.

AI Agents

Modal (platform)

Modal is a serverless cloud computing platform that lets developers run compute-intensive artificial intelligence, machine learning, and data-processing workloads on cloud GPUs by writing ordinary Python, with…

AI Companies

Model Parallelism

Model parallelism is a distributed training and inference technique that splits a single neural network across multiple processing units so that no individual accelerator has to hold the entire model.

Deep LearningMachine Learning

Model deployment

Model deployment is the MLOps process of taking a trained machine learning model and making it available in a production environment so it can serve predictions to applications, users, or downstream systems.

MLOps

MoonEP

MoonEP is an open-source expert-parallel communication library for mixture-of-experts models, released by Moonshot AI on July 27, 2026 under the MIT License.

Chinese AIMixture of Experts

NAND flash memory

NAND flash memory is the non-volatile storage technology that holds the data in solid-state drives, memory cards, USB sticks, phones and the flash tiers of a modern data center.

AI HardwareComputer Science

NLWeb

NLWeb (short for Natural Language Web) is an open project from Microsoft that makes it easy to add a natural-language conversational interface to a website, drawing on the site's own existing structured data…

AI Agents

NVHBM

NVHBM is an announced custom high-bandwidth memory architecture from NVIDIA for custom AI accelerators that participate in the company's NVLink Fusion platform. NVIDIA introduced it on August 26, 2026.

AI HardwareNVIDIA

NVIDIA AI Enterprise

NVIDIA AI Enterprise is an end-to-end, cloud-native software suite sold by Nvidia as a paid subscription for developing and deploying production artificial intelligence and data analytics.

Enterprise AINVIDIA

NVIDIA B200

The NVIDIA B200 is a data center GPU based on the NVIDIA Blackwell microarchitecture, announced by Jensen Huang at GTC 2024 on March 18, 2024.

AI HardwareData Centers

NVIDIA ConnectX

NVIDIA ConnectX is a family of high-speed network adapters and SmartNICs (smart network interface cards) that connect a server to the data center fabric and accelerate networking in hardware.

AI HardwareNVIDIA

NVIDIA DGX Cloud

NVIDIA DGX Cloud is a managed AI-supercomputing-as-a-service offering from Nvidia that rents enterprises access to multi-node clusters of NVIDIA DGX infrastructure plus the NVIDIA AI software stack over the…

NVIDIA

NVIDIA DGX Station

NVIDIA DGX Station is a line of deskside artificial intelligence workstations from NVIDIA, each marketed as a "personal AI supercomputer" that puts data-center-class compute next to a developer's desk rather…

AI HardwareNVIDIA

NVIDIA Deep Learning Institute

NVIDIA Deep Learning Institute (DLI) is the training and education arm of NVIDIA, offering hands-on courses, instructor-led workshops, and professional certifications in artificial intelligence, accelerated…

Developer ToolsNVIDIA

NVIDIA Dynamo

NVIDIA Dynamo is an open-source, low-latency distributed inference serving framework designed to deploy and scale generative AI and reasoning models across large GPU clusters.

AI InferenceDeveloper Tools

NVIDIA Exemplar Cloud

NVIDIA Exemplar Cloud is a validation program run by NVIDIA that certifies cloud providers whose GPU clusters reproduce at least 95% of the training throughput NVIDIA measures on its own reference architecture…

Data CentersNVIDIA

NVIDIA GB300 NVL72

The NVIDIA GB300 NVL72 is a liquid-cooled, rack-scale AI computing system that integrates 72 NVIDIA Blackwell Ultra (B300) GPUs and 36 Arm-based NVIDIA Grace CPUs into a single NVLink fabric, delivering 1.1…

AI HardwareData Centers

NVIDIA H100

NVIDIA H100 (also called the H100 Tensor Core GPU) is a data-center graphics processing unit built by NVIDIA on the Hopper microarchitecture, fabricated with over 80 billion transistors on a custom TSMC 4N (4…

AI HardwareNVIDIA

NVIDIA HGX

NVIDIA HGX is a family of accelerated-server platform designs built around tightly connected data-center GPUs. It is not one immutable board specification.

AI HardwareData Centers

NVIDIA MGX

NVIDIA MGX is a modular reference architecture that NVIDIA publishes so that server makers and contract manufacturers can build accelerated systems around NVIDIA GPUs, CPUs, DPUs and networking without…

AI HardwareNVIDIA

NVIDIA NIM

NVIDIA NIM (NVIDIA Inference Microservices) is a set of containerized, prebuilt-and-optimized model-serving microservices from NVIDIA that package an AI model, an optimized inference engine, and an…

AI InferenceDeveloper Tools

NVIDIA NeMo

NVIDIA NeMo is an open, end-to-end framework from NVIDIA for building, customizing, and deploying generative AI models, described by NVIDIA as "a scalable generative AI framework built for researchers and…

Developer ToolsNVIDIA

NVIDIA Picasso

NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.

AI HardwareAI Inference