Lambda Labs
Lambda Labs (operating as Lambda, Inc.) is an American AI infrastructure company that provides GPU cloud computing, on-premises GPU hardware, and deep learning software for artificial intelligence research and…
Explore AI Infrastructure through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Infrastructure.
Showing 121-180 of 281 articles
Lambda Labs (operating as Lambda, Inc.) is an American AI infrastructure company that provides GPU cloud computing, on-premises GPU hardware, and deep learning software for artificial intelligence research and…
LanceDB is an open-source, developer-friendly vector database and multimodal lakehouse built on the Lance columnar storage format, designed to store vector embeddings, images, video, audio, and structured…
Land, power, and shell (LPS) is an emerging commercial label for the physical site, deliverable electricity, and building structure needed before computing equipment can be installed in a large data center.
LangSmith is a commercial observability, evaluation, and deployment platform for large language model (LLM) applications and AI agents, developed and operated by LangChain Inc. It provides developers and…
As of July 2026, the largest AI funding round on record is OpenAI's $122 billion raise, which closed in March 2026 at an $852 billion post-money valuation, the biggest private financing in history .
Larry Ellison (born August 17, 1944) is an American businessman who co-founded the database company that became Oracle Corporation in 1977.
Lenovo Group Limited is a Chinese multinational technology company that is the world's largest personal computer vendor by unit shipments and one of the largest builders of AI-optimized data center hardware…
Lepton AI was an American AI cloud company, founded in 2023, that built a cloud-native inference platform for serving large language models, generative image models, and other AI workloads on NVIDIA GPUs.
The Linux Foundation is an American 501(c)(6) nonprofit trade association based in San Francisco, California, that supports the development of Linux and other open-source software projects.
An MCP server is a program that implements the Model Context Protocol (MCP) to expose tools, resources, and prompts to MCP clients running inside an AI application or AI agent
MTIA (Meta Training and Inference Accelerator) is a family of custom silicon chips that Meta designs for use in its own data centers rather than for sale.
Google Cloud is the public cloud arm of Google and one of the three dominant providers of machine learning infrastructure, alongside Amazon Web Services and Microsoft Azure.
Mark Zuckerberg is the co-founder, chairman, and chief executive officer of Meta Platforms, and since the early 2010s he has directed one of the largest corporate artificial-intelligence programs in the world
MediaTek Inc. is a Taiwanese fabless semiconductor company headquartered in Hsinchu, Taiwan.
Medusa is a large language model inference acceleration framework that speeds up text generation by adding multiple lightweight decoding heads on top of an existing model to predict several future tokens in…
Meta Compute is a top-level organization that Meta created in January 2026 to plan, build, and run the gigawatt-scale data center capacity behind its push toward artificial general intelligence and what the…
MTIA (Meta Training and Inference Accelerator) is a family of custom AI chips that Meta designs in-house to run its largest artificial intelligence workloads, beginning with the deep learning recommendation…
Meta Superintelligence Labs (MSL) is the artificial intelligence division that Meta created on June 30, 2025 to consolidate its research, foundation models, and AI products under a single organization aimed at…
Microscaling (MX) formats are a family of low-precision number formats for machine learning in which a small block of values, normally 32 of them, shares one common scale factor while each value is stored in a…
Microsoft Azure is the public cloud computing platform operated by Microsoft, and the world's second-largest cloud provider, behind Amazon Web Services (AWS) and ahead of Google Cloud.
Fairwater is Microsoft's name for a class of large AI datacenter built to train and serve frontier artificial intelligence models.
Foundry Local is an on-device artificial intelligence runtime from Microsoft that lets applications run open weight language models entirely on a user's own hardware.
Microsoft Maia 200 is a custom artificial intelligence accelerator designed by Microsoft for inference workloads in the Azure cloud.
Milvus is an open-source vector database built for billion-scale similarity search, developed by Zilliz and governed under the Linux Foundation AI & Data Foundation.
The MoE load balancing loss is an auxiliary training objective used in sparse mixture of experts (MoE) neural networks to keep work spread evenly across the expert sub-networks.
Modal is a serverless cloud computing platform that lets developers run compute-intensive artificial intelligence, machine learning, and data-processing workloads on cloud GPUs by writing ordinary Python, with…
Model parallelism is a distributed training and inference technique that splits a single neural network across multiple processing units so that no individual accelerator has to hold the entire model.
Model deployment is the MLOps process of taking a trained machine learning model and making it available in a production environment so it can serve predictions to applications, users, or downstream systems.
A model hub is an online platform or repository where people discover, share, version, and download machine learning models.
MoonEP is an open-source expert-parallel communication library for mixture-of-experts models, released by Moonshot AI on July 27, 2026 under the MIT License.
Mooncake is a KVCache-centric, disaggregated serving architecture for large language models, built by Moonshot AI together with researchers at Tsinghua University.
Mubadala Investment Company is a state-owned strategic investor based in Abu Dhabi and wholly owned by the government of the emirate.
NAND flash memory is the non-volatile storage technology that holds the data in solid-state drives, memory cards, USB sticks, phones and the flash tiers of a modern data center.
NCCL, the NVIDIA Collective Communications Library, is a library of topology-aware communication primitives for NVIDIA GPU systems.
NLWeb (short for Natural Language Web) is an open project from Microsoft that makes it easy to add a natural-language conversational interface to a website, drawing on the site's own existing structured data…
NVHBM is an announced custom high-bandwidth memory architecture from NVIDIA for custom AI accelerators that participate in the company's NVLink Fusion platform. NVIDIA introduced it on August 26, 2026.
NVIDIA AI Enterprise is an end-to-end, cloud-native software suite sold by Nvidia as a paid subscription for developing and deploying production artificial intelligence and data analytics.
The NVIDIA AI compute infrastructure financing platforms are a set of proposed, independently run financing vehicles that NVIDIA announced on August 10, 2026, together with Apollo, BlackRock, Blackstone…
The NVIDIA B200 is a data center GPU based on the NVIDIA Blackwell microarchitecture, announced by Jensen Huang at GTC 2024 on March 18, 2024.
NVIDIA ConnectX is a family of high-speed network adapters and SmartNICs (smart network interface cards) that connect a server to the data center fabric and accelerate networking in hardware.
The NVIDIA DGX B300 is an 8-GPU AI supercomputer node built around NVIDIA's Blackwell Ultra architecture.
NVIDIA DGX Cloud is a managed AI-supercomputing-as-a-service offering from Nvidia that rents enterprises access to multi-node clusters of NVIDIA DGX infrastructure plus the NVIDIA AI software stack over the…
NVIDIA DGX Station is a line of deskside artificial intelligence workstations from NVIDIA, each marketed as a "personal AI supercomputer" that puts data-center-class compute next to a developer's desk rather…
The NVIDIA DGX SuperPOD is a reference-architecture artificial intelligence supercomputer designed and sold by Nvidia.
NVIDIA Deep Learning Institute (DLI) is the training and education arm of NVIDIA, offering hands-on courses, instructor-led workshops, and professional certifications in artificial intelligence, accelerated…
NVIDIA Dynamo is an open-source, low-latency distributed inference serving framework designed to deploy and scale generative AI and reasoning models across large GPU clusters.
NVIDIA Exemplar Cloud is a validation program run by NVIDIA that certifies cloud providers whose GPU clusters reproduce at least 95% of the training throughput NVIDIA measures on its own reference architecture…
The NVIDIA GB300 NVL72 is a liquid-cooled, rack-scale AI computing system that integrates 72 NVIDIA Blackwell Ultra (B300) GPUs and 36 Arm-based NVIDIA Grace CPUs into a single NVLink fabric, delivering 1.1…
The NVIDIA GH200 Grace Hopper Superchip is a single-module processor from NVIDIA that combines a 72-core Grace Arm CPU with a Hopper-generation H100-class GPU on one package
NVIDIA H100 (also called the H100 Tensor Core GPU) is a data-center graphics processing unit built by NVIDIA on the Hopper microarchitecture, fabricated with over 80 billion transistors on a custom TSMC 4N (4…
NVIDIA HGX is a family of accelerated-server platform designs built around tightly connected data-center GPUs. It is not one immutable board specification.
NVIDIA Holoscan is a domain-agnostic, multimodal AI sensor processing platform and software development kit (SDK) built by NVIDIA for real-time
NVIDIA MGX is a modular reference architecture that NVIDIA publishes so that server makers and contract manufacturers can build accelerated systems around NVIDIA GPUs, CPUs, DPUs and networking without…
NVIDIA NIM (NVIDIA Inference Microservices) is a set of containerized, prebuilt-and-optimized model-serving microservices from NVIDIA that package an AI model, an optimized inference engine, and an…
NVIDIA NeMo is an open, end-to-end framework from NVIDIA for building, customizing, and deploying generative AI models, described by NVIDIA as "a scalable generative AI framework built for researchers and…
NVIDIA NeMo Switchyard is an open-source proxy and Rust library for routing requests among configured large language models.
NVIDIA OSMO is an open-source workflow orchestration platform developed by Nvidia for physical AI and robotics development.
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
NVIDIA Quantum-X Photonics is a family of co-packaged-optics network switches that NVIDIA announced at its GTC conference on March 18, 2025.
NVIDIA Spectrum-6 is an Ethernet switch ASIC and system architecture for large AI infrastructure.