CUDA
CUDA (Compute Unified Device Architecture) is NVIDIA's platform and programming model for general-purpose computation on its graphics processing units.
Explore AI Infrastructure through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Infrastructure.
Showing 1-35 of 35 articles
CUDA (Compute Unified Device Architecture) is NVIDIA's platform and programming model for general-purpose computation on its graphics processing units.
CUTLASS is an open-source library of reusable building blocks for writing high-performance matrix kernels on NVIDIA GPUs.
KAI Scheduler is an open-source Kubernetes scheduler that optimizes the allocation of GPU resources for artificial intelligence and machine learning workloads.
Lepton AI was an American AI cloud company, founded in 2023, that built a cloud-native inference platform for serving large language models, generative image models, and other AI workloads on NVIDIA GPUs.
NCCL, the NVIDIA Collective Communications Library, is a library of topology-aware communication primitives for NVIDIA GPU systems.
NVHBM is an announced custom high-bandwidth memory architecture from NVIDIA for custom AI accelerators that participate in the company's NVLink Fusion platform. NVIDIA introduced it on August 26, 2026.
NVIDIA AI Enterprise is an end-to-end, cloud-native software suite sold by Nvidia as a paid subscription for developing and deploying production artificial intelligence and data analytics.
The NVIDIA AI compute infrastructure financing platforms are a set of proposed, independently run financing vehicles that NVIDIA announced on August 10, 2026, together with Apollo, BlackRock, Blackstone…
The NVIDIA B200 is a data center GPU based on the NVIDIA Blackwell microarchitecture, announced by Jensen Huang at GTC 2024 on March 18, 2024.
NVIDIA ConnectX is a family of high-speed network adapters and SmartNICs (smart network interface cards) that connect a server to the data center fabric and accelerate networking in hardware.
The NVIDIA DGX B300 is an 8-GPU AI supercomputer node built around NVIDIA's Blackwell Ultra architecture.
NVIDIA DGX Cloud is a managed AI-supercomputing-as-a-service offering from Nvidia that rents enterprises access to multi-node clusters of NVIDIA DGX infrastructure plus the NVIDIA AI software stack over the…
NVIDIA DGX Station is a line of deskside artificial intelligence workstations from NVIDIA, each marketed as a "personal AI supercomputer" that puts data-center-class compute next to a developer's desk rather…
The NVIDIA DGX SuperPOD is a reference-architecture artificial intelligence supercomputer designed and sold by Nvidia.
NVIDIA Deep Learning Institute (DLI) is the training and education arm of NVIDIA, offering hands-on courses, instructor-led workshops, and professional certifications in artificial intelligence, accelerated…
NVIDIA Dynamo is an open-source, low-latency distributed inference serving framework designed to deploy and scale generative AI and reasoning models across large GPU clusters.
NVIDIA Exemplar Cloud is a validation program run by NVIDIA that certifies cloud providers whose GPU clusters reproduce at least 95% of the training throughput NVIDIA measures on its own reference architecture…
The NVIDIA GB300 NVL72 is a liquid-cooled, rack-scale AI computing system that integrates 72 NVIDIA Blackwell Ultra (B300) GPUs and 36 Arm-based NVIDIA Grace CPUs into a single NVLink fabric, delivering 1.1…
The NVIDIA GH200 Grace Hopper Superchip is a single-module processor from NVIDIA that combines a 72-core Grace Arm CPU with a Hopper-generation H100-class GPU on one package
NVIDIA H100 (also called the H100 Tensor Core GPU) is a data-center graphics processing unit built by NVIDIA on the Hopper microarchitecture, fabricated with over 80 billion transistors on a custom TSMC 4N (4…
NVIDIA HGX is a family of accelerated-server platform designs built around tightly connected data-center GPUs. It is not one immutable board specification.
NVIDIA Holoscan is a domain-agnostic, multimodal AI sensor processing platform and software development kit (SDK) built by NVIDIA for real-time
NVIDIA MGX is a modular reference architecture that NVIDIA publishes so that server makers and contract manufacturers can build accelerated systems around NVIDIA GPUs, CPUs, DPUs and networking without…
NVIDIA NIM (NVIDIA Inference Microservices) is a set of containerized, prebuilt-and-optimized model-serving microservices from NVIDIA that package an AI model, an optimized inference engine, and an…
NVIDIA NeMo is an open, end-to-end framework from NVIDIA for building, customizing, and deploying generative AI models, described by NVIDIA as "a scalable generative AI framework built for researchers and…
NVIDIA NeMo Switchyard is an open-source proxy and Rust library for routing requests among configured large language models.
NVIDIA OSMO is an open-source workflow orchestration platform developed by Nvidia for physical AI and robotics development.
NVIDIA Picasso is a cloud-based generative AI foundry from NVIDIA for building, training, and deploying visual generative models that produce images, video, and 3D content from text prompts.
NVIDIA Quantum-X Photonics is a family of co-packaged-optics network switches that NVIDIA announced at its GTC conference on March 18, 2025.
NVIDIA Spectrum-6 is an Ethernet switch ASIC and system architecture for large AI infrastructure.
NVIDIA Spectrum-X Photonics is a co-packaged optics (CPO) Ethernet switch platform from Nvidia, announced at the company's GPU Technology Conference (GTC) on March 18, 2025.
NVLink Fusion is a program and silicon technology from NVIDIA that opens its NVLink high speed interconnect to third party chips.
SpaceX Starmind is SpaceX's planned constellation of solar-powered artificial intelligence compute satellites, intended to function as orbital data centers that run AI workloads in space and beam results back…
A Tensor Core is a specialized execution unit inside NVIDIA GPUs that computes a small matrix multiplication and accumulation, D = A x B + C
Raptor is the second-generation AI inference accelerator from d-Matrix, a Santa Clara semiconductor startup, and the first commercial chip built on the company's 3D stacked digital in-memory compute technology