NVIDIA Spectrum-X Photonics
NVIDIA Spectrum-X Photonics is a co-packaged optics (CPO) Ethernet switch platform from Nvidia, announced at the company's GPU Technology Conference (GTC) on March 18, 2025.
Explore AI Infrastructure through related topics and the articles other pages reference most.
Articles that also belong to these categories. Counts cover all of AI Infrastructure.
Showing 181-240 of 281 articles
NVIDIA Spectrum-X Photonics is a co-packaged optics (CPO) Ethernet switch platform from Nvidia, announced at the company's GPU Technology Conference (GTC) on March 18, 2025.
NVIDIA Spectrum-XGS, marketed in full as NVIDIA Spectrum-XGS Ethernet, is a networking technology from NVIDIA designed to link several geographically separated data centers into a single
NVLink Fusion is a program and silicon technology from NVIDIA that opens its NVLink high speed interconnect to third party chips.
Nebius is an AI cloud company incorporated in the Netherlands and listed on the Nasdaq stock exchange under the ticker symbol NBIS
A neocloud is a cloud computing company whose business is renting out GPU capacity for artificial intelligence workloads, rather than selling the broad catalogue of storage, database, networking and…
Neuromorphic computing is brain-inspired computer hardware that processes information with spiking neural networks (SNNs) and event-driven, in-memory computation, co-locating memory and processing the way…
Nscale is a British AI infrastructure company, usually described as a neocloud, that builds GPU-dense data center capacity and rents large-scale compute to hyperscalers, AI labs, and governments.
The Nvidia-OpenAI partnership is a strategic agreement announced on September 22, 2025, under which NVIDIA said it intends to invest up to $100 billion in OpenAI as OpenAI builds out a large amount of new…
OctoAI (originally OctoML) was an American artificial intelligence infrastructure company that operated a generative-AI inference platform and, before its pivot
Oklo Inc. (NYSE: OKLO) is an American advanced-fission company developing the Aurora "powerhouse," a small sodium-cooled fast reactor that Oklo plans to build, own, and operate to sell electricity directly to…
OpenAI AgentKit is a suite of agent-building tools that OpenAI introduced at OpenAI DevDay on 6 October 2025 to take AI agents from prototype to production on OpenAI's hosted models.
The OpenAI-Broadcom collaboration is a strategic partnership, announced on October 13, 2025, under which OpenAI designs its own custom AI inference accelerators and Broadcom co-develops, manufactures, and…
OpenRoboto is a robotics model competition that runs as subnet 80 on the Bittensor network.
Oracle Corporation is an American multinational technology company, founded in 1977 and headquartered in Austin, Texas
PCI Express (PCIe) is a high-speed serial interconnect standard used to attach processors to graphics cards, network adapters, storage devices, and accelerators inside nearly every modern computer and server.
The PORTS Technology Campus, also called the PORTS-Pike Technology Campus, is a planned AI infrastructure and power-generation complex in Pike County, Ohio.
PagedAttention is a KV-cache memory management algorithm for serving large language models that applies the virtual-memory paging technique used by operating systems to the GPU, eliminating memory…
Pallas is an experimental extension to JAX that lets users write custom hardware kernels in Python and lower them to both Tensor Processing Units and NVIDIA GPUs from a single source.
Parallel Web Systems is an American artificial-intelligence infrastructure company founded in early 2024 by Parag Agrawal, the former chief executive of Twitter.
Pathways is the name Google has used for two related but distinct things in its artificial-intelligence work: a research vision for a next-generation AI architecture, first articulated by Jeff Dean in October…
Pinecone is a managed vector database for artificial intelligence applications that lets developers store, index, and query high-dimensional vector embeddings at scale without managing infrastructure.
Pipeline parallelism (often abbreviated PP) is a distributed training strategy that splits the layers of a deep neural network across multiple accelerator devices so that each device holds one contiguous block…
Prefix caching is an inference optimization for large language model serving that stores and reuses the key-value (KV) cache computed for a shared prompt prefix, so that the prefill for that prefix is computed…
Prime Intellect is a San Francisco based artificial intelligence company building infrastructure for decentralized AI development, with a focus on training large models across heterogeneous
Product quantization (PQ) is a vector-compression technique for approximate nearest-neighbor (ANN) search that splits each high-dimensional vector into M equal sub-vectors and quantizes each sub-vector with…
Project Rainier is a distributed artificial intelligence supercomputer built by Amazon Web Services around its in-house AWS Trainium accelerators, created principally to train and serve the Claude models of…
Prometheus is a roughly 1-gigawatt AI supercluster built by Meta at its data center campus in New Albany, Ohio.
Prompt lookup decoding (PLD), also called n-gram speculative decoding, is an inference acceleration method for large language models that speeds up text generation without changing the model or its outputs.
Qdrant (pronounced "quadrant") is an open-source vector database and similarity search engine written in Rust and designed for high-performance retrieval over high-dimensional data.
QuIP (Quantization with Incoherence Processing) is a family of weight-only post-training quantization methods for large language models developed in the RelaxML group at Cornell University
Quantization-aware training (QAT) is a model compression technique in which the effects of quantization are simulated during the training or fine-tuning of a neural network, so that the model learns parameter…
RISC-V (pronounced "risk-five") is an open standard instruction set architecture (ISA), the contract that defines which instructions a processor executes and which registers software can see.
ROCm is AMD's open software stack for GPU computing and the main alternative to NVIDIA's CUDA platform.
RadixAttention is a KV cache management technique introduced in SGLang that uses a radix tree data structure to automatically share and reuse cached key-value tensors across inference requests.
Ray is an open-source distributed computing framework, developed at the University of California, Berkeley's RISELab and commercialized by Anyscale, that lets developers scale Python and artificial…
Ray Serve is a scalable, framework-agnostic model serving library built on top of the Ray (framework) distributed computing system.
The Reliance Jamnagar AI data center is a planned gigawatt-scale artificial intelligence computing campus that Reliance Industries, India's largest company by market value
Replicate is a cloud platform for running, deploying, and sharing machine learning models via a simple API.
Replit is an online integrated development environment (IDE) and AI-powered coding platform that lets users write, run, and deploy software directly from a web browser, including by describing an application…
The Research SuperCluster (RSC) is an AI supercomputer built by Meta AI, the artificial intelligence research division of Meta Platforms (the company formerly known as Facebook).
A robot is a programmable machine that senses its environment, computes a decision, and acts on the physical world through a closed loop of perception, planning, and actuation.
Run:ai (legal name Runai Labs Ltd.) is an Israeli software company that developed a Kubernetes-based orchestration and scheduling platform for graphics processing unit (GPU) resources used in artificial…
RunPod is a globally distributed GPU cloud computing platform that provides on-demand and serverless GPU compute for artificial intelligence training, fine-tuning, and inference.
SRAM, or static random-access memory, is a semiconductor memory that stores each bit in a latch built from cross-coupled inverters.
The SambaNova SN40L is a reconfigurable dataflow AI accelerator designed by SambaNova Systems and unveiled on September 19, 2023.
The SambaNova SN50 is a reconfigurable dataflow accelerator that SambaNova Systems unveiled in late February 2026 as the successor to its SN40L chip.
Sanjay Ghemawat (born 1966) is an American computer scientist and software engineer, best known for co-creating the core distributed-systems infrastructure that powered Google's rise, including the Google File…
Self-speculative decoding is a family of speculative decoding methods that accelerate large language model (LLM) inference by using the target model itself, run in a cheaper reduced-depth mode
SemiAnalysis is an independent research and analysis firm covering the semiconductor and artificial intelligence industries.
Sequence parallelism (SP) is a family of distributed training techniques for transformer-based neural networks that partitions activations along the sequence (token) dimension across multiple accelerators…
Slurm is an open-source workload manager and job scheduler for Linux clusters, developed and maintained by SchedMD, which NVIDIA acquired in December 2025 .
Snowflake AI is the suite of artificial intelligence and machine learning capabilities built into the Snowflake AI Data Cloud, anchored by Cortex AI (managed generative AI services callable in SQL), the…
The Snowflake-AWS chip deal is a five-year, roughly $6 billion infrastructure commitment that the data-cloud company Snowflake made to Amazon Web Services (AWS), announced on May 27, 2026.
Space-based data centers are data centers launched into Earth orbit (or, in some proposals, placed on the Moon), where satellites carrying AI accelerators draw power from solar arrays and reject heat by…
SpaceX Starmind is SpaceX's planned constellation of solar-powered artificial intelligence compute satellites, intended to function as orbital data centers that run AI workloads in space and beam results back…
SpinQuant is a post-training quantization method for large language models that inserts learned rotation matrices into a transformer network to make its weights, activations, and KV cache easier to represent…
Starcloud, Inc. is an American space infrastructure startup that builds space-based data centers: satellites carrying data-center-class GPUs that are powered by solar arrays and cooled by radiating waste heat…
Stargate Argentina is an announced, early-stage plan to build a large artificial intelligence data center of up to 500 megawatts (MW) in Argentina
The Stargate Initiative (formally The Stargate Project, incorporated in Delaware as Stargate LLC) is an American joint venture announced on January 21, 2025 at the White House to build AI infrastructure for…
Stargate Michigan is an artificial intelligence data center campus under construction in Saline Township, Michigan, about 10 miles southwest of Ann Arbor in Washtenaw County.