Citation and evidence

Cerebras Systems

53 min full readUpdated 69 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI HardwareArtificial Intelligence

Cite this article

Cerebras Systems is an American artificial intelligence hardware company that builds wafer-scale processors for AI training and inference, and it operates the fastest commercial large language model inference service measured by tokens per second. Founded in March 2016 and headquartered in Sunnyvale, California, Cerebras is best known for the Wafer-Scale Engine (WSE), the largest computer chip ever made, which fabricates an entire silicon wafer into a single processor with up to 4 trillion transistors and 900,000 AI cores. By keeping a model's weights in 44 GB of on-chip SRAM with 21 petabytes per second of memory bandwidth, roughly 7,000 times the off-chip bandwidth of an NVIDIA H100, the WSE-3 sidesteps the memory bottleneck that throttles GPU inference, which is the technical basis for Cerebras's speed records and its positioning as the leading pure-play challenger to NVIDIA in the AI accelerator market [3]. Cerebras completed an initial public offering on the Nasdaq Global Select Market under the ticker symbol CBRS on May 14, 2026, pricing 30 million shares at $185 to raise $5.55 billion in what became the largest global IPO of the year and the largest U.S. tech IPO since Uber's 2019 debut [18][19]. When the offering closed on May 15, 2026, the underwriters had exercised their option in full, taking the total to 34.5 million shares and gross proceeds of about $6.38 billion [52]. In 2026 Cerebras also began selling its systems as the decode half of heterogeneous inference deployments, pairing them with AWS Trainium in March and with AMD Helios in July, and in August it introduced the CS-4, a three-wafer rack built on a higher-clocked WSE-3 Turbo processor [44][53][63].

History and Founding

Cerebras Systems was incorporated in March 2016 by Andrew Feldman, Gary Lauterbach, Michael James, Sean Lie, and Jean-Philippe Fricker. The five-person founding team had worked together at the energy-efficient microserver company SeaMicro, which Feldman and Lauterbach had previously co-founded. SeaMicro was acquired by AMD in 2012 for $334 million, after which the core engineering group remained inside AMD until breaking away to start Cerebras. The founders brought decades of combined experience in networking, server design, and microprocessor architecture to the new venture [1][2][20].

The core insight behind Cerebras was that AI workloads, particularly deep learning training, could benefit from a processor built at a radically larger scale than conventional chips. Rather than designing a small chip and connecting many of them together (the approach used by GPUs), Cerebras decided to build a single chip spanning an entire 300mm silicon wafer. This approach eliminates the inter-chip communication bottlenecks that slow down distributed AI systems and allows the entire model to reside on-chip during computation.

The company operated in stealth mode for its first three years before unveiling the original Wafer-Scale Engine (WSE) at the Hot Chips conference in August 2019. During the stealth period the team focused on solving five interlocking engineering problems that had blocked wafer-scale integration for decades, including yield, power delivery, thermal management, cross-reticle wiring, and packaging. By the time Cerebras revealed its first chip, the company had already shipped early CS-1 systems to research customers and was working with the U.S. Department of Energy on scientific computing workloads [11].

Who founded Cerebras and who leads it?

The founders of Cerebras have remained closely involved with the company since 2016, an unusual pattern for a hardware startup that scaled past ten years.

FounderRole at CerebrasBackground
Andrew FeldmanCo-founder, CEOBA Stanford 1991, MBA Stanford GSB 1997; previously CEO of SeaMicro, executive roles at Force10 Networks and RiverStone Networks
Gary LauterbachCo-founder, CTOFormer chief architect at Sun Microsystems and CTO of SeaMicro
Sean LieCo-founder, Chief Hardware ArchitectFormer lead architect at SeaMicro and AMD
Michael JamesCo-founder, Chief Software ArchitectVeteran compiler and runtime engineer with prior work at SeaMicro
Jean-Philippe FrickerCo-founder, Chief System ArchitectSystems and packaging engineer with prior work at SeaMicro

Expanded article table

Feldman has been a serial entrepreneur in Silicon Valley, having sold three companies and taken a fourth public before founding Cerebras. He grew up on the campus of Stanford University, where his father was a faculty member, and he completed both undergraduate and graduate degrees there. Following the company's May 2026 IPO, Feldman's stake was valued at roughly $3.2 billion and Sean Lie's at roughly $1.7 billion, making both billionaires on paper [21].

Wafer-Scale Engine Technology

The defining innovation of Cerebras is the Wafer-Scale Engine, a processor that occupies an entire silicon wafer rather than being cut into individual dies like a traditional chip. This approach required Cerebras to solve several engineering challenges that had prevented wafer-scale integration for decades, including managing defects, distributing power evenly, and handling thermal dissipation across such a large area.

Engineering challenges and solutions

Building a chip the size of an entire wafer presented five major technical challenges that Cerebras had to overcome:

Defect tolerance. In traditional chip manufacturing, defective dies are simply discarded. On a wafer-scale chip, defects cannot be avoided but must be managed. Cerebras designed the WSE with approximately 100x the defect tolerance of a conventional GPU. Each AI core on the WSE-3 occupies roughly 0.05 mm2, about 1% the size of an NVIDIA H100 streaming multiprocessor (SM) core. When a defect hits a WSE core, it disables only 0.05 mm2 of silicon, whereas the same defect on an H100 would disable approximately 6 mm2. The WSE-3 contains 970,000 physical cores, with 900,000 active in the shipping product, meaning Cerebras achieves 93% silicon utilization, which is higher than leading GPUs despite building the world's largest chip [10].

Cerebras also developed a sophisticated routing architecture that allows dynamic reconfiguration of connections between cores. When a defect is detected during manufacturing test, the system automatically routes around the disabled core using redundant communication pathways, maintaining the fabric's full bandwidth and connectivity.

Power delivery. Delivering tens of kilowatts of power uniformly across a wafer-sized chip cannot be accomplished with traditional edge-of-die power connections. Cerebras designed a custom engine block that delivers power directly into the face of the wafer, achieving the power density required for hundreds of thousands of active cores. The custom connector and PCB design ensures uniform voltage across the entire wafer surface [11].

Thermal management. With overall power delivery in the mid-teen kilowatt range, the WSE generates substantial heat that must be removed uniformly. Traditional heat sink attachment techniques could not be used because of the thermal expansion mismatch between silicon and copper. Cerebras invented a new material and connector design that allows the wafer to expand and contract while remaining in thermal contact with a copper heat exchanger. Water flows through micro-fins on the backside of the heat exchanger, and the wafer slides against the polished front surface, maintaining thermal coupling despite differing coefficients of thermal expansion [11].

Cross-reticle connectivity. A standard photolithographic reticle can expose only a portion of a wafer at a time. To create a chip that spans the entire wafer, Cerebras had to connect circuits across reticle boundaries with high bandwidth and low latency, a problem unique to wafer-scale integration.

Die-to-die communication. The fabric interconnect that links all 900,000 cores must provide enormous aggregate bandwidth while consuming minimal power and area. The WSE-3's on-chip fabric provides 214 Pbit/s of bandwidth, enabling data to flow between any two cores on the wafer with predictable, low latency.

WSE-1 (2019)

The first-generation Wafer-Scale Engine was announced in August 2019. It featured 400,000 AI-optimized processing cores, 1.2 trillion transistors, and 18 gigabytes of on-chip SRAM. The CS-1 system, which housed the WSE-1, included twelve 100 Gigabit Ethernet connections for data transfer. At the time, it was by far the largest chip ever fabricated, with a die area of 46,225 square millimeters, roughly 56 times larger than the biggest GPU available. Argonne National Laboratory deployed the first CS-1 in 2019, becoming the inaugural national-lab customer for the technology [22].

WSE-2 (2021)

In April 2021, Cerebras announced the second-generation Wafer-Scale Engine (WSE-2), manufactured using TSMC's 7nm process. The WSE-2 represented a major leap in specifications:

SpecificationWSE-1WSE-2
Transistors1.2 trillion2.6 trillion
AI cores400,000850,000
On-chip SRAM18 GB40 GB
Memory bandwidth9.6 PB/s20 PB/s
Fabric bandwidth100 Pb/s220 Pb/s
Process node16nm7nm

Expanded article table

The CS-2 system built around the WSE-2 became the primary commercial product for Cerebras, used by research institutions and enterprises for AI training. Notably, the 40 GB of on-chip SRAM eliminated the need for external high-bandwidth memory (HBM), allowing the entire working set of many AI models to reside directly on the processor.

WSE-3 (2024)

The third-generation Wafer-Scale Engine (WSE-3) was announced in March 2024 at a dedicated launch event and later presented in detail at the Hot Chips 2024 conference. Manufactured on TSMC's 5nm process, the WSE-3 is the most powerful AI chip ever built, containing approximately 4 trillion transistors across its 46,255 mm2 die area. Detailed coverage of the chip lives in the sibling article on the Cerebras WSE-3 [12][13].

SpecificationWSE-2WSE-3
Transistors2.6 trillion4 trillion
AI cores (active)850,000900,000
Physical cores~900,000970,000
On-chip SRAM40 GB44 GB
Memory bandwidth20 PB/s21 PB/s
Fabric bandwidth220 Pb/s214 Pb/s
Peak AI performanceN/A125 PFLOPS
Process node7nm5nm
ManufacturerTSMCTSMC

Expanded article table

The WSE-3 delivers 125 petaflops of peak AI performance, which Cerebras claims is roughly double the performance of the WSE-2 at the same power consumption and cost. The memory bandwidth of 21 PB/s is approximately 7,000 times greater than the NVIDIA H100's off-chip HBM bandwidth, which is the fundamental advantage that enables Cerebras's inference speed records. Because the WSE stores model weights in on-chip SRAM distributed alongside the compute cores, data travels only fractions of a millimeter from memory to compute, rather than crossing package boundaries as it does in GPU architectures [3].

WSE-3 Turbo (2026)

On August 18, 2026, Cerebras announced the Wafer Scale Engine 3 Turbo (WSE-3T), a higher-performance version of the WSE-3 rather than a new generation. According to the company's release, the WSE-3T keeps the WSE-3's physical design, with four trillion transistors, 900,000 AI-optimized cores, 46,225 mm2 of silicon, and 44 GB of on-wafer SRAM, while doubling AI compute to 250 PFLOPS per wafer and memory bandwidth to 43.2 PB/s; on-chip fabric bandwidth and off-chip I/O also double, to 53.5 PB/s and 2.4 Tb/s, and I/O latency falls from five microseconds to as low as two [63]. Cerebras attributes the higher operating frequency to a redesigned power path that moves conversion from roughly 50 mm to about 0.5 mm from the processor, delivering twice as much power to the wafer [63][64]. These are company-published figures. The fourth-generation wafer that trade commentary had expected around the IPO remains unannounced as of September 2026; see Cerebras WSE-4.

SpecificationWSE-3WSE-3 Turbo
Transistors4 trillion4 trillion
AI cores (active)900,000900,000
Die area46,225 mm246,225 mm2
On-chip SRAM44 GB44 GB
Peak AI compute125 PFLOPS250 PFLOPS
Memory bandwidth21 PB/s43.2 PB/s
Fabric bandwidth214 Pb/s53.5 PB/s (stated as double the WSE-3)
Off-chip I/On/a2.4 Tb/s (stated as double the WSE-3)
I/O latency5 microseconds (per the WSE-3T release)As low as 2 microseconds
AnnouncedMarch 2024August 18, 2026

Expanded article table

CS-2 and CS-3 Systems

The Cerebras CS-2 and CS-3 are complete computing systems designed to house the WSE-2 and WSE-3 chips, respectively. Each system integrates the wafer-scale processor with all necessary power delivery, cooling, and I/O connectivity in a single rack-mountable unit. The systems are designed to be straightforward to deploy, requiring only standard datacenter power and cooling.

CS-3 system specifications

The CS-3, powered by the WSE-3, delivers 125 petaflops of AI compute while consuming 23 kW of power per system. Key system-level specifications include:

SpecificationCS-3
ProcessorWSE-3 (4T transistors, 900K cores)
Peak AI compute125 PFLOPS
On-chip memory44 GB SRAM
System memory (with MemoryX)Up to 1.2 PB
Power consumption23 kW
CoolingDirect liquid cooling
Max systems in cluster2,048 (via SwarmX)
Max cluster compute~0.25 zettaFLOPS

Expanded article table

One of the key advantages of the CS systems is that they eliminate the need for complex multi-node distributed computing setups. A single CS-3, for instance, can train models that would otherwise require clusters of hundreds of GPUs, dramatically simplifying the software stack and reducing the engineering effort required to scale AI training.

MemoryX and SwarmX

Cerebras developed two complementary technologies to extend the CS systems beyond single-chip limitations:

MemoryX is an external memory system that extends the available memory for the WSE beyond the on-chip SRAM. With up to 1.2 petabytes of capacity, MemoryX enables the CS-3 to handle models with trillions of parameters by streaming weight data to the processor at high bandwidth. This is large enough to store models with 24 trillion parameters in a single logical memory space without partitioning or refactoring, enabling training of next-generation frontier models 10x larger than GPT-4 and Gemini [3].

SwarmX is a high-bandwidth fabric that connects multiple CS systems together. Up to 2,048 CS-3 systems can be linked via SwarmX to build hyperscale AI supercomputers delivering up to a quarter of a zettaFLOP. SwarmX maintains near-linear scaling efficiency, meaning that doubling the number of CS-3 systems nearly doubles aggregate performance.

CS-4 (announced August 2026)

The CS-4, introduced on August 18, 2026, is the first Cerebras system to put more than one wafer in a rack. Cerebras describes it as a rack-scale system built from three WSE-3T processors and the first product on its Nexus platform architecture, a modular design that separates compute, power, and I/O into self-contained assemblies [63][64]. The compute sits in a rear-mounted "Wafer-Scale Backpack" that packages power conversion, direct liquid cooling, high-speed I/O, and control electronics around each wafer; the company says the backpack has 50 percent fewer components than the prior-generation system, uses 60 percent more automated manufacturing, and cuts deployment time from days to hours [63][64]. A new programmable I/O subsystem supports standards-based RoCE v2 RDMA over Ethernet for connecting to other vendors' hardware and switch-free "Direct Wafer Links" between Cerebras systems [64].

SpecificationCS-3CS-4
Processors1 x WSE-33 x WSE-3 Turbo
Peak AI compute125 PFLOPS750 PFLOPS
On-wafer SRAM (total)44 GB132 GB (three wafers of 44 GB, per The Register's tally [65])
Memory bandwidth21 PB/s129.6 PB/s
Fabric bandwidth214 Pb/s160.5 PB/s
System I/On/a7.2 Tb/s
Wafer-to-wafer latencyn/aAs low as 2 microseconds
Largest supported modelsUp to 24 trillion parameters (with MemoryX)Over 50 trillion parameters (company claim)
First shipments2024Third quarter of 2026 (company plan)

Expanded article table

Cerebras's headline claims are that the CS-4 is up to twice as fast as the CS-3, which it says extends its per-user tokens-per-second advantage over GPU systems to "up to 30x", and that it delivers up to 10 times more throughput per watt than the CS-3; the company's blog post sources the 30x figure to "Artificial analysis and internal benchmarking (August 2026)" and the throughput figure to "internal benchmarking and projections" [63][64]. It also says the CS-4 can deliver more than 1,000 tokens per second on models exceeding 10 trillion parameters, a figure it labels an extrapolation from internal benchmarking [64]. SemiAnalysis founder Dylan Patel is quoted in the release praising the system's deployability, reliability, and networking [63]. The Register's Tobias Mann argued after Hot Chips 2026 that the CS-4 and NVIDIA's Groq 3 LPX figures both appear to be single-request measurements that customers are unlikely to see in production, estimating a maximum batch of about 12 requests at a 100,000-token input length for Gemma 4 31B given the CS-4's 132 GB of SRAM, and that both designs make more sense as decode accelerators beside GPUs [65]. Cerebras said first CS-4 shipments would begin in the third quarter of 2026, and full specifications are in its CS-4 datasheet [63]. The announcement also states that the system has "native support for disaggregated inference" and names AMD Helios and AWS Trainium as prefill platforms it is designed to pair with, discussed under cloud and hyperscaler partnerships below [64].

Software Platform

Cerebras's software stack is exposed to customers through two distinct surfaces. The first is the Cerebras Software Platform, branded CSoft, which provides an ML-framework-compatible path so that models written in PyTorch (and historically TensorFlow) can be compiled to run on the WSE with minimal code modification. The PyTorch integration uses XLA's lazy tensor backend, which lets the Cerebras Graph Compiler (CGC) lower a model's computational graph to wafer-scale executable code, handling core allocation across the WSE's hundreds of thousands of processing elements and minimizing communication latency in the fabric [41]. CSoft supports both single-CS-3 training and distributed training across MemoryX/SwarmX clusters under the same programming model, which is the central selling point versus GPU-based tensor- and pipeline-parallelism pipelines [41][42].

The second surface is the Cerebras SDK and Cerebras Software Language (CSL), a domain-specific C-like language for writing custom kernels directly against the WSE microarchitecture. In CSL, each on-wafer core is exposed as a Processing Element (PE), and developers can place computation and data on individual PEs to maximize fabric throughput. The SDK ships a cycle-accurate simulator so kernels can be developed and debugged without on-prem CS-3 hardware, and the SDK is the path most often used by HPC customers at national labs whose workloads do not map cleanly onto standard deep-learning frameworks [42]. As of the 2.10 SDK release, CSL covers full WSE-3 functionality including the cross-reticle fabric and on-chip SRAM hierarchy [42]. Cerebras also offers AI Model Studio, a managed-cloud experience that exposes CS-3 clusters as a hosted training service for large language model work without requiring customers to operate hardware themselves [41].

Cerebras Inference

Cerebras Inference launched in August 2024 as a cloud-based inference service that leverages the unique architecture of the WSE to deliver what the company describes as the fastest LLM inference available. Because the WSE keeps the entire model on-chip (or streams it efficiently via MemoryX), it avoids the memory bandwidth bottleneck that limits GPU-based inference. At launch Cerebras claimed inference 10x to 20x faster than systems built using NVIDIA H100 Hopper GPUs, and the service has set successive speed records on each major open-source model since [23]. Cerebras frames raw speed as an existential requirement rather than a luxury. "For hard problems, there is no upper bound to how much faster you want to be," Feldman has said of the inference market. "How big is the market for slow search? Zero. How big is the market for dial-up internet? Zero" [43].

Inference performance benchmarks

The CS-3 has achieved remarkable inference speeds across a range of models, consistently setting records for tokens-per-second output:

ModelTokens/secondNotes
Llama 3.1 8B1,800Single-user latency
Llama 3.1 70B2,100Per-user, roughly 8x faster than H200
Llama 3.1 405B969Largest open-source model at launch
Llama 4 Scout2,600+19x faster than fastest GPU at the time
Llama 4 Maverick2,500+Beat NVIDIA Blackwell on independent benchmarks
gpt-oss-120B2,700+Via Core42 partnership
K2 Think2,000Reasoning-optimized model
Kimi K2.6Approaching 1,000First trillion-parameter model served on Cerebras; figure attributed by Cerebras to Artificial Analysis measurements (June 2026 results release) [45]
GPT-5.6 Sol (Ultrafast preview)Up to 750OpenAI's advertised maximum for the Cerebras-powered tier, August 2026 [46][47]

Expanded article table

In head-to-head comparisons, Cerebras claims the CS-3 achieves over 21x faster inference than NVIDIA's flagship Blackwell B200 GPU on the Llama 3 70B model in a reasoning scenario with 1,024 input tokens and 4,096 output tokens. On the gpt-oss-120B model, the CS-3 delivered 2,700+ tokens/second compared to approximately 900 tokens/second on Blackwell B200 [4]. Independent benchmarks from third-party site Artificial Analysis have repeatedly placed Cerebras at the top of token-per-second rankings, with inference speeds up to 75 times faster than the same models served by hyperscalers such as Amazon, Microsoft, and Google [24].

The service supports popular open-source models including Llama 3.1 (8B, 70B, and 405B variants), Llama 4 Scout, Llama 4 Maverick, Mistral models, DeepSeek variants, and the gpt-oss family. In a flagship partnership with Mistral, Cerebras powered the Le Chat AI assistant at over 1,100 tokens per second, roughly ten times faster than ChatGPT at the time [7]. Cerebras and Core42 (a subsidiary of G42) also launched global access to OpenAI's gpt-oss-120B model, serving it at approximately 3,000 tokens per second [9].

The speed advantage of Cerebras Inference has made it particularly attractive for real-time applications, agentic AI workflows, and scenarios where low latency directly impacts user experience. Cerebras planned to increase its inference capacity from 2 million to over 40 million tokens per second by Q4 2025, distributed across eight data centers spanning North America and Europe.

Meta Llama API partnership

In April 2025, Meta announced that Cerebras would power its new Llama API, the first cloud inference service offered directly by Meta for the Llama family of models. Developers selecting the Cerebras backend in the Llama API see generation speeds up to 18 times faster than traditional GPU-based offerings, with Llama 4 Scout running at 2,600 tokens per second compared to roughly 130 tokens per second for ChatGPT and 25 tokens per second for DeepSeek on equivalent endpoints [25]. Announcing the deal, Feldman said: "Developers building agentic and real-time apps need speed. With Cerebras on Llama API, they can build AI systems that are fundamentally out of reach for leading GPU-based inference clouds" [25]. The partnership represented a meaningful endorsement from Meta, the originator of the Llama model family, and channeled significant developer traffic to Cerebras hardware through Meta's own API surface.

Enterprise and developer customers

Throughout 2025 and 2026 Cerebras expanded its roster of paying enterprise customers across multiple industries:

CustomerUse caseNotes
MetaLlama API backendPowers Meta's official Llama inference cloud
MistralLe Chat assistant1,100+ tokens/s, claimed consumer chat speed record
NotionEnterprise searchSub-300 ms results across 100M+ users
AlphaSenseMarket intelligence10x faster insights via Cerebras Inference
CognitionAgentic codingPowers Devin-style autonomous coding workflows
IBMEnterprise AIIntegrates Cerebras into broader AI offerings
Mayo ClinicGenomic foundation modelRheumatoid arthritis treatment prediction
GlaxoSmithKlineDrug discoveryProtein and epigenomic transformer models
AstraZenecaDrug discoveryLarge biomedical training jobs
U.S. Department of DefenseClassified workloadsDisclosed in S-1 customer list

Expanded article table

This breadth of customers, from frontier labs to regulated pharmaceutical firms to U.S. government agencies, became a central pillar of Cerebras's IPO marketing in 2026 [17][26]. The company's second-quarter 2026 results release added to the list: it reported new cloud capacity agreements with the AI coding companies Cognition and Lovable, named Block, Figma, AlphaSense, and GSK as customers running agentic workflows on Cerebras, and described an inline-security deployment with CrowdStrike that runs LLMs on a large share of enterprise traffic [56].

Open-Source Models and Research

Cerebras has contributed several open-source models and research efforts to the AI community:

  • Cerebras-GPT: A family of GPT-3-like models ranging from 111 million to 13 billion parameters, trained on the open-source Pile dataset following DeepMind's Chinchilla scaling rules. These were among the first compute-optimal open-source language models.
  • BTLM-3B-8K: The Bittensor Language Model, a 3 billion parameter model developed in collaboration with Opentensor that achieved performance competitive with 7 billion parameter models while being significantly smaller and faster [8].
  • CrystalCoder-7B: Developed in partnership with Petuum and MBZUAI as part of the LLM360 initiative, which aimed to create fully transparent and reproducible large language models. CrystalCoder was trained on the Condor Galaxy 1 supercomputer.
  • Jais-30B and Med42: Arabic-language foundation models and a medical reasoning model jointly trained with G42 and MBZUAI on Condor Galaxy infrastructure.

G42 Partnership and Condor Galaxy

One of Cerebras's most significant strategic partnerships has been with G42, the Abu Dhabi-based AI technology holding company. The two companies announced a multi-billion-dollar deal in July 2023 to co-build a constellation of AI supercomputers and share infrastructure for training and inference. Together, they have built the Condor Galaxy series of AI supercomputers:

  • Condor Galaxy 1 (CG-1): Located in Santa Clara, California, CG-1 delivers 4 exaFLOPS of AI compute with 54 million AI-optimized compute cores. It was used to train several of Cerebras's open-source models and became available for external workloads.
  • Condor Galaxy 2 (CG-2): Also delivering 4 exaFLOPS with 54 million cores, CG-2 expanded the constellation's total capacity to 8 exaFLOPS and 108 million cores upon completion.
  • Condor Galaxy 3 (CG-3): Located in Dallas, Texas, CG-3 features 64 CS-3 systems powered by the WSE-3. It delivers 8 exaFLOPS of AI compute with 58 million AI-optimized cores, bringing the constellation total to 16 exaFLOPS.
  • Condor Galaxy India: In May 2026, G42 and the Government of India formalized a commercial framework to deploy an 8-exaFLOPS Condor Galaxy supercomputer in India built on Cerebras CS-3 systems, extending the constellation beyond North America for the first time [27].
SystemLocationComputeAI CoresWSE Generation
CG-1Santa Clara, CA4 exaFLOPS54 millionWSE-2
CG-2Undisclosed (US)4 exaFLOPS54 millionWSE-2
CG-3Dallas, TX8 exaFLOPS58 millionWSE-3
CG IndiaIndia8 exaFLOPSTBDWSE-3
Full constellation (planned)Global36 exaFLOPS-Mixed

Expanded article table

The full Condor Galaxy constellation of nine interconnected AI supercomputers is designed to deliver 36 exaFLOPS of AI compute by the end of 2026, making it one of the largest collections of interconnected AI supercomputers in the world.

The G42 partnership provided Cerebras with both a major customer and a development platform for demonstrating the capabilities of its wafer-scale technology at datacenter scale. In 2024, G42 accounted for roughly 85% of Cerebras's $290 million in revenue, a concentration that drew scrutiny from U.S. regulators during the IPO process. By the May 2026 amended S-1 filing, G42's direct share had fallen to 24% of 2025 revenue, with the Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), a related party of G42, accounting for an additional 62%, leaving total UAE-linked revenue at roughly 86% of 2025 sales [17][26].

CFIUS review and restructuring

The relationship with G42 created significant complications for Cerebras's IPO plans because of U.S. regulatory scrutiny of G42's prior business ties to China, including investments related to Huawei. After Cerebras filed its first S-1 in September 2024, the Committee on Foreign Investment in the United States (CFIUS) opened a review of G42's minority stake in Cerebras, focused on the risk that advanced AI chip technology could be transferred to China through UAE intermediaries. Cerebras paused its IPO process until the review concluded.

G42 divested its Chinese investments and signed a parallel partnership with Microsoft during 2024, both of which were designed to address national security concerns. CFIUS cleared the restructured arrangement on March 31, 2025, after G42 agreed to convert its Cerebras stake into non-voting shares, removing the company from Cerebras's voting investor base. This regulatory clearance unlocked Cerebras's ability to refile for an IPO and continue selling chips to G42 and its subsidiaries [16][28].

Cloud and hyperscaler partnerships

AWS partnership

In March 2026, Amazon Web Services signed a multiyear deal with Cerebras to make the WSE-3 wafer-scale chip available to cloud customers through Amazon Bedrock. AWS became the first major cloud provider to offer Cerebras's disaggregated inference solution, which splits a request between AWS Trainium servers handling prompt processing and CS-3 systems handling token generation, connected over Amazon's Elastic Fabric Adapter networking, with plans to launch in the coming months and add open-source LLM and Amazon Nova support later in 2026. "Partnering with AWS to build a disaggregated inference solution will bring the fastest inference to a global customer base," said Andrew Feldman, Founder and CEO of Cerebras Systems [44]. This partnership represented a significant validation of Cerebras's technology by one of the world's largest cloud providers and is described in the S-1 as a binding term sheet for ongoing capacity commitments [5][17].

Cerebras's own announcement post, written by James Wang, added detail the joint release did not. AWS would support both aggregated and disaggregated configurations; in disaggregated mode Trainium would compute the KV cache and send it to the wafer over EFA, with the WSE performing decode exclusively; and Cerebras projected that the arrangement would give AWS customers "5x more high-speed token capacity in the same hardware footprint" [55]. The company's June 2026 results release specified Trainium3 as the prefill chip [45]. The commercial terms followed the March announcement: according to Cerebras's quarterly report for the period ended June 30, 2026, the company entered into an "AWS Commercial Agreement" in June 2026 that includes an initial multi-year lease of Cerebras systems, options for future procurement, pricing commitments, exclusivity, and minimum manufacturing capacity guarantees in AWS's favor, and in the same month it issued Amazon's investment arm a warrant to acquire up to 2,696,678 shares of Class N common stock at $100 per share [57]. The timetable moved as well. The March release said the Bedrock service would launch "in the next couple of months" [44]; by its August 12, 2026 results release Cerebras said it expected to bring disaggregated inference and "the same 5x throughput benefits" to Amazon Bedrock in the first quarter of 2027 [56].

AMD Helios partnership

On July 23, 2026, at AMD's Advancing AI 2026 event in San Francisco, AMD and Cerebras announced a technical partnership to deliver a disaggregated inference system that combines AMD's Helios rack-scale platform, built on Instinct GPUs, with the Wafer-Scale Engine. The joint release describes the two machines operating "as a single disaggregated inference workflow": Helios serves as the high-throughput engine that processes prompts and large context windows, while the WSE handles "the memory-bandwidth-intensive token generation" [53]. The companies said the pairing is "expected to deliver up to 5x higher tokens per second per watt", a figure the release footnotes as based on July 2026 modelling by AMD Performance Labs and Cerebras of tokens per second per kilowatt at a comparable interactivity point on the "Kimi 2.6 1T" model (Kimi K2.6), comparing a Helios-plus-WSE configuration with a WSE-only configuration. It is therefore a modelled comparison against Cerebras's own standalone hardware, not a measured result against GPU systems [53]. "AI inference is becoming one of the largest infrastructure opportunities in AI, and its growing diversity requires a more flexible approach," AMD chair and CEO Lisa Su said in the release. "Together with Cerebras, we are extending that leadership into the most latency-sensitive applications and creating a powerful new platform for real-time agentic AI" [53]. Feldman said the partnership "gives us an incredible opportunity to bring that performance to even more customers" [53]. AMD's own Advancing AI release listed Cerebras alongside Anthropic, OpenAI, and Meta among the partners featured at the event, and AMD had participated in Cerebras's Series H round in February 2026 [34][69].

Cerebras plans to deploy Helios systems in its own data centers, with the joint solution expected to be available first through Cerebras Cloud in the second half of 2026 [53]. CNBC's Kif Leswing reported that Feldman said at the event that the Helios systems would be installed in Cerebras data centers "starting later this year", that server buyers would also be able to configure AMD systems with wafer-scale chips, and that Cerebras shares rose about 5 percent on the day [58]. Axios covered the announcement the same day [62]. In a theCUBE interview published by SiliconANGLE on July 29, Cerebras chief marketing officer Julie Choi called the arrangement "one plus one equals five X" and said joint go-to-market efforts were expected before the end of the year [60]. In its August 12 results release Cerebras said the AMD solution "will be in production in Q4 2026" [56].

Disaggregated inference strategy

Both partnerships rest on the same architectural argument, which Cerebras laid out in a March 26, 2026 blog post by Sarah Chieng titled "The GPU Is Being Split in Half". The post explains disaggregated serving as splitting the compute-bound prefill phase from the memory-bandwidth-bound decode phase onto different machines, argues that the largest gains come from "heterogeneous disaggregation" in which each phase runs on hardware matched to it, and states that Cerebras expects a disaggregated Trainium-Cerebras solution "will increase token throughput by 5x over an aggregated solution" [54]. The CS-4 announcement carries the framing into hardware: Cerebras says the system was designed "to work as part of a heterogeneous AI infrastructure solution", in which "a purpose-built prefill engine" prepares the model state and hands it to the Cerebras system for decode, and it names AMD Helios and AWS Trainium as the prefill platforms it is built to pair with [64]. The company's second-quarter results release went further and described the AMD and AWS deals as establishing Cerebras "as the leader in disaggregated inference", a company characterisation [56].

In both announced deals the split is between prefill and decode: the Trainium or Helios side builds the KV cache for the prompt, and the Cerebras system then runs the entire decode step, attention and feed-forward layers included, with the cache resident on the wafer [44][53][55]. That differs from attention-FFN disaggregation, the finer split NVIDIA describes for its Groq 3 LPX rack, in which GPUs keep the KV cache and run attention while an SRAM accelerator runs only the feed-forward or expert layers, exchanging activations on every layer of every token. Commentary has nonetheless grouped Cerebras with that approach. In a widely shared September 4, 2026 post, Luminal chief executive Joe Fioti wrote that attention-FFN disaggregation is "exactly the approach Nvidia is taking with their SRAM-only Groq LPX system, and Cerebras is targeting with their partnerships with AMD and AWS Trainium" [68]. The public record does not support the Cerebras half of that sentence: neither press release, Cerebras's own blog posts, nor its CS-4 documentation mentions attention-FFN disaggregation, and no Cerebras statement describing such a deployment had been located as of September 6, 2026. The Register's Tobias Mann made a narrower point in August, arguing that SRAM-heavy accelerators such as the WSE "work a lot better as decode accelerators" beside GPUs and noting that no benchmarks of the combined AMD, AWS, or NVIDIA pairings had yet been published [65].

OpenAI deal

In January 2026, Cerebras agreed to provide 750 megawatts of computing power to OpenAI through 2028 [29][30]. The transaction was structured as a Master Relationship Agreement (MRA) valued at more than $20 billion, roughly twice the $10 billion headline figure that initially circulated in press coverage [30][45]. The MRA includes explicit expansion provisions that allow OpenAI to scale the commitment up to 2 gigawatts of capacity by 2030, which would make it among the largest single-customer AI infrastructure deals ever signed [26][35]. As part of the arrangement, OpenAI advanced Cerebras a $1 billion loan at 6% annual interest to help fund the data centers and received warrants for roughly 33 million shares, making OpenAI simultaneously a customer, a lender, and a prospective shareholder [26][35]. The contract accounted for the bulk of the $24.6 billion order backlog Cerebras disclosed in its S-1 [26][35].

The deployment began rolling out in phases in 2026, with Cerebras providing dedicated low-latency inference as part of OpenAI's inference stack [29][46]. Cerebras CEO Andrew Feldman described the deal as a turning point that validated wafer-scale silicon for production-scale frontier AI deployment, and analysts treated it as the centerpiece of the 2026 IPO narrative [29][30].

On August 13, 2026, OpenAI opened a limited preview of Ultrafast, a Cerebras-powered service tier for GPT-5.6 Sol in the OpenAI API. Ultrafast changes how the existing model is processed; OpenAI did not announce a new model or checkpoint. Access was limited to a select group of customers, and OpenAI said it would expand access as capacity grew without giving a general-availability date [46][47][51].

OpenAI advertised up to 750 output tokens per second and up to 14 times the speed of Standard processing. Both figures are vendor maximums rather than guarantees across requests. Cerebras reported that its own GPT-5.6 Sol Ultrafast evaluation completed all 2,500 Humanity's Last Exam questions in 11 hours 11 minutes, compared with 78 hours 27 minutes for Claude Fable 5 at what Cerebras described as comparable accuracy. It also reported a 5.6-times end-to-end speedup over Standard on GDP-Val without quality degradation [47]. Cerebras disclosed the reasoning settings and test dates, but it ran both evaluations itself. In a separate same-prompt dashboard demonstration, Cerebras reported 1 minute 50 seconds for Ultrafast and 12 minutes 20 seconds for Standard, characterizing the result as nearly seven times faster [49]. These tests do not establish generalized throughput or quality parity. Artificial Analysis independently measured the regular GPT-5.6 Sol first-party API at 68.7 output tokens per second with high reasoning, but its page did not evaluate the selected-customer Ultrafast preview [50].

The preview materials did not publish Ultrafast pricing, an API selector or Ultrafast-specific model ID, rate limits, an uptime or latency service-level agreement, a feature and quality parity matrix, or a general-availability date [46][47]. OpenAI's public documentation described Fast mode separately, with up to 2.5 times Standard speed and the selectors service_tier: "fast" or service_tier: "priority". It did not document a corresponding selector for Ultrafast [48]. Standard GPT-5.6 Sol and Fast mode terms therefore cannot be assumed to apply to the Ultrafast preview.

Qualcomm joint inference platform

In March 2024, Cerebras and Qualcomm announced a joint training-and-inference platform that pairs the CS-3 for model training with the Qualcomm Cloud AI 100 Ultra inference accelerator for deployment. The two companies claimed combined price-performance gains of roughly 10x over GPU-only pipelines, driven by hardware-aware training techniques implemented on the CS-3 and consumed on AI 100 Ultra silicon at inference time [40]. Specifically, models trained with 85% unstructured sparsity ran 3-4x faster during training and delivered 2-3x higher inference throughput; speculative decoding paired a small draft model with a larger target model to double token throughput; and the AI 100 Ultra's MX6 micro-exponent format halved the memory footprint relative to FP16 while doubling throughput [40]. The deal was strategically important because it moved Cerebras beyond selling training systems and into a position to influence the inference deployment of models customers had trained on its hardware. G42 also separately committed to deploy Qualcomm's AI 100 inference cards as part of its Cerebras-aligned datacenter buildout in 2024, reinforcing the three-way relationship between training silicon, inference silicon, and Middle East infrastructure capital [40].

Department of Energy collaborations

Cerebras has had an active relationship with the U.S. Department of Energy (DOE) since 2019. Argonne National Laboratory deployed the first CS-1 system that year, and Lawrence Livermore National Laboratory followed shortly after. In 2022, a collaboration between Cerebras and Argonne National Laboratory used WSE-based models to study SARS-CoV-2 genomic dynamics, work that won a Gordon Bell Special Prize.

In December 2025, Cerebras signed a memorandum of understanding (MOU) with DOE to accelerate the White House Genesis Mission, a national AI initiative to apply AI to scientific discovery and national security. The Pittsburgh Supercomputing Center and the National Energy Technology Laboratory have also operated CS systems for HPC and AI workloads [31][32].

Data center footprint

Cerebras shifted from selling individual CS systems to operating its own inference clusters during 2024 and 2025, building out a North American and European footprint to support the Cerebras Inference cloud. Most facilities are operated jointly with regional data center partners.

LocationPartnerStatus (May 2026)
Santa Clara, CaliforniaCerebrasOperational (CG-1, inference)
Stockton, CaliforniaNautilus Data TechnologiesOperational (floating barge facility)
Dallas, TexasG42 / CrusoeOperational (CG-3)
Minneapolis, MinnesotaUndisclosedOperational
Montreal, CanadaBit DigitalOperational
Oklahoma City, OklahomaScale DatacentersOperational (10 MW)
Undisclosed Europe siteUndisclosedComing online 2026
Condor Galaxy IndiaG42Announced May 2026

Expanded article table

The Oklahoma City facility opened in late 2025 as a 10 MW site providing more than 44 exaFLOPS of AI compute. The Stockton barge facility, hosted at Nautilus's water-cooled platform on the San Joaquin River, became Cerebras's first west coast inference cluster outside Santa Clara. The expansion was specifically designed to push aggregate token-generation capacity past 40 million tokens per second by the end of 2025 [33].

Announcements after May 2026 shifted the footprint toward Europe and toward contracted capacity rather than individual sites. On July 9, 2026, at the RAISE Summit in Paris, Cerebras said it would bring its first European data center capacity online by the end of 2026, build out across France and the Nordics, and reach 200 MW of European capacity by the end of 2027, with a portion supporting OpenAI workloads; Feldman named Norway and Finland as slated locations [67]. On September 1, 2026, Cerebras and Compute Nordic Finland announced an AI data centre in Mikkeli, Finland, that is to scale in phases from an initial 50 MW, on which construction was already under way, to 80 MW and then 165 MW of contracted IT capacity, under service orders with seven-year terms; the site is designed for closed-loop cooling and waste-heat recovery [66]. In its second-quarter 2026 results Cerebras said data center capacity live and under contract for delivery by the end of 2027 exceeded 600 MW, that its pipeline of opportunities had reached gigawatts, and that new factory lines at contract manufacturers Flex, Sanmina, and Rocket EMS would raise manufacturing capacity more than tenfold in 2026 [56]. All of these figures are company-reported.

LocationPartnerAnnouncedStatus (September 2026)
France and the Nordics (multiple sites)UndisclosedJuly 9, 2026First capacity due online by end of 2026; 200 MW targeted by end of 2027 [67]
Mikkeli, FinlandCompute Nordic FinlandSeptember 1, 2026Initial 50 MW phase under construction; 165 MW contracted in phases [66]

Expanded article table

Performance Comparison with GPU Clusters

The architectural differences between Cerebras's wafer-scale approach and traditional GPU clusters lead to fundamentally different performance profiles:

MetricCerebras CS-3NVIDIA DGX B200 (8x B200)Advantage
Inference latency (Llama 70B)~2,100 tok/s per user~250 tok/s per userCS-3 ~8x faster
On-chip memory bandwidth21 PB/s~64 TB/s (8 GPUs combined)CS-3 ~300x higher
Programming modelSingle deviceDistributed (tensor/pipeline parallelism)CS-3 simpler
Power consumption23 kW (system)~10 kW (8 GPUs + system)GPU cluster more efficient per watt
Training flexibilityOptimized via MemoryX/SwarmXIndustry-standard frameworksGPUs more flexible
Software ecosystemCerebras SDKCUDA/PyTorch/TensorFlowGPUs much broader

Expanded article table

The CS-3's primary advantage is inference latency, driven by the massive on-chip memory bandwidth that eliminates the memory wall problem. For training, GPU clusters retain advantages in software ecosystem maturity and flexibility, though Cerebras has made significant progress in training support through its compiler and MemoryX/SwarmX infrastructure [39].

Funding and Valuation

Cerebras has raised substantial capital through multiple funding rounds, including a final pre-IPO Series H in February 2026 that took the valuation to $23 billion:

RoundDateAmountPost-money valuationLead investors
Series AMay 2016$27Mn/aBenchmark, Foundation Capital, Eclipse Ventures
Series B2017$60Mn/aBenchmark Capital
Series C2018$112Mn/aBenchmark Capital
Series D2019$272M$1.6B (unicorn)Koch Disruptive Technologies
Series E2021$250Mn/aAlpha Wave Ventures
Series FNov 2021$720M$4BAlpha Wave Ventures, Abu Dhabi Growth Fund
Series GSept 2025$1.1B$8.1BFidelity, Atreides Management
Series HFeb 2026$1.0B$23BTiger Global (with Benchmark, Fidelity, Atreides, Alpha Wave, Altimeter, AMD, Coatue)
IPOMay 2026$5.55B base; $6.38B including the underwriters' option [52]~$48B at $185 listing pricePublic markets

Expanded article table

Total private funding before the IPO exceeded $4 billion. The Series H round, completed shortly after the OpenAI MRA was signed and closed on February 3, 2026, was led by Tiger Global and served as the final crossover round before Cerebras filed its amended S-1 in April 2026. The round nearly tripled the $8.1 billion valuation set by the September 2025 Series G [14][26][34]. Alongside equity, Cerebras arranged a revolving credit facility of up to $850 million from a syndicate of banks; the company said it closed the facility in April 2026, and its quarterly report states that the facility stepped up to $850 million on June 17, 2026 once post-IPO conditions were met [45][57].

IPO Filing and Public Listing

Cerebras first filed for an IPO on the Nasdaq in September 2024 under the reserved ticker symbol CBRS. The original filing faced delays after CFIUS opened its review of the G42 stake, which forced Cerebras to withdraw the prospectus and pause the listing process. After restructuring its investor base and completing the Series G round, Cerebras refiled its S-1 on April 17, 2026, and amended it on May 4, 2026, with concrete pricing terms [6][15][35].

When did Cerebras go public and how did the stock perform?

The IPO priced on May 13, 2026 at $185 per share, more than double the original target range, selling 30 million shares for total gross proceeds of $5.55 billion. Trading opened the following day at $350 and closed at roughly $311, up 68%, valuing Cerebras at around $95 billion on a fully diluted basis at the opening price, with the stock briefly topping a $100 billion market cap intraday. The offering was the largest global IPO of 2026 at the time of pricing, the largest U.S. tech IPO since Uber's 2019 debut, and one of the largest pure-play semiconductor listings in history [18][19][36]. The offering closed on May 15, 2026 with the underwriters' option to buy an additional 4.5 million shares exercised in full, bringing the total to 34.5 million shares and gross proceeds of about $6.38 billion before expenses, a figure Cerebras rounds to $6.4 billion in later filings [52][56]. The company's June 2026 results release called it the largest semiconductor IPO of all time, a company characterisation [45].

Financial performance

The S-1 filings disclosed rapid revenue growth alongside continued operating losses:

YearRevenueGAAP net income (loss)Notes
2022$24.6Mn/aEarly commercial period
2023$78.7Mn/aFirst full year of G42 ramp
2024$290.3M$(481.6M)Loss reflects share-based comp
2025$510.0M$237.8MOperating loss of $145.9M; net income driven by one-time gain

Expanded article table

In the 2025 fiscal year, Cerebras reported $510 million in revenue, up 76% from 2024, and a GAAP net income of $237.8 million. The swing to profitability was driven largely by a one-time, non-cash gain of $363.3 million from extinguishing a forward-contract liability tied to G42's original investment; on an operating basis Cerebras recorded a GAAP operating loss of $145.9 million and a non-GAAP net loss of roughly $75.7 million, meaning the company remained unprofitable on its core operations [17][26]. Customer concentration remained the dominant risk factor in the S-1, with UAE-linked entities (G42 and MBZUAI combined) accounting for roughly 86% of 2025 revenue, split 62% to MBZUAI and 24% to G42, which the prospectus notes are related parties with respect to each other [17][26].

Results as a public company

Cerebras reports both GAAP figures and a set of "core" non-GAAP figures that exclude the amortization of customer warrants, data center pass-through revenue and costs, stock-based compensation, and certain other items; the two differ materially, and the table shows both [45][56].

MetricQ1 2026 (quarter ended March 31)Q2 2026 (quarter ended June 30)
GAAP total revenue$193.4M (up 94% year over year)$180.1M (up 74%)
GAAP hardware revenue$110.6M$54.1M (down from $70.3M a year earlier)
GAAP cloud and other services revenue$82.8M (up 178%)$126.0M (up 281%)
Core total revenue$191.3M (up 92%)$209.9M (up 103%)
GAAP gross margin45%14%
Core gross margin47%41%
GAAP net income (loss)$(14.0M)$(450.5M)
Cash, equivalents, restricted cash, and short-term investments$3.3B$8.6B
Remaining performance obligationsNot stated in the release$25.4B
ReportedJune 23, 2026 [45]August 12, 2026 [56]

Expanded article table

The second quarter was the first reported quarter in which cloud and other services revenue exceeded hardware revenue. The GAAP net loss of $450.5 million compared with net income of $309.5 million in the year-earlier quarter, which had included the one-time gain described above; SiliconANGLE's Mike Wheatley reported that most of the 2026 loss was tied to stock-based compensation of $386.6 million [56][61]. Customer concentration eased but remained high: the quarterly report states that MBZUAI accounted for 34 percent of revenue in the June 2026 quarter, against 70 percent a year earlier, and G42 for 9 percent against 18 percent, while for the first half of 2026 the two were 49 percent and 10 percent [57]. Cerebras guided to core revenue of $214 million to $216 million for the third quarter, raised its full-year 2026 core revenue outlook to $880 million to $890 million from $855 million to $865 million, and said it planned to more than triple revenue in 2027 [45][56]. The market reaction was negative: CNBC reported that the shares fell about 14 percent in extended trading on August 12 after GAAP revenue of $180 million missed an LSEG consensus of $194 million, a comparison LSEG later said should be made on core revenue instead [59]. Cerebras's quarterly report also spells out an obligation the headline backlog carries: under the OpenAI agreement the company must deliver capacity tranches across specified numbers of data centers on time-based milestones, and OpenAI may terminate part or all of the agreement if it does not [57].

Competition and Market Position

Cerebras competes in the AI accelerator market against several established and emerging players:

CompetitorApproachKey products
NVIDIAGPU-based acceleratorsH100, H200, B200, B300
AMDGPU-based accelerators (Helios rack); disaggregated-inference partner since July 2026 [53]MI300X, MI325X, MI350, MI455X (Helios rack) [69]
GoogleCustom ASICsTPU v6 (Trillium), TPU v7 (Ironwood)
GroqLPU inference accelerators; technology licensed to NVIDIA in December 2025 (see Groq) [65]GroqChip1, LPU v2, NVIDIA Groq 3 LPX
SambaNovaReconfigurable dataflow units (RDUs)SN40L, SN50
IntelGaudi accelerators (discontinued)Gaudi 3
AmazonCustom ASICsTrainium2, Trainium3

Expanded article table

Cerebras differentiates itself through the sheer scale of its wafer-scale approach, the simplicity of programming a single massive chip versus a distributed cluster, and its focus on both training and inference performance. The company's inference speed records have been a particularly effective marketing tool, demonstrating tangible performance advantages over GPU-based solutions.

How does Cerebras compare with Groq and SambaNova?

The pure-play inference startups all chase a similar value proposition (orders-of-magnitude lower latency than NVIDIA GPUs on open-source LLMs), but they take very different architectural paths:

VendorArchitectureLlama 4 Maverick (tok/s)Notable model
CerebrasWafer-scale SRAM2,500+gpt-oss-120B at 2,700+ tok/s
SambaNovaReconfigurable Dataflow Units794DeepSeek-V3 at 198 tok/s
GroqLPU streaming architecture549Llama 70B steady at ~275 tok/s
NVIDIA BlackwellGPU clusters~290 (cloud)General-purpose flexibility

Expanded article table

Cerebras has consistently topped Artificial Analysis token-per-second leaderboards for the largest open-source models, while Groq tends to win on consistency across context lengths and SambaNova on cost-per-token in dense reasoning workloads. NVIDIA continues to dominate training and the broader ecosystem, while Cerebras and its peers compete primarily on inference performance and total cost of ownership [37][38].

Strengths and limitations

Cerebras's wafer-scale approach offers several distinct advantages:

  • Elimination of the memory wall: On-chip SRAM bandwidth of 21 PB/s removes the bottleneck that limits GPU inference speed.
  • Programming simplicity: A single CS-3 replaces a cluster of GPUs, eliminating the need for complex distributed computing frameworks.
  • Near-linear scaling: SwarmX fabric enables efficient multi-system scaling without the communication overhead of GPU clusters.
  • Defect tolerance: Fine-grained core architecture provides higher effective yield than traditional chips.

Limitations include:

  • Software ecosystem: The CUDA ecosystem's maturity and breadth remain a significant advantage for NVIDIA GPUs.
  • Total memory capacity: While MemoryX addresses this, the on-chip SRAM of 44 GB is smaller than the HBM capacity of modern GPUs (80-288 GB per chip).
  • Training ecosystem: Most AI researchers are trained on GPU-based workflows, creating inertia against adoption.
  • Cost and availability: CS-3 systems are not yet as widely available as GPU-based alternatives.
  • Customer concentration: UAE-linked customers continue to dominate revenue, creating execution and geopolitical risk that public investors have flagged.

Current State

As of May 2026, Cerebras is in a strong position. The company completed the largest global IPO of 2026 on May 14, 2026, raising $5.55 billion at a $185 price and trading up roughly 68% on the first day to imply a fully diluted market value around $95 billion. Cerebras has secured anchor partnerships with AWS and OpenAI, continued to set inference speed records on every major open-source model, and is broadening its data center footprint to handle Meta's Llama API traffic and the OpenAI Master Relationship Agreement.

The Condor Galaxy constellation continues to grow, with CG-3 fully operational in Dallas and the Condor Galaxy India site announced in May 2026. The company is on track to deliver more than 40 million tokens per second of aggregate inference capacity by the end of 2026 and is rolling out additional facilities in Oklahoma City, Minneapolis, Montreal, and undisclosed European sites. Cerebras remains the most prominent pure-play challenger to NVIDIA in AI inference and one of only a handful of independent companies capable of supplying frontier-scale AI compute to hyperscalers and frontier labs at the gigawatt level.

As of September 2026, the picture is more mixed than the May snapshot. The company's first two quarterly reports as a public company showed cloud revenue overtaking hardware revenue and a raised full-year outlook, but also a GAAP net loss of $450.5 million in the June quarter, driven largely by stock-based compensation, and a share price that CNBC described as volatile, reaching $386.34 on the first day of trading, dipping below $161 in late June, and closing at $262.06 on August 12 [56][58][59][61]. The product roadmap moved to the CS-4 and WSE-3 Turbo, with first shipments due in the third quarter of 2026 [63]. Cerebras's stated plan is to sell its systems both as standalone inference clusters and as the decode half of heterogeneous deployments: the AMD Helios pairing is due in production in the fourth quarter of 2026, the disaggregated Amazon Bedrock service in the first quarter of 2027, and European capacity, led by the Mikkeli site, is to come online from late 2026 [56][66][67]. Remaining performance obligations stood at $25.4 billion at June 30, 2026, and the quarterly report says the OpenAI agreement represents a substantial portion of projected revenue over the next several years [56][57].

References

  1. ^"Cerebras Systems." Wikipedia. en.wikipedia.org/...Cerebras
  2. ^"Report: Cerebras Business Breakdown & Founding Story." Contrary Research. research.contrary.com/...cerebras
  3. ^1 ^2 ^3"Cerebras Systems Unveils World's Fastest AI Chip with Whopping 4 Trillion Transistors." Cerebras Press Release. cerebras.ai/...third-generation-wafer-scale-engine
  4. ^"Cerebras CS-3 vs. Nvidia DGX B200 Blackwell." Cerebras Blog. cerebras.ai/...s-cs-3-vs-nvidia-dgx-b200-blackwell
  5. ^"AWS will bring Cerebras' wafer-size WSE-3 chip to its cloud platform." SiliconANGLE, March 13, 2026. siliconangle.com/...size-wse-3-chip-cloud-platform
  6. ^"Report: AI chipmaker Cerebras Systems rekindles IPO plans, targeting early 2026 listing." SiliconANGLE, December 21, 2025. siliconangle.com/...s-targeting-early-2026-listing
  7. ^"Cerebras powers Mistral's Le Chat to claim AI speed record." Digital Watch Observatory. dig.watch/...rals-le-chat-to-claim-ai-speed-record
  8. ^"BTLM-3B-8K: 7B Performance in a 3 Billion Parameter Model." Cerebras Blog. cerebras.ai/...ance-in-a-3-billion-parameter-model
  9. ^"Core42 and Cerebras Deliver Record-Breaking Performance for OpenAI's gpt-oss-120B." G42. g42.ai/...ring-ai-innovation-enterprises-worldwide
  10. ^"100x Defect Tolerance: How Cerebras Solved the Yield Problem." Cerebras Blog. cerebras.ai/...w-cerebras-solved-the-yield-problem
  11. ^1 ^2 ^3"The five technical challenges Cerebras overcame in building the first trillion-transistor chip." TechCrunch. techcrunch.com/...e-first-trillion-transistor-chip
  12. ^"Cerebras Wafer-Scale AI." Hot Chips 2024 Presentation. hc2024.hotchips.org/...Cerebras.Sean.v03.final.pdf
  13. ^"Cerebras CS-3: the world's fastest and most scalable AI accelerator." Cerebras Blog. cerebras.ai/...cerebras-cs3
  14. ^"Cerebras Systems Raises $1.1 Billion Series G at $8.1 Billion Valuation." BusinessWire, September 30, 2025. businesswire.com/...en
  15. ^"Cerebras Launches World's Fastest Inference for Meta Llama 4." Cerebras Press Release. cerebras.ai/...llama4PR
  16. ^"Cerebras WSE-3: Third Generation Superchip for AI." IEEE Spectrum. spectrum.ieee.org/cerebras-chip-cs3
  17. ^1 ^2 ^3 ^4 ^5"Breaking down AI chipmaker Cerebras' S-1." PitchBook. pitchbook.com/...ng-down-ai-chipmaker-cerebras-s-1
  18. ^1 ^2"Cerebras prices its shares at $185 ahead of biggest tech IPO in years, raising $5.5B." SiliconANGLE, May 13, 2026. siliconangle.com/...5-ahead-biggest-tech-ipo-years
  19. ^1 ^2"Cerebras pops 68% in Nasdaq debut, pushing the AI chipmaker's market cap to $95 billion." CNBC, May 14, 2026. cnbc.com/...cerebras-cbrs-stock-trade-nasdaq-ipo
  20. ^"Cerebras Systems: The Company That Built a Chip the Size of a Dinner Plate." Mukund Mohan. mukundmohan.blog/...-plate-and-is-now-going-public
  21. ^"Cerebras CEO Feldman's Stake Hits $3.2 Billion After Year's Biggest IPO." Bloomberg, May 14, 2026. bloomberg.com/...gest-ipo-into-3-2-billion-fortune
  22. ^"Argonne National Laboratory Deploys Cerebras CS-1." Argonne National Laboratory. anl.gov/...astest-artificial-intelligence-computer
  23. ^"Introducing Cerebras Inference: AI at Instant Speed." Cerebras Blog. cerebras.ai/...ebras-inference-ai-at-instant-speed
  24. ^"Cerebras Shatters Inference Records: Llama 3.1 405B Hits 969 Tokens Per Second." FinancialContent. markets.financialcontent.com/...ining-real-time-ai
  25. ^1 ^2"Meta Collaborates with Cerebras to Drive Fast Inference for Developers in New Llama API." Cerebras Press Release. cerebras.ai/...nce-for-developers-in-new-llama-api
  26. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8"Cerebras S-1/A." SEC, May 2026. sec.gov/...cerebras-sx1a2
  27. ^"G42 and Government of India Formalize Commercial Framework for Condor Galaxy India." HPCwire. hpcwire.com/...ondor-galaxy-india-ai-supercomputer
  28. ^"CFIUS Approves Cerebras Systems Shares Sale to UAE-Based AI Firm G42." GovCon Exec International, April 2025. govconexec.com/...ulatory-approval-for-shares-sale
  29. ^1 ^2 ^3"OpenAI partners with Cerebras." OpenAI, January 14, 2026. openai.com/...cerebras-partnership
  30. ^1 ^2 ^3"Cerebras scores OpenAI deal worth over $10 billion ahead of AI chipmaker's IPO." CNBC, January 14, 2026. cnbc.com/...ores-openai-deal-worth-over-10-billion
  31. ^"Cerebras Systems and U.S. Department of Energy Sign MOU." BusinessWire, December 2025. businesswire.com/...nd-U.S.-National-AI-Initiative
  32. ^"Department of Energy and Cerebras Systems Partner to Accelerate Science." BusinessWire, September 2019. businesswire.com/...-Scale-Artificial-Intelligence
  33. ^"Cerebras Announces Six New AI Datacenters Across North America and Europe." Cerebras Press Release. cerebras.ai/...ca-and-europe-to-deliver-industry-s
  34. ^1 ^2"Cerebras Systems Raises $1 Billion Series H." Cerebras Press Release, February 2026. cerebras.ai/...ystems-raises-usd1-billion-series-h
  35. ^1 ^2 ^3 ^4"Cerebras - S-1 (April 2026)." SEC. sec.gov/...cerebras-sx1april2026
  36. ^"Cerebras IPO mints two billionaires, sets stage for potential AI wave." CNBC, May 14, 2026. cnbc.com/...aires-sets-stage-for-potential-ai-wave
  37. ^"Cerebras vs SambaNova vs Groq: AI Chip Comparison." IntuitionLabs. intuitionlabs.ai/...-vs-sambanova-vs-groq-ai-chips
  38. ^"Cerebras CS-3 vs. Groq LPU." Cerebras Blog. cerebras.ai/...cerebras-cs-3-vs-groq-lpu
  39. ^"A Comparison of the Cerebras Wafer-Scale Integration Technology with Nvidia GPU-based Systems for Artificial Intelligence." arXiv. arxiv.org/...2503.11698v1
  40. ^1 ^2 ^3"Cerebras Selects Qualcomm to Deliver Unprecedented Performance in AI Inference." Cerebras Press Release, March 13, 2024. cerebras.ai/...-announce-10x-inference-performance
  41. ^1 ^2 ^3"Cerebras Software Platform (CSoft)." Cerebras. cerebras.ai/product-software
  42. ^1 ^2 ^3"Cerebras SDK Documentation (CSL Compiler and Language Guide)." Cerebras. sdk.cerebras.net
  43. ^"Beyond GPUs: Cerebras' Wafer-Scale Engine for Lightning-Fast AI Inference." The Data Exchange. thedataexchange.media/cerebras-inference
  44. ^1 ^2 ^3 ^4"AWS and Cerebras Collaborate to Deliver the Fastest Inference." Cerebras Press Release, March 13, 2026. cerebras.ai/...awscollaboration
  45. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8"Cerebras Systems Announces Strong First Quarter 2026 Results." Cerebras Systems, June 23, 2026. sec.gov/...cbrsannouncesfinancialresu
  46. ^1 ^2 ^3 ^4"Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed." OpenAI, August 13, 2026. openai.com/...previewing-ultrafast
  47. ^1 ^2 ^3 ^4"Accelerating GPT-5.6 Sol Ultrafast." Cerebras, August 13, 2026. cerebras.ai/...g-gpt-5-6-sol-ultrafast-with-openai
  48. ^"Fast mode." OpenAI API documentation, accessed August 15, 2026. developers.openai.com/...fast-mode
  49. ^"Ultrafast mode for GPT-5.6 Sol is now in limited preview, powered by Cerebras." Cerebras official X post, August 13, 2026. x.com/...2087961128869748856
  50. ^"GPT-5.6 Sol (high) Intelligence, Performance & Price Analysis." Artificial Analysis, accessed August 15, 2026. artificialanalysis.ai/...gpt-5-6-sol-high
  51. ^"Previewing Ultrafast mode: GPT-5.6 Sol at up to 14x the speed." OpenAI official X post, August 13, 2026. x.com/...2087947721936359705
  52. ^1 ^2 ^3"Cerebras Systems Announces Closing of Initial Public Offering." Cerebras Press Release, May 15, 2026. cerebras.ai/...-closing-of-initial-public-offering
  53. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8"AMD and Cerebras Announce Industry-Leading Ultra-Low-Latency and High Throughput AI Inference Solution." Cerebras Press Release, July 23, 2026. cerebras.ai/...cy-and-high-throughput-ai-inference
  54. ^"The GPU Is Being Split in Half." Cerebras Blog (Sarah Chieng), March 26, 2026. cerebras.ai/...disaggregated-inference
  55. ^1 ^2"Cerebras is coming to AWS." Cerebras Blog (James Wang), March 13, 2026. cerebras.ai/...cerebras-is-coming-to-aws
  56. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13"Cerebras Systems Fast Inference Cloud Business Nearly Quadruples in Second Quarter 2026." Cerebras Systems, Form 8-K Exhibit 99.1, August 12, 2026. sec.gov/...cbrsannouncesfinancialresu
  57. ^1 ^2 ^3 ^4 ^5"Cerebras Systems Inc., Quarterly Report on Form 10-Q for the quarter ended June 30, 2026." SEC, August 12, 2026. sec.gov/...cbrs-20260630
  58. ^1 ^2"Cerebras stock gains on AMD partnership." CNBC (Kif Leswing), July 23, 2026. cnbc.com/...cerebras-stock-gains-on-amd-partnership
  59. ^1 ^2"Cerebras stock plunges 14% after second earnings report following IPO." CNBC (Kif Leswing), August 12, 2026. cnbc.com/...cerebras-cbrs-q2-earnings-report-2026
  60. ^"Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution." SiliconANGLE (Thomas Godwin), July 29, 2026. siliconangle.com/...ce-cerebras-amd-neo4jgraphtalk
  61. ^1 ^2"AI chipmaker Cerebras Systems' stock plunges, despite posting solid earnings and guidance." SiliconANGLE (Mike Wheatley), August 12, 2026. siliconangle.com/...osting-solid-earnings-guidance
  62. ^"AMD, Cerebras team up to split AI workloads across chips." Axios (Ina Fried), July 23, 2026. axios.com/...amd-cerebras-ai-chips
  63. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9"Cerebras Unveils CS-4: Up to 30 Times Faster than GPU-based Solutions." Cerebras Systems press release via GlobeNewswire (Markets Insider), August 18, 2026. markets.businessinsider.com/...olutions-1036472378
  64. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8"Introducing Cerebras CS-4: The Fastest AI Just Got Faster." Cerebras Blog (Angela Yeung and Eric Gardner), August 18, 2026. cerebras.ai/...introducing-cerebras-cs-4
  65. ^1 ^2 ^3 ^4"Nvidia and Cerebras are selling performance their customers will (probably) never see." The Register (Tobias Mann), August 27, 2026. theregister.com/...5293117
  66. ^1 ^2 ^3"Cerebras and Compute Nordic Finland Announce New 165 MW AI Data Centre in Mikkeli, Finland." Cerebras Systems press release via GlobeNewswire (Markets Insider), September 1, 2026. markets.businessinsider.com/...-finland-1036509897
  67. ^1 ^2 ^3"Cerebras Systems Accelerates European Expansion with 200MW of AI Compute Capacity by End of 2027." Cerebras Systems press release via GlobeNewswire (Nasdaq), July 9, 2026. nasdaq.com/...n-200mw-ai-compute-capacity-end-2027
  68. ^"There's only one technique that will save SRAM-only chips: attention-feedforward disaggregation." X post by Joe Fioti (@joefioti), September 4, 2026. x.com/...2095993586320019913
  69. ^1 ^2"AAI 2026: AMD Delivers Full-Stack Compute for the Agentic AI Era." AMD Investor Relations press release, July 23, 2026. ir.amd.com/...stack-compute-for-the-agentic-ai-era

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

9 revisions · v10 · 10,684 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: xf96 Sep 6 2026: verifier V2 checked 96 claims (AMD/AWS releases, CS-4, 10-Q, Q1/Q2 results, IPO closing); Axios quote unverifiable and removed

Cite this page: AI Wiki. "Cerebras Systems." aiwiki.ai, updated 7 Sept 2026, fact-checked 7 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/cerebras

Suggest edit