Citation and evidence

NVIDIA Vera (CPU)

25 min full readUpdated 30 references

This article's verification

Report a problem with this article

More

Use this article

Raw MarkdownExplore connections

Improve this page

Suggest editRevision historyDiscussion

Browse categories

AI AgentsAI HardwareAI InfrastructureNVIDIA

Cite this article

NVIDIA Vera is a custom Arm-based data center central processing unit from NVIDIA, marketed by the company as "the CPU for agents," a processor designed first for agentic AI workloads rather than the human-driven, interactive computing that has shaped server CPUs for decades. NVIDIA first showed Vera and its custom cores in Jensen Huang's GTC keynote in March 2025, formally launched the chip at GTC on March 16, 2026, and reaffirmed at the GTC Taipei keynote at COMPUTEX 2026 on June 1, 2026 that it was in full production, a status first stated in the March release (Taipei time; the press release is dated May 31 in the United States).[23][15][1][3] Vera is built around 88 custom Arm-compatible "Olympus" cores with NVIDIA Spatial Multithreading (176 threads), and it serves as the host processor for the NVIDIA Vera Rubin platform, where it is paired with NVIDIA's Rubin GPUs over a coherent NVLink-C2C link that carries up to 1.8 TB/s. NVIDIA presents Vera as the next step beyond its Grace CPU and claims it completes agentic tasks up to 1.8 times faster than leading x86 processors, although the size of that claimed advantage has varied across NVIDIA's own materials.[1][2][9][13] NVIDIA hand-delivered the first Vera systems to customers in May 2026.[18] On September 28, 2026, NVIDIA called Vera "the first purpose-built CPU for agentic AI" (its own marketing description) when it named the chip as the optimized host for the NVIDIA OpenShell agent runtime in its new NVIDIA Open Agent Safety Platform.[25]

This article covers the Vera CPU specifically. For the full rack, the GPUs, and the broader system, see NVIDIA Vera Rubin.

Key facts

AttributeDetail
DeveloperNVIDIA
TypeData center CPU for agentic AI
First shownGTC keynote, March 2025[23]
LaunchedMarch 16, 2026 (GTC 2026); in full production per the March launch release, reaffirmed June 1, 2026 (GTC Taipei keynote at COMPUTEX 2026)[15][1][3]
CPU cores88 custom NVIDIA "Olympus" cores (a single 88-core SKU)[1][2][28]
Instruction setArm-compatible (Armv9.2)[2][5][11]
ThreadingNVIDIA Spatial Multithreading, 176 threads[2]
Cache2 MB L2 per core; 164 MB unified L3[5][17]
MemoryLPDDR5X on SOCAMM2 modules, up to 1.5 TB capacity[2][17]
Memory bandwidthUp to 1.2 TB/s[1][2]
CPU-GPU linkSecond-generation NVLink-C2C, up to 1.8 TB/s coherent[1][2]
On-chip fabricSecond-generation SCF, 3.4 TB/s bisection bandwidth[2]
TDPConfigurable, 250 W to 450 W[16]
Paired GPUNVIDIA Rubin (in Vera Rubin)[5]
PredecessorNVIDIA Grace CPU[1]
AvailabilityFull production; first systems hand-delivered May 2026; available from system builders and cloud partners from fall 2026[1][18]

Expanded article table

What is the NVIDIA Vera CPU?

Vera is NVIDIA's second-generation data center CPU and its first to use a fully custom CPU core, rather than the licensed Arm Neoverse V2 cores used in the earlier Grace design; NVIDIA describes Olympus as its "first fully custom data center CPU core."[2][5] NVIDIA frames it as a clean break in CPU philosophy. Announcing full production at GTC Taipei, NVIDIA chief executive Jensen Huang said that "AI agents will be the largest users of computing" and that Vera "is the first CPU designed for that future," one "built to run agentic AI at hyperscale with extraordinary performance, efficiency and programmability," the reasoning being that fleets of autonomous AI agents will become the dominant consumers of compute and that they place very different demands on a CPU than a person clicking through an application does.[1][3][7]

In practice that positioning translates into a chip tuned for the parts of an AI workload that a GPU does not handle well: orchestration, control flow, data movement, scheduling, and keeping accelerators fed. NVIDIA targets three workload classes explicitly, agentic AI, reinforcement learning, and large-scale data processing, all of which mix latency-sensitive control logic with heavy memory traffic. The company's headline claim at GTC Taipei was that Vera enables "1.8x faster task completion compared with x86 CPUs" on these workloads; later NVIDIA releases phrase it as "up to 1.8x." That figure is an NVIDIA marketing claim measured on its own agentic benchmarks and should be read as such.[1][13]

Vera is not a consumer part. It is sold in data center systems in several forms: as the host CPU of the Vera Rubin NVL72 and HGX Rubin NVL8 platforms, in a dense liquid-cooled Vera CPU rack, and in standard single- and dual-socket servers. At GTC Taipei NVIDIA said Dell, HPE, Lenovo and Supermicro would offer Vera in standalone CPU server configurations, which it called "the first standard CPU option beyond x86."[1][2][9]

What is the Olympus core architecture?

Each Vera CPU integrates 88 of NVIDIA's custom "Olympus" cores on a single monolithic compute die. NVIDIA designed Olympus for high single-thread performance and energy efficiency while remaining fully compatible with the Armv9.2 instruction set. NVIDIA's January 2026 comparison with Grace lists six 128-bit SVE2 vector units with FP8 support per core, against four 128-bit SVE2 units in Grace.[2][5][11] NVIDIA claims Olympus delivers up to 50% higher instructions per cycle than Grace.[16]

The core uses a wide, deep microarchitecture aimed at the control-heavy, data-movement-intensive code that surrounds modern AI inference. NVIDIA describes a 10-wide instruction fetch and decode frontend and a neural branch predictor capable of evaluating two taken branches per cycle, alongside improved prefetching and load-store performance. Its July 2026 deep dive adds deep out-of-order execution with a large reorder buffer, dependency-breaking features including memory renaming and value prediction, and a graph prefetcher aimed at the pointer-chasing memory accesses common in graph analytics and agent workflows. The intent is to extract high instructions-per-clock on the irregular, branchy code that drives agent orchestration, where a GPU is a poor fit.[2][17] Drawing on NVIDIA's Vera whitepaper, ServeTheHome reported that each Olympus core has 18 execution pipes in total.[30]

Keeping the cores on one compute die is a deliberate choice. NVIDIA pairs that single die with adjacent dielets that implement the memory and I/O subsystems, but the 88 cores and the coherency fabric live together on one piece of silicon. In NVIDIA's Hot Chips 2026 presentation in August, as reported by ServeTheHome, the package comprised six dies on a single interposer: the monolithic compute die plus separate memory and I/O chiplets, with eight 128-bit LPDDR5X memory controllers. The company argues the monolithic compute die avoids cross-chiplet latency and gives the predictable, low-variance response times that agent workloads need.[2][21]

How does Spatial Multithreading work?

The standout architectural feature is what NVIDIA calls Spatial Multithreading, its variant of simultaneous multithreading. Across 88 cores Vera exposes 176 threads, two per core. Where conventional SMT time-shares a single core's execution resources between threads, NVIDIA says its approach runs two hardware threads per core "by physically partitioning resources instead of time-slicing."[2][5]

The payoff NVIDIA emphasizes is determinism. Because the threads are not competing for the same shared structures cycle by cycle, the design is meant to deliver stronger isolation between threads and more predictable tail latency, which matters when many agents share a machine and one slow request can stall a pipeline. NVIDIA also notes that operators can choose between maximizing per-thread performance and maximizing thread count at runtime: a core can act as a high-throughput single-threaded core, with the sibling thread handling management work, or as two more isolated execution contexts.[2][17] At Hot Chips 2026, ServeTheHome reported, NVIDIA described the scheme as statically partitioned and said it achieves the same level of performance as traditional SMT while being less susceptible to noise from threads competing for the same resources.[21]

How much memory bandwidth does Vera have?

Vera uses a second-generation LPDDR5X memory subsystem that delivers up to 1.2 TB/s of total bandwidth and up to 1.5 TB of capacity. NVIDIA reports roughly 14 GB/s of memory bandwidth provisioned per core, which it characterizes as about three times the bandwidth per core of leading x86 CPUs using DDR5, and about twice the bandwidth at half the power of conventional CPU memory.[1][2][9]

A notable packaging change accompanies the move to LPDDR5X. Vera's memory ships on SOCAMM modules (small outline compression-attached memory modules; NVIDIA's later material specifies SOCAMM2), which replace the soldered LPDDR found in earlier low-power designs with detachable, field-replaceable modules. This brings low-power mobile-class memory into a serviceable data center form factor, addressing one of the practical drawbacks of soldered memory at rack scale.[2][17] NVIDIA says the LPDDR5X subsystem typically consumes less than 30 watts, compared with "well over 100 watts" for DDR5 configurations, and that Vera has 40% lower peak memory latency than x86 CPUs.[16]

Binding the cores and memory together is a second-generation Scalable Coherency Fabric (SCF). The SCF connects all 88 Olympus cores to a shared, unified last-level cache and provides 3.4 TB/s of bisection bandwidth, and NVIDIA states the design sustains over 90 percent of peak memory bandwidth under load. NVIDIA's own GTC Taipei live blog gave a different figure, describing a "3.6TB/s on-chip fabric"; its technical blogs and product page use 3.4 TB/s.[2][3][9] NVIDIA's January 2026 Grace-versus-Vera table lists 2 MB of L2 cache per core (double Grace's 1 MB) and a 164 MB unified L3 cache (Grace: 114 MB), figures repeated in its later blog posts.[5][11][17]

How does Vera connect to the Rubin GPU?

Vera connects to NVIDIA's GPUs over NVLink-C2C (chip-to-chip), a coherent interconnect that provides up to 1.8 TB/s of bandwidth between the CPU and GPU. This is roughly double the 900 GB/s NVLink-C2C link used in the earlier Grace Hopper generation.[1][2]

Coherence is the important property. The link lets a Vera CPU and the GPUs it serves share a single unified memory address space, so the GPU can reach the CPU's LPDDR5X and the CPU can reach the GPU's HBM as one coherent pool without explicit copies. For agentic and reasoning workloads this is what makes practices like KV-cache offload, where key-value attention state spills from GPU memory into the larger CPU memory pool, efficient enough to be useful. NVIDIA positions Vera not as a passive host but as a high-bandwidth data-movement engine tightly coupled to GPU execution.[1][2][5]

The same second-generation NVLink-C2C links two Vera CPUs in dual-socket servers at up to 1.8 TB/s, and NVIDIA says each socket presents as a single NUMA domain. For other I/O, Vera supports PCIe Gen 6 (NVIDIA's July 2026 deep dive specifies PCIe 6.4) and CXL 3.1, with 176 PCIe lanes in a dual-socket configuration. NVIDIA also extends confidential computing to Vera, built on Arm CCA/RME with per-VM encryption keys and encrypted C2C links.[2][5][17]

What is the Vera Rubin platform?

Vera is the host CPU of the Vera Rubin platform, NVIDIA's Blackwell successor for AI factories. The basic building block is the Vera Rubin Superchip, which pairs one Vera CPU with two Rubin GPUs over NVLink-C2C. In the flagship rack, the Vera Rubin NVL72, NVIDIA combines 36 Vera CPUs with 72 Rubin GPUs in a single liquid-cooled, NVLink-connected domain, a two-to-one GPU-to-CPU ratio.[5][10]

Within that system Vera handles orchestration, data movement, and coherent memory access across the rack while the Rubin GPUs do the dense math, with BlueField-4 DPUs handling networking, storage and security offload. Vera is also the host CPU for NVIDIA's HGX Rubin NVL8 reference designs, where it connects to Rubin GPUs over PCIe. At GTC in March 2026 NVIDIA described a Vera Rubin POD built from five rack-scale systems: Vera Rubin NVL72, the Vera CPU rack, Groq 3 LPX, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet, the same five-rack lineup Huang said at GTC Taipei was ramping into full production. Full details of the rack, the GPUs, and the networking are covered on the NVIDIA Vera Rubin page.[2][3][5][29]

Separately, NVIDIA introduced a CPU-only Vera CPU rack at GTC in March 2026 for customers who want Vera's general-purpose compute on its own. Built on the NVIDIA MGX modular architecture, it integrates up to 256 liquid-cooled Vera CPUs, which NVIDIA says sustain more than 22,500 concurrent CPU environments, and NVIDIA claims over 4x the capacity and 2x the performance per watt of x86-based server racks. Tom's Hardware's GTC coverage reported the rack pairs the 256 CPUs with 74 BlueField-4 DPUs and ConnectX SuperNIC networking.[2][8][15]

What is the Vera BlueField-4 STX Storage Processor?

On August 3, 2026, NVIDIA published a set of storage microbenchmarks that extend Vera's stated role beyond hosting GPUs. In a developer blog post by Jason Hardy and Ronil Prasad, the company described the NVIDIA Vera BlueField-4 STX Storage Processor, a component of what it calls the NVIDIA STX foundation for AI-native data platforms, as bringing the Vera CPU directly into the storage data path of the BlueField line. The framing follows the same agentic logic as the CPU itself: NVIDIA argues that every agent step can trigger multiple storage operations, retrieving enterprise knowledge, reading and writing persistent memory, and reusing KV cache data, that these operations repeat across thousands of concurrent agents with growing context windows, and that the encryption, checksumming, compression, and redundancy work in the storage path runs on CPUs rather than GPUs.[11]

In the microbenchmarks, run with a purpose-built NVIDIA test framework on memory-resident data using common libraries including OpenSSL, Zstandard, and LZ4, the company reports Vera delivering up to 1.43x higher AES-128 encryption throughput and up to 1.29x higher AES-128 decryption throughput than the x86 CPU used for comparison, up to 3.26x higher Reed-Solomon throughput in an erasure-coding recovery workload, up to 3.67x higher CRC32C integrity-checking throughput, up to 3.29x higher compression throughput and up to 1.72x higher decompression throughput under concurrency, and up to 3.21x higher throughput on a two-stage pipeline that compresses and then encrypts each data buffer, the figure NVIDIA led with when announcing the results on X the same day. These are the vendor's own numbers: the comparison processor is identified only as "the x86 CPU," the tests exclude file I/O, disk performance, and networking, and NVIDIA itself notes that end-to-end testing is still required to quantify complete storage-system or GPU-performance outcomes.[11][12]

NVIDIA attributes the results to the architecture described above: sustained per-core throughput from the Olympus core's wide execution, branch prediction, and vector and cryptographic resources for the serial stages inside each data stream, and the SCF, unified L3 cache, and up to 1.2 TB/s SOCAMM2 LPDDR5X memory subsystem for the many concurrent streams around them, with Spatial Multithreading credited with reducing thread-to-thread interference. The post adds that across selected workloads Vera also delivered higher measured performance per watt, and positions a common Vera architecture and toolchain across compute and storage as one foundation for scaling both agent execution and data infrastructure.[11]

How does Vera differ from Grace?

Vera is the direct successor to NVIDIA's Grace CPU, the Arm-based data center processor introduced earlier in the Grace Hopper and Grace Blackwell generations. NVIDIA said in its May 31, 2026 GTC Taipei release that Grace had nearly 2.5 million shipments to date, and it presents Vera as building on that base while moving to a fully custom core.[1]

The generational jump is substantial on paper. NVIDIA's own comparison table shows Vera increasing the core count from 72 to 88, tripling memory capacity (from up to 480 GB to up to 1.5 TB of LPDDR5X), raising memory bandwidth 2.4 times (from up to 512 GB/s to up to 1.2 TB/s), and doubling the NVLink-C2C link to 1.8 TB/s. The biggest qualitative change is the switch from licensed Arm Neoverse V2 cores to NVIDIA's own Olympus cores, which is what lets the company tune the microarchitecture specifically for agentic workloads.[2][5]

MetricGraceVeraSource basis
CPU cores72 (Arm Neoverse V2)88 (custom Olympus)NVIDIA[5]
Threads72176 (Spatial Multithreading)NVIDIA[5]
L2 cache per core1 MB2 MBNVIDIA[5]
Unified L3 cache114 MB164 MBNVIDIA[5]
Memory capacityUp to 480 GB LPDDR5XUp to 1.5 TB LPDDR5XNVIDIA[5]
Memory bandwidthUp to 512 GB/sUp to 1.2 TB/sNVIDIA[5]
SIMD4x 128-bit SVE26x 128-bit SVE2, FP8NVIDIA[5]
NVLink-C2C900 GB/sUp to 1.8 TB/sNVIDIA[1][5]
PCIe / CXLGen 5Gen 6 / CXL 3.1NVIDIA[5]
Confidential computingNot supportedSupportedNVIDIA[5]
Core sourceLicensed Arm NeoverseNVIDIA custom OlympusNVIDIA[2][5]

Expanded article table

Reported specifications

The table below collects the published figures. Almost all come from NVIDIA's own materials; where NVIDIA's documents disagree, the difference is noted.

SpecificationValueSource basis
Cores88 custom Olympus cores (single 88-core SKU)NVIDIA[1][2]; Tom's Hardware[28]
Threads176 (Spatial Multithreading)NVIDIA[2]
Instruction setArm-compatible; Armv9.2NVIDIA[2][11]
Frontend10-wide fetch and decodeNVIDIA[2]
Branch predictorNeural, 2 taken branches/cycleNVIDIA[2]
L2 cache2 MB per coreNVIDIA[5]; Phoronix[22]
L3 cache164 MB unifiedNVIDIA[5][11]
On-chip fabric2nd-gen SCF, 3.4 TB/s bisection (3.6 TB/s in NVIDIA's GTC Taipei live blog)NVIDIA[2][3]
MemoryLPDDR5X (SOCAMM2), up to 1.5 TBNVIDIA[2][17]
Memory bandwidthUp to 1.2 TB/s (up to 14 GB/s per core, ~3x x86 DDR5 per core)NVIDIA[1][2][9]
CPU-GPU linkNVLink-C2C, up to 1.8 TB/s coherentNVIDIA[1][2]
Host I/OPCIe Gen 6 (PCIe 6.4), CXL 3.1; 176 PCIe lanes in dual-socket systemsNVIDIA[5][17]
PrecisionFP8 supported (6x 128-bit SVE2)NVIDIA[5]
Socket powerConfigurable 250 W to 450 W TDP; 450 W peak socket TDP on the pre-production system Phoronix testedNVIDIA[16]; Phoronix[22]

Expanded article table

What benchmarks has NVIDIA published?

NVIDIA's headline performance claim for Vera has shifted over the year. The March 16, 2026 launch release said Vera delivered "results with twice the efficiency and 50% faster than traditional rack-scale CPUs," and the accompanying technical blog claimed up to 1.5x the agentic sandbox performance of competitive x86 platforms, measured against AMD EPYC Turin and Intel Xeon 6 Granite Rapids. By June 1, 2026, NVIDIA was claiming "more than 1.8x higher agentic sandbox performance than traditional x86 architectures," and its product page and later releases use "up to 1.8x."[15][2][16][9][13] A July 2026 NVIDIA post attributes a 1.8x single-thread advantage to the Olympus cores and illustrates it with a reinforcement-learning example in which Vera completes up to 85% of evaluations within a training window, against 45% for a baseline CPU.[27]

The first public third-party numbers came from Phoronix, which NVIDIA invited to its Santa Clara headquarters (the review was published April 14, 2026). Phoronix wrote that Vera was more competitive with Intel and AMD x86 CPUs than any other Arm or non-x86 processor it had tested, but noted that NVIDIA limited testing to workloads it chose and asked that CPU power consumption not be monitored because power-management tuning was still being upstreamed. NVIDIA's GTC Taipei release cited Phoronix as finding that Vera delivered the fastest overall performance across agentic workloads including code compilation, Python, Java and database processing.[22][1][4]

On July 21, 2026, NVIDIA released a Vera whitepaper with internally measured SPEC CPU 2026 results (the SPECrate integer suite) against a dual-socket AMD EPYC 9755 system. NVIDIA presented the results normalized per physical core under full socket load, which it argued is the relevant metric for agent sandboxes. Tom's Hardware noted that this is not how SPECrate results are normally reported: on overall base score, the dual-socket Vera system scored 925 against 898 for the EPYC 9755, about 3% ahead in total throughput despite having fewer threads, rather than the 70% to 80% margin the per-core chart suggests.[20][17] An August 24, 2026 NVIDIA post estimated that Vera delivers up to 1.5x the per-core performance of AMD's next-generation Venice across four agentic workloads (compiler, static analysis and Python benchmarks), with Venice's figures estimated from a published score and internal Turin measurements rather than measured. The same post cited telemetry from 163,594 agentic sessions, over 97% of which had unique trajectory profiles, and a 33-minute Claude Code session, to argue that one balanced CPU design suits agent fleets better than several specialized ones.[19] Tom's Hardware reported that AMD had countered that its 256-core Zen 6 Venice processors beat Vera by 3.3 times in rack-level performance.[14]

NVIDIA presented Vera at Hot Chips 2026 in August. ServeTheHome reported that NVIDIA's SPEC-based proxies showed agentic workloads close to 1.8x and data processing workloads closer to 1.5x, and that NVIDIA itself stressed these were proxy workloads.[21]

When is Vera available, and who is using it?

NVIDIA's March 16, 2026 launch release said Vera was in full production and would be available from partners in the second half of 2026. It named customers collaborating with NVIDIA to deploy Vera, including Alibaba Cloud, ByteDance, Meta, Oracle Cloud Infrastructure (OCI), CoreWeave, Lambda, Nebius and Nscale, and said Cursor was adopting Vera for its AI coding agents. Streaming-data company Redpanda said in the release that its tests of Kafka-compatible workloads on Vera showed up to 5.5x lower latency, and the Leibniz Supercomputing Centre, Los Alamos National Laboratory, NERSC at Lawrence Berkeley National Laboratory and the Texas Advanced Computing Center (TACC) were named as national laboratories planning to deploy Vera. The infrastructure providers listed in March included Cisco, Dell, HPE, Lenovo and Supermicro, and NVIDIA's March technical blog said Vera systems would be available from those OEMs in the second half of 2026.[15][2]

At GTC Taipei, NVIDIA said Vera and the Vera Rubin platform were in full production and on schedule to ship in the fall of 2026, and that Vera systems would be available from system builders and cloud partners "starting this fall." That release described NYSE, Anthropic, OpenAI, SpaceXAI, ByteDance, CoreWeave, Lambda, Nebius, Nscale and OCI as customers "exploring" Vera; it said Anthropic was "evaluating adding Vera to scale CPU-intensive agentic workloads" and that the NYSE would use Vera with Redpanda and HPE. Cloud providers it said were planning to deploy Vera included Akamai, Cloudflare, Crusoe, Together AI and Vultr alongside those already named. Its OEM list named Dell, HPE, Lenovo and Supermicro but not Cisco.[1][6]

Deliveries began before the Taipei announcement. According to an NVIDIA blog post first published on May 18, 2026, Ian Buck, NVIDIA's vice president of hyperscale and HPC, hand-delivered the first Vera CPU systems in May to Anthropic, OpenAI, SpaceXAI and OCI; OCI's Karan Batta said OCI planned to deploy "hundreds of thousands of NVIDIA Vera CPUs beginning in 2026." An August 27 update to the post said Amazon Web Services had received its first Vera CPU server and Vera Rubin GPU, and described Vera as beginning to ship at scale.[18] The day before, on August 26, 2026, AWS and NVIDIA announced that they were working to bring Vera CPU-based infrastructure to AWS as part of a wider expansion of their partnership.[24] Tom's Hardware reported that Meta had confirmed in February 2026 that it would run Grace-only servers in production, with Vera to follow as soon as 2027.[14]

On Aug. 24, 2026, NVIDIA announced that SpaceXAI would deploy Vera CPUs for the CPU-side work around its next generation of agentic AI applications, including tool orchestration, code execution, data processing, and simulation between model calls. NVIDIA also said SpaceXAI planned to expand the infrastructure behind Grok with the Vera Rubin platform, and that SpaceXAI's planned first-generation Starmind AI satellite would be based on an optimized Vera Rubin NVL72 system.[13] The announcement did not disclose a Vera CPU count, deal value, or start date for the terrestrial rollout; Tom's Hardware reported that the only date attached was the satellite launch window, which Elon Musk put at the fourth quarter of 2027.[14] NVIDIA expressly classified the planned deployment and related benefit claims as forward-looking. NVIDIA's own delivery blog says a Vera system was hand-delivered to SpaceXAI's Palo Alto offices in May 2026 and that SpaceXAI was evaluating Vera for reinforcement learning workloads and agent-based simulation pipelines, which is evidence of evaluation hardware rather than a production deployment.[13][18]

How does Vera fit into NVIDIA's agent safety platform?

On September 28, 2026, NVIDIA announced the NVIDIA Open Agent Safety Platform, which combines the open source OpenShell runtime with NVIDIA Sentry, a reference system design that runs on BlueField-4 DPUs. The press release says OpenShell "provides a secure runtime boundary that traces all actions and enforces policy as agents run on NVIDIA Vera CPUs," and that it "delivers this protection with minimal overhead on NVIDIA Vera, the first purpose-built CPU for agentic AI." The "first" and "minimal overhead" wording is NVIDIA's own, and the release gives no overhead figure. NVIDIA added that, as open source software, OpenShell can be extended to third-party compute platforms, including those from Arm and Intel.[25]

NVIDIA's companion technical blog says the platform "is optimized to run on NVIDIA Vera CPU- and BlueField DPU-based systems, and is also compatible with other hardware systems." In a Vera Rubin POD, it says, a BlueField-4 DPU sits on each compute tray's only path to the model, where organizations can run Sentry as an optional security layer alongside OpenShell, and for anyone already running a Vera system with BlueField-4, "enabling these protections is just a software update."[26] The GTC Taipei release had already positioned the Vera BlueField-4 STX processor as combining Vera with networking, storage acceleration and "in-silicon security."[1]

References

  1. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22NVIDIA Newsroom. "NVIDIA Unveils Vera, the CPU for Agents." nvidianews.nvidia.com/...s-vera-the-cpu-for-agents
  2. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27 ^28 ^29 ^30 ^31 ^32 ^33 ^34 ^35 ^36NVIDIA Technical Blog. "NVIDIA Vera CPU Delivers High Performance, Bandwidth, and Efficiency for AI Factories." developer.nvidia.com/...fficiency-for-ai-factories
  3. ^1 ^2 ^3 ^4 ^5 ^6NVIDIA Blog. "NVIDIA GTC Taipei at COMPUTEX: Live Updates on What's Next in AI." blogs.nvidia.com/...-gtc-taipei-computex-2026-news
  4. ^SQ Magazine. "NVIDIA Vera ARM CPU Outperforms Intel Xeon and AMD EPYC." sqmagazine.co.uk/...a-arm-cpu-intel-amd-benchmarks
  5. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10 ^11 ^12 ^13 ^14 ^15 ^16 ^17 ^18 ^19 ^20 ^21 ^22 ^23 ^24 ^25 ^26 ^27NVIDIA Technical Blog. "Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer." developer.nvidia.com/...chips-one-ai-supercomputer
  6. ^DataCenterKnowledge. "GTC Taipei: Nvidia Says Vera Rubin, Vera CPU on Track." datacenterknowledge.com/...-os-to-run-ai-factories
  7. ^DigiTimes. "Nvidia's Vera CPU is built for agents, not humans; Jensen Huang says it opens a market that never existed before." digitimes.com/...uang-nvidia-gtc-ai-agent-cpu-2026
  8. ^Tom's Hardware. "Nvidia unveils details of new 88-core Vera CPUs positioned to compete with AMD and Intel." tomshardware.com/...to-a-6x-gain-in-cpu-throughput
  9. ^1 ^2 ^3 ^4 ^5 ^6NVIDIA. "Next Gen Data Center CPU | NVIDIA Vera CPU." nvidia.com/...vera-cpu
  10. ^NVIDIA. "Rack-Scale Agentic AI Supercomputer | NVIDIA Vera Rubin NVL72." nvidia.com/...vera-rubin-nvl72
  11. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8Jason Hardy and Ronil Prasad, NVIDIA Technical Blog. "NVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage." August 3, 2026. developer.nvidia.com/...very-for-ai-native-storage
  12. ^NVIDIA AI Infrastructure (@NVIDIAAIInfra) on X. August 3, 2026. x.com/...2084336350715637768
  13. ^1 ^2 ^3 ^4 ^5NVIDIA Newsroom. "SpaceXAI Adopts NVIDIA Vera CPU to Accelerate Agentic AI at Massive Scale." August 24, 2026. nvidianews.nvidia.com/...entic-ai-at-massive-scale
  14. ^1 ^2 ^3Luke James, Tom's Hardware. "SpaceXAI will deploy standalone Nvidia Vera CPUs for Grok's agentic workloads." August 25, 2026. tomshardware.com/...us-for-groks-agentic-workloads
  15. ^1 ^2 ^3 ^4 ^5NVIDIA Newsroom. "NVIDIA Launches Vera CPU, Purpose-Built for Agentic AI." March 16, 2026. nvidianews.nvidia.com/...pose-built-for-agentic-ai
  16. ^1 ^2 ^3 ^4 ^5Praveen Menon, Ivan Goldwasser, Ian Finder and Diana Aung, NVIDIA Technical Blog. "NVIDIA Vera CPU Sets a New Standard for Agentic Workloads in AI Factories." June 1, 2026. developer.nvidia.com/...-workloads-in-ai-factories
  17. ^1 ^2 ^3 ^4 ^5 ^6 ^7 ^8 ^9 ^10Praveen Menon and Eduardo Alvarez, NVIDIA Technical Blog. "NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI." July 21, 2026. developer.nvidia.com/...-performance-in-agentic-ai
  18. ^1 ^2 ^3 ^4NVIDIA Blog. "Delivering Vera: NVIDIA's First CPU Built for Agents Is Shipping Now." Originally published May 18, 2026; updated August 27, 2026. blogs.nvidia.com/...vera-cpu-delivery
  19. ^Eduardo Alvarez and Praveen Menon, NVIDIA Technical Blog. "Solving Agentic AI Fleet Challenges with NVIDIA Vera CPU." August 24, 2026. developer.nvidia.com/...enges-with-nvidia-vera-cpu
  20. ^Jake Roach, Tom's Hardware. "Nvidia deep dives Vera CPU for AI data centers: SPEC CPU 2026 benchmarks revealed, Olympus architecture specifics, and more." July 21, 2026. tomshardware.com/...architecture-detailed-and-more
  21. ^1 ^2 ^3ServeTheHome. "NVIDIA Vera CPU at Hot Chips 2026." August 24, 2026. servethehome.com/nvidia-vera-cpu-at-hot-chips-2026
  22. ^1 ^2 ^3Phoronix. "NVIDIA Vera CPU Benchmarks: Olympus Cores Delivering The Best Performance Ever Seen On ARM." phoronix.com/...nvidia-vera-benchmarks
  23. ^1 ^2Jeremy Laird, PC Gamer. "Nvidia reveals Vera, a new CPU with 'custom' cores which could be very exciting for its upcoming premium PC processor." March 19, 2025. pcgamer.com/...r-its-upcoming-premium-pc-processor
  24. ^NVIDIA Newsroom. "AWS and NVIDIA to Deliver 2 Million Additional GPUs and Next-Generation Infrastructure for Agentic and Physical AI." August 26, 2026. nvidianews.nvidia.com/...r-agentic-and-physical-ai
  25. ^1 ^2NVIDIA Newsroom. "NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment." September 28, 2026. nvidianews.nvidia.com/...open-agent-safety-platform
  26. ^John Myers, Alex Watson, Ali Golshan and Ofir Arkin, NVIDIA Technical Blog. "NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring." September 28, 2026. developer.nvidia.com/...n-silicon-agent-monitoring
  27. ^Praveen Menon, Eduardo Alvarez and Farshad Ghodsian, NVIDIA Technical Blog. "NVIDIA Vera CPU Boosts AI Factory Throughput to Accelerate Agentic Workloads." July 7, 2026. developer.nvidia.com/...celerate-agentic-workloads
  28. ^1 ^2Jake Roach, Tom's Hardware. "Hot Chips 2026: Nvidia breaks down 88-core Vera CPU: spatial multithreading benchmarked, 1.2 TB/s SOCAMM2 memory, agentic workloads detailed, and more." August 25, 2026. tomshardware.com/...ic-workloads-detailed-and-more
  29. ^Rohil Bhargava, Taylor Allison and Harry Petty, NVIDIA Technical Blog. "NVIDIA Vera Rubin POD: Seven Chips, Five Rack-Scale Systems, One AI Supercomputer." March 16, 2026. developer.nvidia.com/...stems-one-ai-supercomputer
  30. ^ServeTheHome. "Diving Deeper on NVIDIA's Vera CPU: New Architectural Details and SPEC CPU 2026 Benchmarks." July 21, 2026. servethehome.com/...s-and-spec-cpu-2026-benchmarks

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

7 revisions · v8 · 4,905 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent verification 28 Sep 2026 (xg11 V5): specs, launch chronology, benchmarks and customers checked against NVIDIA releases/blogs, Phoronix, Tom's Hardware; 6 minor defects fixed

Cite this page: AI Wiki. "NVIDIA Vera (CPU)." aiwiki.ai, updated 28 Sept 2026, fact-checked 28 Sept 2026. CC BY 4.0. https://aiwiki.ai/wiki/nvidia_vera_cpu

Suggest edit