AWS Graviton
AWS Graviton is a family of Arm-based server processors designed by Amazon Web Services for use in its own cloud computing fleet. The first generation launched with the Amazon EC2 A1 instance in November 2018 [3][4], and five generations had shipped by mid-2026. Customers reach the chips through EC2 instance types and through managed AWS services that run on them, such as Amazon SageMaker inference endpoints [23].
The processors come from Annapurna Labs, the Israeli chip design company Amazon acquired in January 2015 [29]. The same group produces the Nitro cards that offload networking, storage, and virtualization from EC2 servers, along with the AWS Trainium and AWS Inferentia machine learning accelerators [29]. That combination is why Graviton appears in AI infrastructure discussions at all: AWS designs the general-purpose CPU, the accelerator, and the offload silicon as one portfolio, and it markets Graviton as the part that runs the code around the model rather than the model itself.
AWS's own framing shifted noticeably with the fifth generation. Announcing Graviton5 at re:Invent 2025, the company tied it directly to agentic AI: "As AI shifts from answering questions to taking actions, running code, using tools, evaluating results, and orchestrating multi-step tasks, the demand for CPU compute is growing rapidly" [13]. In April 2026 Meta agreed to deploy tens of millions of Graviton5 cores for exactly that class of workload [27][28].
Origins: Annapurna Labs and Nitro
Annapurna Labs was founded in Israel in 2011, with Nafea Bshara among its co-founders, and Amazon acquired it in January 2015 [29]. Bshara remained at Amazon and was still speaking for the silicon organization a decade later, as vice president and distinguished engineer, when the Meta agreement was announced [27].
The group's first major contribution was not a CPU but the Nitro System. AWS describes Nitro as "a combination of purpose-built server designs, data processors, system management components, and specialized firmware which provide the underlying platform for all Amazon EC2 instances launched since the beginning of 2018" [17]. It has three parts: Nitro Cards that handle input/output virtualization on separate hardware, a Nitro Security Chip that provides a hardware root of trust and secure boot, and the Nitro Hypervisor, "a deliberately minimized and firmware-like hypervisor designed to provide strong resource isolation, and performance that is nearly indistinguishable from a bare metal server" [17]. Nitro moves input/output virtualization onto dedicated hardware "independent of the main system board with its CPUs and memory" [17], so the host processor carries less of the virtualization work than in a conventional server design. AWS also uses Nitro to offer Nitro Enclaves, isolated compute environments inside an EC2 instance for handling sensitive data [18].
Generations
First generation and the A1 instance
AWS announced the original Graviton and the A1 instance family on 26 November 2018, describing the chip as built from "64-bit Arm Neoverse cores and custom silicon designed by AWS" [4]. A1 shipped in five sizes with 1, 2, 4, 8, and 16 vCPUs, priced from $0.0255 per hour for a1.medium to $0.4080 per hour for a1.4xlarge in the US East (N. Virginia) region [3]. AWS aimed it at scale-out work that was already portable: web servers, containerized microservices, caching fleets, distributed data stores, and development environments [4]. AWS did not publish core-level details at launch; secondary documentation lists the first-generation part as a 16-core Cortex-A72 design running at 2.3 GHz [30].
Graviton2
AWS announced Graviton2 on 3 December 2019, built from 64-bit Arm Neoverse cores on a 7 nm process with "double-sized per-core caches" [5]. AWS said the new instances could deliver up to seven times the performance of A1, including twice the floating-point performance, and published per-vCPU comparisons against the x86-based M5 generation: 24 percent better on HTTPS load balancing with Nginx, 43 percent on Memcached, 26 percent on x264 video encoding, and 54 percent on EDA simulation [5]. The general-purpose M6g instances were advertised as delivering "up to 40% better price performance over M5 instances" [6].
Graviton2 uses Arm's Neoverse N1 core at 2,500 MHz with eight DDR4 channels and a 32 MB shared last-level cache [2]. It went on to power an unusually wide instance range, including T4g burstable instances, X2gd memory-optimized instances, I4g and Im4gn storage instances, and the G5g instances that pair Graviton2 with NVIDIA T4G Tensor Core GPUs for Android game streaming and low-cost machine learning inference [2][26].
Graviton3 and Graviton3E
Graviton3 became generally available with the C7g instance on 23 May 2022 [7]. AWS claimed up to 25 percent higher compute performance, up to twice the floating-point performance, up to twice the cryptographic performance, and 50 percent faster memory access, the last of these coming from DDR5 [7][8]. C7g was described as the first instance type in the cloud to use DDR5 memory [8]. AWS also stated that Graviton3 "uses up to 60 percent less energy for the same performance as comparable EC2 instances" [7].
For AI work, Graviton3 is the pivotal generation. It moved to Arm's Neoverse V1 core with the Scalable Vector Extension (two 256-bit SVE units), and it added bfloat16 and int8 matrix multiplication (MMLA) instructions [2]. AWS claimed "up to 3x better performance compared to Graviton2 processors for ML workloads" [8]. The generation also introduced always-on memory encryption and pointer authentication [7][8].
Graviton3E is a variant tuned for vector-heavy work. AWS says it delivers up to 35 percent higher vector instruction processing performance than Graviton3 and uses it in the Hpc7g instances, which offer 16, 32, or 64 cores with 128 GiB of memory and 200 Gbps of Elastic Fabric Adapter bandwidth, and in the network-optimized C7gn instances [2][9].
Graviton4
AWS previewed Graviton4 at re:Invent on 28 November 2023. Jeff Barr's announcement gave the configuration directly: "96 Neoverse V2 cores, 2 MB of L2 cache per core, and 12 DDR5-5600 channels", plus encrypted high-speed hardware interfaces and Branch Target Identification [10]. Amazon's accompanying press release claimed "up to 30% better compute performance, 50% more cores, and 75% more memory bandwidth than current generation Graviton3 processors" [11].
The memory-optimized R8g instances reached general availability on 9 July 2024 in US East (N. Virginia), US East (Ohio), US West (Oregon), and Europe (Frankfurt) [12]. AWS quoted 30 percent faster performance for web applications, 40 percent for databases, and 45 percent for large Java applications relative to Graviton3, with up to three times the vCPUs (to 48xlarge) and up to 1.5 TB of memory compared with R7g, across 12 sizes including two bare metal options [12]. The 48xlarge size is a two-socket configuration, giving 192 vCPUs from two 96-core packages [2]. Graviton4 also brought the Armv9.0-A architecture revision and SVE2 [2].
Graviton5
AWS introduced Graviton5 at re:Invent 2025, held from 30 November to 4 December, describing it as "the company's most powerful and efficient CPU" [16]. The chip has 192 cores, five times the L3 cache of Graviton4, DDR5-8800 memory, and PCIe Gen 6 support, and AWS calls it the first CPU in its fleet to support PCIe Gen 6 and DDR5-8800 memory [13]. AWS states that Graviton5 "adopts the latest 3nm technology, optimizes the design for AWS use cases, and allows for system-level optimizations such as bare-die cooling" [14], and reports up to 33 percent lower inter-core latency than the previous generation [1].
The M9g and M9gd instances became generally available on 10 June 2026 in the same four launch regions AWS used for R8g [13]. AWS claims up to 25 percent better compute performance than Graviton4-based instances, and specifically up to 35 percent faster web application performance, up to 35 percent faster machine learning inference, and up to 30 percent faster database performance [13]. M9gd, the NVMe SSD variant, is credited with 30 percent higher IOPS than M8gd [13]. M9g scales from m9g.medium at 1 vCPU and 4 GiB to m9g.48xlarge at 192 vCPUs and 768 GiB, with up to 100 Gbps of network bandwidth [13][15]. Customer figures published by AWS include Honeycomb reporting 36 percent better throughput per core versus Graviton4 across a six-month A/B test [13], and Airbnb reporting up to 25 percent over other system architectures of the same generation and up to 20 percent over Graviton4 [14].
Generation summary
| Generation | Availability | Core | Cores | Clock | Arm ISA | Memory | LLC | Representative instances |
|---|---|---|---|---|---|---|---|---|
| Graviton | Nov 2018 | Cortex-A72 (per published summaries) | 16 | 2.3 GHz | Armv8-A | not published | not published | A1 |
| Graviton2 | Dec 2019 | Neoverse N1 | 64 | 2,500 MHz | Armv8.2-A | DDR4, 8 channels | 32 MB | M6g, C6g, R6g, T4g, X2gd, G5g, I4g |
| Graviton3 / 3E | May 2022 | Neoverse V1 | 64 | 2,600 MHz | Armv8.4-A | DDR5, 8 channels | 32 MB | C7g, M7g, R7g, C7gn, Hpc7g |
| Graviton4 | Jul 2024 | Neoverse V2 | 96 per socket | 2,800 MHz | Armv9.0-A | DDR5-5600, 12 channels | 36 MB | C8g, M8g, R8g, X8g, I8g |
| Graviton5 | Jun 2026 | Neoverse V3 | 192 | 3,300 MHz | Armv9.2-A | DDR5-8800, 12 channels | 48 MB per NUMA domain | M9g, M9gd |
Sources: AWS Graviton Technical Guide for core, clock, ISA, cache, and channel counts [2]; the announcement posts cited above for availability dates; published family summaries for the first generation [30].
The vector and machine learning instruction support tracks the same progression [2]:
| Generation | SIMD units | ML-relevant instructions |
|---|---|---|
| Graviton2 | 2 x 128-bit Neon | fp16, dot product |
| Graviton3 / 3E | 4 x 128-bit Neon or 2 x 256-bit SVE | bfloat16, int8 MMLA |
| Graviton4 | 4 x 128-bit Neon/SVE, SVE2 | SVE bfloat16, SVE int8 |
| Graviton5 | 4 x 128-bit Neon/SVE, SVE2 | SVE bfloat16, SVE int8 |
Price, performance, and energy claims
Every headline number attached to Graviton originates with AWS, and the framing has changed over time. The long-running claim is that Graviton instances "deliver up to 40% better price performance over comparable x86-based instances" [19]. The current Graviton product page instead says the instances "cost up to 20% less than comparable x86-based Amazon EC2 instances" and use "up to 60% less energy than comparable EC2 instances" [1]. These are different measures, and neither is an independently audited benchmark. AWS also publishes workload-specific customer figures, such as 20 to 80 percent higher performance on Twitter timelines moving from Graviton2-based C6g to Graviton3-based C7g [19].
The energy figure matters most to data center planning, where power draw per rack sets how much compute fits in a building. AWS first attached the 60 percent number to Graviton3 in 2022 and still uses it for the family [1][7].
Machine learning and AI workloads
CPU inference
Graviton3 and later generations are the target of a sustained AWS effort to make CPU inference competitive for small and mid-sized models. The work runs through the Arm Compute Library and oneDNN, which supply Neon and SVE GEMM kernels to PyTorch [20].
| Stack | Reported result | Baseline | Date |
|---|---|---|---|
| PyTorch 2.0 on Graviton3 (C7g) | ResNet50 up to 3.5x faster, BERT up to 1.4x faster; up to 50% cost saving for inference | previous PyTorch release; comparable x86 instances | May 2023 [20] |
| Amazon SageMaker on Graviton3 | up to 50% cost saving for PyTorch, TensorFlow, XGBoost, and scikit-learn inference; latency up to 50% better | comparable x86 SageMaker instances | May 2023 [23] |
| ONNX Runtime 1.17.0 on Graviton3 | fp32 throughput up to 65% higher; int8 throughput up to 30% higher (BERT, RoBERTa, GPT2) | earlier ONNX Runtime on Graviton3 | May 2024 [22] |
| torch.compile on Graviton3 | up to 2x for Hugging Face model inference; up to 1.35x for TorchBench | PyTorch eager mode | July 2024 [21] |
Two implementation details recur across this documentation. First, bfloat16 fast math is not on by default: PyTorch users are told to set DNNL_DEFAULT_FPMATH_MODE=BF16, and ONNX Runtime requires explicit session options, while int8 kernels are enabled automatically on Graviton3 [22][25]. Second, AWS strongly recommends AMIs based on Linux kernel 5.10 or later for the best inference performance on Graviton, so an out-of-date AMI can quietly give up much of the machine learning gain [24][25].
For large language model work, AWS documents llama.cpp as the reference path, recommending Graviton3(E) (C7g, M7g, R7g, C7gn, Hpc7g), Graviton4 (R8g), and Graviton5 (M9g) instances, and walking through Llama 3 8B in the Q4_0 quantization format from Hugging Face [24]. TensorFlow, ONNX, XGBoost, and scikit-learn are covered by the same optimization program [22][23].
Managed AI services inherit this. Amazon SageMaker exposes Graviton-backed inference instance types such as ml.c7g.xlarge, and AWS's published benchmarks put Graviton3-based c7g up to 50 percent ahead of c5 and c6i on inference latency [23].
Work around the accelerator
AWS positions Graviton as complementary to its AI accelerator line rather than as a substitute for it. Training and high-volume accelerated serving run on Trainium and Inferentia, while Graviton carries the CPU-side work around them. AWS attributes rising CPU demand to systems that take actions, run code, use tools, evaluate results, and orchestrate multi-step tasks [13], and describes Graviton5 as built for processors that "must handle large numbers of concurrent environments and keep accelerators moving" [14].
One instance family puts both in the same box: G5g pairs Graviton2 with NVIDIA T4G Tensor Core GPUs, which AWS markets for cost-effective machine learning inference and Android game streaming [26].
Agentic AI and the Meta agreement
The clearest evidence for the agentic framing is commercial. On 24 April 2026 Meta announced an agreement with AWS to bring tens of millions of Graviton5 cores into its compute portfolio, citing the CPU-intensive character of agent workloads [27][28]. Meta's head of infrastructure, Santosh Janardhan, said that "diversifying our compute sources is a strategic imperative" and that Graviton lets Meta "run the CPU-intensive workloads behind agentic AI with the performance and efficiency we need at our scale" [27]. Meta framed it as a portfolio decision, on the reasoning that no single chip architecture serves every workload efficiently [28]. AWS also names Uber and Snowflake as deploying Graviton for their respective agentic workloads [14]. See Meta's AWS Graviton agreement for detail.
Adoption
AWS has released adoption figures at intervals. At re:Invent 2023 the company said it had built more than two million Graviton processors, that more than 50,000 customers used Graviton-based instances, that every one of the top 100 EC2 customers used Graviton, and that more than 150 Graviton-powered instance types were available globally [10][11]. By the Graviton5 general availability announcement in June 2026, AWS put the figure at more than 120,000 customers [1][14].
Graviton is not the only Arm server line aimed at hyperscaler fleets: Ampere Computing sells merchant Arm processors, and NVIDIA Grace pairs Arm cores with NVIDIA GPUs. What is specific to Graviton is that AWS tunes it for a single fleet, which the company cites as the reason it can make design choices such as the bare-die cooling used on Graviton5 [14]. Because Graviton builds on Arm's Neoverse cores rather than a fully custom core design, each generation tracks the corresponding Neoverse release: N1, V1, V2, and V3 across the second through fifth generations [2].
Limitations and caveats
Software has to be built for aarch64. AWS maintains a substantial porting and tuning guide precisely because moving a fleet is not free: compilers, language runtimes, container images, and third-party dependencies all need Arm builds, and AWS's own advice is that "later versions of compilers, language runtimes, and applications should be used whenever possible" [2]. Closed-source components with no Arm build are the hard case, because a customer cannot recompile them.
For AI specifically, the machine learning gains depend on configuration that is not default. Bfloat16 fast math must be turned on explicitly in PyTorch and ONNX Runtime, and AWS advises recent kernel and framework versions [22][24][25]. A team that lifts a container onto a Graviton instance without changing anything will see the general-purpose improvement but not the inference improvement.
The published comparisons are vendor comparisons. AWS chooses the baselines, the instance sizes, and the benchmarks, and the family carries several different headline numbers across AWS pages: 40 percent better price performance against x86 [19], 20 percent lower cost against x86 [1], and 25 to 30 percent better compute performance against the previous Graviton generation [12][13]. Customer-reported figures such as Honeycomb's 36 percent per-core throughput gain [13] are also selected by AWS for publication.
Finally, Graviton is a general-purpose CPU. It has SVE2, bfloat16, and int8 matrix instructions, but no dedicated matrix engine on the scale of a GPU or of Trainium, so it does not compete for large-model training or high-throughput generative serving. AWS does not claim otherwise; its own positioning puts Graviton around the accelerator, not in place of it [14].
See also
References
- ^AWS, "AWS Graviton Processor" product page. aws.amazon.com/...graviton
- ^AWS, "AWS Graviton Technical Guide" (aws/aws-graviton-getting-started, GitHub). github.com/...aws-graviton-getting-started
- ^Jeff Barr, "New: EC2 Instances (A1) Powered by Arm-Based AWS Graviton Processors," AWS News Blog, 26 November 2018. aws.amazon.com/...rm-based-aws-graviton-processors
- ^AWS, "Introducing Amazon EC2 A1 Instances," What's New with AWS, 26 November 2018. aws.amazon.com/...roducing-amazon-ec2-a1-instances
- ^Jeff Barr, "Coming Soon: Graviton2-Powered General Purpose, Compute-Optimized, & Memory-Optimized EC2 Instances," AWS News Blog, 3 December 2019. aws.amazon.com/...d-memory-optimized-ec2-instances
- ^AWS, "Amazon EC2 M6g Instances." aws.amazon.com/...m6g
- ^AWS News Blog, "New Amazon EC2 C7g Instances, Powered by AWS Graviton3 Processors," 23 May 2022. aws.amazon.com/...ered-by-aws-graviton3-processors
- ^AWS, "Amazon EC2 C7g Instances." aws.amazon.com/...c7g
- ^AWS, "Amazon EC2 Hpc7g Instances." aws.amazon.com/...hpc7g
- ^Jeff Barr, "Join the preview for new memory-optimized, AWS Graviton4-powered Amazon EC2 instances (R8g)," AWS News Blog, 28 November 2023. aws.amazon.com/...powered-amazon-ec2-instances-r8g
- ^Amazon, "AWS Unveils Next Generation AWS-Designed Chips," press release, 28 November 2023. press.aboutamazon.com/...ration-aws-designed-chips
- ^AWS, "Amazon EC2 R8g instances powered by AWS Graviton4 now generally available," What's New with AWS, 9 July 2024. aws.amazon.com/...ws-graviton4-generally-available
- ^Esra Kayabali, "Now available: Amazon EC2 M9g and M9gd instances powered by new AWS Graviton5 processors," AWS News Blog, 10 June 2026. aws.amazon.com/...-by-new-aws-graviton5-processors
- ^About Amazon, "AWS Graviton5 is now generally available, delivering purpose-built performance for the agentic AI era," originally published 4 December 2025, updated 10 June 2026. aboutamazon.com/...aws-graviton-5-cpu-amazon-ec2
- ^AWS, "Amazon EC2 M9g Instances." aws.amazon.com/...m9g
- ^AWS News Blog, "Top announcements of AWS re:Invent 2025." aws.amazon.com/...nouncements-of-aws-reinvent-2025
- ^AWS, "The Security Design of the AWS Nitro System," whitepaper, 15 February 2024. docs.aws.amazon.com/...-design-of-aws-nitro-system
- ^AWS, "AWS Nitro System." aws.amazon.com/...nitro
- ^AWS, "AWS Silicon Innovation." aws.amazon.com/silicon-innovation
- ^AWS Machine Learning Blog, "Optimized PyTorch 2.0 inference with AWS Graviton processors," 3 May 2023. aws.amazon.com/...nce-with-aws-graviton-processors
- ^AWS Machine Learning Blog, "Accelerated PyTorch inference with torch.compile on AWS Graviton processors," 2 July 2024. aws.amazon.com/...mpile-on-aws-graviton-processors
- ^AWS Machine Learning Blog, "Accelerate NLP inference with ONNX Runtime on AWS Graviton processors," 15 May 2024. aws.amazon.com/...ntime-on-aws-graviton-processors
- ^AWS Machine Learning Blog, "Reduce Amazon SageMaker inference cost with AWS Graviton," 10 May 2023. aws.amazon.com/...inference-cost-with-aws-graviton
- ^AWS Graviton Technical Guide, "LLM inference on Graviton CPUs with llama.cpp." github.com/...llama.cpp.md
- ^AWS Graviton Technical Guide, "PyTorch inference on Graviton." github.com/...pytorch.md
- ^AWS, "Amazon EC2 G5g Instances." aws.amazon.com/...g5g
- ^About Amazon, "Meta expands Amazon partnership with AWS Graviton chips for AI," April 2026. aboutamazon.com/...meta-aws-graviton-ai-partnership
- ^Meta Newsroom, "Meta Partners With AWS on Graviton Chips to Power Agentic AI," 24 April 2026. about.fb.com/...graviton-chips-to-power-agentic-ai
- ^Wikipedia, "Annapurna Labs." en.wikipedia.org/...Annapurna_Labs
- ^Wikipedia, "AWS Graviton." en.wikipedia.org/...AWS_Graviton
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
v1 · 3,251 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Independent adversarial fact-check at creation (wanted175 campaign, 2026-07-24): every claim verified against primary sources by a dedicated verification agent; corrections applied before publication.
Cite this page: AI Wiki. "AWS Graviton." aiwiki.ai, updated 24 Jul 2026, fact-checked 24 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/aws_graviton