AWS Graviton

RawGraph

AWS Graviton is a family of Arm-based server processors designed by Amazon Web Services for use in its own cloud computing fleet. The first generation launched with the Amazon EC2 A1 instance in November 2018 [3][4], and five generations had shipped by mid-2026. Customers reach the chips through EC2 instance types and through managed AWS services that run on them, such as Amazon SageMaker inference endpoints [23].

The processors come from Annapurna Labs, the Israeli chip design company Amazon acquired in January 2015 [29]. The same group produces the Nitro cards that offload networking, storage, and virtualization from EC2 servers, along with the AWS Trainium and AWS Inferentia machine learning accelerators [29]. That combination is why Graviton appears in AI infrastructure discussions at all: AWS designs the general-purpose CPU, the accelerator, and the offload silicon as one portfolio, and it markets Graviton as the part that runs the code around the model rather than the model itself.

AWS's own framing shifted noticeably with the fifth generation. Announcing Graviton5 at re:Invent 2025, the company tied it directly to agentic AI: "As AI shifts from answering questions to taking actions, running code, using tools, evaluating results, and orchestrating multi-step tasks, the demand for CPU compute is growing rapidly" [13]. In April 2026 Meta agreed to deploy tens of millions of Graviton5 cores for exactly that class of workload [27][28].

Origins: Annapurna Labs and Nitro

Annapurna Labs was founded in Israel in 2011, with Nafea Bshara among its co-founders, and Amazon acquired it in January 2015 [29]. Bshara remained at Amazon and was still speaking for the silicon organization a decade later, as vice president and distinguished engineer, when the Meta agreement was announced [27].

The group's first major contribution was not a CPU but the Nitro System. AWS describes Nitro as "a combination of purpose-built server designs, data processors, system management components, and specialized firmware which provide the underlying platform for all Amazon EC2 instances launched since the beginning of 2018" [17]. It has three parts: Nitro Cards that handle input/output virtualization on separate hardware, a Nitro Security Chip that provides a hardware root of trust and secure boot, and the Nitro Hypervisor, "a deliberately minimized and firmware-like hypervisor designed to provide strong resource isolation, and performance that is nearly indistinguishable from a bare metal server" [17]. Nitro moves input/output virtualization onto dedicated hardware "independent of the main system board with its CPUs and memory" [17], so the host processor carries less of the virtualization work than in a conventional server design. AWS also uses Nitro to offer Nitro Enclaves, isolated compute environments inside an EC2 instance for handling sensitive data [18].

Generations

First generation and the A1 instance

AWS announced the original Graviton and the A1 instance family on 26 November 2018, describing the chip as built from "64-bit Arm Neoverse cores and custom silicon designed by AWS" [4]. A1 shipped in five sizes with 1, 2, 4, 8, and 16 vCPUs, priced from $0.0255 per hour for a1.medium to $0.4080 per hour for a1.4xlarge in the US East (N. Virginia) region [3]. AWS aimed it at scale-out work that was already portable: web servers, containerized microservices, caching fleets, distributed data stores, and development environments [4]. AWS did not publish core-level details at launch; secondary documentation lists the first-generation part as a 16-core Cortex-A72 design running at 2.3 GHz [30].

Graviton2

AWS announced Graviton2 on 3 December 2019, built from 64-bit Arm Neoverse cores on a 7 nm process with "double-sized per-core caches" [5]. AWS said the new instances could deliver up to seven times the performance of A1, including twice the floating-point performance, and published per-vCPU comparisons against the x86-based M5 generation: 24 percent better on HTTPS load balancing with Nginx, 43 percent on Memcached, 26 percent on x264 video encoding, and 54 percent on EDA simulation [5]. The general-purpose M6g instances were advertised as delivering "up to 40% better price performance over M5 instances" [6].

Graviton2 uses Arm's Neoverse N1 core at 2,500 MHz with eight DDR4 channels and a 32 MB shared last-level cache [2]. It went on to power an unusually wide instance range, including T4g burstable instances, X2gd memory-optimized instances, I4g and Im4gn storage instances, and the G5g instances that pair Graviton2 with NVIDIA T4G Tensor Core GPUs for Android game streaming and low-cost machine learning inference [2][26].

Graviton3 and Graviton3E

Graviton3 became generally available with the C7g instance on 23 May 2022 [7]. AWS claimed up to 25 percent higher compute performance, up to twice the floating-point performance, up to twice the cryptographic performance, and 50 percent faster memory access, the last of these coming from DDR5 [7][8]. C7g was described as the first instance type in the cloud to use DDR5 memory [8]. AWS also stated that Graviton3 "uses up to 60 percent less energy for the same performance as comparable EC2 instances" [7].

For AI work, Graviton3 is the pivotal generation. It moved to Arm's Neoverse V1 core with the Scalable Vector Extension (two 256-bit SVE units), and it added bfloat16 and int8 matrix multiplication (MMLA) instructions [2]. AWS claimed "up to 3x better performance compared to Graviton2 processors for ML workloads" [8]. The generation also introduced always-on memory encryption and pointer authentication [7][8].

Graviton3E is a variant tuned for vector-heavy work. AWS says it delivers up to 35 percent higher vector instruction processing performance than Graviton3 and uses it in the Hpc7g instances, which offer 16, 32, or 64 cores with 128 GiB of memory and 200 Gbps of Elastic Fabric Adapter bandwidth, and in the network-optimized C7gn instances [2][9].

Graviton4

AWS previewed Graviton4 at re:Invent on 28 November 2023. Jeff Barr's announcement gave the configuration directly: "96 Neoverse V2 cores, 2 MB of L2 cache per core, and 12 DDR5-5600 channels", plus encrypted high-speed hardware interfaces and Branch Target Identification [10]. Amazon's accompanying press release claimed "up to 30% better compute performance, 50% more cores, and 75% more memory bandwidth than current generation Graviton3 processors" [11].

The memory-optimized R8g instances reached general availability on 9 July 2024 in US East (N. Virginia), US East (Ohio), US West (Oregon), and Europe (Frankfurt) [12]. AWS quoted 30 percent faster performance for web applications, 40 percent for databases, and 45 percent for large Java applications relative to Graviton3, with up to three times the vCPUs (to 48xlarge) and up to 1.5 TB of memory compared with R7g, across 12 sizes including two bare metal options [12]. The 48xlarge size is a two-socket configuration, giving 192 vCPUs from two 96-core packages [2]. Graviton4 also brought the Armv9.0-A architecture revision and SVE2 [2].

Graviton5

AWS introduced Graviton5 at re:Invent 2025, held from 30 November to 4 December, describing it as "the company's most powerful and efficient CPU" [16]. The chip has 192 cores, five times the L3 cache of Graviton4, DDR5-8800 memory, and PCIe Gen 6 support, and AWS calls it the first CPU in its fleet to support PCIe Gen 6 and DDR5-8800 memory [13]. AWS states that Graviton5 "adopts the latest 3nm technology, optimizes the design for AWS use cases, and allows for system-level optimizations such as bare-die cooling" [14], and reports up to 33 percent lower inter-core latency than the previous generation [1].

The M9g and M9gd instances became generally available on 10 June 2026 in the same four launch regions AWS used for R8g [13]. AWS claims up to 25 percent better compute performance than Graviton4-based instances, and specifically up to 35 percent faster web application performance, up to 35 percent faster machine learning inference, and up to 30 percent faster database performance [13]. M9gd, the NVMe SSD variant, is credited with 30 percent higher IOPS than M8gd [13]. M9g scales from m9g.medium at 1 vCPU and 4 GiB to m9g.48xlarge at 192 vCPUs and 768 GiB, with up to 100 Gbps of network bandwidth [13][15]. Customer figures published by AWS include Honeycomb reporting 36 percent better throughput per core versus Graviton4 across a six-month A/B test [13], and Airbnb reporting up to 25 percent over other system architectures of the same generation and up to 20 percent over Graviton4 [14].

Generation summary

GenerationAvailabilityCoreCoresClockArm ISAMemoryLLCRepresentative instances
GravitonNov 2018Cortex-A72 (per published summaries)162.3 GHzArmv8-Anot publishednot publishedA1
Graviton2Dec 2019Neoverse N1642,500 MHzArmv8.2-ADDR4, 8 channels32 MBM6g, C6g, R6g, T4g, X2gd, G5g, I4g
Graviton3 / 3EMay 2022Neoverse V1642,600 MHzArmv8.4-ADDR5, 8 channels32 MBC7g, M7g, R7g, C7gn, Hpc7g
Graviton4Jul 2024Neoverse V296 per socket2,800 MHzArmv9.0-ADDR5-5600, 12 channels36 MBC8g, M8g, R8g, X8g, I8g
Graviton5Jun 2026Neoverse V31923,300 MHzArmv9.2-ADDR5-8800, 12 channels48 MB per NUMA domainM9g, M9gd

Sources: AWS Graviton Technical Guide for core, clock, ISA, cache, and channel counts [2]; the announcement posts cited above for availability dates; published family summaries for the first generation [30].

The vector and machine learning instruction support tracks the same progression [2]:

GenerationSIMD unitsML-relevant instructions
Graviton22 x 128-bit Neonfp16, dot product
Graviton3 / 3E4 x 128-bit Neon or 2 x 256-bit SVEbfloat16, int8 MMLA
Graviton44 x 128-bit Neon/SVE, SVE2SVE bfloat16, SVE int8
Graviton54 x 128-bit Neon/SVE, SVE2SVE bfloat16, SVE int8

Price, performance, and energy claims

Every headline number attached to Graviton originates with AWS, and the framing has changed over time. The long-running claim is that Graviton instances "deliver up to 40% better price performance over comparable x86-based instances" [19]. The current Graviton product page instead says the instances "cost up to 20% less than comparable x86-based Amazon EC2 instances" and use "up to 60% less energy than comparable EC2 instances" [1]. These are different measures, and neither is an independently audited benchmark. AWS also publishes workload-specific customer figures, such as 20 to 80 percent higher performance on Twitter timelines moving from Graviton2-based C6g to Graviton3-based C7g [19].

The energy figure matters most to data center planning, where power draw per rack sets how much compute fits in a building. AWS first attached the 60 percent number to Graviton3 in 2022 and still uses it for the family [1][7].

Machine learning and AI workloads

CPU inference

Graviton3 and later generations are the target of a sustained AWS effort to make CPU inference competitive for small and mid-sized models. The work runs through the Arm Compute Library and oneDNN, which supply Neon and SVE GEMM kernels to PyTorch [20].

StackReported resultBaselineDate
PyTorch 2.0 on Graviton3 (C7g)ResNet50 up to 3.5x faster, BERT up to 1.4x faster; up to 50% cost saving for inferenceprevious PyTorch release; comparable x86 instancesMay 2023 [20]
Amazon SageMaker on Graviton3up to 50% cost saving for PyTorch, TensorFlow, XGBoost, and scikit-learn inference; latency up to 50% bettercomparable x86 SageMaker instancesMay 2023 [23]
ONNX Runtime 1.17.0 on Graviton3fp32 throughput up to 65% higher; int8 throughput up to 30% higher (BERT, RoBERTa, GPT2)earlier ONNX Runtime on Graviton3May 2024 [22]
torch.compile on Graviton3up to 2x for Hugging Face model inference; up to 1.35x for TorchBenchPyTorch eager modeJuly 2024 [21]

Two implementation details recur across this documentation. First, bfloat16 fast math is not on by default: PyTorch users are told to set DNNL_DEFAULT_FPMATH_MODE=BF16, and ONNX Runtime requires explicit session options, while int8 kernels are enabled automatically on Graviton3 [22][25]. Second, AWS strongly recommends AMIs based on Linux kernel 5.10 or later for the best inference performance on Graviton, so an out-of-date AMI can quietly give up much of the machine learning gain [24][25].

For large language model work, AWS documents llama.cpp as the reference path, recommending Graviton3(E) (C7g, M7g, R7g, C7gn, Hpc7g), Graviton4 (R8g), and Graviton5 (M9g) instances, and walking through Llama 3 8B in the Q4_0 quantization format from Hugging Face [24]. TensorFlow, ONNX, XGBoost, and scikit-learn are covered by the same optimization program [22][23].

Managed AI services inherit this. Amazon SageMaker exposes Graviton-backed inference instance types such as ml.c7g.xlarge, and AWS's published benchmarks put Graviton3-based c7g up to 50 percent ahead of c5 and c6i on inference latency [23].

Work around the accelerator

AWS positions Graviton as complementary to its AI accelerator line rather than as a substitute for it. Training and high-volume accelerated serving run on Trainium and Inferentia, while Graviton carries the CPU-side work around them. AWS attributes rising CPU demand to systems that take actions, run code, use tools, evaluate results, and orchestrate multi-step tasks [13], and describes Graviton5 as built for processors that "must handle large numbers of concurrent environments and keep accelerators moving" [14].

One instance family puts both in the same box: G5g pairs Graviton2 with NVIDIA T4G Tensor Core GPUs, which AWS markets for cost-effective machine learning inference and Android game streaming [26].

Agentic AI and the Meta agreement

The clearest evidence for the agentic framing is commercial. On 24 April 2026 Meta announced an agreement with AWS to bring tens of millions of Graviton5 cores into its compute portfolio, citing the CPU-intensive character of agent workloads [27][28]. Meta's head of infrastructure, Santosh Janardhan, said that "diversifying our compute sources is a strategic imperative" and that Graviton lets Meta "run the CPU-intensive workloads behind agentic AI with the performance and efficiency we need at our scale" [27]. Meta framed it as a portfolio decision, on the reasoning that no single chip architecture serves every workload efficiently [28]. AWS also names Uber and Snowflake as deploying Graviton for their respective agentic workloads [14]. See Meta's AWS Graviton agreement for detail.

Adoption

AWS has released adoption figures at intervals. At re:Invent 2023 the company said it had built more than two million Graviton processors, that more than 50,000 customers used Graviton-based instances, that every one of the top 100 EC2 customers used Graviton, and that more than 150 Graviton-powered instance types were available globally [10][11]. By the Graviton5 general availability announcement in June 2026, AWS put the figure at more than 120,000 customers [1][14].

Graviton is not the only Arm server line aimed at hyperscaler fleets: Ampere Computing sells merchant Arm processors, and NVIDIA Grace pairs Arm cores with NVIDIA GPUs. What is specific to Graviton is that AWS tunes it for a single fleet, which the company cites as the reason it can make design choices such as the bare-die cooling used on Graviton5 [14]. Because Graviton builds on Arm's Neoverse cores rather than a fully custom core design, each generation tracks the corresponding Neoverse release: N1, V1, V2, and V3 across the second through fifth generations [2].

Limitations and caveats

Software has to be built for aarch64. AWS maintains a substantial porting and tuning guide precisely because moving a fleet is not free: compilers, language runtimes, container images, and third-party dependencies all need Arm builds, and AWS's own advice is that "later versions of compilers, language runtimes, and applications should be used whenever possible" [2]. Closed-source components with no Arm build are the hard case, because a customer cannot recompile them.

For AI specifically, the machine learning gains depend on configuration that is not default. Bfloat16 fast math must be turned on explicitly in PyTorch and ONNX Runtime, and AWS advises recent kernel and framework versions [22][24][25]. A team that lifts a container onto a Graviton instance without changing anything will see the general-purpose improvement but not the inference improvement.

The published comparisons are vendor comparisons. AWS chooses the baselines, the instance sizes, and the benchmarks, and the family carries several different headline numbers across AWS pages: 40 percent better price performance against x86 [19], 20 percent lower cost against x86 [1], and 25 to 30 percent better compute performance against the previous Graviton generation [12][13]. Customer-reported figures such as Honeycomb's 36 percent per-core throughput gain [13] are also selected by AWS for publication.

Finally, Graviton is a general-purpose CPU. It has SVE2, bfloat16, and int8 matrix instructions, but no dedicated matrix engine on the scale of a GPU or of Trainium, so it does not compete for large-model training or high-throughput generative serving. AWS does not claim otherwise; its own positioning puts Graviton around the accelerator, not in place of it [14].

See also

References

  1. ^AWS, "AWS Graviton Processor" product page. aws.amazon.com/...graviton
  2. ^AWS, "AWS Graviton Technical Guide" (aws/aws-graviton-getting-started, GitHub). github.com/...aws-graviton-getting-started
  3. ^Jeff Barr, "New: EC2 Instances (A1) Powered by Arm-Based AWS Graviton Processors," AWS News Blog, 26 November 2018. aws.amazon.com/...rm-based-aws-graviton-processors
  4. ^AWS, "Introducing Amazon EC2 A1 Instances," What's New with AWS, 26 November 2018. aws.amazon.com/...roducing-amazon-ec2-a1-instances
  5. ^Jeff Barr, "Coming Soon: Graviton2-Powered General Purpose, Compute-Optimized, & Memory-Optimized EC2 Instances," AWS News Blog, 3 December 2019. aws.amazon.com/...d-memory-optimized-ec2-instances
  6. ^AWS, "Amazon EC2 M6g Instances." aws.amazon.com/...m6g
  7. ^AWS News Blog, "New Amazon EC2 C7g Instances, Powered by AWS Graviton3 Processors," 23 May 2022. aws.amazon.com/...ered-by-aws-graviton3-processors
  8. ^AWS, "Amazon EC2 C7g Instances." aws.amazon.com/...c7g
  9. ^AWS, "Amazon EC2 Hpc7g Instances." aws.amazon.com/...hpc7g
  10. ^Jeff Barr, "Join the preview for new memory-optimized, AWS Graviton4-powered Amazon EC2 instances (R8g)," AWS News Blog, 28 November 2023. aws.amazon.com/...powered-amazon-ec2-instances-r8g
  11. ^Amazon, "AWS Unveils Next Generation AWS-Designed Chips," press release, 28 November 2023. press.aboutamazon.com/...ration-aws-designed-chips
  12. ^AWS, "Amazon EC2 R8g instances powered by AWS Graviton4 now generally available," What's New with AWS, 9 July 2024. aws.amazon.com/...ws-graviton4-generally-available
  13. ^Esra Kayabali, "Now available: Amazon EC2 M9g and M9gd instances powered by new AWS Graviton5 processors," AWS News Blog, 10 June 2026. aws.amazon.com/...-by-new-aws-graviton5-processors
  14. ^About Amazon, "AWS Graviton5 is now generally available, delivering purpose-built performance for the agentic AI era," originally published 4 December 2025, updated 10 June 2026. aboutamazon.com/...aws-graviton-5-cpu-amazon-ec2
  15. ^AWS, "Amazon EC2 M9g Instances." aws.amazon.com/...m9g
  16. ^AWS News Blog, "Top announcements of AWS re:Invent 2025." aws.amazon.com/...nouncements-of-aws-reinvent-2025
  17. ^AWS, "The Security Design of the AWS Nitro System," whitepaper, 15 February 2024. docs.aws.amazon.com/...-design-of-aws-nitro-system
  18. ^AWS, "AWS Nitro System." aws.amazon.com/...nitro
  19. ^AWS, "AWS Silicon Innovation." aws.amazon.com/silicon-innovation
  20. ^AWS Machine Learning Blog, "Optimized PyTorch 2.0 inference with AWS Graviton processors," 3 May 2023. aws.amazon.com/...nce-with-aws-graviton-processors
  21. ^AWS Machine Learning Blog, "Accelerated PyTorch inference with torch.compile on AWS Graviton processors," 2 July 2024. aws.amazon.com/...mpile-on-aws-graviton-processors
  22. ^AWS Machine Learning Blog, "Accelerate NLP inference with ONNX Runtime on AWS Graviton processors," 15 May 2024. aws.amazon.com/...ntime-on-aws-graviton-processors
  23. ^AWS Machine Learning Blog, "Reduce Amazon SageMaker inference cost with AWS Graviton," 10 May 2023. aws.amazon.com/...inference-cost-with-aws-graviton
  24. ^AWS Graviton Technical Guide, "LLM inference on Graviton CPUs with llama.cpp." github.com/...llama.cpp.md
  25. ^AWS Graviton Technical Guide, "PyTorch inference on Graviton." github.com/...pytorch.md
  26. ^AWS, "Amazon EC2 G5g Instances." aws.amazon.com/...g5g
  27. ^About Amazon, "Meta expands Amazon partnership with AWS Graviton chips for AI," April 2026. aboutamazon.com/...meta-aws-graviton-ai-partnership
  28. ^Meta Newsroom, "Meta Partners With AWS on Graviton Chips to Power Agentic AI," 24 April 2026. about.fb.com/...graviton-chips-to-power-agentic-ai
  29. ^Wikipedia, "Annapurna Labs." en.wikipedia.org/...Annapurna_Labs
  30. ^Wikipedia, "AWS Graviton." en.wikipedia.org/...AWS_Graviton

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 3,251 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent adversarial fact-check at creation (wanted175 campaign, 2026-07-24): every claim verified against primary sources by a dedicated verification agent; corrections applied before publication.

Cite this page: AI Wiki. "AWS Graviton." aiwiki.ai, updated 24 Jul 2026, fact-checked 24 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/aws_graviton

Suggest edit