AWS made Graviton5 generally available on June 10. Eight days later, Bloomberg reported that Amazon was in talks to sell Trainium AI accelerators directly to outside data centers, without requiring customers to use AWS.

The two developments point to a broader role for Amazon's chip business. Graviton5 improves the infrastructure AWS already sells. External Trainium sales would put Amazon into the hardware market alongside companies whose products it has spent years buying. As of June 20, 2026, those sales remain under discussion, but the scale behind them is substantial: Amazon's custom silicon business reached a $20 billion annual revenue run rate in Q1.

What Graviton5 adds

Graviton5 has 192 ARM Neoverse V3 cores in a four-chiplet package manufactured on TSMC's 3nm process. Graviton4 topped out at 96 cores. The new processor also has a reported 192 MB of L3 cache, five times its predecessor's capacity, plus DDR5-8800 memory and PCIe Gen 6 connectivity. Inter-core latency is reported to be 33% lower.

Amazon claims 25% higher overall compute throughput than Graviton4, with gains of 35% for web applications, 35% for machine-learning inference, and 30% for databases. Those are vendor performance claims, rather than a guarantee for every workload. Tom's Hardware described the chip as competitive with high-end AMD EPYC and Intel Xeon parts in cloud configurations. That puts Amazon's in-house CPU in competition on performance, beyond its original appeal as a lower-cost option.

Amazon describes Graviton5 as purpose-built for agentic AI, where software agents can generate many concurrent requests and carry out several tasks in sequence. The larger L3 cache, a fast pool of memory close to the CPU cores, and lower inter-core latency could help workloads that repeatedly move or reuse data.

That is relevant to transformer-based inference. Attention calculations can put heavy pressure on memory bandwidth, and long context windows increase the amount of data a system needs to retain and access. The KV cache stores previously calculated attention data so it doesn't have to be recomputed for each new token. A larger CPU cache may help with some of that data movement, though the benefit depends on the model and workload. With 192 cores, an instance also has more room to handle concurrent agent sessions before CPU threads become a constraint.

Graviton5 is available in M9g and M9gd instances. Compute-optimized C9g and memory-optimized R9g variants are expected later in 2026. Coverage of the Graviton5 launch also highlights the higher core count and claimed AI performance gains.

Trainium could become a product outside the cloud

Trainium currently comes through AWS as a cloud service. Customers can't simply buy the accelerators and install them in their own racks. The discussions reported by Bloomberg would change that arrangement by making Amazon a direct supplier to external data center operators.

Peter DeSantis, identified in the report as Amazon's AI chief, told Bloomberg that AI infrastructure was evolving rapidly and that Amazon was continually looking for ways to reach more customers. CEO Andy Jassy had already raised the possibility on the Q1 2026 earnings call in April. He said there was a good chance Amazon would offer Trainium beyond AWS within the next couple of years. The June reporting suggests concrete discussions are underway, although it doesn't establish a launch date.

Potential customers could include data center operators, regional hyperscalers, and sovereign cloud projects that want their own infrastructure. Selling to them would place Amazon more directly in competition with Nvidia, which supplies accelerators such as the H100 and H200 to outside operators.

DeSantis argued that external sales wouldn't hurt AWS revenue because AI demand remains far below its potential level. That is a plausible short-term argument: demand could support both cloud consumption and hardware purchases. Over time, however, selling chips would also give customers a way to use Amazon-designed compute without buying the surrounding AWS services. Amazon would be accepting that tradeoff in exchange for a larger hardware market.

The scale behind the sales discussions

Amazon's processor strategy started with a narrower economic goal. The first Graviton, launched in 2018, offered a way to run commodity workloads more cheaply and reduce exposure to Intel's pricing. Graviton2 in 2020 strengthened the performance case. By Graviton4, AWS could credibly market its processors against x86 alternatives on performance as well as cost.

Trainium has followed a similar progression. Trainium1 gave Amazon an alternative to Nvidia for machine-learning training and a way to reduce dependence on an outside supplier. Trainium2 improved the throughput case. Trainium3 is now reportedly nearly sold out.

Amazon has signed commitments totaling more than $225 billion in Trainium revenue. OpenAI has agreed to roughly two gigawatts of Trainium capacity through AWS, while Anthropic has committed to up to five gigawatts of current and future Trainium chips. These are commitments rather than revenue already collected, but their size suggests substantial planned workloads rather than small experimental purchases.

The custom silicon business includes Trainium, Graviton, and the Nitro security chip. Its annual revenue run rate crossed $20 billion in Q1 2026, growing more than 100% year over year. Amazon also deployed more than 2.1 million AI chips over the preceding 12 months.

Jassy said on the earnings call that a standalone version of the business, selling chips to outside customers in the manner of a traditional chip company, would have a $50 billion annual revenue run rate. That is his hypothetical estimate, not a separate stream of hardware revenue Amazon currently earns.

What infrastructure teams should evaluate

Graviton5 deserves benchmarking for agentic workloads, particularly multi-agent pipelines with many concurrent sessions. Its core count, cache capacity, and DDR5-8800 memory provide reasons to test M9g instances against existing infrastructure. They don't, by themselves, establish that a CPU instance will replace GPU-backed inference economically.

GPUs generally have an advantage at high request rates when requests can be batched efficiently. CPU inference can offer a better price and latency tradeoff for some lower-concurrency, latency-sensitive endpoints. The relevant comparison is the cost of meeting a workload's latency and throughput requirements, rather than core count alone. Graviton5 could improve that CPU option, but workload-specific measurements are still necessary.

For training, the OpenAI and Anthropic commitments are meaningful signs that Trainium is becoming a more mature alternative to Nvidia. Teams comparing H100 or H200 infrastructure with Trainium still need to account for software compatibility. Trainium's ecosystem remains behind CUDA, Nvidia's widely used software platform for GPU computing, even if the gap appears to be closing faster than it was 18 months ago. Large capacity commitments don't remove the work of adapting a training pipeline.

External Trainium sales would also create a possible option for financial services, government, and healthcare organizations that can't or won't put sensitive training workloads in a public cloud. That option depends on Amazon making hardware available for customers to operate themselves. The next 12 to 18 months should provide a clearer view of whether the reported discussions become a usable product offering.

Nvidia's advantages remain substantial

None of this establishes an immediate crisis for Nvidia. H100 and H200 demand remains backlogged, Blackwell is shipping, and CUDA remains the dominant software ecosystem for AI compute. Competitive hardware takes time to displace an established platform, particularly when customers have built their tools and training pipelines around it.

Amazon's silicon, along with Google's TPUs, Microsoft's Maia, and Meta's MTIA, raises the potential scale of the alternatives. It doesn't establish a near-term shift in market share. The large cloud providers have the research budgets, manufacturing relationships, and internal deployment volumes needed to support expensive chip development.

Amazon's $20 billion run rate supports the argument that internal deployment can provide enough scale to justify those fixed costs. External sales could spread the costs across more customers and potentially fund faster development. The path from reducing supplier dependence to building competitive products is already visible in Graviton and Trainium. Whether Amazon can turn that into a durable outside hardware business is the next test.

For infrastructure investments planned for 2026 to 2028, Amazon's silicon roadmap belongs in the evaluation. Graviton5 can be tested now. Trainium can be evaluated through AWS today, while direct hardware purchases remain a possibility rather than an available procurement option. Plans for training clusters, inference capacity, and regulated on-premises workloads should keep that distinction clear.