AWS Debuts Graviton5 Pro: A 'Processing-in-Memory' Chip to Break Nvidia's Stranglehold
Amazon Web Services just revealed Project Cerebrum: a new "processing-in-memory" chip called Graviton5 Pro. This radical design, available in new P1 instances, directly targets the memory bottlenecks that plague large AI models—and Nvidia's market dominance.

The Shot Heard 'Round the Cloud
In a move that sent shockwaves through the technology industry, Amazon Web Services today, October 5, 2026, unveiled a radical new piece of silicon that challenges decades of computer architecture and takes direct aim at Nvidia’s lucrative AI empire. At a packed keynote, AWS CEO Adam Selipsky introduced the Graviton5 Pro, a chip born from a multi-year, top-secret initiative codenamed 'Project Cerebrum'. This is no mere incremental update. The Graviton5 Pro is AWS's first commercial 'Processing-in-Memory' (PIM) chip, a design that merges logic and memory to dissolve the single greatest bottleneck in modern computing: the costly shuffle of data between processor and RAM. These new chips will power a brand-new EC2 instance family, the P1 instances, which AWS claims will offer an order-of-magnitude improvement in performance-per-watt for large-scale AI inference workloads.
What Just Happened?
The announcement effectively fires the starting gun on a new architectural war. For years, the industry has relied on a standard model: powerful processors (CPUs and GPUs) are paired with vast pools of separate memory (DRAM). AWS is betting that this model is broken, at least for the future of AI. The Graviton5 Pro chip physically integrates computational logic directly within its memory arrays. This means data can be processed where it lives, eliminating the immense energy and latency costs of moving petabytes of data across physical buses.
"We've spent years chasing performance by adding more cores and faster clocks, but we've always been shackled by the 'memory wall'," Selipsky stated during the keynote. "With Project Cerebrum and the Graviton5 Pro, we're not just climbing the wall; we're demolishing it. This is a fundamental rethinking of how a computer should work in the age of generative AI."
The new P1 instances will launch in preview next week. The flagship `p1.32xlarge` instance will reportedly feature 8 Graviton5 Pro chips, offering a staggering 4 terabytes of PIM-enabled memory. While pricing is not final, initial estimates suggest it could be 30-40% cheaper for certain large language model (LLM) inference tasks compared to an equivalent instance powered by Nvidia's latest H2000 GPUs.
The End of the Von Neumann Bottleneck?
To understand the significance of PIM, think of a master chef in a kitchen. The traditional von Neumann architecture is like a chef whose ingredients are all stored in a pantry down the hall. For every step of the recipe, they must walk to the pantry, grab one ingredient, walk back, use it, and repeat. It's inefficient, and most of their time is spent walking, not cooking. A PIM architecture, by contrast, is like a chef with a hyper-advanced workstation where every ingredient, spice, and tool is within arm's reach and can even perform simple tasks on its own. The chef's focus shifts from fetching to orchestrating.
"Nvidia sells you a world-class engine; AWS is now trying to sell you a teleportation device. If it works, the entire map of AI infrastructure changes overnight."
This "memory wall" or "von Neumann bottleneck" is the primary performance limiter for today's enormous AI models. Models with trillions of parameters are too large to fit in a single GPU's memory, forcing developers into complex and slow distributed computing arrangements. The Graviton5 Pro's architecture is designed to make a massive, unified memory space addressable and computationally active, potentially allowing for models of unprecedented scale to be run with far greater efficiency.
A Direct Assault on Nvidia's Throne
Make no mistake: this is a declaration of war on Nvidia. While AWS will continue to be one of Nvidia's largest customers, the Graviton5 Pro is a clear strategy to reduce its dependence and capture the colossal margins of the AI hardware market. Nvidia's GPUs are masters of parallel mathematics, but their performance is increasingly gated by memory bandwidth—how fast data can be fed to the cores. Their flagship H2000 Tensor Core GPU, with its advanced HBM4 memory, is an engineering marvel, but it still operates within the classical architectural paradigm.
"AWS is playing a different game," commented one analyst from SemiAnalysis. "Nvidia is optimizing the existing paradigm to its absolute physical limits. AWS is attempting to sidestep the paradigm entirely. For inference workloads on massive foundation models, where memory access patterns are often sparse and unpredictable, the PIM approach could be a killer app. The latency and energy savings could be enormous."
Nvidia's stock is expected to face significant pressure as the market digests this news. Their moat has always been their CUDA software ecosystem and the sheer performance of their chips. AWS is attacking not on performance alone, but on architectural efficiency and cloud integration.
The Software Challenge: A New Programming Paradigm
A revolutionary chip is useless without a way to program it. AWS was quick to address this, announcing the 'Cerebrum SDK', a new software development kit with deep integrations for popular frameworks like PyTorch and JAX. The SDK includes a specialized compiler that analyzes AI models and automatically offloads certain operations to the memory-integrated logic. However, this is not a free lunch. Unlocking the full potential of the P1 instances will require developers to think differently about data locality and model design.
Early partners, including Anthropic and Perplexity, have reportedly been working with the SDK for months. However, the broader community will face a learning curve. This also introduces a new form of vendor lock-in. Models and code highly optimized for AWS's PIM architecture may not be easily portable to Google's TPU-based infrastructure or Microsoft's Azure offerings. AWS is betting the performance gains will be worth the commitment.
Who Wins, Who Loses, and What's Next?
The implications of this launch are far-reaching. AWS is the clear potential winner, positioning itself to capture a larger slice of the AI value chain and create a powerful new competitive moat. AI startups and researchers could also win big, as the potential for more efficient and affordable access to massive model inference could democratize capabilities previously reserved for tech giants.
The most immediate loser is Nvidia, which now faces a credible, architectural threat to its data center dominance for the first time in years. Microsoft and Google are also on the back foot; while both have their own custom AI silicon (Maia and TPUs, respectively), neither has a commercial PIM offering, and they must now race to respond. Even traditional memory manufacturers like Samsung and Hynix should be nervous, as AWS's custom-fabricated PIM chips disrupt the standard commodity DRAM market.
The coming months will be critical. The industry will be laser-focused on the first independent benchmarks to emerge from the P1 preview. All eyes will turn to Nvidia's next GTC conference for Jensen Huang's response, and to Microsoft and Google for hints of their own post-von Neumann strategies. AWS has thrown down the gauntlet. The race to define the next era of computer architecture has just begun, and the battlefield is the cloud.
Frequently asked questions
Can I use Graviton5 Pro for my regular web server or database?+
No, not effectively. The Graviton5 Pro's PIM architecture is highly specialized for massively parallel workloads with specific data access patterns, like AI inference and scientific computing. For general-purpose tasks like running a web server, traditional CPUs like the standard Graviton4 or Intel Xeon chips will remain far more efficient and cost-effective. Using a P1 instance for a simple database would be like using a Formula 1 car for a grocery run.
Is this actually faster than Nvidia's best H2000 GPU?+
It's more complicated than 'faster'. For raw graphics rendering or certain dense matrix multiplication tasks, Nvidia's H2000 will likely retain the crown. The Graviton5 Pro's advantage lies in performance-per-watt and latency for specific AI workloads, particularly inference on models too large to fit in a single GPU's memory. AWS's chip aims to be more *efficient* by eliminating the memory transfer bottleneck, which can make it 'faster' in practice for those targeted use cases.
How does this affect Google's TPUs or Microsoft's Maia chips?+
It puts immense pressure on them. While Google's TPUs and Microsoft's Maia are also custom AI accelerators, they largely operate within the traditional processor-plus-memory paradigm. AWS's PIM is a more radical architectural leap. This announcement forces Google and Microsoft to accelerate their own roadmaps for post-GPU architectures. Expect to see them publicize their own research into PIM, optical interconnects, or other novel designs to signal they aren't being left behind.
What does 'Processing-in-Memory' (PIM) really mean for a developer?+
Ideally, not much at first. AWS's Cerebrum SDK is designed to abstract away most of the complexity. A developer using PyTorch should be able to run their model on a P1 instance and let the compiler handle offloading tasks to the memory. However, for power users seeking maximum performance, it will mean thinking about data locality. They may need to refactor algorithms to take advantage of the ability to compute where data is stored, minimizing data movement.
When can I actually start using these new P1 instances?+
AWS has announced that P1 instances will be available in a limited public preview starting next week in the `us-east-1` (N. Virginia) and `eu-west-1` (Ireland) regions. However, they have cautioned that preview capacity will be extremely constrained and subject to an approval process. General availability across more regions is tentatively scheduled for Q1 2027, but this timeline may shift based on production yields and initial customer feedback.
Liked this story?
Share it with a colleague, or explore more in the Artificial Intelligence section.
More stories

The Empire Strikes Back: DOJ Sues Coherence AI in Landmark Antitrust Case
The Justice Department just fired the opening salvo in what could be the defining antitrust battle of the decade, suing AI juggernaut Coherence AI. At stake: the very architecture of the AI economy and who gets to build its future.

Cerebrum’s New ‘LPU’ Chip Aims to End NVIDIA’s AI Inference Stranglehold
After years in stealth, Cerebrum Systems just launched the Synapse-1, a specialized "Language Processing Unit." The chip promises to slash the cost of running large language models, posing the most credible threat yet to NVIDIA’s data center empire.

DOJ Sues Nexus AI, Seeks to Force Open-Sourcing of Its Flagship AI
In an unprecedented move that could reshape the future of artificial intelligence, the Department of Justice is suing market leader Nexus AI. The goal? To break its alleged monopoly by forcing its crown-jewel foundation model into the open.