Back to News Hub
📐SiliconANGLE AI
July 29, 2026
Product Updates

Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution

Overview

Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world. The recent collaboration pairs AMD's Helios rack-scale architecture for the compute-intensive pre-fill phase with [... ] The post Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution appeared first on SiliconANGLE.

Key Takeaways

  • Cerebras CMO Julie Choi says the AMD Helios and Cerebras Wafer-Scale Engine partnership delivers 5X higher throughput for disaggregated AI inference.
  • "It's a one plus one equals five X in this case.

    " Choi spoke with theCUBE's John Furrier at the Neo4j GraphTalk event during an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio.

  • "On the decode portion, this is a memory bandwidth constrained problem," Choi said.

    "The Cerebras Wafer-Scale Engine has the largest amount of memory bandwidth.

  • Cerebras will also bring AMD Helios systems into its own data centers this year to power the pre-fill layer, making the partnership both a product collaboration and a production infrastructure commitment.

    "Our vision is to really provide this speed and max intelligence, no trade-off, to every developer on Earth," Choi said.

  • 4k+ theCUBE alumni - Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network.

Stats & Key Facts

  • #The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions.
  • #2,000 times more than Nvidia GPUs.
  • #The 5x throughput gain means the same infrastructure can serve dramatically more concurrent users, with joint go-to-market efforts expected before the end of the year, Choi noted.
Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution

Cerebras CMO Julie Choi says the AMD Helios and Cerebras Wafer-Scale Engine partnership delivers 5X higher throughput for disaggregated AI inference. UPDATED 18:09 EDT / JULY 29 2026 AI Cerebras and AMD partner to build the world's fastest disaggregated AI inference solution by Thomas Godwin Disaggregated AI inference is proving to be more than a complementary answer to the prefill and decode bottleneck slowing enterprise AI at scale, and Cerebras and AMD just announced a partnership to build the fastest version of it in the world. The recent collaboration pairs AMD's Helios rack-scale architecture for the compute-intensive pre-fill phase with the Cerebras Wafer-Scale Engine for ultra-low-latency decode, according to Julie Choi (pictured), chief marketing officer at Cerebras.

The resulting combination delivers 5x higher tokens per second per watt compared to existing solutions. Later this year, Cerebras will bring AMD Helios systems into its own data centers to power the pre-fill layer of the production deployment. "Lisa and Andrew both announced how AMD and Cerebras are collaborating on the world's most powerful disaggregated inference solution," Choi said.

"It's a one plus one equals five X in this case. " Choi spoke with theCUBE's John Furrier at the Neo4j GraphTalk event during an exclusive broadcast on theCUBE, SiliconANGLE Media's livestreaming studio. They discussed the technical architecture behind disaggregated AI inference, along with what workloads are driving the fastest demand and why the partnership extends to deploying AMD Helios inside Cerebras data centers before the end of 2026.

For more details please read the original article at SiliconANGLE AI.

Continue Learning

Originally published by SiliconANGLE AI
Read the original

Comments

Sign in to join the conversation