Skip to main content
Back to News Hub
⚙️IEEE Spectrum AI
June 1, 2026
Research

New Server Hopes to Break Through AI's "Memory Wall"

Overview

AI hardware startup Majestic Labs is building a new server called Prometheus with up to 128 terabytes of memory, over 60 times more than Nvidia's DGX B300, to address what the industry calls the memory wall in large language model inference. The company uses a DRAM-centric architecture with a proprietary copper-cable memory interface and custom aggregation chips. Prometheus pairs this memory with a custom AI processor called Ignite and supports common frameworks without code changes.

Key Takeaways

  • Memory is arguably the most serious constraint on modern AI large language models (LLMs).

    According to one influential paper , LLM token generation is an inherently memory-bound task, meaning the rate at which models output text is limited by how quickly data can be read in from memory.

  • While he acknowledges that "Nvidia's done a phenomenal job creating a system that can scale out," he argues that it becomes less economical as models grow and "ends up greatly over-provisioning on compute and starving on memory."

    DRAM-Centric Architecture for LLM Memory Majestic Labs plans to surmount the "memory wall" with an architecture that fundamentally differs from competitors'.

  • To solve that, Majestic uses a proprietary memory interface constructed from miniature copper cables that's effective up to a meter.

    This is paired with custom memory aggregation chips that sit physically next to memory modules and coordinate memory across the server.

  • The ARM cores act as an on-chip host processor to orchestrate the AI model.
  • Up to four servers can fit in a server rack; power draw is expected to total up to 120 kilowatts per rack; and heat will be managed with cold-plate liquid cooling .

Stats & Key Facts

  • #AI hardware startup Majestic Labs is building a new server called Prometheus with up to 128 terabytes of memory, over 60 times more than Nvidia's DGX B300, to address what the industry calls the memory wall in large language model inference.
  • #That's over 60 times more than Nvidia's DGX B300 server , a cutting-edge AI processing rack.
New Server Hopes to Break Through AI's "Memory Wall"

Memory is arguably the most serious constraint on modern AI large language models (LLMs). According to one influential paper , LLM token generation is an inherently memory-bound task, meaning the rate at which models output text is limited by how quickly data can be read in from memory. The severity of this bottleneck grows with model size.

This creates a "memory wall" that holds back LLM inference performance. AI hardware startup Majestic Labs is taking a direct-and comprehensive-approach to solving this problem. It's developing a new AI server, Prometheus, with up to 128 terabytes of memory.

That's over 60 times more than Nvidia's DGX B300 server , a cutting-edge AI processing rack. Sha Rabii , co-founder and president of Majestic Labs, believes that this drastic increase in memory will provide his company an edge. While he acknowledges that "Nvidia's done a phenomenal job creating a system that can scale out," he argues that it becomes less economical as models grow and "ends up greatly over-provisioning on compute and starving on memory."

For more details please read the original article at IEEE Spectrum AI.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by IEEE Spectrum AI
Read the original