Back to News Hub
🟩NVIDIA Blog
July 21, 2026
Tech

NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

Overview

NVIDIA Vera Rubin is here, and it's going gigascale. Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350-plus factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand.

Key Takeaways

  • Backed by 300 global partners, Vera Rubin is ramping up worldwide.

    NVIDIA partners CoreWeave, Google Cloud, Microsoft Azure and Mistral are among many deploying Vera Rubin, which delivers benchmark leadership on performance per watt and lowest token costs.

  • The NVIDIA Vera CPU is at its center.
  • 6T ConnectX-9 SuperNICs, adaptive routing, advanced congestion control, telemetry and open operating system support, enabling 1.
  • And NVLink Fusion opens the NVIDIA infrastructure platform to third-party XPUs, giving partners a faster path to market on the proven NVLink scale-up stack and ecosystem.

    Saving Setup Time, Water NVIDIA's three generations of rack-scale co-design produced a Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute.

  • Underpinning the partnership is a new multibillion-dollar agreement focused on expanding AI infrastructure in Europe.

Stats & Key Facts

  • #Spanning 350-plus factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand.
  • #Spanning 350-plus factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand.
  • #CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72 - landing directly on the metric that matters most for power-constrained AI factories.
  • #Designed and built for the agent era, its custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs, making it the most efficient single-threaded CPU for the agentic workloads that matter most.
NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

Vera Rubin NVL72 production is ramping up with racks running at partners CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. Spanning 350-plus factory sites in 30 countries, Vera Rubin has the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand. The Vera Rubin platform is built from chip to grid to deliver the highest performance per watt and the lowest token cost.

CoreWeave's first benchmark on DeepSeek-R1 says it all: 10x more throughput per megawatt than Grace Blackwell NVL72 - landing directly on the metric that matters most for power-constrained AI factories. Advancing Performance With Extreme Co-Design What makes this possible is extreme co-design across seven chips and five rack trays - Vera Rubin NVL72, Vera CPU rack, Groq 3 LPX, Spectrum-6 SPX and Vera BlueField-4 STX - all engineered as a single system rather than assembled from separate off-the-shelf products. The NVIDIA Vera CPU is at its center.

It redefines what an AI factory CPU can be. Designed and built for the agent era, its custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs, making it the most efficient single-threaded CPU for the agentic workloads that matter most. Accelerating AI Factories With Purpose-built Networking For networking, the platform's sixth-generation NVLink scale-up delivers more than 2x throughput on complex workloads, 3x lower latency and 10x higher packet rates than off-the-shelf Ethernet.

6x higher RDMA bandwidth than off-the-shelf Ethernet. The world's leading AI infrastructure builders - including CoreWeave, Microsoft, SpaceXAI and Tesla - are among the first to bring in Spectrum-6 switches to accelerate their AI factories. NVIDIA Photonics with co-packaged optics for scale-out - the industry's first such switch in volume manufacturing - adds 5x lower power and 10x higher MTBI versus pluggable transceivers, with CoreWeave , Lambda and OCI among the first adopters.

Spectrum-XGS Ethernet extends performance across sites with 1. 9x multi-site throughput because gigascale AI isn't a single building problem. And NVLink Fusion opens the NVIDIA infrastructure platform to third-party XPUs, giving partners a faster path to market on the proven NVLink scale-up stack and ecosystem.

Saving Setup Time, Water NVIDIA's three generations of rack-scale co-design produced a Vera Rubin NVL72 system with no cables, fans or hoses in the tray, cutting compute tray assembly time from hours to one minute. A 45-degree Celsius liquid cooling inlet temperature design enables chiller-free dry-cooler operation. For new AI factories, this higher temperature dry cooling along with the closed-loop liquid cooling system saves millions of gallons of water per megawatt annually.

PT 🔗 NVIDIA Vera Rubin Powers Europe's Open-Model Era Vera Rubin is delivering next-generation performance to Europe's AI infrastructure. It's the foundation for a newly expanded Microsoft and Mistral partnership that brings frontier AI to the region, combining open European models with cloud and customer-controlled environments so governments and regulated industries can adopt it on their own terms. Underpinning the partnership is a new multibillion-dollar agreement focused on expanding AI infrastructure in Europe.

For more details please read the original article at NVIDIA Blog.

Continue Learning

Originally published by NVIDIA Blog
Read the original

Comments

Sign in to join the conversation