AI News

Nvidia unveils Vera Rubin platform to enhance AI infrastructure efficiency

Nvidia highlights its innovative Vera Rubin platform, designed to enhance AI workloads and improve efficiency in data center operations.

Nvidia showcased its Vera Rubin platform during a recent briefing aimed at demonstrating the physical infrastructure behind artificial intelligence (AI). Ian Buck, general manager of Nvidia’s hyperscale and HPC computing business, emphasized the tangible nature of this infrastructure, stating, “Infrastructure is physical, it’s real. You can touch it, you can see it, and it’s what helps bring AI to life.” The event included a tour of a mini-datacenter located in Sunnyvale, California, although attendees were instructed to keep the location confidential.

The briefing comes at a time of heightened scrutiny over the environmental impact of data centers. Just a day after Nvidia’s event, Dutch activists protested against a Microsoft data center, highlighting concerns about climate and political issues related to tech operations. Despite such protests, Nvidia aims to mitigate neighborhood concerns regarding power and water usage through advancements in its data center technology.

Andrew Bell, senior vice president of hardware engineering at Nvidia, demonstrated that the new Vera Rubin NVL72 compute tray can be assembled in one minute, a significant reduction from the 90 minutes required for the previous GB200 compute tray. This automation represents a 90-fold improvement in installation efficiency.

The Vera Rubin platform, first announced at Computex 2024 and further detailed at CES in January, consists of six key components: the Vera CPU, Rubin GPU, NVLink 6 switch, ConnectX-9 network interface, BlueField-4 data processing unit (DPU), and Spectrum-6 Ethernet switch. Initial testing indicates that the Vera Rubin platform can deliver ten times more tokens per watt than the previous GB200 NVL72, based on results from CoreWeave on DeepSeek-R1.

Nvidia’s architecture for the Vera Rubin platform is designed to support AI workloads more effectively. The company argues that agentic workloads, which involve running AI agents, require low latency and sustained inference across multiple reasoning steps. The platform aims to unify compute capabilities throughout the data center to optimize performance.

The Vera CPU, featuring the Olympus Core, is designed to accelerate AI agents by two times, provide three times the core-to-core bandwidth, and reduce latency by 40 percent using LPDDR5X memory. Buck stated, “Every AI factory is power-constrained. And frankly the most important metric is your delivered performance in a fixed watt data center.” This focus on efficiency is essential for Nvidia’s customers, who stand to gain by moving more tokens within their AI inference operations.