
Nvidia has released detailed SPEC CPU 2026 benchmark results for its Vera CPU, revealing performance against AMD's Turin-based Epyc 9755 in dual-socket configurations. According to the latest white paper, Vera achieved an overall score of 901 compared to the Epyc 9755's 898 score, representing a 3% advantage for Nvidia's custom Olympus core design. The benchmarks were conducted using GNU 15.2 compiler with both systems running in fully-loaded socket configurations. Nvidia's Ian Buck explained that per-core performance under fully-loaded socket conditions is crucial for agentic AI and reinforcement learning systems, where many sandboxes run concurrently with sequential, latency-sensitive operations.
Recent benchmark results from DeepInfra, an early access participant in Nvidia's open AI ecosystem, demonstrate the Vera CPU's superiority in agentic AI workloads. The cloud platform, which processes nearly 5 trillion tokens a week with about 30% driven by agentic systems, independently conducted benchmarks using its production AI agent infrastructure. DeepInfra's benchmarks show the Vera CPU is more than twice as fast and can support up to 1.6x more concurrent AI agents at the same quality of service compared to other CPUs. The results also indicate up to 2.2x faster orchestration than alternative CPUs, while improving infrastructure utilization and cost efficiency. These benchmarks show the Vera CPU delivers the cost efficiency, low latency and throughput that production agentic AI demands, with per-core performance under fully-loaded socket conditions being crucial for these applications.
The Vera CPU sits at the heart of Nvidia's new Vera Rubin platform, built on the company's custom Olympus core architecture. The processor delivers twice the single-threaded performance, three times the core-to-core bandwidth and 40% lower memory latency compared to competing chiplet-based designs. According to the white paper analysis, Vera provides more than four times the per-core bandwidth of AMD's 9755, with loaded memory latency showing significantly better performance than the Turin chip. The chip features Sixth-generation NVLink delivering more than twice the throughput, three times lower latency and ten times higher packet rates than conventional Ethernet, while Spectrum-X Ethernet enhances large-scale networking performance. The platform combines seven chips and five rack trays into a single integrated system with liquid cooling system designed to lower water consumption in AI factories. The Vera Rubin NVL72 system delivers up to 10x more tokens per megawatt and one-tenth the cost per million tokens compared to NVIDIA GB200 NVL72, providing more intelligence within the same power footprint.
Nvidia said Vera Rubin is ramping up worldwide with support from more than 300 partners. Production of the Vera Rubin NVL72 platform is underway across over 350 factory sites in 30 countries, representing the largest, most mature rack-scale supply chain ever assembled to meet customer compute demand. Cloud providers including CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure are among those deploying the platform, which Nvidia claims delivers industry-leading performance per watt and the lowest token cost. CoreWeave became the first AI cloud to bring up and validate Vera Rubin NVL72, sharing the first measured performance numbers from live hardware. The company has shared that general availability is on track for the back half of this year, with the Vera CPU representing Nvidia's first custom core design specifically targeting the expanding AI infrastructure market.
Nvidia said Vera Rubin will serve as the computing foundation for its expanded partnership with Microsoft and Mistral in Europe. The platform is expected to power the next generation of Microsoft's European AI infrastructure and Mistral Compute, supporting sovereign AI deployments across public cloud, private cloud and customer-controlled environments using tens of thousands of GPUs. The European expansion represents a key component of Nvidia's strategy to establish end-to-end AI infrastructure capabilities. A new multibillion-dollar agreement focuses on expanding AI infrastructure in Europe, with Mistral adding its GPU capacity drawing on thousands of the latest NVIDIA Vera Rubin GPUs. Mistral Medium 3.5 and OCR 4 are now available in Microsoft Foundry, with Mistral models integrated into Microsoft Copilot Studio. Through Azure Local and Foundry Local, customers can use the same models, tools and operating patterns across cloud and customer-controlled environments, enabling government agencies, healthcare organizations, and manufacturers to deploy within regional requirements.