NVIDIA Vera Rubin NVL72 on Lambda

The next compute platform for frontier AI training and inference. Deployed as a production-ready Lambda Supercluster.

Available H2 2026.

NVIDIA_NVL72-2

One platform for training and serving at scale

10x more inference throughput per watt.
One-tenth the token cost. Six purpose-built chips, designed to operate as one system.

Compute

Train frontier models with fewer GPUs. Serve inference at a fraction of the cost per token.

  • 72 NVIDIA Rubin GPUs: 50 PFLOPS NVFP4 per GPU
  • 36 NVIDIA Vera CPUs: 88 custom NVIDIA Olympus cores each
  • 288GB HBM4 per GPU at ~22 TB/s bandwidth (~2.7x over HBM3e)
  • 20.7TB HBM4 and 1,580 TB/s aggregate memory bandwidth per rack

Networking

5x better power efficiency. Every collective stays fast from GPU to supercluster.

  • NVIDIA NVLink 6 Switch: unifies all 72 GPUs in a rack into a single high-bandwidth domain, keeping collective operations local
  • NVIDIA ConnectX-9 SuperNIC: GPU-direct scale-out networking with programmable RDMA for low-latency inter-rack communication
  • NVIDIA BlueField-4 DPU: offloads infrastructure services and enforces zero-trust security isolation, freeing GPUs for compute
  • NVIDIA Spectrum-6 Ethernet switch (co-packaged optics): scale-out switching with integrated silicon photonics

Storage

Eliminate GPU stalls for long-context and agentic workloads.

  • NVIDIA STX: 5x higher token throughput and 2x faster ingestion than traditional storage

End-to-end network performance

Predictable collectives and fault isolation at every layer, from GPU to supercluster.

Scale-up

NVLink 6 unifies all 72 GPUs in a rack. 3.6 TB/s per GPU. 260 TB/s per rack.

Scale-out (InfiniBand)

NVIDIA Quantum-X800 InfiniBand with SHARP-accelerated collectives, up to 41,000 GPUs

Scale-out (Ethernet)

NVIDIA Spectrum-X Ethernet (RoCE) scales to 128,000 GPUs 

Built for agentic workloads

Large-scale training

Train frontier models with fewer GPUs. Vera Rubin delivers the same MoE training throughput with one-fourth the GPUs compared to Blackwell, compressing cluster size and reducing infrastructure cost for runs at 10T+ tokens.

Mixture-of-Experts (MoE)

All-to-all token routing stays within the NVLink domain, keeping expert dispatch fast and reducing step-time variance across multi-rack deployments.

Long-context inference

Million-token context windows run without offloading to slower memory tiers.

Agentic execution

Vera CPU handles scheduling, orchestration, and KV-cache management without competing for GPU cycles. BlueField-4 isolates infrastructure services from application workloads, maintaining deterministic performance under multi-step agent execution.

Datasheet

Get the full NVIDIA Vera Rubin NVL72 technical specifications, performance comparisons, and networking reference architectures.