Direct answer
Vera Rubin is NVIDIA’s next rack-scale AI platform after Grace Blackwell. The flagship Vera Rubin NVL72 combines 72 Rubin GPUs, 36 Vera CPUs, sixth-generation NVLink, ConnectX-9 networking, BlueField-4 DPUs, and HBM4 memory in one liquid-cooled system. NVIDIA’s July 2026 update says production is ramping and early systems are running at CoreWeave, Google Cloud, Microsoft Azure, Oracle Cloud Infrastructure, and Nebius.
For a buyer, the important distinction is between a platform entering production, a provider announcing deployment, a dated allocation, and capacity that can be contracted with a defined start date and acceptance plan. Those are four different evidence states.
Vera Rubin is a platform, not just a new GPU
The name combines the NVIDIA Vera CPU and NVIDIA Rubin GPU, but the commercial unit is a tightly codesigned rack-scale system. NVIDIA describes Vera Rubin as a set of compute, scale-up, scale-out, networking, storage, security, and management components engineered together rather than assembled from independent parts.
The core platform includes:
- Rubin GPU: the accelerator, with 288 GB of HBM4 and up to 22 TB/s of memory bandwidth in NVIDIA’s preliminary specification;
- Vera CPU: NVIDIA’s custom Arm-compatible CPU with 88 Olympus cores per superchip configuration;
- NVLink 6: the scale-up fabric connecting the 72-GPU rack-scale domain;
- ConnectX-9: high-bandwidth scale-out network interfaces;
- BlueField-4: infrastructure processing for networking, storage, and security;
- Spectrum-6 and Quantum-X800: Ethernet and InfiniBand paths for scaling beyond one rack;
- Mission Control: NVIDIA’s management layer for configuration, facilities integration, and cluster operations.
This matters because a quote for “Rubin GPUs” is not enough. The buyer needs to know whether the offer is Vera Rubin NVL72, Vera Rubin NVL4, another Rubin system configuration, or future cloud instances abstracting the underlying hardware.
The headline Vera Rubin NVL72 specifications
NVIDIA labels the public specifications as preliminary and subject to change. They should be used to understand the architecture, not copied into a contract without confirming the provider’s final bill of materials.
| System area | Vera Rubin NVL72 preliminary specification | Buyer implication |
|---|---|---|
| Compute | 72 Rubin GPUs and 36 Vera CPUs | The procurement unit is an integrated rack-scale system, not a loose pool of cards. |
| GPU memory | 20.7 TB HBM4 total; 288 GB per Rubin GPU | Larger models, longer contexts, and memory-heavy inference may require less partitioning. |
| Memory bandwidth | Up to 1,580 TB/s across the rack; 22 TB/s per GPU | Potentially material for memory-bound training and inference, but workload measurement still matters. |
| Scale-up | 260 TB/s NVLink 6 switch bandwidth across the NVL72 system | Parallelism strategy should be designed around the 72-GPU scale-up domain. |
| CPU | 3,168 Olympus cores and 54 TB LPDDR5X across 36 Vera CPUs | Agent orchestration, data preparation, simulation, and CPU-heavy control paths are first-order parts of the system. |
| Scale-out | 28.8 TB/s system bandwidth; ConnectX-9 and BlueField-4 interfaces | Multi-rack design still requires a specified fabric, topology, congestion plan, and acceptance test. |
| Operations | Mission Control, NVIDIA AI Enterprise, and DGX OS in the DGX configuration | The provider’s actual software and managed-service boundary must match the buyer’s operating model. |
What changes from Grace Blackwell NVL72
The clearest architectural changes are HBM4 memory, the Vera CPU, sixth-generation NVLink, new network interfaces, and a platform designed around agentic AI, mixture-of-experts models, long-context reasoning, and more CPU-intensive orchestration. NVIDIA and CoreWeave report that an early DeepSeek-R1 benchmark delivered ten times more throughput per megawatt than Grace Blackwell NVL72.
That is an important launch benchmark, but it is not a universal workload guarantee. It was produced on a specific model, software stack, configuration, and early provider system. A buyer should ask what portion of the improvement comes from accelerator throughput, memory, precision, CPU orchestration, software maturity, batching, or power-envelope differences—and then benchmark the intended workload.
Vera Rubin does not make Blackwell obsolete on arrival. An available and accepted B200, B300, GB200, or GB300 cluster can be commercially superior to a future Rubin system if it starts earlier, fits the workload, has a proven operating stack, or avoids the risk and premium of a launch-generation deployment.
“Ramping” is not the same as broadly available
NVIDIA’s July update names early cloud partners and says production is ramping across a large supply chain. It also says racks are running at several partners. Those facts establish that Vera Rubin has moved beyond a paper roadmap. They do not establish that a specific provider has unallocated capacity in a buyer’s preferred region, count, date, term, or service model.
Omega Gradient would classify a Vera Rubin opportunity using the same evidence ladder applied to other reserved clusters:
- Live and testable: the provider can identify the system, site, available block, and a scoped validation path.
- Installed or commissioning: hardware is on site, but network, software, burn-in, or acceptance remains.
- Allocated and scheduled: equipment and facility milestones support a documented future date.
- Conditional deployment: capacity depends on another customer, financing, equipment delivery, power, cooling, or facility work.
- Announcement only: a roadmap or partnership exists, but no buyer-specific allocation has been evidenced.
A provider announcement can be credible and still sit in the fifth category for a particular buyer.
Who should evaluate Vera Rubin first?
Vera Rubin is most relevant when the workload can exploit more than a simple generational speed increase:
- large mixture-of-experts training or inference with demanding communication patterns;
- long-context and reasoning workloads constrained by memory capacity or bandwidth;
- agentic systems with substantial CPU orchestration, tool execution, sandboxing, or state management;
- power-constrained deployments optimizing useful tokens per megawatt;
- large scientific and HPC programs needing native FP64 paths or tightly integrated AI and simulation;
- multi-rack AI factories willing to adopt a new system and software generation at scale.
Smaller fine-tuning jobs, conventional inference, early-stage experiments, and teams without stable utilization may be better served by mature Hopper or Blackwell capacity. The operational and commercial premium of a launch platform has to create measurable workload value.
The facility is part of the Vera Rubin product
Vera Rubin NVL72 is a liquid-cooled rack-scale system. A credible delivery plan needs more than a manufacturer allocation. It needs a site with appropriate power, cooling, water-loop responsibilities, rack integration, network, security, commissioning, spares, software, and staff.
Before accepting a future start date, buyers should ask:
- Which facility and hall will host the system?
- Is power committed to the quoted racks, and what milestone makes it usable?
- Which party owns the CDU, facility-water loop, and cooling remediation obligations?
- Is the scale-out network Quantum-X800 InfiniBand, Spectrum-X Ethernet, or another design?
- What is the blocking ratio, oversubscription boundary, and multi-rack topology?
- When do installation, energization, networking, commissioning, burn-in, and acceptance occur?
- Who controls the equipment, who operates the cloud, and who signs the buyer contract?
A Vera Rubin procurement checklist
- Name the exact system. NVL72, NVL4, cloud instance, or another configuration; final GPU, CPU, memory, NIC, DPU, storage, and software specification.
- Define the capacity. Rack count, GPU count, region, start window, ramp, term, tenancy, and whether the block is exclusive.
- Map the delivery chain. Manufacturer or OEM, equipment owner, facility operator, cloud operator, reseller, contracting entity, and support owner.
- Verify the facility path. Power, liquid cooling, network, installation, commissioning, and unresolved dependencies.
- Confirm software readiness. Drivers, CUDA, frameworks, orchestration, Mission Control responsibility, observability, and workload compatibility.
- Normalize the economics. GPU-hour or rack economics, deposit, prepayment, storage, bandwidth, support, software, ramp, renewal, and termination.
- Write the acceptance plan. Inventory, firmware, burn-in, NVLink, scale-out fabric, storage, representative workload, pass criteria, cure period, and remedies.
- Protect the delivery date. Milestone payments, refund or termination rights, partial delivery, substitution rules, and what happens if final specifications change.
Should buyers reserve Vera Rubin now or wait?
Begin sourcing now if the program has a known 2026–27 start window, the workload can plausibly benefit from the architecture, and early access is strategically valuable. The purpose of an early process is to learn the real provider, facility, allocation, and commercial landscape—not to treat every announcement as supply.
Wait or benchmark first if utilization is uncertain, software compatibility is untested, a mature Blackwell system meets the requirement, or the commercial case depends on unverified performance assumptions. An H200 or Blackwell benchmark can establish the workload baseline that a future Rubin proposal must beat.
The procurement question is not “Is Vera Rubin faster?” It is “Will this specific Vera Rubin system deliver more useful work per dollar, megawatt, and month of deployment risk for our workload?”
Omega Gradient can structure a Vera Rubin market scan while keeping Hopper, Blackwell, and scheduled or on-demand alternatives in the same decision frame.
NVIDIA, Vera Rubin NVL72 architecture and preliminary specifications; NVIDIA, DGX Vera Rubin NVL72; NVIDIA, July 2026 production-ramp and partner update; NVIDIA, DGX SuperPOD and Rubin platform announcement; NVIDIA Newsroom, Vera Rubin for scientific computing. NVIDIA marks specifications as preliminary and subject to change; confirm the final provider configuration before procurement.