Direct answer
A request for a “B200 cluster” is incomplete. It could mean eight-GPU HGX nodes connected over a scale-out fabric, a GB200 NVL72 rack-scale system with 72 Blackwell GPUs and 36 Grace CPUs, or capacity marketed before the facility and system are ready. A buyer should specify the workload and required system boundary, then compare only options that make the exact architecture, site, delivery state, and acceptance path explicit.
Omega Gradient prioritizes reserved Blackwell mandates where the buyer needs a coherent cluster, a future delivery window, or a commercial comparison across different provider structures. We do not publish private operator identities or uncontracted capacity as “available.”
Decode the names before comparing quotes
B200 and B300
B200 is a Blackwell data-center GPU; B300 is the Blackwell Ultra successor. Providers may offer these accelerators in HGX systems or cloud instance products. The accelerator tells you the generation, but it does not by itself define GPU count per node, host architecture, NVLink/NVSwitch domain, network, cooling, or storage.
GB200 and GB300 NVL72
GB200 and GB300 pair Grace CPUs with Blackwell GPUs in rack-scale NVL72 systems. NVIDIA describes GB300 NVL72 as a fully liquid-cooled rack-scale architecture with 72 Blackwell Ultra GPUs and 36 Grace CPUs. That 72-GPU scale-up domain is materially different from a collection of independent eight-GPU nodes. It also carries different facility, installation, commissioning, and operational requirements.
Cloud product names
Cloud providers wrap the underlying platform in their own instance, cluster, or capacity-block products. The product can add an orchestration layer, storage, network, and commercial rules. Compare the underlying system and service boundary, not a provider SKU alone.
Choose HGX nodes or rack-scale NVL72 from the workload
| Question | HGX B200 / B300 cluster | GB200 / GB300 NVL72 |
|---|---|---|
| Scale-up domain | Typically centered on the GPUs within each HGX node, then scales out across the cluster fabric | Rack-scale 72-GPU NVLink domain |
| Deployment unit | Node-based; can support more incremental cluster shapes | Integrated rack-scale system with Grace CPUs, compute trays, switch trays, and liquid cooling |
| Facility burden | High-density GPU infrastructure; exact cooling and power depend on system | Facility integration and liquid-cooling readiness are central to delivery |
| Workload fit | Broad training, fine-tuning, inference, and HPC where node-based scaling is suitable | Large AI training and reasoning workloads that benefit from the rack-scale scale-up domain |
| Diligence focus | Node consistency, scale-out fabric, storage, orchestration, and available count | Rack configuration, liquid cooling, power, NVLink domain, scale-out network, commissioning, and system software |
Neither architecture is automatically better. NVL72’s integrated scale-up domain can be decisive for some model architectures and parallelism strategies. Node-based clusters can be easier to source incrementally, fit existing facilities, or align with an established Slurm or Kubernetes operating model. Benchmark the intended software stack and communication pattern.
Teams planning beyond Blackwell should also review the Vera Rubin buyer guide. Vera Rubin introduces HBM4, Vera CPUs, NVLink 6, and a new facility and delivery cycle, but an announced provider deployment should not be treated as contractable capacity without buyer-specific evidence.
Facility readiness is part of the product
Blackwell supply discussions often focus on manufacturer allocation. The actual delivery chain is longer: equipment, rack integration, facility space, power, cooling, network, installation, firmware, cluster software, burn-in, and buyer acceptance. Each dependency needs an owner and a date.
Facility questions for a material reservation
- Which site will host the system, and is the quoted power allocated to that hall and deployment?
- What cooling architecture serves the exact rack configuration? If liquid-cooled, who owns the CDU and facility-water responsibilities?
- What are the installation, energization, network, and commissioning milestones?
- Is the equipment live, on site, in transit, allocated, ordered, or dependent on a future purchase?
- Which party controls the equipment and which party controls the site?
- What redundancy and maintenance design applies to the buyer’s service boundary?
A provider can have a legitimate allocation and still miss the desired start date because the facility is the critical path. Omega Gradient records that distinction rather than treating all future capacity as equivalent.
Scale-up fabric, scale-out fabric, and storage all need acceptance criteria
NVLink and NVSwitch describe communication inside the scale-up domain. A cluster with multiple nodes or NVL72 racks still needs a scale-out network. NVIDIA positions Quantum-X800 InfiniBand and Spectrum-X Ethernet with its GB300 NVL72 platform; the provider must specify the actual network deployed, topology, link speed, RDMA design, isolation, and expected collective behavior.
Storage can be just as limiting. The brief should define dataset ingress, aggregate read throughput, metadata behavior, checkpoint frequency and size, restore path, local scratch, durable storage, and the cost and time to move data out. A provider benchmark that runs entirely in GPU memory does not validate the buyer’s end-to-end training pipeline.
Separate live, near-term, and development capacity
Omega Gradient classifies Blackwell options by delivery state:
- Live and testable: system is operating and can support a scoped technical validation.
- Installed or commissioning: equipment is on site but acceptance dependencies remain.
- Allocated and scheduled: equipment and site path are documented, with future milestones.
- Conditional development: delivery depends on financing, power, equipment, customer anchor, or another unresolved condition.
- Indicative only: market signal without sufficient evidence for a procurement decision.
That classification does not automatically exclude a future project. It determines the evidence, contract structure, deposit protection, milestone payments, and remedies the buyer should require.
Blackwell cluster acceptance
A suitable plan is architecture-specific. It should cover inventory and firmware, component health, sustained burn-in, thermals and power, scale-up link health, scale-out bandwidth and collective tests, storage throughput, orchestration, observability, tenant access, and a representative workload. For NVL72, acceptance should confirm the expected rack-scale system boundary rather than treating each compute tray as an unrelated node.
The contract should state the test versions and conditions, acceptance window, partial-delivery treatment, cure period, component replacement process, service-start adjustment, and the remedy if the delivered topology or performance materially differs from the agreed design.
Minimum Blackwell buyer brief
- Target platform: HGX B200, HGX B300, GB200 NVL72, GB300 NVL72, or architecture-flexible;
- GPU or rack count, target start window, ramp, term, and acceptable regions;
- Model, framework, precision, parallelism strategy, sequence/context characteristics, and expected utilization;
- Required scale-up and scale-out topology, network preference, and orchestration model;
- Dataset size, ingress, storage throughput, checkpoint pattern, and egress constraints;
- Security, tenancy, compliance, support, and managed-service requirements;
- Budget range, deposit or prepayment constraints, and acceptance requirements.
With that brief, Omega Gradient can compare credible paths instead of collecting generic “B200 available” claims.
NVIDIA, GB300 NVL72 overview and specifications; NVIDIA, GB200/GB300 data-center architecture; NVIDIA, HGX AI Factory components; AWS, B200 and B300 Capacity Blocks documentation. Verify current product, region, and availability details with the provider.