Direct answer

Choose an H100 or H200 cluster by useful workload performance, memory fit, delivery confidence, and total commitment—not by accelerator name or headline GPU-hour rate alone. H200’s 141 GB HBM3e footprint can reduce memory pressure for large models and memory-bound inference, while H100 remains a broadly deployed and often easier-to-source platform. In either case, node topology, fabric, storage, and operational consistency determine whether a distributed workload scales.

Omega Gradient sources and compares reserved Hopper clusters for buyers that need a defined count, region, start date, topology, or term. On-demand options remain available for benchmarking, experiments, and flexible workloads.

H100 vs H200: start with memory behavior and workload economics

NVIDIA’s DGX documentation describes eight-GPU systems with 640 GB total H100 memory or 1,128 GB total H200 memory. At the accelerator level, NVIDIA lists H100 SXM with 80 GB HBM3 and H200-based systems with 141 GB per GPU. That larger memory footprint can reduce model partitioning, enable larger batches or contexts, or keep more of an inference workload resident. It does not guarantee that every H200 workload is faster or cheaper.

Decision areaH100H200
MemoryCommon H100 SXM systems provide 80 GB per GPU; H100 NVL is a different PCIe product141 GB HBM3e per GPU in common H200 systems
Best reason to chooseBroad deployment, mature operating experience, potentially wider supplyWorkloads that benefit from larger memory and higher memory bandwidth
Validation focusExact variant, node topology, fabric, and system consistencyMeasured memory benefit, software readiness, node topology, and fabric
Commercial questionDoes availability or rate outweigh the value of more memory?Does the workload save enough nodes, time, or engineering complexity to justify the premium?

A useful comparison runs the buyer’s model or a representative benchmark under the intended precision, parallelism, sequence length, batch behavior, and storage path. The economic unit is cost per completed training objective, fine-tuning run, inference token at the target latency, or scientific result—not the GPU-hour in isolation.

“H100” is not a complete hardware description

H100 appears in SXM systems and in H100 NVL/PCIe configurations. NVIDIA lists different memory and NVLink characteristics for those products. An eight-GPU HGX or DGX-style node with NVSwitch is not interchangeable with independent PCIe GPUs connected only through host PCIe and the scale-out network.

The quote should state exact accelerator variant, GPUs per node, CPU, host memory, local NVMe, NICs, NVLink/NVSwitch topology, and the network between nodes. If the seller cannot provide that information, the buyer does not yet have a cluster design to compare.

Fabric and storage determine distributed efficiency

NVIDIA cites 900 GB/s GPU-to-GPU interconnect for H100 SXM through fourth-generation NVLink, while DGX H100/H200 systems use NVSwitch inside the eight-GPU node. Multi-node scaling still depends on the external fabric. The provider should specify InfiniBand or Ethernet, generation and link speed, topology, blocking ratio, RDMA, network isolation, and the expected behavior of NCCL or other collectives.

Storage should be qualified against the buyer’s data and checkpoint pattern. Important facts include usable capacity, aggregate and per-node read/write throughput, metadata performance, local scratch, object-store path, checkpoint frequency, restore behavior, ingress, and egress. A cluster can pass a GPU microbenchmark and fail the actual training pipeline because the data path is undersized.

Normalize the whole quote

Hopper is available through hyperscalers, neoclouds, private-cloud operators, bare-metal providers, scheduled-capacity products, marketplaces, and resellers. Each channel packages the system differently. Omega Gradient compares:

  • effective GPU-hour economics at the contracted count and billed hours;
  • term, ramp, deposit, prepayment, taxes, currency, and minimums;
  • included storage, bandwidth, public IPs, support, management, and software;
  • exact node and fabric design, orchestration, access, and tenant isolation;
  • delivery state, dependencies, acceptance window, service start, and remedies;
  • renewal, price reset, termination, portability, and end-of-term data handling.

Public on-demand pricing is useful as a market reference, not as a substitute for a reserved quote. Omega Gradient’s daily GPU price index and provider pricing report help establish the visible market; reserved pricing still depends on count, term, start date, system, region, and commercial structure.

Availability needs a date and an evidence status

A Hopper cluster may be live, allocated to another customer until a future date, assembled from inventory across sites, or planned. The diligence record should state what is available, where, from when, in what topology, and under whose control.

A provider representation can be commercially useful before full validation, but it should remain attributed. For a material deposit or long commitment, the buyer may need current system evidence, a controlled test, chain-of-control documentation, facility confirmation, or milestone-based payment and termination protection.

H100/H200 acceptance should test the contracted system

  1. Inventory: count, exact variant, memory, node configuration, NICs, storage, firmware, and access.
  2. Health: component diagnostics, error history, thermals, sustained burn-in, and spare/replacement process.
  3. Scale-up: NVLink/NVSwitch health and topology inside each node.
  4. Scale-out: link state, bandwidth, latency, RDMA, and collective performance across the reserved block.
  5. Storage: read/write and metadata behavior under realistic concurrency, plus checkpoint and restore.
  6. Workload: an agreed benchmark or representative job using versioned drivers, libraries, and configuration.

The contract should define pass criteria, test duration, partial failure, cure period, replacement, service-start adjustment, and the buyer remedy if the system does not match the agreed design.

Minimum Hopper cluster brief

  • H100, H200, or flexible; acceptable form factor and GPUs per node;
  • total GPU count, region, target start window, ramp, and term;
  • training, post-training, inference, or HPC workload and current benchmark data;
  • required scale-up topology and scale-out fabric;
  • dataset, storage throughput, checkpointing, ingress, and egress;
  • bare metal, private cloud, Slurm, Kubernetes, or managed-service preference;
  • security, compliance, support, budget, and commercial constraints.

If the team does not yet know the right platform, begin with a measured on-demand benchmark and use the result to size the reservation.

Primary references
NVIDIA, H100 specifications; NVIDIA, DGX H100/H200 system architecture; NVIDIA, H200 NVL reference architecture; AWS, H100/H200 Capacity Blocks documentation. Verify current product, region, and availability details with the provider.