Direct answer

Before signing a reserved GPU cluster agreement, a buyer should be able to answer four questions without relying on inference: What exact system is being delivered? Who controls the equipment and facility? What evidence supports the start date and performance? What contractual remedy applies when delivery, acceptance, uptime, or support misses the promise?

A defensible procurement process covers nine connected areas: workload requirements, capacity provenance, facility readiness, node and fabric architecture, storage and software, operations, normalized economics, acceptance, and the final decision record. Missing facts do not automatically disqualify a provider, but they must be labeled as open, conditional, or represented—not silently upgraded into validated capacity.

Working method: print or save this page, mark every item as validated, represented, conditional, or open, and attach the evidence location. Re-score whenever the provider, facility, delivery date, or bill of materials changes.

Use a four-state evidence model

Binary “yes/no” diligence hides too much. A provider may give a credible answer that is not yet evidenced, or show an allocation that still depends on facility work. The following status model preserves that difference throughout sourcing and negotiation.

ValidatedPrimary evidence supports the claim and matches the offered system.
RepresentedThe accountable party states it, but documentary or technical proof is pending.
ConditionalThe outcome depends on a named milestone, third party, or unresolved dependency.
OpenThe answer, owner, evidence, or remedy has not been established.

1. Workload and system definition

Start with the useful work. A provider cannot design or price the right cluster when the brief stops at a GPU model and count.

  1. What workload must the cluster run? Record training, post-training, inference, rendering, simulation, or HPC; model scale; frameworks; precision; expected concurrency; and data behavior.
  2. What result defines success? Establish target time-to-train, tokens per second, latency, throughput, checkpoint time, or another workload-level measure.
  3. Which requirements are hard constraints? Separate mandatory GPU, memory, region, security, topology, start date, and term from preferences that can trade against price or availability.
  4. What utilization will actually be sustained? Model ramp, idle periods, maintenance, data preparation, experimentation, and demand variability before comparing reservation economics with on-demand or blended access.

2. Capacity control and delivery evidence

The commercial seller, equipment owner, facility operator, cloud operator, and support team may be different entities. Map the chain before relying on an availability claim.

  1. Who is the contracting entity? Confirm the legal name, jurisdiction, authority to sell the service, and the balance sheet standing behind refunds, credits, and damages.
  2. Who owns or controls the equipment? Ask whether the seller owns it, leases it, has a binding allocation, resells another operator, or is still assembling supply.
  3. What is the capacity state today? Classify the block as live and testable, installed, commissioning, allocated, ordered, financed, forecast, or dependent on another customer.
  4. What evidence matches this exact block? Request redacted serial inventory, rack or cluster identifiers, purchase or allocation support, facility confirmation, commissioning records, or a scoped test path—not a generic screenshot.

3. Facility and delivery readiness

An equipment allocation is not usable capacity until the site can power, cool, network, secure, commission, and operate the promised system.

  1. Where will the cluster run? Identify country, metro, facility, hall or deployment boundary, data-sovereignty implications, and whether relocation is permitted.
  2. Is power committed to these racks? Distinguish contracted utility capacity, energized capacity, reserved distribution, and power that still depends on construction or interconnection.
  3. Is the cooling design ready? For liquid-cooled systems, assign responsibility for CDUs, facility-water loops, water quality, leak response, monitoring, and remediation.
  4. What milestones lead to service? Put equipment arrival, installation, energization, network turn-up, burn-in, customer access, acceptance, and production start on one dated critical path.

4. Node, scale-up, and scale-out architecture

“H100,” “B200,” or “Rubin” is not a system specification. The useful cluster is defined by the entire communication path around the accelerator.

  1. What is the exact node or rack-scale bill of materials? Record accelerator form factor, GPUs per node or rack, GPU memory, CPU, host memory, local NVMe, NICs, DPUs, firmware, and OEM.
  2. What is the scale-up domain? Confirm NVLink and NVSwitch generation, bandwidth, GPU boundaries, and whether the offered partition preserves the topology assumed by the workload.
  3. What is the scale-out fabric? Specify InfiniBand or Ethernet generation, link speed per GPU or node, topology, rail design, RDMA configuration, adaptive routing, and congestion controls.
  4. Where is the oversubscription boundary? Ask for blocking ratio and topology diagrams across compute, storage, management, and external connectivity; “non-blocking” should name the measured boundary.

5. Storage, security, and software platform

GPU utilization often fails outside the GPU. Storage, identity, orchestration, observability, and image support need measurable service boundaries.

  1. Can storage feed the workload? Define usable capacity, aggregate and per-client throughput, metadata performance, checkpoint behavior, local scratch, durability, backup, and restore.
  2. How does data enter and leave? Document network paths, private connectivity, transfer windows, bandwidth caps, egress pricing, encryption, deletion, and end-of-term export.
  3. What isolation and access model applies? Confirm single tenancy or partitioning, identity provider integration, roles, audit logs, key management, vulnerability process, remote-hands controls, and compliance evidence.
  4. Who owns software compatibility? Freeze or bound driver, firmware, CUDA, framework, scheduler, container, image, observability, and upgrade responsibilities; define who cures regressions.

6. Operations, support, and service level

A large cluster is an operating relationship. The buyer needs named response paths and measurable responsibilities, not only an uptime percentage.

  1. What does the provider operate? Draw the boundary across facility, hardware, network, storage, base software, scheduler, Kubernetes or Slurm, images, security, and application support.
  2. How are failed components handled? Set spares strategy, GPU and node replacement targets, drain procedures, maintenance coordination, and the treatment of chronically degraded hardware.
  3. What is measured for the SLA? Define availability denominator, excluded events, planned maintenance, degraded performance, network and storage failures, reporting source, and dispute process.
  4. Who responds at 2:00 AM? Establish severity definitions, 24/7 channels, acknowledgement and restoration targets, escalation names, incident reports, and executive review cadence.

7. Commercial normalization and total obligation

Normalize the whole commitment. Two identical GPU-hour rates can produce different effective costs once minimums, deposits, ramp, support, and data movement are included.

  1. What exactly is the billing unit? Confirm whether the quote is per GPU, node, rack, cluster, reserved block, or consumed hour; define calendar hours, partial availability, and rounding.
  2. What cash moves before acceptance? Record deposit, prepayment, milestone payment, letter of credit, security, refund conditions, and the party holding buyer funds.
  3. Which costs sit outside the headline rate? Include storage, bandwidth, egress, IP addresses, support, licenses, remote hands, setup, taxes, power adjustment, and cross-connects.
  4. What changes during the term? Model ramp, take-or-pay, usage minimums, renewal, indexation, hardware substitution, expansion, contraction, assignment, and early termination.

8. Acceptance, cure, and remedies

Acceptance is where sales language becomes an operational obligation. The test plan should be agreed before delivery and attached to the contract.

  1. What converts delivery into acceptance? Require inventory and configuration verification, burn-in, GPU health, NVLink, fabric, storage, security, scheduler, and access tests.
  2. Does the test reflect the workload? Add representative collectives, framework tests, checkpointing, or application benchmarks with dataset, configuration, duration, and reproducible pass criteria.
  3. What happens when only part passes? Define partial acceptance, replacement, retest, cure period, billing start, schedule relief, and whether a smaller topology can be forced on the buyer.
  4. What remedy matches each failure? Separate late delivery, failed acceptance, chronic degradation, downtime, security events, specification changes, and abandonment; pair each with refund, credit, replacement, or termination rights.

9. Decision record and red flags

The final memo should make the trade-off legible to technical, finance, security, and legal reviewers—and preserve what was known when the commitment was approved.

  1. Which facts are still unresolved? Carry every open and conditional item into the approval record with an owner, evidence request, deadline, and consequence.
  2. What is the credible alternative? Compare the preferred reservation with a second provider, a mature GPU generation, delayed start, smaller base plus burst, and on-demand capacity.
  3. What would cause a no-go? Predefine red lines such as unclear equipment control, unverifiable site, payment before evidence, refusal to attach specifications, no acceptance rights, or remedies that cannot be collected.
  4. Who signs each risk? Require explicit approval from the owners of technical architecture, security, finance, legal, operations, and the business workload—not a blended “team approved” status.

How to compare shortlisted clusters

Do not reduce the scorecard to one total. A high aggregate score can hide a fatal dependency. First apply hard gates: legal seller identified, capacity path evidenced, site and architecture defined, economics normalized, acceptance attached, and remedies enforceable. Then compare the surviving options across workload performance, delivery confidence, operating model, total cost, flexibility, and counterparty exposure.

Red flags deserve asymmetric weight. A provider may have an excellent technical design but still be unbuyable if it cannot demonstrate control of the capacity. Another may control live equipment but fail the workload because the fabric or storage boundary is wrong. The purpose of diligence is not to make every option look comparable; it is to reveal when they are not.

A reserved cluster is ready to buy when the system, evidence, obligations, and remedies describe the same capacity.

What should be attached to the contract?

At minimum, attach the final bill of materials, topology, facility region, capacity and tenancy definition, delivery milestones, acceptance plan, service description, support matrix, SLA measurement method, pricing schedule, excluded charges, data return or deletion procedure, security responsibilities, substitution rules, and remedies. If a material diligence answer lives only in email or a sales deck, convert it into the agreement or explicitly accept that it is not an obligation.

Omega Gradient uses this framework to structure demand briefs, qualify providers, normalize offers, and produce an evidence-marked shortlist. For the underlying method, read how Omega Gradient qualifies reserved GPU clusters. For platform-specific considerations, see the H100 and H200 guide, Blackwell guide, and Vera Rubin buyer guide.

Scope note
This is a buyer-side operational framework, not legal, engineering, security, tax, or financial advice. Adapt the questions, evidence threshold, tests, and contract language to the workload, jurisdiction, risk level, and provider structure.