The core principle: claims, evidence, and obligations are different things
A capacity market moves faster than traditional enterprise procurement. That speed creates useful access, but it also creates ambiguity. A provider may control a live cluster, hold an equipment allocation, have a right to resell another operator’s capacity, or be marketing a build that still depends on financing, power, cooling, networking, and delivery. All four can be described casually as “available.” They are not the same risk.
Omega Gradient does not turn an indicative supply signal into a fact. We keep the opportunity visible and label the evidence status.
The diligence depth should match the commitment. An eight-GPU on-demand test does not require the same evidence package as a multi-year reserved cluster. But as deposits, term, GPU count, deployment complexity, and business dependency increase, the buyer needs a clearer chain of control and a stronger acceptance plan.
1. Capacity provenance and chain of control
The first question is not “what is the price?” It is “who controls the system being quoted?” We map the relevant parties: equipment owner, facility owner, data-center operator, cloud operator, managed-service provider, reseller, contracting entity, billing entity, and support owner. Sometimes one company fills every role. Sometimes four companies sit between the buyer and the hardware.
For live capacity, useful evidence can include a current inventory export, management-plane view, serial-number sample, facility-specific topology, test access, or other artifacts that tie the represented system to a location and operator. For future capacity, the evidence changes: equipment purchase or allocation, delivery schedule, facility space, power and cooling readiness, network plan, installation milestones, and the conditions that could move the date.
Questions we want answered
- Is the capacity live, allocated, under installation, financed, ordered, or forecast?
- Who owns the GPUs, who operates the service, and who signs the buyer contract?
- Does the seller have a direct right to allocate the capacity, or is another approval required?
- Is the quote exclusive to one channel, or is the same capacity being marketed through several intermediaries?
- Which dependency—equipment, power, cooling, networking, software, or financing—sets the delivery date?
2. Facility, power, cooling, and network readiness
Modern GPU systems are facility projects. This becomes especially important for Blackwell rack-scale infrastructure, where power density, liquid cooling, rear-door heat rejection, network cabling, and commissioning can become the critical path. A purchase order for GPUs does not prove a site can energize and operate them.
We ask for the site and readiness facts that materially affect delivery: region and data sovereignty, available power, redundancy design, cooling architecture, rack density, network carriers, cross-connect timing, physical security, relevant certifications, and commissioning milestones. The goal is not to perform an engineering audit from a PDF. It is to identify whether the provider’s schedule and service claim are consistent with the facility plan.
3. System architecture and workload fit
The accelerator name is the beginning of the technical description. A decision-ready cluster brief should identify the exact platform and the boundaries that govern workload scaling:
- Node: GPU form factor, GPUs per node, CPU architecture, host memory, local storage, NICs, and firmware baseline.
- Scale-up domain: NVLink and NVSwitch connectivity inside a node or rack-scale system.
- Scale-out fabric: InfiniBand or Ethernet generation, link speed, topology, blocking ratio, RDMA configuration, and network isolation.
- Storage: usable capacity, read and write throughput, metadata behavior, checkpoint target, object storage, and data-transfer constraints.
- Software: driver, CUDA and container baseline, orchestration, scheduler, observability, and supported operating model.
We compare these facts to the workload. Training needs can be communication-heavy and checkpoint-sensitive. Large-model inference may be constrained by memory capacity, memory bandwidth, latency targets, batching, or inter-node traffic. HPC workloads can depend on precision, host-to-device movement, and a specific MPI or storage pattern. “Eight GPUs” is not a portable unit across those cases.
4. Operating model, support, security, and service boundary
A buyer needs to know where provider responsibility stops. The service may be bare metal with buyer-managed orchestration, a private cloud, managed Kubernetes or Slurm, or a fully managed cluster. That choice changes the staffing burden, incident path, patch responsibility, observability, and time to remediate failed components.
Operational diligence covers access model, tenant isolation, identity and key management, network controls, logging, monitoring, maintenance windows, spares, replacement process, support hours, escalation contacts, and security evidence. Certifications can help, but they do not replace clarity on the actual service boundary.
5. Commercial normalization
Reserved quotes often use incompatible units and include different services. Omega Gradient converts each option into a comparable view without discarding the terms that matter. At minimum, the matrix records:
- GPU count, rate basis, billed hours, ramp schedule, term, and effective GPU-hour cost;
- deposit, prepayment, credit support, taxes, currency, and payment timing;
- included storage, bandwidth, IP space, support, management, and software;
- delivery dependencies, acceptance window, renewal, termination, and assignment;
- SLA measurement, exclusions, maintenance, credits, replacement, and chronic-failure remedies.
The cheapest normalized rate can still be the wrong choice. A higher-priced cluster may have a stronger fabric, a cleaner chain of control, a better acceptance remedy, or a start date with fewer external dependencies. The buyer output should make that trade-off explicit instead of hiding it inside a single score.
6. Acceptance turns delivery into a measurable event
A contract should define what the buyer will receive and how both sides will determine that it works. “GPUs online” is too vague for a material reservation. The acceptance plan should be agreed before the service start date, not invented after a problem appears.
Typical acceptance layers
- Inventory and access: contracted node count, accelerator model, memory, CPU, NICs, storage, tenant access, and management-plane visibility.
- Health and burn-in: component health, error logs, thermals, power behavior, firmware consistency, and a defined burn-in window.
- Fabric: topology validation, link state, bandwidth, latency, collective performance, and failure isolation.
- Storage: capacity, throughput, metadata performance, checkpoint and restore behavior, and data path.
- Workload-relevant test: an agreed benchmark or representative job with versioned software and repeatable conditions.
- Remediation: cure period, replacement process, partial acceptance, service-start adjustment, credit, or termination right.
NVIDIA’s own system guidance demonstrates why system boundaries matter: its DGX H100/H200 documentation describes eight-GPU systems with NVSwitch, while newer rack-scale Blackwell systems extend the scale-up domain across an entire rack. Acceptance must test the architecture that was sold, not a generic GPU benchmark.
Evidence classification: what we know and how we know it
| Status | Meaning | How it appears in a buyer brief |
|---|---|---|
| Validated | Supported by current, relevant evidence or direct technical validation. | Presented as a fact with date, source, and scope. |
| Represented | Stated by the provider or channel but not independently validated to the same depth. | Attributed to the representing party. |
| Conditional | Depends on a future milestone, third party, buyer action, or commercial close. | Shown with the dependency and target date. |
| Open | Material question remains unanswered or evidence is incomplete. | Listed as an explicit diligence item, not silently filled. |
| Contradicted | Available evidence conflicts with the representation. | Escalated and excluded from the validated view until resolved. |
What the buyer receives
The useful output is a decision package, not a directory. Depending on the mandate, that can include a normalized option matrix, provider and chain-of-control map, technical gap list, evidence register, delivery dependency map, acceptance checklist, commercial issues list, and a recommendation framed by the buyer’s priorities.
Omega Gradient does not publish private provider evidence, operator names, buyer requirements, or live commercial terms as SEO content. Public intelligence explains the framework. Confidential diligence supports the actual transaction.
NVIDIA, DGX H100/H200 system architecture; NVIDIA, GB300 NVL72 architecture; AWS, Capacity Blocks for ML. These sources establish product and reservation mechanics; the diligence classifications and workflow are Omega Gradient’s methodology.