Choose a GPU cluster network against the workload, server configuration and scale you need to operate. InfiniBand and Ethernet can both support distributed GPU communication, but they require different equipment and operating choices. Compare a complete, tested fabric covering adapters, switches, interconnects, software and support rather than selecting by port speed alone.

An offer for GPU servers can leave the network almost entirely outside its scope. That omission matters when a workload must exchange data across several machines. The buying question is whether the supplied system can deliver the required service with the network you can deploy and operate.

Use SecondWatt's GPU and accelerator catalogue to identify the compute platform, then specify its external network as a connected purchasing decision.

Key Takeaways

  • Select the network against the actual distributed workload and server configuration.
  • Separate compute, storage, application access and management traffic.
  • Count logical ports, connector cages, uplinks and complete cable assemblies separately.
  • Require a supported software and congestion-management design with an operating owner.
  • Accept the delivered cluster through reproducible workload and failure-response tests.

Start with the traffic your workload creates#

Describe how work is divided between GPUs and servers. A model running within one machine can have a different external-network requirement from a training job or inference service that communicates across many machines. Identify which operations cross the server boundary and which remain inside it.

NVIDIA's NCCL overview describes collective and point-to-point communication across GPUs, including communication within and between nodes. That is a useful starting point for an NVIDIA workload: the relevant measure is how the application's communication behaves on the offered system.

Record model size, precision, parallelization method, batch or concurrency settings and the software stack used for evaluation. Then choose an acceptance measure such as job completion time or throughput at an agreed latency and quality target. A network benchmark can help explain a result, but it is a different measure from application delivery.

Questions that distinguish the two purchasing paths#

Decision InfiniBand proposal should explain Ethernet proposal should explain
Workload support Supported adapters, fabric and communication software Supported adapters, transport and communication software
Traffic management Routing, congestion behavior and fabric management Routing, congestion behavior and relevant Ethernet features
Existing environment Integration with storage, management and outside networks Integration with existing Ethernet and any dedicated GPU fabric
Operations Required expertise, tools and support ownership Required expertise, tools and support ownership
Acceptance Application result and behavior during agreed events The same application requirement and event cases

The InfiniBand Trade Association's RoCE initiative identifies RDMA over Converged Ethernet as an Ethernet approach to remote direct memory access. Therefore, RDMA capability alone does not select InfiniBand. The actual implementation, supported combination and measured workload result need to determine the shortlist.

Ask bidders to explain why their proposed fabric fits the requirement. Avoid a universal rule that one transport wins every training or inference deployment.

Separate the networks before counting equipment#

Create a traffic map before requesting switch quantities. Compute communication, storage access, application traffic and out-of-band management have different purposes. Combining them may be a deliberate design choice, but the offer should identify the consequences and any separation required.

NVIDIA's published DGX H100 SuperPOD component reference provides a named example. It describes separate compute, storage, in-band management and out-of-band management fabrics. That reference architecture should not be assigned automatically to every H100 server or another GPU generation.

A named architecture reference to interpret carefully#

Reference field Published DGX H100 SuperPOD arrangement Procurement boundary
Compute fabric connection Eight NDR400 connections per system Specific reference architecture, not a universal requirement for H100
Compute fabric switch NVIDIA Quantum QM9700 Match the actual switch and supported configuration
Storage fabric Separate InfiniBand fabric Include the required storage-side equipment
In-band management Ethernet network Define access and services separately
Out-of-band management Separate management network Preserve required administrative access

Draw the external endpoints around the servers: storage systems, management servers, security boundaries and connections to the rest of the facility. Allocate responsibility for each interface. A compute-fabric quotation can be complete within its own scope while still leaving necessary site equipment unpriced.

Also distinguish the internal GPU interconnect from the external fabric. An internal NVLink arrangement does not make an external switch, adapter or optical module unnecessary. Ask the server supplier to identify which ports are present, which are usable for the intended design and which components remain to be supplied.

Count logical ports and bandwidth in the same direction#

Switch names, front-panel cages and logical ports can describe different quantities. Begin with a port map that identifies each endpoint, speed, connector, breakout arrangement and role. Count the connections required by the offered server configuration, including infrastructure endpoints beyond the GPU servers.

Illustrative leaf-switch bandwidth calculation#

Item Assumption One-direction aggregate line rate
Downlinks 32 links at 400 Gb/s 12.8 Tb/s
Uplinks, case A 32 links at 400 Gb/s 12.8 Tb/s
Uplinks, case B 16 links at 400 Gb/s 6.4 Tb/s

The arithmetic is 32 multiplied by 400, or 12,800 Gb/s. Case A has a 1:1 downlink-to-uplink line-rate ratio at this leaf boundary. Case B has a 2:1 ratio because 12.8 divided by 6.4 equals 2.

2:1 — the illustrative downlink-to-uplink line-rate ratio when 32 equal-speed downlinks share 16 equal-speed uplinks.

This is capacity bookkeeping for a hypothetical leaf, not a complete fabric design or a prediction of application throughput. Traffic pattern, routing, contention, protocol overhead and the rest of the topology still matter. Do not add transmit and receive directions on one side while using only one direction on the other.

For another unit check, a nominal 400 Gb/s divided by eight is 50 GB/s before accounting for the limits of the actual transfer. Keep bits and bytes explicit when comparing adapter, storage and application figures.

The current Arista 7060X6 datasheet gives a concrete example of why the front-panel specification matters.

Named switch connector references#

Model Published high-speed connector arrangement Buyer check
Arista 7060X6-64PE 64 OSFP 800G connectors Supported port modes and complete interconnect bill of materials
Arista 7060X6-32PE 32 OSFP 800G connectors Same checks for this specific model

Those connector counts do not establish a populated optical configuration, a supported breakout or compatibility with every server interface. Have the supplier reconcile the switch ordering codes with the end-to-end port map.

Specify every interconnect and its receiving endpoint#

Require an interconnect schedule detailed enough to assemble the delivered network. For each link, identify both endpoints, the required operating mode, the connector at each end, the cable medium and the installed route length. Include patching where the design uses it.

NVIDIA's DGX H100 SuperPOD network-fabric documentation maps its compute-fabric transceiver arrangement to the system connections. This illustrates why a visible cage should not automatically be counted as one logical network connection.

Interconnect purchasing schedule#

Field Evidence to request
Endpoint A and B Exact system, port and connector designation
Link mode Required speed, lanes and supported breakout
Assembly Direct-attach cable, active cable or specified optical arrangement
Reach and routing Installed path, patching and manufacturer-supported reach
Compatibility Supported part numbers at both endpoints
Quantity Deployed count, installation allowance and agreed spares
Identification Labels matched to the final port map

A detachable optical link may require modules at both ends plus the fiber path. A cable assembly can package the endpoints differently. Count the actual ordered objects so a quotation does not omit half of a connection or price the same component twice.

Confirm the airflow direction and service access of switches beside the cable layout. Cable routing should leave the equipment maintainable in its installed position. A configuration that works on a test bench may need different interconnect lengths, management accessories or rack placement at the receiving facility.

For used network equipment, obtain an inventory of installed components and supported software versions. Ask whether an offered chassis includes the power supplies, fans, licenses and management features required by the proposed design.

Make software and operational ownership part of the offer#

A switch and adapter inventory is only part of the system. Request the operating-system, driver, firmware and communication-library versions used for validation. Record the settings that materially affect the result and identify who will maintain the supported combination after handover.

The joint Arista and Broadcom RoCE deployment guide covers switch and adapter configuration, congestion notification, traffic management and verification for its documented combinations. It demonstrates that an Ethernet GPU fabric requires a deliberate configuration. Its example settings should not be copied into an unrelated installation without review.

For InfiniBand, ask who supplies and operates fabric management, routing policy and monitoring. For Ethernet, ask who owns the corresponding routing, traffic-management and observability design. Each proposal should explain how the operators will detect degraded links and distinguish network problems from server or workload problems.

Software and service handover#

Deliverable Acceptance question
Version manifest Which complete combination was validated?
Configuration record Which settings must remain consistent?
Management access Can the buyer administer the delivered equipment?
Monitoring Which link, congestion and error signals are visible?
Support entitlement Which organization supports each component and integration issue?
Change process What must be retested after a material update?

Check license and service rights for the actual transaction. Original equipment ownership does not establish every software entitlement the new buyer needs. Keep the seller's commitments, OEM terms and integration support scope distinct in the purchase record.

Arrange an operational handover using the delivered configuration. The team receiving the equipment should have usable records, access and a support path, including a way to recover the documented configuration after replacement or maintenance.

Test the assembled cluster and the facility interfaces#

Agree an acceptance plan before the demonstration. Identify the tested servers, adapters, switches, interconnects and software. State the workload, data, settings, duration and acceptance threshold. Keep a hardware and configuration manifest with the results.

Test the communication patterns the production workload will use. Retain application results alongside network telemetry so reviewers can explain whether a limit comes from compute, storage, communication or configuration. A single successful link test cannot establish the behavior of an assembled multi-server system.

For a phased deployment, define the acceptance boundary for each delivery. Record which switches, servers and links participated, and what remains to be tested when the next phase joins. Evidence from an initial fabric should not silently become evidence for a later topology with additional endpoints or different routing.

If the delivered batch contains different adapter, optical-module or firmware revisions, identify the supported combinations rather than treating a shared marketing name as equivalence. Agree which changes require renewed testing. Retain the final inventory with the accepted results so replacement decisions can start from the configuration that was actually demonstrated.

Acceptance cases for the delivered fabric#

Case Evidence to retain
Initial discovery All required endpoints and links visible at the agreed modes
Baseline workload Reproducible application result and system configuration
Sustained operation Throughput, latency where relevant, and error or congestion telemetry
Concurrent activity Behavior with the agreed storage and background traffic
Controlled failure case Observed effect and recovery under an approved test plan
Final handover Exceptions, remedies, records and accepted configuration

Include the facility in the readiness review. Switches, optical modules, storage and management equipment add electrical load and heat beyond the GPU servers. Request complete equipment input and environmental requirements; avoid estimating the network's facility demand from switch throughput.

The transformer-sizing guide supports the upstream load review, while the UPS sizing guide helps frame the protected-load requirement. Neither replaces the offered network equipment's own electrical documentation.

Connect rack readiness and the network acceptance date to the site's power procurement roadmap. The intended service requires compute, network, storage and facility interfaces to be ready together.

Request a complete GPU cluster network#

Normalize offers around the same deployment boundary. Separate server adapters, switches, interconnects, storage connectivity, management systems, licenses, installation, testing and support. Require the respondent to list exclusions and explain proposed substitutions.

GPU cluster networking RFQ#

Requirement Buyer input
Compute Exact server configurations and deployment phases
Workload Application, parallelization and measurable service target
Fabric options Accepted transports and constraints from existing infrastructure
Topology Required endpoints, expansion plan and failure cases
Interconnects Rack layout, link routes and supported connector requirements
Facility Power, cooling, rack and operating constraints
Acceptance Workload, evidence, ownership and handover requirements

Ask each bidder to price its complete recommended configuration and explain the tradeoffs. Where the workload is not yet fixed, identify what is provisional and request a validation step before the final equipment schedule is committed.

SecondWatt acts as an independent intermediary. To request GPU cluster equipment sourcing, provide the server platform, quantity, workload, intended network arrangement, destination and deployment date. State whether you need compute hardware alone or help sourcing the surrounding network equipment, with the proposed configuration subject to confirmation.

FAQ: InfiniBand and Ethernet for GPU clusters#

Is InfiniBand always the right choice for AI training?#

No universal choice follows from the training label. Evaluate the actual communication pattern, scale, supported server configuration and operating requirements. Compare complete proposals against the same workload and acceptance measure. An established reference architecture can reduce uncertainty for its documented scope, but it does not decide every purchasing case.

Does an Ethernet connection automatically support the required RDMA workload?#

No. Confirm the adapter, transport, switch features, software and configuration required by the intended deployment. RoCE provides an Ethernet path to RDMA, but the end-to-end system still needs validation. Request the supported combination and measured result instead of treating a nominal link speed as evidence that the workload will perform.

Can I count one switch cage as one server connection?#

Only after checking the supported port mode and interconnect arrangement. Some configurations expose multiple logical connections through one physical cage. Others require specific optical modules, breakouts or adapters. Build the quantity schedule from the documented endpoint map and actual part numbers, including the components needed at both ends.

What does a 2:1 oversubscription ratio tell me?#

At a stated boundary, it means the summed downlink capacity is twice the summed uplink capacity when measured on the same basis. It does not predict that every job will run at half speed. Traffic patterns, concurrency, routing and the rest of the system determine how that boundary affects a workload.

Is a network benchmark enough to accept the cluster?#

It provides evidence for the test performed. Acceptance should also include the intended application, relevant storage behavior, monitoring and agreed failure-response cases. Preserve the hardware and software manifest with the results. The buyer needs to know which delivered configuration produced the measured service and which conditions remain outside the test.

What should a used network-equipment offer include?#

Request exact switch and adapter identities, installed components, supported versions, interconnect quantities and condition evidence. Confirm access, licenses and service terms for the new owner. Require a complete configuration and test plan so missing optics, software rights or integration work remain visible before the purchasing decision.