Choose an H200, B200 or B300 server by testing a complete configuration against your workload, memory needs and facility constraints. Compare measured service performance and total deployment scope. GPU generation, theoretical compute and memory capacity can narrow the search, but they cannot independently establish which offered system delivers the required result.
The useful buying question is specific: which documented server can run the intended workload, meet its acceptance target and fit the receiving facility on the required schedule? That question can produce a different answer from simply selecting the accelerator with the largest published specification.
SecondWatt's GPU catalogue provides the model research. This comparison concentrates on complete servers. It does not treat loose accelerators, baseboards and integrated NVL72 racks as interchangeable purchases, and it does not claim that these three choices exhaust every current platform option.
Key Takeaways
- Compare named server configurations instead of transferring specifications between OEM implementations.
- Size memory for the workload's full operating footprint, including runtime requirements.
- Evaluate throughput together with latency, quality, software and test conditions.
- Use complete-system electrical and cooling requirements for facility planning.
- Request a scoped offer with acceptance evidence, support terms and documented exclusions.
Compare named systems on a defined boundary#
For a transparent reference, the table below uses NVIDIA's DGX implementations. These are specific systems, not generic requirements for every H200, B200 or B300 server. The figures were checked against the manufacturer guides in September 2026.
Named DGX reference configurations#
| Reference system | GPUs per system | Published aggregate GPU memory | Published system maximum electrical input |
|---|---|---|---|
| DGX H200 | 8 | 1,128 GB | 10.2 kW |
| DGX B200 | 8 | 1,440 GB | 14.3 kW |
| DGX B300 | 8 | 8 × 288 GB; advertised as 2.3 TB | 15 kW |
Sources: the DGX H100/H200 guide, DGX B200 guide and DGX B300 guide. Use the complete offered configuration and its applicable manual when preparing the purchase specification.
The B300 guide separately lists power consumption and a higher system maximum in its power-specification section. This table uses the expressly labeled system maximum. Neither field is a prediction of average production demand, and neither should be replaced with the sum of power-supply nameplates.
Aggregate GPU memory also needs interpretation. It is distributed across devices; using it for one workload requires the appropriate software and partitioning strategy. The total should not be presented as one automatically available memory space for an unmodified application.
Use the H200 dossier, B200 dossier and B300 dossier to identify variants. Do not assign the DGX table to a different OEM server without checking that system's documentation.
Define the workload before choosing the generation#
Write the intended service in terms the testing team can reproduce. For inference, specify the model revision, precision, context length, concurrency, batching and required output quality. For training or other accelerated computing, specify the dataset, software, precision, parallelism and completion requirement relevant to that application.
Separate a hard requirement from a preference. A model that must fit within a particular device arrangement imposes a different constraint from a workload that can be divided across more systems. Likewise, an existing software qualification can matter more than a theoretical gain that requires an unplanned migration.
Workload brief for comparable server proposals#
| Field | Information to specify |
|---|---|
| Application | Model or software version and intended production use |
| Quality | Required accuracy or output-quality criterion |
| Performance | Throughput, latency or completion-time requirement |
| Memory | Expected model, runtime and concurrency footprint |
| Parallelism | Accepted device and multi-system arrangement |
| Software | Required operating system, framework and deployment stack |
| Operations | Monitoring, support and recovery expectations |
Treat the software environment as part of the result. A benchmark produced with a different precision or optimized implementation may be relevant, but the buyer needs to know whether that implementation is acceptable and maintainable for the intended service.
If production data cannot be shared, define a representative test case and document its limits. A restricted trial can still resolve memory fit, deployment effort or a known bottleneck. It should not be described as complete production validation when it leaves important workloads or operating conditions untested.
Estimate memory, then validate the complete footprint#
Start with the model's storage requirement, then account for the runtime and workload arrangement. The Hugging Face Accelerate memory estimator offers a way to examine model-memory estimates by data type. Such an estimate is a planning input, not a guarantee that an application will fit and perform under production concurrency.
Consider a deliberately simplified illustration. An assumed model with 70 billion parameters stored at two bytes per parameter requires 140 billion bytes for those parameters alone. In decimal units, that is 140 GB. It excludes runtime buffers, activations, cache, framework overhead and any other application state.
Illustrative parameter-storage arithmetic#
| Assumption | Calculation | Result |
|---|---|---|
| 70 billion parameters, two bytes each | 70 billion × 2 bytes | 140 GB, decimal |
| Same parameter count, one byte each | 70 billion × 1 byte | 70 GB, decimal |
These are storage calculations, not validated model configurations or promises about model quality at a lower precision. A GPU memory figure close to the first result does not establish adequate operating headroom. The implementation, workload and memory distribution must be tested.
For inference, document how context and concurrency affect the runtime footprint. For training, include the relevant optimizer, gradient and activation requirements. Ask the application team which values come from measurement and which remain estimates so the hardware comparison does not hide uncertainty behind a single memory number.
Where a larger-memory system reduces the number of devices required, evaluate the resulting workload arrangement directly. It may simplify one deployment, while another application remains limited by compute, communication or input data. Memory capacity can remove a constraint without guaranteeing the largest overall performance improvement.
Compare useful performance under reproducible conditions#
Use a buyer-defined acceptance target to evaluate the offered systems. The MLCommons data-center inference framework provides structured benchmark scenarios and quality conditions. Published results are useful when their configuration and scenario match the question being asked; they do not automatically predict another server's production result.
Request the actual test configuration, software versions, workload settings, measurement period and relevant telemetry. Keep the quality criterion constant when comparing throughput. A faster result obtained by changing the service requirement answers a different buying question.
Performance evidence to compare#
| Evidence field | Why the buyer needs it |
|---|---|
| Exact hardware | Connect the result to the offered configuration |
| Software and settings | Explain reproducibility and optimization differences |
| Workload definition | Establish what the measured result represents |
| Quality and latency | Prevent throughput from hiding a failed service requirement |
| Sustained operation | Reveal behavior beyond a short demonstration |
| Power and thermal conditions | Explain limits or throttling during the run |
| Exceptions | Identify failures, retries or excluded observations |
If two systems need different optimized settings, document them and compare the delivered service against the same acceptance requirement. Do not force identical settings when that would misrepresent a supported architecture, but do make the implementation differences visible to the buyer.
A short test should also have a stated purpose. It might screen an unsuitable configuration before a longer trial. It should not be promoted into a reliability claim or a promise that the same performance will persist across all models, input distributions and future software versions.
Record the amount of work that actually meets the service requirement. For an inference test, exclude neither failed requests nor latency exceptions without explaining the treatment. For a batch workload, state whether setup and data-loading time are included. These choices can change the apparent purchasing result even when the underlying hardware is unchanged.
Repeat the agreed workload enough to investigate material variability. The goal is to understand whether the observed difference is stable under the stated conditions, not to run an arbitrary number of benchmark repetitions. If the result changes with thermal state, background work or software settings, resolve that cause before treating a small difference as a reason to choose one system.
Check power, cooling and rack fit together#
The named-system table provides reference maximum input figures. The facility team needs the offered server's actual installation documentation, redundancy behavior, connector requirements, environmental limits and expected operating profile before approving deployment.
GPU power, server input and facility input describe different boundaries. Distribution losses, network equipment and cooling auxiliaries can sit outside the server figure. Keep those loads separate when building the project estimate so they are neither omitted nor counted twice.
Supermicro's accelerator portfolio includes different air- and liquid-cooled system implementations. That is why a GPU generation alone cannot determine the cooling architecture. Ask for the particular server model and its required facility interfaces.
Facility screen for the offered system#
| Interface | Required confirmation |
|---|---|
| Rack space and loading | Dimensions, weight and service access |
| Electrical input | Voltage, maximum input and connection arrangement |
| Power resilience | Behavior under the project's defined supply-loss case |
| Cooling | Air or liquid implementation and operating conditions |
| Heat rejection | Responsibility for heat leaving the server boundary |
| Network | External switches, optics, cables and topology |
SecondWatt's power procurement roadmap connects the server decision to wider infrastructure dependencies. Where new electrical capacity is needed, coordinate the transformer-sizing review and UPS procurement scope with the actual equipment schedule.
An illustrative server-only planning subtotal can show the boundary clearly. Four systems at an assumed maximum of 15 kW each produce a 60 kW sum. That arithmetic excludes switches, cooling and distribution losses and does not specify a circuit, UPS or transformer size. The facility engineer must apply the actual design requirements.
Compare complete deployment cost and support#
Ask for a complete bill of materials and separate hardware, networking, software entitlement, support, installation and acceptance work. Record the offer date, currency, condition and validity period. This article provides no current market price range because an unsourced range would obscure those differences.
A useful comparison can normalize cost to the measured service requirement after all candidates pass the minimum technical screen. If the systems require different quantities to meet that requirement, compare the resulting complete deployments, including their infrastructure and support scope.
Commercial comparison schedule#
| Cost or obligation | What to keep visible |
|---|---|
| Server hardware | Exact configuration and accepted condition |
| External infrastructure | Network, rack, power and cooling additions |
| Software and support | Included rights, duration and service provider |
| Testing | Pre-shipment and receiving acceptance obligations |
| Logistics | Packaging, freight, delivery point and installation |
| Operating estimate | Workload assumptions, measured power and excluded costs |
| Remedies | Response to missing components or failed acceptance |
For used equipment, require configuration and condition evidence in addition to workload performance. A generation comparison does not replace transaction due diligence. Confirm the seller's authority to supply the identified equipment and the written warranty or service terms that will apply to the buyer and destination.
Avoid assuming the newest configuration is automatically the lowest-cost deployment or the only commercially credible choice. An existing qualification, a facility constraint or a documented offer can change the decision. Equally, a lower acquisition price is not enough if the system cannot meet the required service with an acceptable operating arrangement.
Define the comparison period for any operating-cost estimate. Use measured or explicitly assumed workload demand, operating hours and electricity charges, and identify the infrastructure boundary. Do not substitute maximum system input for average consumption without labeling that conservative assumption. Keep the estimate separate from the quoted hardware price so changes in utilization or energy cost can be reviewed without rewriting the equipment offer.
The same discipline applies to software migration. Record any required porting, qualification or operator training as project work, with an owner and a completion criterion. A performance advantage becomes commercially useful when the buyer can deploy and maintain the implementation that produced it.
Choose against a decision matrix and request offers#
Use the evaluation to explain which constraint each candidate resolves. Keep unresolved evidence separate from an unfavorable result. One means more information is needed; the other means the candidate has been tested against the requirement and does not satisfy it.
Decision matrix for H200, B200 and B300 server offers#
| Buyer situation | Evidence that should drive the choice |
|---|---|
| Existing qualified deployment | Compatibility and measured benefit of changing systems |
| Memory-constrained workload | Full runtime footprint and validated partitioning |
| Throughput-constrained service | Sustained performance at the required quality and latency |
| Limited facility capacity | Complete-system and external infrastructure requirements |
| Fixed deployment date | Documented availability and unresolved site dependencies |
| Secondary-market purchase | Condition, acceptance and support evidence for the offered units |
Keep B200 and GB200 naming separate in the enquiry. A B200 server comparison is not an NVL72 rack-system comparison. If the intended purchase is an integrated rack, use the GB200 NVL72 dossier or GB300 NVL72 dossier and request the complete rack boundary.
SecondWatt acts as an independent intermediary. To request a GPU server comparison, share the workload, required quantity, preferred systems and facility power and cooling constraints. Add accepted condition, destination, support expectations and target date so the response can address the deployment you intend to buy.
FAQ: H200, B200 and B300 buying decisions#
Which GPU server should I buy?#
Choose the documented configuration that meets the workload, facility and commercial requirements with an acceptable support arrangement. Use memory and reference specifications to narrow the candidates, then compare measured service performance and complete deployment scope. This article does not assign a universal winner or assume that one generation fits every buyer.
Can I use total GPU memory as one memory pool?#
Not automatically. The published total adds memory across devices. An application needs the appropriate software and workload arrangement to use those devices effectively, and runtime overhead still matters. Validate the intended model and parallelism strategy rather than assuming that the aggregate capacity behaves like a single unpartitioned device.
Does the largest memory specification guarantee faster inference?#
No. Additional memory can remove a capacity constraint or support a different deployment arrangement, but the result also depends on compute, communication, software and service settings. Measure the actual workload at the required quality and latency. A memory comparison alone cannot establish the performance improvement for the offered systems.
Do all B200 and B300 servers require liquid cooling?#
No universal cooling requirement follows from the accelerator name alone. OEM implementations differ. Obtain the exact server's installation documentation and confirm its air or liquid interfaces with the facility team. Do not transfer the requirements of an integrated liquid-cooled rack to a different standalone server configuration.
Can I size facility power from the GPU power rating?#
No. Use the complete server's documented electrical requirements and add the external equipment within the project boundary. Distinguish maximum input from measured operating demand and power-supply capacity. The facility design must also address the required distribution, protection and redundancy arrangement for the actual system and installation.
Is a B200 server the same purchase as a GB200 NVL72 rack?#
No. They describe different system scopes. Confirm the platform and bill of materials before requesting or comparing a price. A rack-system offer needs its own compute, interconnect, power, cooling and commissioning boundaries. A per-server or per-GPU price cannot establish what is included in an integrated rack purchase.