NVIDIA H200 SXM 141GB

141 GB HBM3e at 4.8 TB/s for the same 700 W as H100 SXM5, inside the same 10.2 kW node and 40.8 kW rack.
NVIDIA H200 SXM 141GB is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.
NVIDIA H200 SXM 141GB specs and verified power: 141 GB HBM3e at 4.8 TB/s for 700 W, 10.2 kW per eight-GPU node and 40.8 kW per rack at four nodes.
The H200 SXM 141GB is H100 compute with a rebuilt memory subsystem. Every arithmetic throughput figure NVIDIA publishes is identical to H100 SXM5, down to 989 teraFLOPS of TF32 and 3,958 teraFLOPS of FP8 with sparsity. What changed is 80 GB to 141 GB and 3.35 TB/s to 4.8 TB/s, at the same up-to-700 W configurable TDP. That is a 76 percent capacity gain and a 43 percent bandwidth gain for zero additional watts, and it is the entire buying argument. It is also why the part holds a $30,000 to $35,000 used band against H100 SXM5 at $22,000 to $27,000, with Compute Exchange signalling supply as Low rather than Very High. The facility numbers are the H100's, because the chassis is the same. NVIDIA's DGX user guide covers H100 and H200 together and does not distinguish their power, rating both at 10.2 kW maximum with six 3300 W supplies of which four must be energized. Eight accelerators at 700 W is 5,600 W, so the GPUs are 54.9 percent of node power. One caution about a figure that circulates. NVIDIA's DGX SuperPOD reference architecture for H200 states that power consumption per rack exceeds 25 kW for its example layout and does not say how many systems that rack holds. Because the node power is identical to H100, a four-node H200 rack is 40.8 kW on the same arithmetic, and the 25 kW figure describes a differently composed rack rather than a lower-power node.
Why it matters
This is the clearest capacity-per-watt step in the Hopper generation and the one that makes an existing power envelope go further. Identical arithmetic, identical 700 W, identical 10.2 kW node, and 76 percent more memory. A facility that already carries H100 racks can adopt it without touching electrical or cooling design.
Who buys it
Operators running memory-bound inference or long-context serving, where 141 GB per accelerator removes a sharding step that 80 GB forces. Also training buyers who want more memory per watt inside an unchanged 40 kW air-cooled rack, and hyperscaler-fleet buyers taking H200 boards as those operators move to Blackwell.
Role in the data center
Training and memory-bound inference, plus long-context serving where 141 GB an accelerator removes a sharding step that 80 GB forces. Density band: 40.8 kW electrical and 45.2 kW of cooling at four eight-GPU nodes a rack, air-cooled, identical to H100 because the chassis is identical. Accelerators are only 54.9 percent of the 10.2 kW node, so size from the node, never from 8 x 700 W. Transients: telemetry on DGX-H100 racks sharing this chassis shows site power swinging by tens to hundreds of MW at 0.2 to 3 Hz during synchronised training, so raise step-load response with the UPS vendor.
Power envelope
Specifications verified 2026-09-10.
- Per accelerator
- 700 W — Verified: Maximum configurable board power, manufacturer specification (configurable). source
- Per 8-accelerator node
- 10.2 kW electrical — Verified: NVIDIA DGX H100 published maximum system power — 10.2 kW against 5.6 kW of accelerator (1.82x). Sizing from accelerator watts alone under-provisions the building. source
- Heat load per node
- 11.3 kW thermal — Verified: NVIDIA DGX H100 published heat load. Stated separately from the 10.2 kW electrical figure — size cooling on this one. source
- Per rack
- ~41 kW — Derived: 4 nodes per rack at the published node power, bound by NVIDIA DGX H100 per-rack guidance of 4 systems (power alone would allow 12). State your own budget and rack height and the count changes.
- Cooling class
- DLC recommended — Derived: Per-accelerator power at or above 700 W — air cooling possible at reduced rack density.
- Accelerators per IT MW
- ~784 — Derived: 1 MW IT load / (10.2 kW per 8-accelerator NVIDIA DGX H100 node). IT load only — excludes cooling and distribution losses.
Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.
Key specifications
| GPU memory | 141 GB HBM3e |
|---|---|
| Memory bandwidth | 4.8 TB/s |
| Max thermal design power | Up to 700 (configurable) W |
| Form factor | SXM |
| NVLink interconnect | 900 GB/s |
| PCIe interconnect | Gen5, 128 GB/s |
| FP8 Tensor Core, with sparsity | 3,958 TFLOPS |
| FP64 Tensor Core | 67 TFLOPS |
| Multi-Instance GPU | Up to 7 at 18 GB each instances |
| Confidential Computing | Supported |
| Node maximum power, 8-GPU DGX H200 | 10.2 kW |
| Node GPU memory, 8-GPU DGX H200 | 1,128 GB |
| Node heat output, shared DGX chassis | 38,557 BTU/hr |
| Rack power at 4 nodes | 40.8 kW |
| Server options | HGX H200 partner and NVIDIA-Certified Systems with 4 or 8 GPUs |
| Cooling | Air or liquid depending on server |
Technical summary
Architecture: Hopper, TSMC 4N process. Form factor: SXM, HGX H200 baseboard or DGX H200 only, not a PCIe card. Memory: 141 GB HBM3e at 4.8 TB/s. Max TDP: up to 700 W, configurable. Node power: 10.2 kW maximum for an eight-GPU DGX H200 (NVIDIA user guide). Node GPU memory: 1,128 GB total across eight accelerators. Rack power: 40.8 kW at four nodes per rack, the same ceiling as H100. Interconnect: NVLink 900 GB/s GPU-to-GPU; PCIe Gen5 at 128 GB/s to host. Multi-Instance GPU: up to 7 instances at 18 GB each. Dense throughput: FP8 1,979 teraFLOPS (NVIDIA publishes 3,958 with sparsity). Confidential Computing: supported. Baseboard part number seen in channel: 935-24287-0040-000 (HGX H200 8-GPU).
Major variations
HGX H200 baseboards ship in four-GPU and eight-GPU configurations, and NVIDIA's DGX H200 uses the eight-GPU form. Lenovo catalogues both an eight-GPU board and a four-GPU board as separate options. Cooling is air or liquid depending on the server, and liquid remains optional rather than required. TDP is configurable below the 700 W maximum, which matters where a hall cannot deliver 40.8 kW per rack. Note what did not change from H100 SXM5. Every published arithmetic throughput figure is identical, including 989 teraFLOPS TF32 and 3,958 teraFLOPS FP8 with sparsity. Only the memory subsystem and the MIG slice size differ, at 18 GB per instance against 10 GB.
Configurations and options
The two configurations that matter commercially are the eight-GPU DGX H200 at 8U and 10.2 kW, and an OEM eight-GPU HGX H200 chassis such as the Supermicro SYS-821GE-TNHR, also 8U, with six 3000 W supplies in a four-plus-two group giving 12 kW usable. Ambient ratings differ between them and the system rating governs. NVIDIA rates the DGX chassis from 5 to 30 C while the Supermicro chassis is rated 10 to 35 C. In a hall running warm the 30 C DGX ceiling is the binding constraint.
Compatibility and dependencies
Host platform: an HGX H200 four- or eight-GPU baseboard or an NVIDIA DGX H200, not a PCIe part. Secondary trade is usually the baseboard, part 935-24287-0040-000. Node power, from NVIDIA's user guide, which covers H100 and H200 together and does not distinguish their power: 10.2 kW maximum. Eight accelerators at 700 W is 5,600 W, so GPUs are 54.9 percent of node power, and sizing from GPU nameplate alone under-provisions by 45 percent. Two NVIDIA figures for the same node disagree; do not average. The guide states a 10.2 kW electrical maximum and 38,557 BTU/hr, which is 11.30 kW. Size electrical distribution from 10.2 kW a node and cooling from 11.30 kW, or 45.2 kW a rack at four nodes. Rack density carries the H100 ceiling because the chassis is identical: four systems a rack, 40.8 kW electrical and three rPDU feeds at N+1. Plan alongside liquid cooling, busway, PDUs, racks and CRAC units.
Pricing and availability
Evidence: channel observation · used condition · Eight-GPU baseboard basis. Reference band for another model: 19000–27000 $ per GPU (NVIDIA H100 SXM5 80GB comparable). It is not a price for this model. Trend: stable H200 began shipping in Q2 2024, so the series is short. Compute Exchange placed used units at $30,000 to $35,000 in its Q3 2026 snapshot with supply signalled as Low, the only accelerator in its table carrying that signal. The notable feature is the absence of a used discount. A new eight-GPU baseboard works out at $31,009.18 per accelerator against a used band topping out at $35,000, which is what a Low supply signal predicts. Contract rental tells the same story, with H200 agreements renewing at the rates they were signed at two to three years earlier. Lead time, new: New eight-GPU baseboards quoted by integrators on 2026-09-10, with price validity dated to the following day rather than a stated ship time. Lead time, used or refurbished: Brokered rather than browsable; Compute Exchange signals supply as Low, the tightest in its Q3 2026 table. Warranty, used: Not established for this part from any channel opened; confirm whether cover is NVIDIA-backed or reseller-backed in writing.
Lifecycle and maintenance
Lifecycle status as of 2026-09-10: supported and purchasable, no longer merchandised. No Hopper part appears on NVIDIA's enterprise-software end-of-life notice list, which carries Volta, Turing, Ampere and Ada Lovelace. NVIDIA's HGX platform page now lists only Vera Rubin, Rubin, B300 and B200 baseboards, so this generation is off the page NVIDIA merchandises while both product pages stay live with full specifications. H200 shipped from Q2 2024, making it the youngest SXM part in this generation and the one with the most remaining life on any depreciation schedule. No manufacturer service-life figure is published. Operator schedules of 5 to 6 years and an observed 7 to 9 year Azure retirement history are the available proxies.
Common failure points
- HBM3e memory stack — Correctable ECC counts climb, then uncorrectable errors and retired pages, and long-context inference workers restart
- SXM module on a non-serviceable baseboard — One accelerator drops from enumeration and total memory reads short of 1,128 GB
- Power supply group — A supply drops out of the six-unit group and the node derates or refuses to start
- NVLink and NVSwitch fabric — Links fail to train or retrain repeatedly and collective operations stall with rising per-lane CRC counters
- Cooling path and fan set — Fans pinned near 100 percent at moderate load and clocks dropping under sustained load
- Fleet-level power oscillation — Aggregate site load swings by tens to hundreds of MW with spectral energy between 0.2 and 3 Hz during synchronised training
Inspection checklist
Memory capacity and bandwidth: each module must report 141 GB and benchmark near the published 4.8 TB/s; a board reading 80 GB or 3.35 TB/s is H100 silicon sold as H200. Board enumeration and total memory: all 8 accelerators must enumerate and total GPU memory must read 1,128 GB; a short count means a dead module on a non-serviceable board. HBM3e row-remapping history: run nvidia-smi -q -d ECC on all 8 modules and record retired page counts; HBM3e is not repairable and remap growth caps the board's life. Uncorrectable ECC event log: pull volatile and aggregate counters; 0 is the only acceptable uncorrectable count on a board offered as tested and recertified. Hours in service: request accumulated power-on hours per module; H200 shipped from Q2 2024, so any board claiming pre-2024 service history is misrepresented.
Procurement channels
Broker and integrator channel. Compute Exchange publishes the only quotable indicative band and signals supply as Low. Integrators such as Exxact list new eight-GPU HGX H200 baseboards with dated price validity. Supply is the tightest of the current-generation accelerators tracked here, which shows up as an absent used discount rather than as unavailability. Expect to pay close to new-baseboard equivalent for used units, and expect broker quotes rather than published stock. Exercise more diligence here than on H100.
Regional notes
This is the part BIS named, the rule citing the H200 and AMD MI325X as its examples. Classification: ECCN 3A090.a and.b, with 4A090 for computers incorporating them. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. This module publishes 4,800 GB/s (4.8 TB/s), under the gate, with no TPP figure published.