NVIDIA GH200 Grace Hopper Superchip

Grace CPU and Hopper GPU on one module, programmable across a 450 W to 1000 W envelope that covers CPU, GPU and memory.
NVIDIA GH200 Grace Hopper Superchip is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.
NVIDIA GH200 Grace Hopper Superchip specs and power: 72 Arm cores, 96GB HBM3 or 144GB HBM3e, and a 450 W to 1000 W programmable power envelope.
The GH200 Grace Hopper Superchip puts a 72-core Arm Neoverse V2 CPU and a Hopper GPU on a single module joined by NVLink-C2C at 900 GB/s, memory-coherent in both directions. It carries up to 480 GB of LPDDR5X with ECC at up to 512 GB/s alongside its GPU memory, and exposes up to four PCIe Gen5 x16 links. NVIDIA rates the module as programmable from 450 W to 1000 W covering CPU, GPU and memory together. That combined figure is why per-accelerator comparisons fail here. A GH200 at its 1000 W ceiling is closer in scope to one H100 plus its host than to one H100, because the number already contains a 72-core CPU and up to 480 GB of LPDDR5X. Comparing it to a 700 W H100 SXM5 on a watts-per-GPU basis produces a misleading answer. Two memory configurations shipped and both are real orderable parts. NVIDIA's datasheet gives up to 96 GB of HBM3 at up to 4 TB/s, or 144 GB of HBM3e at up to 4.9 TB/s. Supermicro listed 96 GB HBM3 as the standard configuration with 144 GB HBM3e following, and the HBM3e variant now appears as a catalogued reseller SKU. One reseller states plainly that the module is available only as part of a Supermicro GPU server assembly, and no channel opened on 2026-09-10 displayed a price for a discrete superchip. The secondary market for GH200 is a market for 1U MGX servers, not for modules.
Why it matters
This is the only part in the Hopper generation whose power figure covers a CPU, a GPU and 480 GB of system memory at once, and the only one with no discrete purchase channel. Both facts change how it must be quoted: never per-GPU against an H100, and never as a module line item, because it ships inside a server.
Who buys it
Buyers of workloads that need coherent CPU-to-GPU memory rather than raw accelerator throughput: recommender systems, graph analytics, very large embedding tables and HPC codes that move data across the boundary constantly. Also operators wanting high compute density per rack unit who can accept an Arm software transition and liquid cooling.
Role in the data center
Workloads needing coherent CPU-to-GPU memory rather than peak throughput: recommender systems, graph analytics, large embedding tables and HPC codes that cross the CPU-GPU boundary constantly. Density band: 450 W to 1000 W a superchip in 1U single- or dual-superchip nodes, and no rack-level kW is published, so no rack band is claimed. Cooling is air or liquid at module level, but the dual-superchip 1U is liquid-only, so density forces the choice. Never express power or price per GPU against a discrete accelerator: the 1000 W ceiling covers a 72-core Arm CPU and up to 480 GB of LPDDR5X.
Power envelope
- Per superchip
- 450–1,000 W — Verified: Configurable power envelope published by the manufacturer — scope: CPU plus GPU plus memory. The upper endpoint is what downstream sizing uses; the lower endpoint is a configured setting, not a lighter part. source
- Cooling class
- DLC strongly preferred — Derived: Per-accelerator power at or above 1000 W. Air-cooled OEM systems exist at this level, at reduced rack density and higher airflow; liquid cooling is the practical choice above roughly 40 kW per rack. Configuration decides, not the accelerator alone.
No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.
Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.
Key specifications
| CPU cores | 72 Arm Neoverse V2 cores |
|---|---|
| CPU L3 cache | 117 MB |
| CPU memory | Up to 480 LPDDR5X with ECC GB |
| CPU memory bandwidth | Up to 512 GB/s |
| GPU memory, HBM3 configuration | Up to 96 GB HBM3 |
| GPU memory bandwidth, HBM3 | Up to 4 TB/s |
| GPU memory, HBM3e configuration | 144 GB HBM3e |
| GPU memory bandwidth, HBM3e | Up to 4.9 TB/s |
| NVLink-C2C bandwidth | 900 GB/s |
| Thermal design power, programmable | 450 to 1000 (CPU plus GPU plus memory) W |
| PCIe | Up to 4x x16, Gen5 |
| Cooling | Air cooled or liquid cooled |
| NVLink Switch System scale | Up to 256 GPUs, up to 144 TB memory |
| Node power supplies, single-superchip 1U | 2x 2000 redundant Titanium W |
| Node power supplies, dual-superchip 1U | 2x 2700 redundant Titanium W |
| Nameplate part number | 900-2G530-0060-000 |
Technical summary
CPU: 72 Arm Neoverse V2 cores, 1 MB L2 per core, 117 MB L3. CPU memory: up to 480 GB LPDDR5X with ECC, up to 512 GB/s. GPU memory, configuration 1: up to 96 GB HBM3 at up to 4 TB/s. GPU memory, configuration 2: 144 GB HBM3e at up to 4.9 TB/s. NVLink-C2C: 900 GB/s, memory-coherent, bidirectional CPU to GPU. TDP: programmable from 450 W to 1000 W covering CPU, GPU and memory together. PCIe: up to 4x PCIe x16, Gen5. Form factor: superchip module, sold inside a server assembly. Cooling: air cooled or liquid cooled per NVIDIA's datasheet. NVLink Switch System: up to 256 connected GPUs and up to 144 TB of memory. Node power supplies: 2x 2000 W for a single-superchip 1U (Supermicro MGX). Part numbers seen: 900-2G530-0060-000, 900-2G530-0070-000, 900-2G535-0020-000.
Major variations
Two GPU memory configurations, both in NVIDIA's datasheet and both shipped. Up to 96 GB of HBM3 at up to 4 TB/s came first, and 144 GB of HBM3e at up to 4.9 TB/s followed; Supermicro listed the 96 GB part as standard, and the HBM3e version now appears as a catalogued reseller SKU. This is one dossier rather than two because the memory change moves neither the power envelope nor the host platform. Both configurations sit in the same 450 W to 1000 W programmable band on the same module socket. Part numbers observed in channel: 900-2G530-0060-000, described as 12V with 480 GB supporting HBM3 or HBM3e and carrying Supermicro SKU GPU-NVGH480-12V; 900-2G530-0070-000, described as CG1 12V with 144 GB HBM3e; and 900-2G535-0020-000. Confirm which you are quoted, because the memory configuration is not always in the title.
Configurations and options
Supermicro's MGX line offers three relevant nodes. ARS-111GL-NHR is a 1U single-superchip air-cooled node with two 2000 W redundant Titanium supplies. ARS-111GL-NHR-LCC is the same 1U single-superchip node liquid-cooled, also two 2000 W. ARS-111GL-DNHR-LCC is a 1U dual-superchip node, liquid-cooled only, with two 2700 W redundant Titanium supplies. Each node carries 480 GB of integrated LPDDR5X with ECC per superchip. The pattern worth noting is that the dual-superchip variant is offered liquid-cooled only, with no air option, which follows directly from putting two modules in a single rack unit.
Compatibility and dependencies
Host platform: none in the usual sense. The Grace CPU is the host and the module ships inside a server assembly; one reseller lists it only inside a Supermicro server assembly, so there is no discrete-module channel. Power envelope: programmable from 450 W to 1000 W covering CPU, GPU and memory together, so it is not comparable per accelerator to a 700 W H100 SXM5. Node power from the OEM: Supermicro provisions two 2000 W supplies for a single-superchip 1U and two 2700 W for the dual-superchip 1U. NVIDIA and Supermicro publish no rack-level kW for GH200, so none is asserted here. At the 1000 W ceiling the dual-superchip 1U places 2 kW of superchip in one rack unit before host overhead, which is why that node is liquid-cooled only. Interconnect: NVLink-C2C at 900 GB/s memory-coherent. The host is Arm, not x86, so validate the stack. Plan alongside liquid cooling, racks, PDUs and busway.
Pricing and availability
Trend: stable No purchase price series exists for this part in the open record, and none can be constructed. The module is not sold discretely, so there is no unit whose price could be tracked over time. Compute Exchange, which publishes dated indicative bands for six other accelerators in its Q3 2026 snapshot, carries no band for GH200. Lead time, new: Available only as part of a Supermicro GPU server assembly per one reseller, so lead time follows the server rather than the module. Lead time, used or refurbished: No used-condition channel displayed stock or a lead time for a discrete GH200 module on 2026-09-10. Warranty, new: One reseller states three years with 30 days advance replacement for dead-on-arrival units, and overnight replacement for 30 days then two weeks afterwards.
Lifecycle and maintenance
Lifecycle status as of 2026-09-10: supported and purchasable, no longer merchandised. No Hopper part appears on NVIDIA's enterprise-software end-of-life notice list, which carries Volta, Turing, Ampere and Ada Lovelace. NVIDIA's HGX platform page now lists only Vera Rubin, Rubin, B300 and B200 baseboards, so this generation is off the page NVIDIA merchandises while both product pages stay live with full specifications. NVIDIA announced GH200 in full production on 28 May 2023, and the 144 GB HBM3e configuration followed. No service-life figure is published. Note a structural consequence: CPU, GPU, HBM and LPDDR5X share one module, so none of the four is separately serviceable and any single failure writes down the whole superchip.
Common failure points
- HBM stack on the module — Correctable ECC counts climb, then uncorrectable errors and retired pages, and the whole module becomes suspect
- On-module LPDDR5X system memory — ECC events against CPU memory, or reported capacity below the expected 480 GB
- NVLink-C2C coherent link — The CPU-to-GPU link trains below 900 GB/s or renegotiates under load, and coherent-memory workloads collapse in throughput
- Liquid cooling loop on the LCC variants — Coolant leak staining around cold plates or fittings, rising die temperatures at unchanged load, flow alarms
- Power envelope misconfiguration — Sustained performance well below expectation with no thermal or error indication
- Arm software stack compatibility — Workloads or drivers built for x86 hosts fail to run or run without acceleration
Inspection checklist
Memory configuration on the module: confirm whether it is the 96 GB HBM3 or the 144 GB HBM3e variant, because part numbers differ and the two are priced and quoted differently. LPDDR5X capacity and ECC: verify up to 480 GB of system memory reports correctly with ECC enabled, since the memory is on-module and cannot be replaced. Programmed power ceiling: read the configured TDP within the 450 W to 1000 W band, because the same module in 2 servers can be strapped to very different envelopes. NVLink-C2C link health: confirm the coherent CPU-to-GPU link trains at 900 GB/s, as a degraded C2C link removes the only reason to buy this part over a discrete GPU. HBM row-remapping history: read retired page counts via nvidia-smi -q -d ECC; HBM is not repairable and on this module a failed stack condemns the CPU alongside the GPU.
Procurement channels
System integrators only, in practice. The module is catalogued by resellers with NVIDIA part numbers but is stated to be available only as part of a Supermicro GPU server assembly, so the purchase is a 1U MGX node rather than a component. No broker price discovery exists. Compute Exchange publishes indicative bands for six other accelerators and none for GH200, and neither reseller page opened on 2026-09-10 displayed a price for the module. Practical consequence for sourcing: quote and compare whole nodes, specify the memory configuration explicitly because part titles do not always state it, and establish whether an air-cooled or liquid-cooled node is being conveyed.
Regional notes
Classification: ECCN 3A090.a and.b, with 4A090 for computers incorporating them. Because it ships inside a server assembly, classify the assembly under 4A090 rather than the module alone. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. This module publishes up to 4,900 GB/s (4.9 TB/s), under the gate; no TPP figure.