AMD Instinct MI350X

AMD Instinct MI350X — gpu, 288 GB HBM3E
AMD Instinct MI350X

CDNA 4 OAM module with 288 GB of HBM3E and 8 TB/s of bandwidth, air cooled at 1000 W in eight-module nodes.

AMD Instinct MI350X is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

AMD Instinct MI350X dossier: 288 GB HBM3E, 8 TB/s, 1000 W air-cooled OAM, node and rack kW, and why it sits above the 6,500 GB/s export gate.

The AMD Instinct MI350X is a CDNA 4 accelerator on an OAM module, launched 12 June 2025 alongside the MI355X. AMD publishes 288 GB of HBM3E, 8 TB/s of memory bandwidth and a typical board power of 1000 W, on a TSMC 3nm and 6nm process with 256 compute units. Cooling is passive OAM. This is the air-cooled half of the MI350 series and that is the reason to choose it. It carries the same 288 GB, the same 8 TB/s and the same 256 compute units as the MI355X, but runs at 1000 W rather than 1400 W. The throughput gap is roughly 10 percent, with 9.2 PFLOPs of MXFP4 against 10.1. A facility that cannot deliver direct-to-chip liquid should buy this part and lose that 10 percent rather than buy the MI355X and run it de-rated. Note that the compute unit count went down against CDNA 3, not up. MI300X and MI325X carry 304 compute units and 19,456 stream processors; CDNA 4 delivers roughly double the FP8 throughput from 256 units against 304, each faster and on a smaller node. FP64 also fell slightly, from 81.7 to 72.1 TFLOPs, which matters to a high performance computing buyer and not at all to an inference buyer.

Why it matters

This part is the answer to the assumption that CDNA 4 requires liquid cooling. Identical silicon, memory and bandwidth to the MI355X for 400 W less and roughly 10 percent less throughput, on air. It also sits above the 6,500 GB/s export gate at 8,000 GB/s, so it is harder to ship than the MI325X it succeeds.

Who buys it

Operators who want CDNA 4 memory capacity and throughput without rebuilding a hall for liquid cooling. At 1000 W and eight to a baseboard the accelerator load alone is 8.0 kW, on a platform rated 15.75 kW usable, which a well-contained air hall can deliver. Buyers must accept ROCm and must be purchasing a whole eight-module system.

Role in the data center

Distributed training and large-model inference on CDNA 4, plus double-precision work at 72.1 TFLOPs FP64. Lands in a 50 to 65 kW rack density band as an air-cooled OAM node, which is the highest density in this category still achievable without direct-to-chip liquid.

Power envelope

Per accelerator
1,000 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
DLC strongly preferred — Derived: Per-accelerator power at or above 1000 W. Air-cooled OEM systems exist at this level, at reduced rack density and higher airflow; liquid cooling is the practical choice above roughly 40 kW per rack. Configuration decides, not the accelerator alone.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureCDNA 4
Form factorOAM module
Memory capacity288 GB HBM3E
Memory bandwidth8 TB/s
Typical board power1000 W TBP
Cooling modePassive OAM, air
LithographyTSMC 3nm and 6nm FinFET
Compute units256 CUs
Dense MXFP4 and MXFP69.2 PFLOPs
Dense FP84.6 PFLOPs
Dense FP162.3 PFLOPs
FP6472.1 TFLOPs
Platform shared memory2.3 TB across 8 modules
Infinity Fabric link153.6 GB/s bidirectional
Accelerator-only load, 8 modules8.0 kW
Total DRAM bandwidth vs export gate8,000 above 6,500 GB/s

Technical summary

Architecture: CDNA 4, TSMC 3nm and 6nm FinFET. Memory: 288 GB HBM3E, 8 TB/s peak bandwidth. Typical board power: 1000 W, passive OAM cooling, air. Compute units: 256. Dense compute: MXFP4 and MXFP6 9.2 PFLOPs, FP8 4.6 PFLOPs, FP16 2.3 PFLOPs. FP32 144.2 TFLOPs, FP64 72.1 TFLOPs. Platform: 8 modules on UBB 2.0, 2.3 TB shared memory, 8 PCIe Gen 5 x16. Infinity Fabric: 153.6 GB/s bidirectional links to each of seven peers. Accelerator-only load, 8-module node: 8.0 kW. AMD publishes no node power rating; Supermicro's AS-8126GS-TNMR platform provides 15.75 kW usable from six 5,250 W supplies at 3+3, which is the figure to size against. Export: 8,000 GB/s total DRAM bandwidth, above the 6,500 GB/s gate. Launched 12 June 2025, alongside the 1400 W MI355X.

Major variations

MI355X is the same silicon at 1400 W with an active cooling option, reaching 10.1 PFLOPs of MXFP4 against 9.2 here. Memory, bandwidth and compute unit count are identical at 288 GB, 8 TB/s and 256 units. The entire difference is 400 W of power budget and the cooling that carries it. MI325X is the CDNA 3 predecessor at 256 GB HBM3E and 6 TB/s, also at 1000 W, but with roughly half the FP8 throughput and no MXFP4 or MXFP6 datatypes. It sits below the 6,500 GB/s export gate where this module sits above it.

Configurations and options

The purchasable unit is an eight-module platform on a UBB 2.0 baseboard carrying 2.3 TB of coherent shared memory, with eight PCIe Gen 5 x16 host links and 153.6 GB/s bidirectional Infinity Fabric to each of seven peers. Supermicro's AS-8126GS-TNMR 8U system lists support for this module and the MI325X, with six 5250 W Titanium supplies in a 3+3 arrangement for 15.75 kW usable. AMD's MI350X platform brochure does not state a cooling mode, and the MI350 platform page returned an error on 10 September 2026. The module page itself states passive OAM. Whether a liquid-cooled MI350X configuration is offered was not established from AMD documentation.

Compatibility and dependencies

Requires an eight-module OAM chassis on a UBB 2.0 baseboard and cannot be retrofitted into a PCIe server. Host attachment is PCIe Gen 5 x16 per module; peer traffic runs on Infinity Fabric at 153.6 GB/s bidirectional per link. Cooling is passive OAM, so the chassis supplies all airflow and the module has no fan. Accelerator-only load for eight modules is 8.0 kW. AMD publishes no node power rating, so size against the platform instead: Supermicro's AS-8126GS-TNMR provides 15.75 kW usable from six 5,250 W supplies at 3+3, and the balance above the 8.0 kW accelerator load covers CPUs, memory, storage and fans. Plan alongside racks, busway, PDUs and CRAC or chiller capacity, and note that this is the CDNA 4 part that does not require liquid cooling.

Pricing and availability

Evidence: analyst estimate. Reference band for another model: 10000–15000 $ per accelerator (AMD Instinct MI300X comparable, Feb 2024). It is not a price for this model. Trend: stable No price series exists for this module. The part launched 12 June 2025, so any depreciation curve would cover fifteen months at most even if a channel published one. Direction is flat rather than falling: this remains a current AMD offering, no successor generation has reached general availability, and no upgrade cycle has released volume into resale. Lead time, new: Ordered as an eight-module OAM system through Supermicro, Dell or another OEM. Supermicro's AS-8126GS-TNMR 8U system lists support for this module.. Lead time, used or refurbished: This is a 2025 part and no upgrade cycle has released volume into resale.. Warranty, new: OEM system warranty set by the server vendor rather than by AMD for the module alone.. Warranty, used: No manufacturer warranty transfer was documented. Every warranty observed in this category is a reseller warranty written by the dealer..

Lifecycle and maintenance

Launched 12 June 2025 and part of the MI350 series that AMD's Instinct landing page presents as a current offering. It is AMD's current air-cooled CDNA 4 module, with the MI400 series announced but not generally available. AMD also publishes no configurable-TDP range for this module, and the MI350 platform page returned an error on that date, so whether it can be power-capped in firmware was not established. Operator depreciation practice runs 4 years at Nebius, 5 at Lambda Labs and 6 at CoreWeave.

Common failure points

Inspection checklist

Driver-reported memory capacity must read 288 GB: 256 GB indicates an MI325X and 192 GB an MI300X sold as this part. Stream benchmark bandwidth within 10 percent of 8 TB/s: 6 TB/s is the MI325X figure and a red flag on identity. Sustained power draw within 5 percent of 1000 W TBP: a module drawing toward 1400 W is an MI355X, not this part. Compute units must enumerate 256, not the 304 of CDNA 3: a 304 count means the module is a previous generation. MXFP4 and MXFP6 datatype support confirmed in the driver: CDNA 3 does not carry these and their absence identifies the die. HBM3E correctable and uncorrectable ECC counters read before purchase and after a 24-hour soak: HBM is not field-serviceable. All Infinity Fabric links train at 153.6 GB/s bidirectional to each of 7 peers: a dead link is a baseboard return.

Procurement channels

New: OEM system orders through Supermicro, Dell, HPE and comparable vendors, priced as an eight-module node rather than per accelerator. Supermicro's AS-8126GS-TNMR is a documented host. Secondary, as of 10 September 2026: Compute Exchange publishes indicative secondary bands for seven NVIDIA parts and lists no AMD or Intel accelerator, and no L4. The market structure explains it: OAM modules ship on a baseboard inside an OEM system rather than discretely, so what trades second-hand is the 8U chassis rather than the module, and no open-channel Instinct listing appears at the brokers surveyed on that date.

Regional notes

AMD's CDNA 4 parts are harder to ship than the CDNA 3 module BIS named, an unusual direction of travel. Confirm classification with counsel before any restricted-destination quote. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Published bandwidth is 8,000 GB/s (8 TB/s), above the gate, so the relief does not reach it and presumption of denial holds.