AMD Instinct MI355X

AMD Instinct MI355X — gpu, 288 GB HBM3E
AMD Instinct MI355X

CDNA 4 OAM module at 1400 W with 288 GB of HBM3E, the highest-power part in this category and not reachable on air.

AMD Instinct MI355X is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

AMD Instinct MI355X dossier: 288 GB HBM3E, 8 TB/s, 1400 W OAM needing liquid cooling, with node and rack kW, export status and observed pricing.

The AMD Instinct MI355X is a CDNA 4 accelerator on an OAM module, launched 12 June 2025. AMD publishes 288 GB of HBM3E, 8 TB/s of memory bandwidth, 256 compute units, 16,384 stream processors and a typical board power of 1400 W on a TSMC 3nm and 6nm process. Cooling is listed as passive and active. 1400 W makes this the highest-power part in this category and the top of a range that starts at 72 W for an L4, a span of 19.4 times. One of these modules draws more than nineteen L4 cards. AMD lists cooling for this module as passive and active at a 1400 W typical board power, and its platform brief presents direct liquid cooling as the option that carries the full 1400 W. Air-cooled platforms are published for the same module — Supermicro lists a 10U air-cooled 8-GPU MI355X system (AS-A126GS-TNMR) alongside its 4U liquid-cooled AS-4126GS-NMR-LCC — so cooling mode is a platform decision, and the sustained per-module power an air platform supports should be confirmed with the OEM in writing. That makes the real decision MI350X against MI355X rather than AMD against anyone else. The two are identical on memory, bandwidth and compute unit count. What a buyer purchases with the extra 400 W is roughly 10 percent more throughput, 10.1 PFLOPs of MXFP4 against 9.2. Eight modules draw 11.2 kW of accelerator load and put a rack near 91 kW, which is a different building from the 26 to 40 kW halls that take the legacy parts in this category.

Why it matters

This part defines the top of the power range covered here, and the cooling decision attached to it is unusually clean. AMD states that liquid cooling is what enables 1400 W, so a facility without it should buy the MI350X and lose roughly 10 percent rather than run this module de-rated. At 8,000 GB/s it also fails the export gate the MI325X clears.

Who buys it

Operators building or retrofitting for direct-to-chip liquid cooling who want the highest memory capacity and throughput AMD ships. The 288 GB pool and 10.1 PFLOPs of MXFP4 suit frontier training and high-throughput inference. Buyers must have a coolant distribution unit, a rack able to carry roughly 91 kW, and tolerance for ROCm rather than CUDA.

Role in the data center

Frontier training and high-throughput low-precision inference, plus double-precision work at 78.6 TFLOPs FP64. Lands near a 91 kW rack density band with direct liquid cooling required, the highest density in this category outside the rack-scale NVLink systems.

Power envelope

Per accelerator
1,400 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
DLC strongly preferred — Derived: Per-accelerator power at or above 1000 W. Air-cooled OEM systems exist at this level, at reduced rack density and higher airflow; liquid cooling is the practical choice above roughly 40 kW per rack. Configuration decides, not the accelerator alone.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureCDNA 4
Form factorOAM module
Memory capacity288 GB HBM3E
Memory bandwidth8 TB/s
Typical board power1400 W TBP
Cooling modePassive and active; liquid enables full 1400 W
Compute units256 CUs
Stream processors16,384
Dense MXFP410.1 PFLOPs
Dense FP85 PFLOPs
FP6478.6 TFLOPs
Infinity Fabric scale-up link153 GB/s
Host interfacePCIe 5.0 x16
Accelerator-only load, 8 modules11.2 kW
Estimated rack density, four nodes91 kW
Total DRAM bandwidth vs export gate8,000 above 6,500 GB/s

Technical summary

Architecture: CDNA 4, TSMC 3nm and 6nm FinFET. Memory: 288 GB HBM3E, 8 TB/s peak bandwidth. Typical board power: 1400 W TBP; AMD lists cooling as passive and active. Compute units: 256; stream processors: 16,384. Dense compute: MXFP4 10.1 PFLOPs, MXFP6 10.1 PFLOPs, FP8 5 PFLOPs. With structured sparsity (vendor figure): FP8 10.1 PFLOPs, FP16 and BF16 5 PFLOPs. FP16 and BF16 matrix 2.5 PFLOPs, FP32 157.3 TFLOPs, FP64 78.6 TFLOPs, INT8 5 POPs. Infinity Fabric: scale-up 153 GB/s, scale-out 128 GB/s; host PCIe 5.0 x16. Platform: 8 modules on UBB 2.0, 2.3 TB shared memory, 153.6 GB/s bidirectional links. Accelerator-only load, 8-module node: 11.2 kW; proxy-estimated node draw, not an OEM figure, 22.7 kW. Rack density near 91 kW at four nodes, with direct liquid cooling required. Export: 8,000 GB/s total DRAM bandwidth, above the 6,500 GB/s gate.

Major variations

MI350X is the same silicon at 1000 W with passive OAM cooling, reaching 9.2 PFLOPs of MXFP4 against 10.1 here. Memory, bandwidth and compute unit count are identical at 288 GB, 8 TB/s and 256 units. The whole difference is 400 W of power budget and the cooling that carries it, so the choice is a facility decision rather than a specification one. MI325X is the CDNA 3 predecessor at 256 GB HBM3E and 6 TB/s at 1000 W, with roughly half the FP8 throughput and no MXFP4 datatypes. It sits below the 6,500 GB/s export gate where this module sits above it.

Configurations and options

The purchasable unit is an eight-module platform on a UBB 2.0 baseboard carrying 2.3 TB of coherent shared memory, 1400 W per module, 153.6 GB/s bidirectional Infinity Fabric to each of seven peers, and eight PCIe Gen 5 x16 host links. AMD's platform brief states that an air-cooled option is available and that the direct liquid cooling option carries the full 1400 W per module. Supermicro publishes both an air-cooled 10U 8-GPU platform (AS-A126GS-TNMR) and a 4U liquid-cooled one (AS-4126GS-NMR-LCC) for this module. AMD does not state the air-cooled power ceiling, so the de-rated figure for an air configuration is not established. No configurable-TDP range is published either. Coolant flow rate, inlet temperature and coolant distribution unit sizing are absent from the brief and must come from the OEM.

Compatibility and dependencies

Requires an eight-module OAM chassis on a UBB 2.0 baseboard with direct-to-chip liquid cooling to reach rated power, and cannot be retrofitted into a PCIe server. Host attachment is PCIe 5.0 x16 per module; peer traffic runs on Infinity Fabric at 153 GB/s scale-up and 128 GB/s scale-out. Accelerator-only load for eight modules is 11.2 kW, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a cross-vendor proxy because the manufacturer publishes no node figure, suggests roughly 22.7 kW. Four such nodes put a rack near 91 kW. That requires a coolant distribution unit, liquid-rated racks and busway sized well above the nameplate assumption, so plan alongside liquid cooling, racks, busway and PDUs. A 26 to 40 kW air hall cannot host this part.

Pricing and availability

Evidence: analyst estimate. Reference band for another model: 10000–15000 $ per accelerator (AMD Instinct MI300X comparable, Feb 2024). It is not a price for this model. Trend: stable No price series exists for this module. The part launched 12 June 2025, so any curve would cover fifteen months at most. Direction is flat rather than falling: this is AMD's current flagship, the MI400 series has not reached general availability, and no upgrade cycle has released volume into resale. A depreciation view on this part is a 2027 question. Lead time, new: Ordered as an eight-module liquid-cooled OAM system through an OEM rather than as a module; no discrete module lead time is published.. Lead time, used or refurbished: This is a 2025 flagship and no upgrade cycle has released volume into resale.. Warranty, new: OEM system warranty set by the server vendor rather than by AMD for the module alone.. Warranty, used: No manufacturer warranty transfer was documented. Every warranty observed in this category is a reseller warranty written by the dealer..

Lifecycle and maintenance

Launched 12 June 2025 and AMD's current shipping flagship accelerator, with the MI350 series presented as available on AMD's Instinct landing page. The MI400 series was announced and its specifications are published as engineering projections, with volume deployment future-dated, so this module remains the top of AMD's generally available line. Operator depreciation practice runs 4 to 6 years across Nebius, Lambda Labs and CoreWeave.

Common failure points

Inspection checklist

Driver-reported memory capacity must read 288 GB: 256 GB indicates an MI325X sold as this part. Sustained power draw must reach toward 1400 W TBP under load: a module capping near 1000 W is an MI350X, not this part. Stream benchmark bandwidth within 10 percent of 8 TB/s: 6 TB/s is the CDNA 3 figure and a red flag on identity. Compute units must enumerate 256 with 16,384 stream processors: a 304 and 19,456 count means a CDNA 3 module. FP64 measured near 78.6 TFLOPs: a reading near 72.1 TFLOPs indicates an MI350X and near 81.7 a CDNA 3 part. MXFP4 and MXFP6 datatype support confirmed in the driver: dense MXFP4 should measure toward 10.1 PFLOPs. HBM3E correctable and uncorrectable ECC counters read before purchase and after a 24-hour soak: HBM is not field-serviceable. Cold plate flatness, seating torque and coolant fitting integrity inspected: at 1400 W this is the primary failure path.

Procurement channels

New: OEM system orders for liquid-cooled eight-module nodes through Supermicro, Dell, HPE and comparable vendors, priced as a node rather than per accelerator. Cloud rental is available, with Oracle at $8.60 per accelerator-hour. Secondary, as of 10 September 2026: Compute Exchange publishes indicative secondary bands for seven NVIDIA parts and lists no AMD or Intel accelerator, and no L4. The market structure explains it: OAM modules ship on a baseboard inside an OEM system rather than discretely, so what trades second-hand is the 8U chassis rather than the module, and no open-channel Instinct listing appears at the brokers surveyed on that date.

Regional notes

AMD's current flagship is harder to ship than the CDNA 3 module BIS named, an unusual direction of travel and the most consequential export fact about the MI350 series. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Published bandwidth is 8,000 GB/s (8 TB/s), above the gate on published specifications, so the January 2026 relief does not reach it.