AMD Instinct MI325X

AMD Instinct MI325X — gpu, 256 GB HBM3E
AMD Instinct MI325X

CDNA 3 OAM module with 256 GB of HBM3E at 1000 W, named by BIS in the January 2026 rule alongside the H200.

AMD Instinct MI325X is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

AMD Instinct MI325X dossier: 256 GB HBM3E, 6 TB/s, 1000 W OAM, node and rack kW, cooling, and the export gate that names this module by name.

The AMD Instinct MI325X is a CDNA 3 accelerator on an OAM module, launched 10 October 2024. AMD publishes 256 GB of HBM3E on an 8192-bit interface, 6 TB/s of memory bandwidth and a typical board power of 1000 W peak. External power comes in on a 54V UBB connection, and cooling is passive OAM. The procurement argument is unusual and worth stating plainly. Compute is identical to the MI300X on every precision AMD publishes, with the same 304 compute units, the same 19,456 stream processors and the same TSMC 5nm and 6nm process. What changed is 192 GB of HBM3 at 5.3 TB/s becoming 256 GB of HBM3E at 6 TB/s, and 750 W becoming 1000 W. That is a third more memory capacity and a third more power for zero throughput gain. It also carries the clearest export position of any part in this category. Its 6 TB/s is 6,000 GB/s, under the 6,500 GB/s gate in 91 FR 1684, and BIS names this exact module as an example of a commodity eligible for case-by-case review. Facility-wise, eight modules draw 8.0 kW of accelerator load and land a node in a 36 to 65 kW rack band, still on air.

Why it matters

This is the part BIS named, which makes it the reference point for reading the January 2026 threshold rule. It is also the clearest case in the catalog of paying power for memory: a third more capacity and a third more watts for identical compute. At 1000 W it sits at the ceiling of what an air-cooled OAM baseboard carries.

Who buys it

Operators running large-model inference where memory capacity per accelerator, not throughput, sets the node count. The 256 GB pool holds models that need two accelerators elsewhere. Buyers must be acquiring a whole eight-module OAM system, and must accept ROCm rather than CUDA. Anyone weighing MI300X against this part is buying memory, because the compute is the same.

Role in the data center

Large-model inference where 256 GB per accelerator sets the node count, plus training and double-precision work at 81.7 TFLOPs FP64. Lands in a 36 to 65 kW rack density band as an air-cooled OAM node, the wide range reflecting the unresolved gap between usable PSU capacity and the estimated node draw.

Power envelope

Per accelerator
1,000 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
DLC strongly preferred — Derived: Per-accelerator power at or above 1000 W. Air-cooled OEM systems exist at this level, at reduced rack density and higher airflow; liquid cooling is the practical choice above roughly 40 kW per rack. Configuration decides, not the accelerator alone.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureCDNA 3
Form factorOAM module
Memory capacity256 GB HBM3E
Memory bandwidth6 TB/s
Memory interface8192 bit
Last-level cache256 MB
Typical board power1000 W peak
External power connection54V UBB
Cooling modePassive OAM, air
Compute units304 CUs
Matrix cores1216
Peak engine clock2100 MHz
Dense FP82.61 PFLOPs
Accelerator-only load, 8 modules8.0 kW
Total DRAM bandwidth vs export gate6,000 below 6,500 GB/s
Launch date10 October 2024

Technical summary

Architecture: CDNA 3, TSMC 5nm and 6nm FinFET, 153 billion transistors. Memory: 256 GB HBM3E, 6 TB/s peak, 8192-bit interface, 6 GHz memory clock. Last-level cache: 256 MB; peak engine clock 2100 MHz. Typical board power: 1000 W peak, passive OAM, 54V UBB external power. Compute units: 304; stream processors: 19,456; matrix cores: 1216. Dense compute: FP8 2.61 PFLOPs, FP16 and BF16 1.3 PFLOPs, TF32 653.7 TFLOPs. With structured sparsity (vendor figure): FP8 5.22 PFLOPs, INT8 5.22 POPs. FP64 81.7 TFLOPs, FP32 163.4 TFLOPs; identical to MI300X on every precision. Host interface: PCIe 5.0 x16; 8 Infinity Fabric links at 128 GB/s peak. Accelerator-only load, 8-module node: 8.0 kW; proxy-estimated node draw, not an OEM figure, 16.2 kW. Export: 6,000 GB/s total DRAM bandwidth, below the 6,500 GB/s gate.

Major variations

MI300X is the same CDNA 3 silicon at 192 GB HBM3, 5.3 TB/s and 750 W. Compute is identical between the two on every precision AMD publishes, so the entire difference is 64 GB of capacity, 0.7 TB/s of bandwidth and 250 W of power. MI350X and MI355X are the CDNA 4 successors at 288 GB HBM3E and 8 TB/s on a 3nm node, with roughly double the FP8 throughput from fewer compute units. Both sit above the 6,500 GB/s export gate that this module clears, which is a rare case of a newer part being harder to ship than the one it replaces.

Configurations and options

The purchasable unit is an eight-module platform on a UBB 2.0 baseboard. Supermicro's 8U H14 system carries eight of these modules with 2 TB of HBM3E in a single node, 400 Gbps of networking dedicated to each accelerator, and six or eight 3000 W N+N Titanium power supplies, which is 9 kW or 12 kW usable after redundancy. Cooling is dual-zone air with five front and five rear counter-rotating fans. Supermicro's AS-8126GS-TNMR 8U system supports this module and the MI350X, with six 5250 W supplies in a 3+3 arrangement for 15.75 kW usable and up to fourteen heavy-duty fans. Read the redundancy arithmetic carefully: N+N on six 3000 W units means three active and three redundant, so 9 kW usable rather than 18 kW installed.

Compatibility and dependencies

Requires an eight-module OAM chassis on a UBB 2.0 baseboard and cannot be retrofitted into a PCIe server. Host attachment is PCIe 5.0 x16 per module, and accelerator-to-accelerator traffic runs over eight Infinity Fabric links at 128 GB/s each. External power arrives on a 54V UBB connection rather than a card-style connector. Cooling is passive OAM, so the chassis supplies all airflow. Accelerator-only load for eight modules is 8.0 kW and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a cross-vendor proxy because the manufacturer publishes no node figure, suggests roughly 16.2 kW, against 9 kW to 15.75 kW of usable PSU capacity depending on configuration. That gap is unresolved and is the reason the rack band below is stated as a range. Plan alongside racks, busway, PDUs and CRAC or chiller capacity.

Pricing and availability

Evidence: analyst estimate. Reference band for another model: 10000–15000 $ per accelerator (AMD Instinct MI300X comparable, Feb 2024). It is not a price for this model. Trend: falling No price series exists for this module. The nearest dated anchor is a sibling estimate: Citi put MI300X at roughly $10,000 to $15,000 per unit on 2 February 2024. One estimate for a different part, more than two years old, is not a trend line. Direction is downward on generational grounds, because CDNA 4 shipped in June 2025 and the MI400 series was announced in 2026. Lead time, new: Ordered as an eight-module OAM system through Supermicro, Dell or another OEM rather than as a module; no discrete module lead time is published.. Lead time, used or refurbished: Resale of this generation is a market for whole 8U chassis, not for modules.. Warranty, new: OEM system warranty set by the server vendor rather than by AMD for the module alone.. Warranty, used: No manufacturer warranty transfer was documented. Every warranty observed in this category is a reseller warranty written by the dealer..

Lifecycle and maintenance

Launched 10 October 2024 and superseded within AMD's own line by the MI350 series in June 2025, which moved to CDNA 4 and 8 TB/s. The product page remains live with a full specification table, and BIS named this module in a rule effective 15 January 2026, so it retains current regulatory relevance. The usable service-life evidence in this category is observed operator practice: depreciation runs 4 years at Nebius, 5 at Lambda Labs and 6 at CoreWeave, against an observed 7.5-year service life on an earlier NVIDIA generation.

Common failure points

Inspection checklist

Driver-reported memory capacity must read 256 GB: 192 GB indicates an MI300X sold as this part. Stream benchmark bandwidth within 10 percent of 6 TB/s: 5.3 TB/s is the MI300X figure and a red flag on identity. Sustained power draw within 5 percent of 1000 W peak TBP: a module capped near 750 W may be misidentified. Compute units must enumerate 304 and stream processors 19,456, with 1216 matrix cores: a low count means a re-marked die. Peak engine clock reaches 2100 MHz and memory clock 6 GHz under load: sustained shortfall indicates power or thermal limiting. HBM3E correctable and uncorrectable ECC counters read before purchase and after a 24-hour soak: HBM is not field-serviceable. 54V UBB power connection inspected for discolouration, pitting or melted keying: a 1000 W module is unforgiving here.

Procurement channels

New: OEM system orders through Supermicro, Dell, HPE and comparable vendors, priced as an eight-module node rather than per accelerator. Cloud rental is available at a median near $3.10 per accelerator-hour. Secondary, as of 10 September 2026: Compute Exchange publishes indicative secondary bands for seven NVIDIA parts and lists no AMD or Intel accelerator, and no L4. The market structure explains it: OAM modules ship on a baseboard inside an OEM system rather than discretely, so what trades second-hand is the 8U chassis rather than the module, and no open-channel Instinct listing appears at the brokers surveyed on that date.

Regional notes

The module BIS named: the rule gives its example as the H200 or the MI325X. Reexport and transfer licences stay under presumption of denial regardless. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Published bandwidth is 6,000 GB/s (6 TB/s), 500 GB/s under the gate, and no TPP figure is published, so bandwidth alone does not establish eligibility.