AMD Instinct MI300X

AMD Instinct MI300X — gpu, 192 GB HBM3
AMD Instinct MI300X

CDNA 3 OAM accelerator carrying 192 GB of HBM3 at 750 W, eight to a UBB 2.0 baseboard inside an OEM node.

AMD Instinct MI300X is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

AMD Instinct MI300X dossier: 192 GB HBM3, 5.3 TB/s, 750 W OAM, node and rack kW, cooling mode, export status and what the secondary market shows.

The AMD Instinct MI300X is a CDNA 3 accelerator on an OAM module, launched 6 December 2023. AMD publishes 192 GB of HBM3 and 5.3 TB/s of memory bandwidth against a typical board power of 750 W peak. It is not a card and cannot be added to an existing server. Eight modules mount to a UBB 2.0 baseboard inside a purpose-built chassis such as Dell's 6U PowerEdge XE9680 or Supermicro's 8U H14. Memory capacity is the whole procurement argument. 192 GB per module against 80 GB on an H100 SXM5 decides how many accelerators a given model needs at all, which is a different question from how fast each one runs. AMD also publishes more openly than its competitors here, so every figure on this page traces to AMD's own product page rather than to an analyst estimate. For a facility, 750 W is the middle of this category. Eight modules draw 6.0 kW of accelerator load before the host, and applying NVIDIA's DGX A100 node ratio of 2.03x as a cross-vendor proxy, since the manufacturer publishes no node figure, suggests roughly 12.2 kW. Air-cooled OAM at roughly 36 to 50 kW a rack is achievable in a well-contained air hall, which is worth knowing before anyone assumes OAM implies liquid cooling.

Why it matters

MI300X is the reference point for what an OAM accelerator does to a facility. At 750 W it sits between a 350 W PCIe card and a 1,400 W liquid-cooled module, and it is still air-coolable at eight to a baseboard. It also anchors the memory argument in this category at 192 GB, which is 2.4 times an H100 SXM5.

Who buys it

Operators who want large-memory training and inference capacity outside the NVIDIA supply queue and can accept the ROCm software stack rather than CUDA. The 192 GB memory pool suits large-model inference where capacity per accelerator decides the node count. Buyers must be purchasing a whole OAM system, because the module is not sold or deployed discretely.

Role in the data center

Large-model inference and distributed training, plus genuine double-precision work at 163.4 TFLOPs FP64 matrix. Lands in a 36 to 50 kW rack density band as an air-cooled OAM node, which is inside what a well-contained air hall can deliver.

Power envelope

Per accelerator
750 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
DLC recommended — Derived: Per-accelerator power at or above 700 W — air cooling possible at reduced rack density.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureCDNA 3
Form factorOAM module
Memory capacity192 GB HBM3
Memory bandwidth5.3 TB/s
Typical board power750 W peak
Cooling modePassive OAM, air
Compute units304 CUs
Stream processors19,456
Dense FP82.61 PFLOPs
Dense FP16 and BF161.3 PFLOPs
FP64 matrix163.4 TFLOPs
Host interfacePCIe 5.0 x16
Infinity Fabric link bandwidth128 GB/s
Platform aggregate peer-to-peer896 GB/s
Accelerator-only load, 8 modules6.0 kW
Launch date6 December 2023

Technical summary

Architecture: CDNA 3, TSMC 5nm and 6nm FinFET. Memory: 192 GB HBM3, 5.3 TB/s peak bandwidth. Typical board power: 750 W peak, passive OAM cooling. Compute units: 304; stream processors: 19,456. Dense compute: FP8 2.61 PFLOPs, FP16 and BF16 1.3 PFLOPs, TF32 653.7 TFLOPs. With structured sparsity (vendor figure): FP8 5.22 PFLOPs, FP16 and BF16 2.61 PFLOPs. FP64: 81.7 TFLOPs vector, 163.4 TFLOPs matrix; FP32 163.4 TFLOPs. Host interface: PCIe 5.0 x16; Infinity Fabric links 128 GB/s. Platform: 8 modules on UBB 2.0, 1.5 TB HBM3 total, 896 GB/s aggregate peer-to-peer. Accelerator-only load, 8-module node: 6.0 kW; proxy-estimated node draw, not an OEM figure, 12.2 kW.

Major variations

MI325X is the memory-and-power upgrade within the same CDNA 3 generation: 256 GB HBM3E at 6 TB/s and 1000 W against 192 GB HBM3 at 5.3 TB/s and 750 W. Compute is identical between them on every precision AMD publishes, with the same 304 compute units and 19,456 stream processors. MI308 is the China-market derivative of this generation. AMD received US export licences to resume MI308 shipments in 2025 alongside a reported agreement to pay 15 percent of China revenue. It has no comparable public specification page and is named here only as lineage.

Configurations and options

The purchasable unit is an eight-module platform on a UBB 2.0 baseboard: 8 modules, 1.5 TB of HBM3, 896 GB/s aggregate bi-directional peer-to-peer bandwidth, baseboard 417 mm by 553 mm. Documented host systems are Dell's 6U PowerEdge XE9680, which lists this module at 192 GB and 750 W and is air-cooled only, and Supermicro's 8U H14 line. AMD publishes no configurable-TDP range for this module, unlike the Grace-paired NVIDIA superchip that is documented as programmable across a power band. Whether it can be power-capped in firmware was not established from AMD documentation.

Compatibility and dependencies

Requires an eight-module OAM chassis on a UBB 2.0 baseboard and cannot be retrofitted into a PCIe server. Host attachment is PCIe 5.0 x16 per module; accelerator-to-accelerator traffic runs over Infinity Fabric at 128 GB/s per link, eight links per module. Cooling is passive OAM, so the chassis supplies all airflow and the module has no fan. Accelerator-only load for an eight-module node is 6.0 kW, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a cross-vendor proxy because the manufacturer publishes no node figure, suggests roughly 12.2 kW. The related facility systems are racks, busway, PDUs and CRAC or chiller capacity, since air cooling holds at this power level where it does not at 1,400 W.

Pricing and availability

Trend: falling No price series can be written for this module because no channel publishes one. The only dated points are a Citi analyst estimate of roughly $10,000 to $15,000 per unit from 2 February 2024 and an undated aggregator figure of $18,000. Two points more than two years apart, one an estimate and one undated, do not describe a trend. The direction is downward on generational grounds alone, since two newer AMD generations have shipped since launch. Lead time, new: Ordered as an eight-module OAM system through Dell, Supermicro or another OEM rather than as a module; no discrete module lead time is published.. Lead time, used or refurbished: Resale of this part is a market for whole 6U and 8U chassis, not for modules.. Warranty, new: OEM system warranty, set by the server vendor rather than by AMD for the module alone.. Warranty, used: No manufacturer warranty transfer was documented. Every warranty observed in this category is a reseller warranty written by the dealer..

Lifecycle and maintenance

Launched 6 December 2023 and two generations behind AMD's MI400 series, with the MI350 series shipping between them. AMD's Instinct landing page still presents the MI300 series as a current offering and the product page remains live with a full specification table. The best available service-life evidence in this category is observed hyperscaler practice rather than MTBF: operator depreciation schedules run 4 years at Nebius, 5 at Lambda Labs and 6 at CoreWeave, while Azure ran V100-based instances about 7.5 years before retirement.

Common failure points

Inspection checklist

Driver-reported memory capacity must read 192 GB: a short reading indicates a failed HBM stack or a re-marked module. Stream benchmark bandwidth within 10 percent of 5.3 TB/s: a low result points to a degraded stack or clock limiting. Sustained power draw within 5 percent of 750 W peak TBP: a module reading well under rating is throttling. Compute unit count must enumerate 304 CUs and 19,456 stream processors: a reduced count means a re-marked die. HBM correctable and uncorrectable ECC counters read before purchase and again after a 24-hour soak: HBM is not field-serviceable. Infinity Fabric link training across all 8 links at 128 GB/s: a dead link cannot be repaired in the field. Host PCIe link trains at Gen 5 x16: negotiating Gen 4 or x8 halves host bandwidth and indicates seating or firmware fault.

Procurement channels

New: OEM system orders through Dell, Supermicro, HPE and comparable vendors, priced as an eight-module node rather than per accelerator. Cloud rental runs $1.71 to $6.00 per accelerator-hour across TensorWave, Runpod, DigitalOcean, Crusoe, Vultr, Azure and Oracle. Secondary, as of 10 September 2026: Compute Exchange publishes indicative secondary bands for seven NVIDIA parts and lists no AMD or Intel accelerator, and no L4. The market structure explains it: OAM modules ship on a baseboard inside an OEM system rather than discretely, so what trades second-hand is the 8U chassis rather than the module, and no open-channel Instinct listing appears at the brokers surveyed on that date.

Regional notes

BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. This module publishes 5,300 GB/s (5.3 TB/s), under the gate, and no TPP figure is published. BIS named the MI325X and not this module, which sits further below the same gate; that is a reading of a rule naming the marginal case, not a determination.