Intel Gaudi 3

The one accelerator here with a published list price, and 24 ports of 200 GbE on the die for scale-out.
Intel Gaudi 3 is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.
Intel Gaudi 3 dossier: 128 GB HBM2E, 3.7 TB/s, 900 W OAM or 600 W PCIe, on-die 200 GbE, rack kW, export status and the published list price.
Intel Gaudi 3 is a third-generation accelerator built on a 5nm process, shipping in two form factors from one design. The HL-325L is an OCP Accelerator Module V2.0 part rated up to 900 W; the HL-338 is a full-height dual-slot PCIe card 10.5 inches long rated 600 W. Both carry 128 GB of HBM2E at 3.7 TB/s, 96 MB of on-die SRAM, eight matrix engines and 64 programmable tensor processor cores. Two things make it genuinely different from everything else in this category. First, the fabric is on the die: 24 ports of 200 GbE running RoCE v2, 9.6 terabits per second bidirectional, so scale-out uses standard Ethernet switching rather than a proprietary interconnect. Second, Intel published a list price, which no other vendor here does. An eight-accelerator universal baseboard kit including networking was offered to system providers at $125,000 at Computex 2024, which is $15,625 per accelerator on a stated divisor of eight. Intel labelled that guidance as being for modeling purposes only. The market position, covered below, is where the honest limits sit. Gaudi 3 is fully documented, purchasable and deployed on IBM Cloud, and Intel's own forward roadmap has moved past it. Both facts belong in a buying decision.
Why it matters
This is the only accelerator in the category with a manufacturer-published price, which makes it the one honest anchor in a market where every other figure is reseller-reported or analyst-estimated. It is also the clearest case of Ethernet scale-out replacing a proprietary fabric, and the only part here where liquid cooling buys thermal margin rather than extra watts.
Who buys it
Operators buying inference or training throughput on price rather than peak performance, who can build scale-out on Ethernet they already understand. The PCIe card suits adding accelerators to conventional servers with a 600 W slot budget. Buyers must weigh the software support horizon, which Intel's own executives identified as the reason uptake was slower than the company planned.
Role in the data center
Inference and distributed training in Ethernet-scale-out clusters, with the PCIe card serving inference and fine-tuning in conventional servers. The OAM configuration lands near a 58 kW rack density band; the 600 W PCIe card lands four to eight per node in a standard rack.
Power envelope
- Per accelerator
- 900 W — Verified: Maximum configurable board power, manufacturer specification. source
- Cooling class
- DLC recommended — Derived: Per-accelerator power at or above 700 W — air cooling possible at reduced rack density.
No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.
Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.
Key specifications
| Process | 5 nm |
|---|---|
| Memory capacity | 128 GB HBM2E |
| Memory bandwidth | 3.7 TB/s |
| On-die SRAM | 96 MB |
| OAM variant TDP (HL-325L) | 900 W |
| PCIe variant TDP (HL-338) | 600 W |
| OAM cooling | Passive or liquid, both at 900 W |
| PCIe form factor | Full-height dual-slot, 10.5 inches |
| Compute engines | 8 matrix engines plus 64 tensor processor cores |
| FP8 and BF16 throughput | 1.8 PFLOPS each |
| On-die networking | 24 ports of 200 GbE RoCE v2 |
| Network bandwidth | 9.6 Tb/s bidirectional |
| Host interface | PCIe Gen 5.0 x16 128 GB/s |
| Accelerator-only load, 8 OAM modules | 7.2 kW |
| Published list price, 8-accelerator kit | 125,000 USD, Computex 2024 |
| Total DRAM bandwidth vs export gate | 3,700 below 6,500 GB/s |
Technical summary
Architecture: Intel Gaudi, third generation, 5nm process. Memory: 128 GB HBM2E, 3.7 TB/s peak bandwidth, 96 MB on-die SRAM. Compute engines: 8 matrix engines plus 64 programmable tensor processor cores. Throughput: 1.8 PFLOPS of FP8 and 1.8 PFLOPS of BF16, equal on both formats. Datatypes: FP8, BF16, FP16, TF32, FP32. OAM variant HL-325L: up to 900 W, OCP Accelerator Module V2.0, passive or liquid, both at 900 W. PCIe variant HL-338: 600 W, full-height dual-slot, 10.5 inches, passive air. Networking: 24 ports of 200 GbE RoCE v2, 9.6 Tb/s bidirectional, on the die. Host interface: PCIe Gen 5.0 x16, 128 GB/s bidirectional. Baseboard: HLB-325 carries eight accelerators in a non-blocking all-to-all arrangement. Accelerator-only load, 8 OAM modules: 7.2 kW; proxy-estimated node draw, not an OEM figure, 14.6 kW. Published list price: $125,000 per eight-accelerator kit, Computex 2024.
Major variations
HL-325L is the OAM form factor at up to 900 W, OCP Accelerator Module V2.0 compliant, mounting eight to an HLB-325 baseboard. HL-338 is the PCIe add-in card at 600 W, full-height dual-slot and 10.5 inches long. Memory, bandwidth, SRAM and compute engines are identical between them, so the only differences are power rating, form factor and how the RoCE ports are presented. HL-328 and HL-388 are China-market designations, reported as launching in June and September 2024. Both are reported with the same 128 GB HBM2E, 3.7 TB/s, 96 MB cache and PCIe 5.0 x16 as the standard part, so whatever compliance margin they carry is not visible in the memory subsystem. They are named here as lineage only.
Configurations and options
The OAM part is bought as an eight-accelerator kit on a universal baseboard, the HLB-325, in a non-blocking all-to-all arrangement. Intel documents a 512-node cluster topology reaching 4,096 accelerators, networked with 48 spine switches of 64 ports at 800 Gbps. The kit price includes the baseboard and the on-die Ethernet, which removes the separate InfiniBand adapters a proprietary fabric would need. The PCIe card is deployed in conventional servers with a 600 W dual-slot thermal budget. Intel's PCIe brief does not state the port count for that variant, nor its power connector type or availability date. No configurable-TDP range is published for either form factor.
Compatibility and dependencies
The OAM part requires an HLB-325 baseboard and an eight-module chassis, and cannot be retrofitted into a PCIe server. The PCIe card fits any PCIe Gen 5 x16 slot that can supply and cool 600 W in a dual-slot full-height space, which is a demanding budget for a general-purpose server. Scale-out is the distinguishing dependency and it is a network requirement rather than a fabric one. The die carries 24 ports of 200 GbE running RoCE v2, so a cluster needs 200 GbE and 800 Gbps Ethernet switching instead of a proprietary switch. Accelerator-only load for eight OAM modules is 7.2 kW, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a cross-vendor proxy because the manufacturer publishes no node figure, suggests roughly 14.6 kW and a four-node rack near 58 kW. Plan alongside racks, busway, PDUs and liquid cooling.
Pricing and availability
Published price range: $125,000 per eight-accelerator UBB kit, which is $15,625 per accelerator on a divisor of eight (Intel list price, Computex 2024, guidance for modeling purposes only). Evidence: list price · Kit price divided by eight accelerators. Comparable secondary-market band: 8125–15625 $ per accelerator (Intel Gaudi 2 and Gaudi 3 eight-accelerator kits comparable, Jun 2024). Trend: falling Only two dated points exist, both Intel list prices from one announcement. At Computex 2024 Intel offered an eight-accelerator Gaudi 2 kit at $65,000 and an eight-accelerator Gaudi 3 kit at $125,000, which is $8,125 and $15,625 per accelerator on a divisor of eight. Intel positioned the Gaudi 3 kit at roughly two-thirds the cost of comparable platforms and the Gaudi 2 kit at one-third. No later price is published in any channel as of 10 September 2026, so no market series can be written. Direction is downward on generational grounds: Intel missed its $500 million Gaudi revenue target for 2024, and its roadmap has moved to Crescent Island and Jaguar Shores rather than a Gaudi 4. Lead time, new: Intel sold eight-accelerator UBB kits to system providers rather than direct.. Lead time, used or refurbished: Compute Exchange does not publish a band for this part.. Warranty, new: OEM system warranty set by the server vendor. Intel sold the accelerator to system providers as a baseboard kit rather than as a retail part.. Warranty, used: No manufacturer warranty transfer was documented. Every warranty observed in this category is a reseller warranty written by the dealer..
Lifecycle and maintenance
Launched 2024 and available on IBM Cloud from 4 April 2025 in Frankfurt and Washington DC. Intel's forward data center accelerator roadmap has moved past this line: Falcon Shores was redirected to an internal test chip rather than brought to market in January 2025, and Intel's current direction is Crescent Island, a 160 GB LPDDR5X inference part on Xe3P sampling in the second half of 2026, plus Jaguar Shores at rack scale. The software support horizon is the risk that matters more than the hardware.
Common failure points
- HBM2E memory stack — Rising correctable ECC counts then uncorrectable errors and row retirement under sustained memory pressure
- On-die 200 GbE ports — One or more of the 24 RoCE v2 ports fails to link, collapsing collective throughput across the cluster
- Passive cooling airflow dependency — Throttling under full load in a chassis not qualified for the card's rating
- Thermal interface and cold plate — Clocks decline at constant load as delta-T rises over weeks
- Gaudi software stack support — Part functions but the current framework release will not target it
- PCIe link degradation — Link trains below Gen 5.0 x16 or reports corrected-error storms on the host link
Inspection checklist
Driver-reported memory capacity must read 128 GB HBM2E across 8 stacks: a short reading indicates a failed stack. Stream benchmark bandwidth within 10 percent of 3.7 TB/s: a low result points to a degraded stack or clock limiting. Sustained power draw within 5 percent of rating: 900 W for an HL-325L OAM module, 600 W for an HL-338 PCIe card. On-die SRAM enumerates 96 MB: the driver must also report 8 matrix engines and 64 tensor processor cores. All 24 ports of 200 GbE must link up, totalling 9.6 Tb/s bidirectional: a dead port is not field-repairable. HBM2E correctable and uncorrectable ECC counters read before purchase and after a 24-hour soak: HBM is not serviceable. Host PCIe link trains at Gen 5.0 x16 for 128 GB/s: negotiating Gen 4 or x8 halves host bandwidth. OAM revision confirmed as OCP Accelerator Module V2.0: UBB revision checked on the HLB-325 baseboard silkscreen.
Procurement channels
New: Intel sold eight-accelerator UBB kits to system providers, so the buying route is an OEM system rather than a discrete accelerator. Cloud access is available on IBM Cloud, which deployed Gaudi 3 in Frankfurt and Washington DC from 4 April 2025. Secondary, as of 10 September 2026: Compute Exchange publishes indicative bands for seven NVIDIA parts and none for any Intel or AMD accelerator, and no dealer, broker or ITAD inventory page publishes a Gaudi 3 price. There is no listed secondary price to quote from, so expect a quoted OEM or ITAD route.
Regional notes
A licence is required to ship it to China. Intel notified Chinese customers in the week of 7 April 2025 of licences at 1,400 GB/s DRAM, 1,100 GB/s I/O or 1,700 GB/s combined; at 3.7 TB/s it is 2.6x over. HL-328 and HL-388 are China-market SKUs. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Its 3,700 GB/s is under the gate, the H200 and MI325X tier.