NVIDIA L4

NVIDIA L4 — gpu, 24 GB
NVIDIA L4

The 72 W floor of this category, single-slot low-profile with no supplemental power connector

NVIDIA L4 is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

NVIDIA L4 dossier: 24 GB, 300 GB/s, 72 W single-slot low-profile PCIe, rack kW at ten cards a node, disputed export status and observed pricing.

The NVIDIA L4 is an Ada Lovelace accelerator on a single-slot low-profile PCIe card. NVIDIA publishes 24 GB of memory at 300 GB/s, a maximum thermal design power of 72 W, PCIe Gen4 x16 at 64 GB/s, and media engines comprising two encoders, four decoders and four JPEG decoders. 72 W is the headline and it is a facility number rather than a specification. The card needs no supplemental power connector, occupies one low-profile slot, and can be deployed in 1U and edge servers that cannot host any other part in this category. It sits at the bottom of a range topping out at 1,400 W for an MI355X, a span of 19.4 times, and one MI355X module draws more than nineteen of these cards. Ten of them in a 5U chassis is 720 W of card load. One specification trap is worth naming because it is the most common L4 error in circulation. NVIDIA's table footnote states the sparsity figures are one-half lower without sparsity, so the widely quoted 485 TFLOPS of FP8 is a sparsity figure and dense FP8 is roughly 242 TFLOPS. A reseller page observed on 10 September 2026 listed this part number with 48 GB and 300 W, wrong by two times on memory and 4.2 times on power.

Why it matters

This card anchors the low end of the power comparison that makes this category worth publishing. An L4 at 72 W in a PCIe slot and an MI355X at 1,400 W in an OAM module are not the same kind of problem, and a planner who sizes from FLOPS will miss that entirely. It is also the only part here with no secondary channel presence at all.

Who buys it

Operators deploying inference or video transcode at density or at the edge, where slot count and watts decide the design rather than throughput. It suits 1U servers, telco cabinets and retail back-offices, and ten fit a 5U chassis. Buyers cross-shopping it against an L40S are choosing a fifth of the power and a single slot over 48 GB and 864 GB/s of bandwidth.

Role in the data center

Inference serving and video transcode, at density in a data hall or at the edge. Four media engines make it a transcode part as much as an accelerator. Lands near a 12 kW rack density band at eight 5U nodes of ten cards, the lowest in this category.

Power envelope

Per accelerator
72 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
Air-coolable — Derived: Per-accelerator power below 400 W — standard air-cooled halls.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureAda Lovelace
Memory capacity24 GB
Memory bandwidth300 GB/s
Maximum thermal design power72 W
Form factorSingle-slot low-profile PCIe
InterconnectPCIe Gen4 x16 64 GB/s
Dense FP3230.3 TFLOPS
FP8 with sparsity485 TFLOPS, dense roughly 242
INT8 with sparsity485 TOPS
Video encoders2 NVENC
Video decoders4 NVDEC
JPEG decoders4
Card load, ten cards in 5U0.72 kW
Estimated rack density12 kW at eight nodes of ten
Total DRAM bandwidth vs export gate300 below 6,500 GB/s
Named in NVIDIA's 17 October 2023 export filingNo

Technical summary

Architecture: Ada Lovelace. Memory: 24 GB, 300 GB/s bandwidth; NVIDIA does not state the memory technology. Maximum thermal design power: 72 W, with no supplemental power connector required. Form factor: single-slot low-profile PCIe. Interconnect: PCIe Gen4 x16 at 64 GB/s. Dense compute: FP32 30.3 TFLOPS; dense FP8 roughly 242 TFLOPS. With sparsity (vendor figures): TF32 120, FP16 242, BF16 242, FP8 485 TFLOPS, INT8 485 TOPS. Media engines: 2 encoders, 4 decoders, 4 JPEG decoders. Ten cards in a 5U chassis: 0.72 kW of card load; proxy-estimated node draw, not an OEM figure, 1.5 kW. Rack density near 12 kW at eight 5U nodes of ten cards. Lowest-power part in this category, against 1,400 W at the top, a 19.4 times span.

Major variations

This is a single SKU rather than a family, so there is no memory or power variant to enumerate. The nearest relatives are the T4 16GB, its direct predecessor in the same low-power inference role, and the L40S at 48 GB and 350 W, which is the same architecture scaled up nearly five times in power. NVIDIA does not state the memory technology for this card, only its 24 GB capacity and 300 GB/s bandwidth.

Configurations and options

Deployed in conventional PCIe servers and in edge and 1U chassis. Supermicro's 5U SYS-521GE-TNRT lists support for this card and the L40S, taking up to ten single-width cards, with four 2700 W Titanium redundant supplies. Ten cards in that chassis is 0.72 kW of card load, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a proxy because the manufacturer publishes no node figure, suggests roughly 1.5 kW. Eight such nodes fill a 40U span at roughly 12 kW a rack, which is the lowest rack density in this category by a wide margin. NVIDIA publishes no configurable-TDP range, so 72 W should be treated as fixed.

Compatibility and dependencies

Fits any PCIe Gen4 x16 slot with single-slot low-profile clearance. At 72 W it requires no supplemental power connector, which is what lets it go into 1U and edge servers that cannot host any other part in this category. There is no NVLink, so multi-card work scales over PCIe Gen4 at 64 GB/s and over the network. The constraint that catches people is airflow at the slot rather than power at the rack. Low-profile slots often sit in the worst-cooled part of a chassis, and ten cards concentrate 720 W into one span. Validate airflow at the slot, not for the chassis. Plan alongside racks and PDUs; at 12 kW a rack this part has no meaningful cooling-plant implication, which is precisely the point of buying it.

Pricing and availability

Evidence: channel observation · used condition. Reference band for another model: 2200–2800 $ per accelerator (NVIDIA T4 16GB PCIe comparable). It is not a price for this model. Trend: stable No series can be written for this card. Only two priced points exist, one new dealer price of $3,410 and one undated aggregator figure of $2,800, and Compute Exchange does not list the part at all. The direction is flat rather than falling, and an aggregator reckoning the market figure at 12 percent above original list supports that. A part still bought new for inference does not get flooded by an upgrade cycle. For the role, the predecessor T4 16GB moved from $2,500 to $3,500 in mid-2025 down to $2,200 to $2,800 in mid-2026, a 15 percent decline, which is the closest available proxy for how this card will age. Lead time, new: In production and available from resellers such as Dihuni at $3,410. Also configured into new servers by Supermicro, Dell and comparable vendors.. Lead time, used or refurbished: Compute Exchange does not publish a band for this part.. Warranty, new: Reseller or OEM system warranty.. Warranty, used: No manufacturer warranty transfer was documented anywhere in this category. Every warranty observed is a reseller warranty written by the dealer..

Lifecycle and maintenance

Still in production and sold new in the $2,800 to $3,410 range as of 10 September 2026. Because it is often deployed at the edge, environmental conditions rather than electrical stress are the likely life-limiting factor. Operator depreciation practice across the category runs 4 to 6 years.

Common failure points

Inspection checklist

Driver-reported memory must read 24 GB: a 48 GB reading means the card is not this part, whatever the listing says. Stream benchmark bandwidth within 10 percent of 300 GB/s: a low result points to degraded memory or clock limiting. Sustained power draw at or under 72 W under full load: a card drawing 300 W is a different part entirely. Confirm the card is single-slot low-profile with no supplemental power connector: any auxiliary connector is wrong. Part number matches 900-2G193-0000-001 or 900-2G193-0000-000: reseller pages for this number carry wrong specifications. Memory ECC error counters and retired-page logs read before purchase: repeat after a 24-hour soak at full load. Host PCIe link trains at Gen4 x16 for 64 GB/s: negotiating x8 halves host bandwidth in a dense node. Media engines enumerate 2 encoders, 4 decoders and 4 JPEG decoders: a short count indicates a re-marked die.

Procurement channels

New: reseller listings such as Dihuni at $3,410, and configuration into new servers by Supermicro, Dell and comparable vendors. Cloud access is widely available from providers offering low-cost inference instances. Secondary, as of 10 September 2026: Compute Exchange publishes indicative bands for seven NVIDIA parts and does not list this one, and no dealer, broker or ITAD inventory page publishes a price for it.

Regional notes

Control status is disputed: NVIDIA's Form 8-K of 17 October 2023 does not name this card and a July 2026 restrictions compilation does not list it, while one trade outlet reports it added alongside the RTX 6000 Ada and RTX A6000. Get a classification check. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Its 300 GB/s is far under the gate.