NVIDIA L40S

NVIDIA L40S — gpu, 48 GB GDDR6 with ECC
NVIDIA L40S

Ada Lovelace PCIe card with 48 GB GDDR6 at 350 W, and a real used market at $5,500 to $9,000

NVIDIA L40S is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.

NVIDIA L40S dossier: 48 GB GDDR6, 864 GB/s, 350 W passive PCIe, rack kW, export status, and a used range of $5,500 to $9,000 from three channels.

The NVIDIA L40S is an Ada Lovelace accelerator on a passive dual-slot PCIe card measuring 4.4 inches high by 10.5 inches long. NVIDIA publishes 48 GB of GDDR6 with ECC, 864 GB/s of memory bandwidth and 350 W maximum power, with 18,176 CUDA cores, 568 fourth-generation tensor cores and 142 third-generation ray tracing cores. It carries four DisplayPort 1.4a outputs and a 16-pin power connector. Two facts a buyer should not miss. It uses GDDR6 rather than HBM, so 864 GB/s against an A100 SXM4's 2,039 GB/s is a 2.4 times bandwidth deficit that makes it a poor fit for memory-bandwidth-bound work regardless of its FP8 throughput. And it is PCIe Gen4, not Gen5, with no NVLink and no multi-instance partitioning, so multi-card scaling runs over PCIe and the network only. The used discount is only about 25 percent off new, which is far shallower than the A100 alongside it. That is what a Moderate supply signal looks like: the part is still in production and still bought new, so no upgrade cycle is flooding resale.

Why it matters

L40S is the price-discovery anchor of the inference tier: three independent channels agree inside $5,500 to $9,000, which almost nothing else in this category manages. It is also export-controlled, named in NVIDIA's own October 2023 filing alongside H100, which surprises most buyers of a mid-priced inference card and changes how it can be moved.

Who buys it

Operators adding inference, fine-tuning, graphics, virtual desktop or transcode capacity to an existing server fleet rather than building a new hall. At 350 W passive in a standard dual-slot PCIe space it drops into conventional chassis, eight to a 5U node. Buyers cross-shopping it against a used A100 80GB PCIe are choosing FP8 and a newer architecture over memory bandwidth.

Role in the data center

Inference, fine-tuning, virtual desktop, rendering and video transcode rather than large-scale training. Ray tracing cores and four DisplayPort outputs make it a graphics part as much as an accelerator. Lands near a 40 kW rack density band at seven 5U nodes of eight cards.

Power envelope

Per accelerator
350 W — Verified: Maximum configurable board power, manufacturer specification. source
Cooling class
Air-coolable — Derived: Per-accelerator power below 400 W — standard air-cooled halls.

No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.

Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.

Key specifications

ArchitectureAda Lovelace
Memory capacity48 GB GDDR6 with ECC
Memory bandwidth864 GB/s
Maximum power350 W
Thermal solutionPassive, server airflow
Form factor4.4 in H x 10.5 in L, dual slot
InterconnectPCIe Gen4 x16 64 GB/s bidirectional
NVLink supportNo
Multi-instance partitioningNo
CUDA cores18,176
Dense FP8733 TFLOPS
Dense FP3291.6 TFLOPS
Power connector16-pin
Accelerator-only load, 8 cards2.8 kW
Used and refurbished range5,500 to 9,000 USD per card, three channels
Total DRAM bandwidth vs export gate864 below 6,500 GB/s

Technical summary

Architecture: Ada Lovelace; 18,176 CUDA cores, 568 tensor cores, 142 ray tracing cores. Memory: 48 GB GDDR6 with ECC, 864 GB/s bandwidth. Maximum power: 350 W; thermal solution passive, server airflow required. Form factor: 4.4 inch high by 10.5 inch long, dual slot; 16-pin power connector. Interconnect: PCIe Gen4 x16 at 64 GB/s bidirectional; no NVLink; no multi-instance partitioning. Dense compute: FP32 91.6 TFLOPS, TF32 183, BF16 and FP16 362.05, FP8 733, INT8 733 TOPS. With sparsity (vendor figure): TF32 366, BF16 and FP16 733, FP8 1,466 TFLOPS. Display outputs: 4 x DisplayPort 1.4a; virtual GPU supported. Accelerator-only load, 8 cards: 2.8 kW; proxy-estimated node draw, not an OEM figure, 5.7 kW. Rack density near 40 kW at seven 5U nodes of eight cards. Export: named in NVIDIA's Form 8-K of 17 October 2023 as licence-controlled.

Major variations

This is a single SKU rather than a family, so there is no memory or power variant to enumerate. The related L40 is a distinct part and is named separately in NVIDIA's October 2023 export filing. The card also ships as PNY SKU NVL40STCGPU-KIT.

Configurations and options

Deployed in conventional PCIe servers. Supermicro's 5U SYS-521GE-TNRT lists support for this card and the L4, taking up to two quad-width, two triple-width, eight double-width or ten single-width cards, with four 2700 W Titanium redundant supplies and up to ten heavy-duty fans plus an optional direct-to-chip cold plate. Eight cards in that chassis is 2.8 kW of card load, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a proxy because the manufacturer publishes no node figure, suggests roughly 5.7 kW. Seven such nodes fill a 35U span at roughly 40 kW a rack. NVIDIA publishes no configurable-TDP range for this card, so 350 W should be treated as fixed.

Compatibility and dependencies

Fits any PCIe Gen4 x16 slot with a dual-slot 4.4 by 10.5 inch space, a 16-pin power feed and airflow rated for a 350 W passive card. There is no NVLink and no multi-instance partitioning, so multi-card work scales over PCIe Gen4 at 64 GB/s and over the network, which is a real limit for training. The bandwidth constraint is the one to design around. 864 GB/s of GDDR6 against 1,935 GB/s on a used A100 80GB PCIe means bandwidth-bound inference belongs on the A100 even though this card has FP8 and the A100 does not. Plan alongside racks, PDUs and CRAC or chiller capacity; the optional cold plate on the 5U chassis makes liquid cooling relevant at high card counts.

Pricing and availability

Published price range: $5,500 to $9,000 per card used or refurbished, three independent channels, Q3 2026. Evidence: channel observation · used or refurbished condition · Single card. Trend: falling The curve here is shallow and that is the finding. sell-server.com records $8,000 to $10,000 in mid-2025 falling to $7,000 to $9,000 in mid-2026, a 12 percent decline over the year. Compute Exchange puts used stock at $5,500 to $7,500 for Q3 2026 with a Moderate supply signal, and an aggregator reckons the $9,000 market figure at 13 percent above the original list price. Set that against A100 SXM4 at roughly 20 percent a year and the difference is a product in its selling life against one in its disposal life. A part still bought new does not get flooded by an upgrade cycle. Do not cite the PC Server and Parts range as independent corroboration; it restates sell-server.com exactly. Lead time, new: In production. Wiredzone showed the card out of stock and available only as part of a Supermicro server assembly; Dihuni listed it new at $9,100.. Lead time, used or refurbished: 2 to 4 weeks from a secondary dealer per sell-server.com, or in stock from a refurbisher. This is among the most liquid items in the category.. Warranty, new: 3 years with 30 days advance replacement for dead-on-arrival units, per the Wiredzone listing, plus 2-week replacement thereafter.. Warranty, used: 3-year reseller warranty on refurbished stock per Network Outlet. No manufacturer warranty transfer was documented anywhere in this category..

Lifecycle and maintenance

Still in production and still sold new at $9,100 to $9,294 as of 10 September 2026, with Compute Exchange rating secondary supply as Moderate rather than High. That combination puts the card in its selling life rather than its disposal life, and it is why the used discount is roughly 25 percent rather than the 45 percent an A100 carries. The card is passive, so its service life is a function of host airflow more than of anything on the board. Operator depreciation practice across the category runs 4 to 6 years, against an observed 7.5-year service life on an earlier generation.

Common failure points

Inspection checklist

Driver-reported memory must read 48 GB GDDR6 with ECC enabled: a short reading indicates a failed module. Stream benchmark bandwidth within 10 percent of 864 GB/s: a low result points to degraded GDDR6 or clock limiting. Sustained power draw within 5 percent of 350 W under load: a card capped well below is throttling. Core counts must enumerate 18,176 CUDA, 568 tensor and 142 ray tracing cores: a low count means a re-marked die. GDDR6 ECC error counters and retired-page logs read before purchase: repeat after a 24-hour soak at full load. Host PCIe link trains at Gen4 x16 for 64 GB/s: this card is Gen4, so a Gen5 reading indicates a different part. Confirm the card reports no NVLink and no multi-instance support: a listing claiming either is misdescribed. 16-pin power connector inspected for discolouration, pitting or melted keying: a 350 W card is unforgiving here.

Procurement channels

New: Dihuni at $9,100, reduced from $11,975, and Wiredzone at $9,294 though out of stock and sold only within a Supermicro server assembly. Also configured into new servers by Supermicro, Dell and comparable vendors. Secondary: Compute Exchange as broker with a published band, Network Outlet as a refurbisher holding stock with a 3-year warranty, and sell-server.com with a 2 to 4 week lead time. This is one of the few parts in the category with genuine open-channel resale liquidity.

Regional notes

Export-controlled: NVIDIA's Form 8-K of 17 October 2023 names L40S and extends the licence to any system incorporating it, for China and Country Groups D1, D4 and D5. Resale is itself an export. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Its 864 GB/s is far under the gate, which sets review policy rather than licence need; no TPP figure.