NVIDIA L40S

Ada Lovelace PCIe card with 48 GB GDDR6 at 350 W, and a real used market at $5,500 to $9,000
NVIDIA L40S is one of 24 data center GPUs tracked in the SecondWatt catalogue, each with specifications, lead times and indicative secondary-market pricing.
NVIDIA L40S dossier: 48 GB GDDR6, 864 GB/s, 350 W passive PCIe, rack kW, export status, and a used range of $5,500 to $9,000 from three channels.
The NVIDIA L40S is an Ada Lovelace accelerator on a passive dual-slot PCIe card measuring 4.4 inches high by 10.5 inches long. NVIDIA publishes 48 GB of GDDR6 with ECC, 864 GB/s of memory bandwidth and 350 W maximum power, with 18,176 CUDA cores, 568 fourth-generation tensor cores and 142 third-generation ray tracing cores. It carries four DisplayPort 1.4a outputs and a 16-pin power connector. Two facts a buyer should not miss. It uses GDDR6 rather than HBM, so 864 GB/s against an A100 SXM4's 2,039 GB/s is a 2.4 times bandwidth deficit that makes it a poor fit for memory-bandwidth-bound work regardless of its FP8 throughput. And it is PCIe Gen4, not Gen5, with no NVLink and no multi-instance partitioning, so multi-card scaling runs over PCIe and the network only. The used discount is only about 25 percent off new, which is far shallower than the A100 alongside it. That is what a Moderate supply signal looks like: the part is still in production and still bought new, so no upgrade cycle is flooding resale.
Why it matters
L40S is the price-discovery anchor of the inference tier: three independent channels agree inside $5,500 to $9,000, which almost nothing else in this category manages. It is also export-controlled, named in NVIDIA's own October 2023 filing alongside H100, which surprises most buyers of a mid-priced inference card and changes how it can be moved.
Who buys it
Operators adding inference, fine-tuning, graphics, virtual desktop or transcode capacity to an existing server fleet rather than building a new hall. At 350 W passive in a standard dual-slot PCIe space it drops into conventional chassis, eight to a 5U node. Buyers cross-shopping it against a used A100 80GB PCIe are choosing FP8 and a newer architecture over memory bandwidth.
Role in the data center
Inference, fine-tuning, virtual desktop, rendering and video transcode rather than large-scale training. Ray tracing cores and four DisplayPort outputs make it a graphics part as much as an accelerator. Lands near a 40 kW rack density band at seven 5U nodes of eight cards.
Power envelope
- Per accelerator
- 350 W — Verified: Maximum configurable board power, manufacturer specification. source
- Cooling class
- Air-coolable — Derived: Per-accelerator power below 400 W — standard air-cooled halls.
No published node-level power rating exists for this part, so node, rack and per-MW figures are not stated. Multiplying accelerator watts by an assumed overhead under-provisions a building by roughly 45% against measured systems.
Node figures are published OEM system ratings; rack counts are arithmetic against a stated rack budget and move with your own cooling design. Assumption set version 2026-09-11a.
Key specifications
| Architecture | Ada Lovelace |
|---|---|
| Memory capacity | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| Maximum power | 350 W |
| Thermal solution | Passive, server airflow |
| Form factor | 4.4 in H x 10.5 in L, dual slot |
| Interconnect | PCIe Gen4 x16 64 GB/s bidirectional |
| NVLink support | No |
| Multi-instance partitioning | No |
| CUDA cores | 18,176 |
| Dense FP8 | 733 TFLOPS |
| Dense FP32 | 91.6 TFLOPS |
| Power connector | 16-pin |
| Accelerator-only load, 8 cards | 2.8 kW |
| Used and refurbished range | 5,500 to 9,000 USD per card, three channels |
| Total DRAM bandwidth vs export gate | 864 below 6,500 GB/s |
Technical summary
Architecture: Ada Lovelace; 18,176 CUDA cores, 568 tensor cores, 142 ray tracing cores. Memory: 48 GB GDDR6 with ECC, 864 GB/s bandwidth. Maximum power: 350 W; thermal solution passive, server airflow required. Form factor: 4.4 inch high by 10.5 inch long, dual slot; 16-pin power connector. Interconnect: PCIe Gen4 x16 at 64 GB/s bidirectional; no NVLink; no multi-instance partitioning. Dense compute: FP32 91.6 TFLOPS, TF32 183, BF16 and FP16 362.05, FP8 733, INT8 733 TOPS. With sparsity (vendor figure): TF32 366, BF16 and FP16 733, FP8 1,466 TFLOPS. Display outputs: 4 x DisplayPort 1.4a; virtual GPU supported. Accelerator-only load, 8 cards: 2.8 kW; proxy-estimated node draw, not an OEM figure, 5.7 kW. Rack density near 40 kW at seven 5U nodes of eight cards. Export: named in NVIDIA's Form 8-K of 17 October 2023 as licence-controlled.
Major variations
This is a single SKU rather than a family, so there is no memory or power variant to enumerate. The related L40 is a distinct part and is named separately in NVIDIA's October 2023 export filing. The card also ships as PNY SKU NVL40STCGPU-KIT.
Configurations and options
Deployed in conventional PCIe servers. Supermicro's 5U SYS-521GE-TNRT lists support for this card and the L4, taking up to two quad-width, two triple-width, eight double-width or ten single-width cards, with four 2700 W Titanium redundant supplies and up to ten heavy-duty fans plus an optional direct-to-chip cold plate. Eight cards in that chassis is 2.8 kW of card load, and NVIDIA's DGX A100 node ratio of 2.03x, used here only as a proxy because the manufacturer publishes no node figure, suggests roughly 5.7 kW. Seven such nodes fill a 35U span at roughly 40 kW a rack. NVIDIA publishes no configurable-TDP range for this card, so 350 W should be treated as fixed.
Compatibility and dependencies
Fits any PCIe Gen4 x16 slot with a dual-slot 4.4 by 10.5 inch space, a 16-pin power feed and airflow rated for a 350 W passive card. There is no NVLink and no multi-instance partitioning, so multi-card work scales over PCIe Gen4 at 64 GB/s and over the network, which is a real limit for training. The bandwidth constraint is the one to design around. 864 GB/s of GDDR6 against 1,935 GB/s on a used A100 80GB PCIe means bandwidth-bound inference belongs on the A100 even though this card has FP8 and the A100 does not. Plan alongside racks, PDUs and CRAC or chiller capacity; the optional cold plate on the 5U chassis makes liquid cooling relevant at high card counts.
Pricing and availability
Published price range: $5,500 to $9,000 per card used or refurbished, three independent channels, Q3 2026. Evidence: channel observation · used or refurbished condition · Single card. Trend: falling The curve here is shallow and that is the finding. sell-server.com records $8,000 to $10,000 in mid-2025 falling to $7,000 to $9,000 in mid-2026, a 12 percent decline over the year. Compute Exchange puts used stock at $5,500 to $7,500 for Q3 2026 with a Moderate supply signal, and an aggregator reckons the $9,000 market figure at 13 percent above the original list price. Set that against A100 SXM4 at roughly 20 percent a year and the difference is a product in its selling life against one in its disposal life. A part still bought new does not get flooded by an upgrade cycle. Do not cite the PC Server and Parts range as independent corroboration; it restates sell-server.com exactly. Lead time, new: In production. Wiredzone showed the card out of stock and available only as part of a Supermicro server assembly; Dihuni listed it new at $9,100.. Lead time, used or refurbished: 2 to 4 weeks from a secondary dealer per sell-server.com, or in stock from a refurbisher. This is among the most liquid items in the category.. Warranty, new: 3 years with 30 days advance replacement for dead-on-arrival units, per the Wiredzone listing, plus 2-week replacement thereafter.. Warranty, used: 3-year reseller warranty on refurbished stock per Network Outlet. No manufacturer warranty transfer was documented anywhere in this category..
Lifecycle and maintenance
Still in production and still sold new at $9,100 to $9,294 as of 10 September 2026, with Compute Exchange rating secondary supply as Moderate rather than High. That combination puts the card in its selling life rather than its disposal life, and it is why the used discount is roughly 25 percent rather than the 45 percent an A100 carries. The card is passive, so its service life is a function of host airflow more than of anything on the board. Operator depreciation practice across the category runs 4 to 6 years, against an observed 7.5-year service life on an earlier generation.
Common failure points
- 16-pin power connector — Intermittent resets under transient load, with discolouration or melted keying visible at the connector
- GDDR6 memory and ECC — Rising correctable error counts then retired pages and job crashes under sustained memory pressure
- Passive cooling airflow dependency — Throttling only under full load in a chassis whose fan curve was not qualified for 350 W passive cards
- PCIe link degradation — Link trains below Gen4 x16 or reports corrected-error storms on the host link
- Thermal interface material — Clocks decline gradually at constant load as die temperature climbs over months
- Driver branch obsolescence — Card functions but a current framework release drops the architecture from its support matrix
Inspection checklist
Driver-reported memory must read 48 GB GDDR6 with ECC enabled: a short reading indicates a failed module. Stream benchmark bandwidth within 10 percent of 864 GB/s: a low result points to degraded GDDR6 or clock limiting. Sustained power draw within 5 percent of 350 W under load: a card capped well below is throttling. Core counts must enumerate 18,176 CUDA, 568 tensor and 142 ray tracing cores: a low count means a re-marked die. GDDR6 ECC error counters and retired-page logs read before purchase: repeat after a 24-hour soak at full load. Host PCIe link trains at Gen4 x16 for 64 GB/s: this card is Gen4, so a Gen5 reading indicates a different part. Confirm the card reports no NVLink and no multi-instance support: a listing claiming either is misdescribed. 16-pin power connector inspected for discolouration, pitting or melted keying: a 350 W card is unforgiving here.
Procurement channels
New: Dihuni at $9,100, reduced from $11,975, and Wiredzone at $9,294 though out of stock and sold only within a Supermicro server assembly. Also configured into new servers by Supermicro, Dell and comparable vendors. Secondary: Compute Exchange as broker with a published band, Network Outlet as a refurbisher holding stock with a 3-year warranty, and sell-server.com with a 2 to 4 week lead time. This is one of the few parts in the category with genuine open-channel resale liquidity.
Regional notes
Export-controlled: NVIDIA's Form 8-K of 17 October 2023 names L40S and extends the licence to any system incorporating it, for China and Country Groups D1, D4 and D5. Resale is itself an export. BIS guidance of 31 May 2026 turns on the counterparty: a licence is required for any entity headquartered in Country Group D:5 or Macau, or with an ultimate parent there, from every destination outside the US, regardless of location. 91 FR 1684 of 15 January 2026 gates case-by-case China and Macau review at both total processing performance under 21,000 and DRAM bandwidth under 6,500 GB/s. Its 864 GB/s is far under the gate, which sets review policy rather than licence need; no TPP figure.