NVIDIA L40S

RawGraph

The NVIDIA L40S is a dual-slot, passively cooled data-center GPU based on the Ada Lovelace architecture. NVIDIA announced it at SIGGRAPH on August 8, 2023, for AI training and inference, generative AI, 3D graphics, rendering, video processing, and industrial digitalization with NVIDIA Omniverse.[1][3] The card has 48 GB of GDDR6 memory with error-correcting code (ECC), 18,176 CUDA cores, 568 fourth-generation Tensor Cores, and 142 third-generation ray-tracing cores. NVIDIA rates its peak FP8 Tensor Core throughput at 733 TFLOPS without structured sparsity and 1,466 TFLOPS with sparsity.[1]

The L40S is closely related to the earlier NVIDIA L40. They are Ada Lovelace PCIe cards with 48 GB of GDDR6 ECC memory and 864 GB/s of memory bandwidth, but the L40S has a 350 W maximum power rating and substantially higher published Tensor Core throughput than the 300 W L40.[1][4][5][6] Unlike the H100, the L40S has neither NVLink nor Multi-Instance GPU (MIG) support.[1][9]

What is the NVIDIA L40S?

NVIDIA markets the L40S as a "universal" data-center processor because it combines AI compute, RTX graphics, and media engines on one card. The product supports transformer training and inference, image generation, 3D visualization, video encoding and decoding, and virtual GPU software.[1] This positioning describes the range of supported workloads; it does not imply that the L40S matches a specialized accelerator on every workload.

At launch, NVIDIA made the L40S part of its OVX server platform. It said OVX systems could contain up to eight L40S GPUs, each with 48 GB of memory, and named ASUS, Dell Technologies, GIGABYTE, Hewlett Packard Enterprise, Lenovo, QCT, and Supermicro among the planned system suppliers. NVIDIA said availability would begin in fall 2023.[3]

What architecture and design does the L40S use?

The L40S uses NVIDIA's Ada Lovelace architecture and an AD102 GPU. NVIDIA's NIM benchmarking documentation identifies its tested L40S as AD102 and lists a 2,520 MHz GPU clock for that test system.[2] The card exposes 18,176 CUDA cores, 568 fourth-generation Tensor Cores, and 142 third-generation RT Cores.[1]

Ada Lovelace Tensor Cores support structured sparsity, TF32, BF16, FP16, and FP8. The L40S also includes NVIDIA's Transformer Engine, which can recast transformer-network layers between FP8 and FP16 precision. Its graphics hardware includes third-generation RT Cores, while three NVENC and three NVDEC engines provide video encode and decode support, including AV1.[1]

The memory subsystem consists of 48 GB of GDDR6 with ECC and 864 GB/s of bandwidth. For comparison, NVIDIA specifies 80 GB of HBM3 and 3.35 TB/s of memory bandwidth for the H100 SXM.[1][9] The L40S connects to its host through PCIe Gen4 x16, rated by NVIDIA at 64 GB/s bidirectional, and has no NVLink. NVIDIA lists vGPU support but no MIG support.[1]

The card measures 4.4 by 10.5 inches, occupies two slots, uses passive cooling, and has a 350 W maximum power rating with a 16-pin power connector. Those characteristics require a server chassis designed to supply airflow and power for a passive full-height, full-length accelerator.[1][8]

What are the L40S specifications?

The following table lists NVIDIA's published peak specifications. Tensor values are shown without and with structured sparsity, in that order.[1]

SpecificationNVIDIA L40S
ArchitectureAda Lovelace (AD102)
AnnouncedAugust 8, 2023
CUDA cores18,176
Tensor Cores568 (fourth generation)
RT Cores142 (third generation)
GPU memory48 GB GDDR6 with ECC
Memory bandwidth864 GB/s
FP3291.6 TFLOPS
TF32 Tensor Core183 | 366 TFLOPS
BF16 Tensor Core362.05 | 733 TFLOPS
FP16 Tensor Core362.05 | 733 TFLOPS
FP8 Tensor Core733 | 1,466 TFLOPS
INT8 Tensor Core733 | 1,466 TOPS
INT4 Tensor Core733 | 1,466 TOPS
RT Core performance212 TFLOPS
Host interfacePCIe Gen4 x16, 64 GB/s bidirectional
NVLinkNo
MIGNo
Form factor4.4 x 10.5 inches, dual slot, passive
Maximum power350 W
Media engines3 NVENC and 3 NVDEC, including AV1

The 1,466 TFLOPS figure is the sparse FP8 peak, not dense FP8 throughput and not a workload benchmark.[1] At launch, NVIDIA also reported up to 1.2 times the generative-AI inference performance and up to 1.7 times the training performance of the A100 for the workloads in its comparison. These are vendor-reported, workload-specific results rather than general performance ratios.[3]

How does the L40S compare to the L40?

The NVIDIA L40 and L40S share their Ada Lovelace architecture, core counts, 48 GB GDDR6 ECC memory capacity, 864 GB/s memory bandwidth, dual-slot form factor, and lack of NVLink.[1][5][6] NVIDIA presents the L40 as a visual-computing GPU for rendering, 3D graphics, virtual workstations, video, and AI, while it presents the L40S as a broader AI, graphics, and media accelerator.[1][5]

Their most visible published differences are power and Tensor Core throughput. The L40 is rated at 300 W and the L40S at 350 W. NVIDIA's L40 data sheet lists 362 TFLOPS of dense FP8 throughput and 724 TFLOPS with sparsity, while the L40S is rated at 733 and 1,466 TFLOPS respectively. Their FP32 and ray-tracing figures are much closer.[1][6]

AttributeNVIDIA L40NVIDIA L40S
Product focusVisual computing, rendering, video, and AI inferenceAI training and inference, graphics, and video
FP3290.5 TFLOPS91.6 TFLOPS
FP8 Tensor Core362 | 724 TFLOPS733 | 1,466 TFLOPS
RT Core performance209 TFLOPS212 TFLOPS
Maximum power300 W350 W
Memory48 GB GDDR6 ECC, 864 GB/s48 GB GDDR6 ECC, 864 GB/s
NVLinkNoNo

These peak figures do not determine application performance on their own. Software, precision, batching, memory use, and the rest of the server configuration affect measured throughput and latency.

How does the L40S compare to the H100?

The L40S was discussed as an alternative during the 2023 shortage of H100 accelerators. In November 2023, The Next Platform reported that H100 allocations extended into 2024 and that buyers could face waits of about six months. In the same report, George Wagner, then executive director of product and technical marketing at Liqid, estimated that H100 delivered two to three times L40S performance across a range of AI training and inference workloads, consumed about twice the power, and cost about three times as much.[7] Those estimates describe one source's assessment during a particular supply shortage, not a universal benchmark or a current price comparison.

The two products also differ substantially in their official specifications. H100 SXM has more memory, nearly four times the memory bandwidth, a higher sparse FP8 peak, NVLink, and hardware FP64 intended for high-performance computing. L40S is a PCIe card with GDDR6 and no NVLink or MIG.[1][9]

AttributeNVIDIA L40SNVIDIA H100 SXM
ArchitectureAda Lovelace (AD102)NVIDIA Hopper (GH100)
Memory48 GB GDDR6 ECC80 GB HBM3
Memory bandwidth864 GB/s3.35 TB/s
Sparse FP8 Tensor Core peak1,466 TFLOPS3,958 TFLOPS
Maximum power350 WUp to 700 W, configurable
GPU interconnectNo NVLink; PCIe Gen4 x16NVLink at 900 GB/s; PCIe Gen5
MIGNoUp to seven 10 GB instances

Peak arithmetic rates and maximum power limits should not be converted into a general performance-per-dollar or performance-per-watt claim. A meaningful comparison requires the same model, software stack, precision, batch size, latency target, server, and price basis.

How was the L40S adopted, and why does it matter?

NVIDIA's launch announcement named CoreWeave as an early cloud provider and listed seven system builders planning OVX systems with the card.[3] Public pricing pages from RunPod, Nebius, and Crusoe continued to list L40S instances in August 2026.[10][11][12] NVIDIA has also published TensorRT-LLM inference results produced on a one-GPU L40S test system, with the model, precision, input and output lengths, concurrency, and software version stated alongside the results.[2]

Within NVIDIA's Ada data-center products, the L40S has more memory and a higher maximum power rating than the 24 GB, 72 W, low-profile L4.[1][13] Its 48 GB memory capacity, FP8 support, graphics hardware, media engines, and vGPU support allow one model to cover several classes of data-center work.[1] Its limits are equally explicit: 864 GB/s of memory bandwidth, a PCIe Gen4 host interface with no NVLink, no MIG, and a 350 W maximum power rating. Whether it is suitable depends on the workload and system configuration rather than a single peak-throughput figure.

References

  1. ^NVIDIA, "NVIDIA L40S." nvidia.com/...l40s
  2. ^NVIDIA, "NVIDIA NIM LLMs Benchmarking: Performance," last updated April 1, 2026. docs.nvidia.com/...performance
  3. ^NVIDIA Newsroom, "NVIDIA, Global Data Center System Manufacturers to Supercharge Generative AI and Industrial Digitalization," August 8, 2023. nvidianews.nvidia.com/...industrial-digitalization
  4. ^ServeTheHome, "NVIDIA L40S GPU for Data Center Visualization Launched," August 8, 2023. servethehome.com/...-center-visualization-launched
  5. ^NVIDIA, "NVIDIA L40 GPU for Data Center." nvidia.com/...l40
  6. ^NVIDIA, "NVIDIA L40 Data Sheet." images.nvidia.com/...vgpu-L40-datasheet.pdf
  7. ^*The Next Platform*, "What To Do When You Can't Get Nvidia H100 GPUs," November 17, 2023. nextplatform.com/...1656657
  8. ^Lenovo Press, "ThinkSystem NVIDIA L40S 48GB PCIe Gen4 Passive GPU Product Guide," last updated July 1, 2025. lenovopress.lenovo.com/...gb-pcie-gen4-passive-gpu
  9. ^NVIDIA, "NVIDIA H100 Tensor Core GPU." nvidia.com/...h100
  10. ^RunPod, "GPU Cloud Pricing," accessed August 17, 2026. runpod.io/pricing
  11. ^Nebius, "AI Cloud Pricing," accessed August 17, 2026. nebius.com/prices
  12. ^Crusoe, "GPU Pricing for AI Compute and Inference," accessed August 17, 2026. crusoe.ai/...pricing
  13. ^NVIDIA, "NVIDIA L4 Tensor Core GPU." nvidia.com/...l4

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

3 revisions · v4 · 1,649 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Fact-checked 2026-08-17 against NVIDIA L40S, L40, H100, L4, and NIM documentation, Lenovo's product guide, and the cited 2023 historical report. Corrected sparse-throughput labeling, attributed and dated the H100 comparison, and removed unsupported availability, packaging, cost-efficiency, and performance-per-watt claims.

Cite this page: AI Wiki. "NVIDIA L40S." aiwiki.ai, updated 17 Aug 2026, fact-checked 17 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/nvidia_l40s

Suggest edit