NVIDIA L40S
The NVIDIA L40S is a dual-slot, passively cooled data-center GPU based on the Ada Lovelace architecture. NVIDIA announced it at SIGGRAPH on August 8, 2023, for AI training and inference, generative AI, 3D graphics, rendering, video processing, and industrial digitalization with NVIDIA Omniverse.[1][3] The card has 48 GB of GDDR6 memory with error-correcting code (ECC), 18,176 CUDA cores, 568 fourth-generation Tensor Cores, and 142 third-generation ray-tracing cores. NVIDIA rates its peak FP8 Tensor Core throughput at 733 TFLOPS without structured sparsity and 1,466 TFLOPS with sparsity.[1]
The L40S is closely related to the earlier NVIDIA L40. They are Ada Lovelace PCIe cards with 48 GB of GDDR6 ECC memory and 864 GB/s of memory bandwidth, but the L40S has a 350 W maximum power rating and substantially higher published Tensor Core throughput than the 300 W L40.[1][4][5][6] Unlike the H100, the L40S has neither NVLink nor Multi-Instance GPU (MIG) support.[1][9]
What is the NVIDIA L40S?
NVIDIA markets the L40S as a "universal" data-center processor because it combines AI compute, RTX graphics, and media engines on one card. The product supports transformer training and inference, image generation, 3D visualization, video encoding and decoding, and virtual GPU software.[1] This positioning describes the range of supported workloads; it does not imply that the L40S matches a specialized accelerator on every workload.
At launch, NVIDIA made the L40S part of its OVX server platform. It said OVX systems could contain up to eight L40S GPUs, each with 48 GB of memory, and named ASUS, Dell Technologies, GIGABYTE, Hewlett Packard Enterprise, Lenovo, QCT, and Supermicro among the planned system suppliers. NVIDIA said availability would begin in fall 2023.[3]
What architecture and design does the L40S use?
The L40S uses NVIDIA's Ada Lovelace architecture and an AD102 GPU. NVIDIA's NIM benchmarking documentation identifies its tested L40S as AD102 and lists a 2,520 MHz GPU clock for that test system.[2] The card exposes 18,176 CUDA cores, 568 fourth-generation Tensor Cores, and 142 third-generation RT Cores.[1]
Ada Lovelace Tensor Cores support structured sparsity, TF32, BF16, FP16, and FP8. The L40S also includes NVIDIA's Transformer Engine, which can recast transformer-network layers between FP8 and FP16 precision. Its graphics hardware includes third-generation RT Cores, while three NVENC and three NVDEC engines provide video encode and decode support, including AV1.[1]
The memory subsystem consists of 48 GB of GDDR6 with ECC and 864 GB/s of bandwidth. For comparison, NVIDIA specifies 80 GB of HBM3 and 3.35 TB/s of memory bandwidth for the H100 SXM.[1][9] The L40S connects to its host through PCIe Gen4 x16, rated by NVIDIA at 64 GB/s bidirectional, and has no NVLink. NVIDIA lists vGPU support but no MIG support.[1]
The card measures 4.4 by 10.5 inches, occupies two slots, uses passive cooling, and has a 350 W maximum power rating with a 16-pin power connector. Those characteristics require a server chassis designed to supply airflow and power for a passive full-height, full-length accelerator.[1][8]
What are the L40S specifications?
The following table lists NVIDIA's published peak specifications. Tensor values are shown without and with structured sparsity, in that order.[1]
| Specification | NVIDIA L40S |
|---|---|
| Architecture | Ada Lovelace (AD102) |
| Announced | August 8, 2023 |
| CUDA cores | 18,176 |
| Tensor Cores | 568 (fourth generation) |
| RT Cores | 142 (third generation) |
| GPU memory | 48 GB GDDR6 with ECC |
| Memory bandwidth | 864 GB/s |
| FP32 | 91.6 TFLOPS |
| TF32 Tensor Core | 183 | 366 TFLOPS |
| BF16 Tensor Core | 362.05 | 733 TFLOPS |
| FP16 Tensor Core | 362.05 | 733 TFLOPS |
| FP8 Tensor Core | 733 | 1,466 TFLOPS |
| INT8 Tensor Core | 733 | 1,466 TOPS |
| INT4 Tensor Core | 733 | 1,466 TOPS |
| RT Core performance | 212 TFLOPS |
| Host interface | PCIe Gen4 x16, 64 GB/s bidirectional |
| NVLink | No |
| MIG | No |
| Form factor | 4.4 x 10.5 inches, dual slot, passive |
| Maximum power | 350 W |
| Media engines | 3 NVENC and 3 NVDEC, including AV1 |
The 1,466 TFLOPS figure is the sparse FP8 peak, not dense FP8 throughput and not a workload benchmark.[1] At launch, NVIDIA also reported up to 1.2 times the generative-AI inference performance and up to 1.7 times the training performance of the A100 for the workloads in its comparison. These are vendor-reported, workload-specific results rather than general performance ratios.[3]
How does the L40S compare to the L40?
The NVIDIA L40 and L40S share their Ada Lovelace architecture, core counts, 48 GB GDDR6 ECC memory capacity, 864 GB/s memory bandwidth, dual-slot form factor, and lack of NVLink.[1][5][6] NVIDIA presents the L40 as a visual-computing GPU for rendering, 3D graphics, virtual workstations, video, and AI, while it presents the L40S as a broader AI, graphics, and media accelerator.[1][5]
Their most visible published differences are power and Tensor Core throughput. The L40 is rated at 300 W and the L40S at 350 W. NVIDIA's L40 data sheet lists 362 TFLOPS of dense FP8 throughput and 724 TFLOPS with sparsity, while the L40S is rated at 733 and 1,466 TFLOPS respectively. Their FP32 and ray-tracing figures are much closer.[1][6]
| Attribute | NVIDIA L40 | NVIDIA L40S |
|---|---|---|
| Product focus | Visual computing, rendering, video, and AI inference | AI training and inference, graphics, and video |
| FP32 | 90.5 TFLOPS | 91.6 TFLOPS |
| FP8 Tensor Core | 362 | 724 TFLOPS | 733 | 1,466 TFLOPS |
| RT Core performance | 209 TFLOPS | 212 TFLOPS |
| Maximum power | 300 W | 350 W |
| Memory | 48 GB GDDR6 ECC, 864 GB/s | 48 GB GDDR6 ECC, 864 GB/s |
| NVLink | No | No |
These peak figures do not determine application performance on their own. Software, precision, batching, memory use, and the rest of the server configuration affect measured throughput and latency.
How does the L40S compare to the H100?
The L40S was discussed as an alternative during the 2023 shortage of H100 accelerators. In November 2023, The Next Platform reported that H100 allocations extended into 2024 and that buyers could face waits of about six months. In the same report, George Wagner, then executive director of product and technical marketing at Liqid, estimated that H100 delivered two to three times L40S performance across a range of AI training and inference workloads, consumed about twice the power, and cost about three times as much.[7] Those estimates describe one source's assessment during a particular supply shortage, not a universal benchmark or a current price comparison.
The two products also differ substantially in their official specifications. H100 SXM has more memory, nearly four times the memory bandwidth, a higher sparse FP8 peak, NVLink, and hardware FP64 intended for high-performance computing. L40S is a PCIe card with GDDR6 and no NVLink or MIG.[1][9]
| Attribute | NVIDIA L40S | NVIDIA H100 SXM |
|---|---|---|
| Architecture | Ada Lovelace (AD102) | NVIDIA Hopper (GH100) |
| Memory | 48 GB GDDR6 ECC | 80 GB HBM3 |
| Memory bandwidth | 864 GB/s | 3.35 TB/s |
| Sparse FP8 Tensor Core peak | 1,466 TFLOPS | 3,958 TFLOPS |
| Maximum power | 350 W | Up to 700 W, configurable |
| GPU interconnect | No NVLink; PCIe Gen4 x16 | NVLink at 900 GB/s; PCIe Gen5 |
| MIG | No | Up to seven 10 GB instances |
Peak arithmetic rates and maximum power limits should not be converted into a general performance-per-dollar or performance-per-watt claim. A meaningful comparison requires the same model, software stack, precision, batch size, latency target, server, and price basis.
How was the L40S adopted, and why does it matter?
NVIDIA's launch announcement named CoreWeave as an early cloud provider and listed seven system builders planning OVX systems with the card.[3] Public pricing pages from RunPod, Nebius, and Crusoe continued to list L40S instances in August 2026.[10][11][12] NVIDIA has also published TensorRT-LLM inference results produced on a one-GPU L40S test system, with the model, precision, input and output lengths, concurrency, and software version stated alongside the results.[2]
Within NVIDIA's Ada data-center products, the L40S has more memory and a higher maximum power rating than the 24 GB, 72 W, low-profile L4.[1][13] Its 48 GB memory capacity, FP8 support, graphics hardware, media engines, and vGPU support allow one model to cover several classes of data-center work.[1] Its limits are equally explicit: 864 GB/s of memory bandwidth, a PCIe Gen4 host interface with no NVLink, no MIG, and a 350 W maximum power rating. Whether it is suitable depends on the workload and system configuration rather than a single peak-throughput figure.
References
- ^NVIDIA, "NVIDIA L40S." nvidia.com/...l40s
- ^NVIDIA, "NVIDIA NIM LLMs Benchmarking: Performance," last updated April 1, 2026. docs.nvidia.com/...performance
- ^NVIDIA Newsroom, "NVIDIA, Global Data Center System Manufacturers to Supercharge Generative AI and Industrial Digitalization," August 8, 2023. nvidianews.nvidia.com/...industrial-digitalization
- ^ServeTheHome, "NVIDIA L40S GPU for Data Center Visualization Launched," August 8, 2023. servethehome.com/...-center-visualization-launched
- ^NVIDIA, "NVIDIA L40 GPU for Data Center." nvidia.com/...l40
- ^NVIDIA, "NVIDIA L40 Data Sheet." images.nvidia.com/...vgpu-L40-datasheet.pdf
- ^*The Next Platform*, "What To Do When You Can't Get Nvidia H100 GPUs," November 17, 2023. nextplatform.com/...1656657
- ^Lenovo Press, "ThinkSystem NVIDIA L40S 48GB PCIe Gen4 Passive GPU Product Guide," last updated July 1, 2025. lenovopress.lenovo.com/...gb-pcie-gen4-passive-gpu
- ^NVIDIA, "NVIDIA H100 Tensor Core GPU." nvidia.com/...h100
- ^RunPod, "GPU Cloud Pricing," accessed August 17, 2026. runpod.io/pricing
- ^Nebius, "AI Cloud Pricing," accessed August 17, 2026. nebius.com/prices
- ^Crusoe, "GPU Pricing for AI Compute and Inference," accessed August 17, 2026. crusoe.ai/...pricing
- ^NVIDIA, "NVIDIA L4 Tensor Core GPU." nvidia.com/...l4
Improve this article
Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.
3 revisions · v4 · 1,649 words · full history
Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify
Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here
Reviewer note: Fact-checked 2026-08-17 against NVIDIA L40S, L40, H100, L4, and NIM documentation, Lenovo's product guide, and the cited 2023 historical report. Corrected sparse-throughput labeling, attributed and dated the H100 comparison, and removed unsupported availability, packaging, cost-efficiency, and performance-per-watt claims.
Cite this page: AI Wiki. "NVIDIA L40S." aiwiki.ai, updated 17 Aug 2026, fact-checked 17 Aug 2026. CC BY 4.0. https://aiwiki.ai/wiki/nvidia_l40s