# NVIDIA Blackwell

> Source: https://aiwiki.ai/wiki/nvidia_blackwell
> Summary: NVIDIA Blackwell is a family of graphics processing unit architectures and computing platforms developed by NVIDIA. NVIDIA announced the data center Blackwell platform on March 18, 2024 as the successor to NVIDIA Hopper, then extended the name to GeForce and professional graphics products in 2025.
> Updated: 2026-07-30
> Fact-checked: 2026-07-30
> Categories: AI Hardware, Deep Learning, NVIDIA
> License: CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/) - attribute to "AI Wiki (aiwiki.ai)"
> Cite as: AI Wiki. "NVIDIA Blackwell." aiwiki.ai, 30 Jul 2026. https://aiwiki.ai/wiki/nvidia_blackwell
> From AI Wiki (https://aiwiki.ai), the free encyclopedia of artificial intelligence. Reuse freely with attribution.

![Nvidia blackwell1.jpg](https://qqcb8dyk5bp2il4c.public.blob.vercel-storage.com/images/300px-nvidia_blackwell1.jpg)
![Nvidia blackwell2.jpg](https://qqcb8dyk5bp2il4c.public.blob.vercel-storage.com/images/300px-nvidia_blackwell2.jpg)

**NVIDIA Blackwell** is a family of [graphics processing unit](https://aiwiki.ai/wiki/gpu) architectures and computing platforms developed by [NVIDIA](https://aiwiki.ai/wiki/nvidia). NVIDIA announced the data center Blackwell platform on March 18, 2024 as the successor to [NVIDIA Hopper](https://aiwiki.ai/wiki/nvidia_hopper), then extended the name to GeForce and professional graphics products in 2025.[1][17][19] The family spans distinct implementations. The GB100 data center design used by B200 and GB200 combines two large dies in one logical GPU, while RTX Blackwell chips such as GB202 are graphics-oriented designs with GDDR7 memory, ray-tracing hardware, and video engines.[1][4][18]

Blackwell is therefore not one processor with one specification. It is an architecture and platform umbrella covering data center accelerators, Grace Blackwell superchips, rack-scale systems, consumer GPUs, professional GPUs, and later Blackwell Ultra products. Statements about 208 billion transistors, two reticle-limited dies, HBM3e memory, and a 10 TB/s internal link apply to the original GB100-class data center GPU. They do not describe every GeForce or RTX PRO Blackwell chip.[1][18]

The architecture's principal data center changes include fifth-generation [Tensor Cores](https://aiwiki.ai/wiki/tensor_core), a second-generation Transformer Engine, support for lower-precision microscaling formats, fifth-generation [NVLink](https://aiwiki.ai/wiki/nvlink), expanded on-chip memory resources, a hardware decompression engine, and additional reliability and confidential-computing functions.[1][3][9][12] These features are intended to improve [deep learning](https://aiwiki.ai/wiki/deep_learning), [large language model](https://aiwiki.ai/wiki/large_language_model), data analytics, and high-performance computing workloads. Their practical benefit depends on the exact GPU, numerical format, model, software stack, interconnect topology, and power envelope.

## Scope and naming

This article covers Blackwell as an architecture and platform family. The [NVIDIA Blackwell B200](https://aiwiki.ai/wiki/nvidia_blackwell_b200) article covers the accelerator product in more detail. [NVIDIA Blackwell Ultra](https://aiwiki.ai/wiki/nvidia_blackwell_ultra) covers the later B300 and GB300 generation, and [NVIDIA RTX PRO 6000 Blackwell](https://aiwiki.ai/wiki/nvidia_rtx_pro_6000) covers a professional GB202 product. Keeping these identities separate prevents specifications from one implementation being assigned to another.

NVIDIA named the architecture for David Harold Blackwell, an American mathematician and statistician. Blackwell made foundational contributions to probability, statistics, game theory, information theory, and dynamic programming. He became the first Black member elected to the United States National Academy of Sciences in 1965.[1][2] The name is an honorific; it does not indicate a technical connection between his research and the GPU design.

## Development and release

NVIDIA introduced Blackwell at its GTC conference on March 18, 2024. Its announcement described a GB100-class GPU made on a custom TSMC 4NP process, with 208 billion transistors across two reticle-limited dies joined by a 10 TB/s chip-to-chip link. The same announcement introduced B200, the GB200 Grace Blackwell Superchip, HGX B200, and GB200 NVL72. It said partner systems were expected later in 2024.[1]

The production transition was not immediate. In its quarterly filing for the period ended July 28, 2024, NVIDIA said it had shipped customer samples and had changed the Blackwell GPU mask to improve production yield. The filing also recorded inventory provisions for low-yielding material and scheduled the production ramp for the fourth quarter of fiscal 2025 and into fiscal 2026.[13] NVIDIA's next quarterly commentary said the mask change had been completed successfully and that production shipments were scheduled to begin in the fourth quarter.[14]

By February 2025, NVIDIA reported that it had ramped mass production and achieved billions of dollars in Blackwell's first quarter of sales. This is company-reported commercial evidence, not an independent count of installed GPUs.[15] In May 2025, NVIDIA said GB200 NVL72 was in full-scale production across system makers and cloud service providers.[16]

The Blackwell name expanded beyond the initial data center platform:

| Date | Family milestone |
|---|---|
| March 18, 2024 | NVIDIA announced the GB100-class data center platform, B200, GB200, and GB200 NVL72.[1] |
| August 2024 | NVIDIA disclosed the production-yield mask change and customer sampling in a regulatory filing.[13] |
| January 6, 2025 | NVIDIA announced GeForce RTX 50-series desktop and laptop GPUs based on RTX Blackwell.[17] |
| March 18, 2025 | NVIDIA announced RTX PRO Blackwell workstation and server products.[19] |
| March 18, 2025 | NVIDIA announced Blackwell Ultra, including B300 and GB300-based systems, as an evolution of the platform.[20] |
| May 28, 2025 | NVIDIA reported that GB200 NVL72 had entered full-scale production.[16] |

This timeline distinguishes announcements, sampling, production, and availability. A roadmap statement or launch-day projection is not evidence that every listed system was generally available on the announcement date.

## Data center architecture

### GB100 package design

The original data center Blackwell GPU is a multi-die package. NVIDIA describes two reticle-limited GPU dies connected by a 10 TB/s internal link and presented to software as one GPU. Together the dies contain 208 billion transistors and are manufactured on a custom TSMC 4NP process.[1] The design lets NVIDIA build a logical GPU larger than a single lithographic reticle field, but it does not make the package identical to a monolithic die. Cross-die communication, placement, and software scheduling can still matter to measured behavior.

Public specifications differ by product configuration. NVIDIA's supported-GPU table identifies both B200 and the GPU in GB200 as GB100 microarchitecture with CUDA compute capability 10.0. It lists B200 with 180 GB of memory and GB200 with 186 GB.[4] NVIDIA's later Blackwell data sheet lists 7.7 TB/s of memory bandwidth and configurable power up to 1,000 W for B200, compared with 8 TB/s and up to 1,200 W for a GB200 GPU.[5] These shipping-oriented figures supersede early material that described a 192 GB configuration.

The [CUDA](https://aiwiki.ai/wiki/cuda) 12.8 tuning guide documents a maximum of 64 resident warps per streaming multiprocessor for compute capability 10.0, 228 KB of shared-memory capacity per SM, and 227 KB available to one opted-in thread block. It also documents a 256 KB combined L1, texture, and shared-memory capacity for B200.[3] These limits describe programming resources, not guaranteed application occupancy. Register use, block dimensions, shared-memory allocation, and thread-block clustering can reduce the number of active blocks or warps.

### Tensor processing and numerical formats

Blackwell's fifth-generation Tensor Cores add native paths for lower-precision matrix operations. The second-generation Transformer Engine combines those paths with scaling and precision-management software. NVIDIA's launch material emphasized 4-bit floating-point inference and micro-tensor scaling, while the PTX instruction set exposes Blackwell-specific `tcgen05` operations, Tensor Memory allocation and movement, and optional decompression of 4-bit or 6-bit values into wider types.[1][9]

Microscaling uses a scale shared by a small block of values together with narrow per-element formats. Sharing a scale reduces storage and arithmetic cost, but also couples the representable range of values in that block. Research preceding and accompanying Blackwell found that microscaling formats can be practical for training and inference, while the achieved accuracy depends on format, block size, model, calibration, and quantization method.[10][11]

This qualification is important. "FP4 support" is a hardware capability, not a promise that any model can be converted to four bits without loss. A 2024 PMLR study obtained negligible loss for tested models with 4-bit weights and 8-bit activations only after combining post-training methods such as SmoothQuant, AWQ, and GPTQ.[11] Production deployments may keep sensitive layers, accumulators, activations, or calibration statistics at higher precision. The throughput advertised for a sparse or low-precision Tensor Core path should not be compared directly with dense FP16 or FP32 work without stating the precision and sparsity assumptions.

Blackwell also introduces Tensor Memory, or TMEM, as an on-chip storage path used by fifth-generation Tensor Core instructions. Software must allocate, move, and release this resource through the relevant instruction sequence.[9] TMEM can reduce register and shared-memory pressure for suitable matrix pipelines, but older kernels do not gain that benefit automatically. Libraries and compilers need Blackwell-specific kernels and scheduling.

### Memory and interconnect

B200 uses [High Bandwidth Memory](https://aiwiki.ai/wiki/high_bandwidth_memory) rather than the GDDR7 used by RTX Blackwell. NVIDIA documents up to 180 GB of HBM3 or HBM3e for B200, while the GB200 configuration exposes 186 GB per GPU.[3][4] Capacity and bandwidth are product attributes. They should not be generalized to every Blackwell processor.

Fifth-generation NVLink provides up to 1.8 TB/s of bidirectional GPU link bandwidth per GB100-class GPU. In GB200 NVL72, 72 GPUs and nine switch trays form a 130 TB/s scale-up domain.[7][8] NVLink is not a replacement for every data center network. It connects GPUs within supported scale-up topologies, while [InfiniBand](https://aiwiki.ai/wiki/infiniband) or [Ethernet](https://aiwiki.ai/wiki/ethernet) carries traffic across racks and larger clusters. End-to-end [distributed training](https://aiwiki.ai/wiki/distributed_training) performance depends on both layers, collective algorithms, topology, congestion, and the ratio of communication to computation.

The GB200 Superchip connects two Blackwell GPUs to one Grace CPU through NVLink-C2C. NVIDIA specifies 900 GB/s of bidirectional coherent CPU-to-GPU link bandwidth.[1][7] Coherent addressing simplifies some data movement and permits the processors to access a larger combined memory space, but CPU memory does not become HBM. Latency, bandwidth, placement, and page migration remain distinct performance considerations.

### Decompression, reliability, and security

The data center design includes a hardware decompression engine for formats including LZ4, Deflate, and Snappy.[1][7] Its purpose is to reduce work and traffic in analytics pipelines that read compressed data. Benefit depends on compression ratio, chunk size, concurrency, and whether compressed input or decompressed output is the bottleneck. It does not accelerate arbitrary compression algorithms or every database query.

NVIDIA also describes a dedicated reliability, availability, and serviceability engine that monitors hardware and software signals and supports diagnosis of potential faults.[1][5] Public material establishes that the functions exist, but it does not provide a workload-independent uptime improvement. Availability is a system property that also depends on firmware, cooling, power, networking, orchestration, and repair procedures.

Blackwell continues GPU confidential-computing support and adds protected multi-GPU paths in supported configurations. NVIDIA's security white paper describes attestation, encrypted CPU-to-GPU transfers, and encrypted NVLink traffic for Blackwell pass-through configurations.[12] These protections require a compatible CPU, host platform, firmware, driver, virtual-machine environment, and attestation policy. They do not protect data outside the defined trusted execution environment, and enabling a GPU feature alone does not make an application confidential.

## Product and platform variants

The following table summarizes major Blackwell variants without treating them as interchangeable:

| Product or platform | GPU implementation | Memory and interconnect | System boundary |
|---|---|---|---|
| B200 Tensor Core GPU | GB100, compute capability 10.0 | 180 GB HBM3e, 7.7 TB/s memory bandwidth, fifth-generation NVLink up to 1.8 TB/s, configurable up to 1,000 W.[3][4][5] | One data center GPU module; normally purchased in HGX, DGX, or partner systems |
| GB200 Grace Blackwell Superchip | Two GB100-class GPUs plus one Grace CPU | 186 GB HBM3e per GPU, 8 TB/s per-GPU memory bandwidth, and 900 GB/s coherent NVLink-C2C between Grace and the GPUs.[1][4][5] | A compute building block, not a 72-GPU rack by itself |
| [NVIDIA DGX](https://aiwiki.ai/wiki/nvidia_dgx) B200 | Eight B200 GPUs with two fifth-generation NVSwitches | 1,440 GB aggregate GPU memory and 14.4 TB/s aggregate NVSwitch bandwidth.[6] | Air-cooled 10U server with six 3.3 kW power supplies and a documented 14.3 kW maximum input.[6] |
| GB200 NVL72 | 36 GB200 Superchips, totaling 72 GPUs and 36 Grace CPUs | About 13.4 TB HBM3e and 130 TB/s aggregate NVLink bandwidth across the 72-GPU domain.[5][8] | Liquid-cooled rack-scale system with 18 compute trays and nine NVLink switch trays |
| RTX Blackwell | Graphics-oriented GB20x chips, including GB202 | GDDR7 rather than HBM; RTX PRO GB202 products use compute capability 12.0 and can provide up to 96 GB.[4][18][19] | Consumer, workstation, and server graphics products with fourth-generation RT Cores, fifth-generation Tensor Cores, neural-rendering features, and video engines |
| Blackwell Ultra | B300 and GB300 platform generation | Higher-capacity HBM3e and updated system configurations | Later platform evolution announced in March 2025; detailed specifications belong to the Blackwell Ultra article.[20] |

The table uses released documentation rather than launch-day peak charts where the two disagree. For example, the current B200 capacity is 180 GB, not the 192 GB figure repeated in some early briefs and third-party summaries.[3][4][5]

### B200 and DGX B200

B200 is the basic GB100 accelerator product. DGX B200 is a complete server containing eight B200 GPUs, two Intel Xeon 8570 processors, two NVSwitches, storage, networking, and management hardware. NVIDIA's user guide lists 1,440 GB of aggregate GPU memory, 14.4 TB/s of aggregate switch bandwidth, a 10U air-cooled chassis, and 14.3 kW maximum AC input.[6]

These figures correct two common category errors. The 14.3 kW value is a server maximum, not one GPU's steady workload draw. Conversely, the GPU's configurable 1,000 W ceiling does not include CPUs, memory, storage, networking, fans, or power-conversion losses. Actual consumption varies with firmware limits and workload.

### GB200 NVL72

GB200 NVL72 is a rack-scale system, not a single accelerator. It uses 18 compute trays, each with two Grace CPUs and four Blackwell GPUs, plus nine NVLink switch trays.[7] The 130 TB/s figure is the aggregate bidirectional bandwidth of the 72-GPU NVLink domain.[5][8] Dividing or comparing this number with a single PCIe link without preserving the aggregation boundary is misleading.

The rack uses direct liquid cooling for dense compute trays.[7] This does not mean all Blackwell systems require liquid cooling: DGX B200 is an air-cooled 10U server, and workstation and consumer products use their own cooling designs.[6][18]

### RTX Blackwell

RTX Blackwell is related architecturally but differs substantially from GB100. NVIDIA's RTX architecture paper describes GB202 as the flagship RTX Blackwell GPU and documents graphics processing clusters, fourth-generation ray-tracing cores, fifth-generation Tensor Cores, neural shaders, GDDR7, ninth-generation NVENC, and sixth-generation NVDEC.[18] NVIDIA's compute-capability table places current GB202 professional products at 12.0, rather than GB100's 10.0.[4]

The distinction matters for software distribution. A CUDA binary compiled only for `sm_100` is not automatically a native `sm_120` binary, even though both products carry the Blackwell name. Applications should ship appropriate architecture targets or forward-compatible PTX where supported, then test the actual driver and library combination.

## Software model

Blackwell uses the existing CUDA programming model, with architecture-specific additions exposed through CUDA libraries, PTX, and optimized frameworks. Existing CUDA source can often be rebuilt for a Blackwell target, but peak features require new kernels. Fifth-generation Tensor Core operations use the `tcgen05` instruction family and TMEM. B200 also supports an opted-in nonportable thread-block cluster size of 16, beyond CUDA's portable maximum of eight.[3][9]

The compute-capability split is a practical compatibility boundary:

- B200 and GB200 are GB100 products with compute capability 10.0.[4]
- The RTX PRO Blackwell products listed in NVIDIA's supported-GPU table are compute capability 12.0.[4]

Source-level CUDA portability does not imply identical instruction support or performance. Build systems should select correct `sm_` targets, retain a tested fallback where needed, and avoid using the word "Blackwell" as if it identified one instruction set.

Higher-level software such as TensorRT-LLM, NCCL, cuBLAS, CUTLASS, and framework-specific backends supplies kernels and communication algorithms that expose much of the hardware. This makes benchmark results inseparable from software versions. A driver, CUDA toolkit, inference engine, kernel-selection policy, precision recipe, and topology should be reported alongside the GPU model.

## Performance evidence

### Vendor projections

At launch, NVIDIA projected that GB200 NVL72 could provide up to 30 times the real-time inference throughput and up to 25 times lower cost and energy than an H100-based comparison.[1][7] The accompanying methodology used a projected 1.8-trillion-parameter mixture-of-experts model, 50 ms token-to-token latency, 5,000 ms first-token latency, 32,768 input tokens, 1,024 output tokens, nine eight-GPU HGX H100 systems, and one liquid-cooled GB200 NVL72. NVIDIA explicitly marked the result as projected and subject to change.[7]

Those numbers are not general GPU speedups. They combine a new GPU, FP4, a different scale-up topology, more memory, updated software, and a workload chosen to stress very large-model inference. They should not be applied to small models, FP32 scientific kernels, graphics rendering, or a one-GPU comparison.

### MLPerf Inference

MLPerf Inference provides measured, accuracy-constrained system results with disclosed configurations. The suite is open source, peer reviewed, architecture neutral, and organized by workload and scenario.[21] In version 5.1, Nebius submitted both an eight-GPU B200 system and an eight-GPU H200 system using TensorRT 10.11 and CUDA 12.9. For the Llama 2 70B 99-percent-accuracy benchmark, the B200 system reported 101,611 tokens/s in the Server scenario and 101,246 tokens/s Offline, while the H200 system reported 34,029.4 and 34,812.1 tokens/s respectively.[22] The ratios are about 2.99 and 2.91.

This is a useful system-level comparison, but it is not a controlled measure of the silicon alone. The B200 submission used FP4, the H200 submission used FP8, the host CPUs differed, and the systems had different GPU power limits. The result shows what those disclosed hardware and software configurations achieved while meeting the benchmark's accuracy requirements. It does not isolate how much of the difference came from Tensor Cores, memory, precision, kernels, or host configuration.

MLPerf Training uses time to reach a defined quality target rather than raw operations per second. Its peer-reviewed benchmark paper explains why fixed epoch counts or peak arithmetic rates do not provide a fair training comparison across systems.[23] Blackwell systems appeared in later MLPerf Training rounds, including version 6.0 in June 2026, but individual scores remain tied to a named model, quality target, system size, and submitted software stack.[26]

### Independent microbenchmarking

An independent University of Delaware study, accepted and presented at IEEE IPDPS 2026, characterized B200 Tensor Core instructions, TMEM, decompression, memory behavior, inference, and training.[24][25] Its reported tests found 1.85 times the ResNet-50 mixed-precision training throughput and 1.55 times the GPT-1.3B training throughput of its H200 comparison, with 32 percent higher training energy efficiency in the reported GPT run.[24]

The study is valuable because it measures features not fully described in vendor peak tables, but it also states limits. Its experiments primarily exercised single-die behavior, leaving cross-die latency and NV-HBI interactions for future work. The authors used particular CUDA, library, model, batch, and cloud configurations. Their numbers should be read as reproducible observations for those tests, not universal application speedups.[24][25]

Together, the three evidence types answer different questions:

| Evidence | What it can support | Main limitation |
|---|---|---|
| NVIDIA launch projections | Intended architecture, topology, and a modeled target workload | Vendor-selected assumptions; some figures were projections |
| MLPerf submissions | Measured end-to-end performance under common rules and accuracy constraints | Results include the whole submitted system and software stack |
| Independent microbenchmarks | Behavior of instructions, memory paths, and selected applications | Narrow configurations and incomplete coverage of rack-scale or cross-die behavior |

## Power, cooling, and deployment

Blackwell's data center performance comes with substantial infrastructure requirements. A B200 can be configured up to 1,000 W, and an eight-GPU DGX B200 has a 14.3 kW maximum input specification.[5][6] GB200 NVL72 places 72 GPUs in a liquid-cooled rack-scale design.[7] Facilities must account for electrical delivery, cooling-water or air-handling capacity, rack weight, network cabling, service access, and redundancy as well as accelerator power.

Energy efficiency should be reported as work completed per unit of energy under a stated workload. A lower time to solution can offset higher instantaneous power, but peak Tensor Core operations per watt do not determine total facility energy. CPU work, memory, network switches, cooling, idle time, failed runs, and model-quality requirements contribute to the result.

Cooling is product-specific. DGX B200 is documented as an air-cooled server, while GB200 NVL72 uses liquid-cooled compute trays.[6][7] Describing "Blackwell cooling" without naming the system merges incompatible deployment requirements.

## Limitations and interpretation

### Product fragmentation

The Blackwell label covers GB100 data center GPUs, GB20x graphics GPUs, Grace Blackwell systems, and Blackwell Ultra. They differ in compute capability, memory technology, power envelope, graphics features, and packaging. Specifications must be attached to a full product name.

### Low-precision accuracy

FP4 and FP6 increase potential throughput and reduce memory traffic, but quantization error is model and layer dependent. Calibration, scaling, clipping, fine-tuning, or mixed-precision exceptions may be necessary. Benchmark accuracy thresholds show that one submitted recipe passed one test; they do not prove exact equivalence for every downstream task.[10][11][21]

### Peak versus sustained performance

Published Tensor Core rates assume specified formats and, in some cases, sparsity. Real applications include data movement, attention operations, synchronization, nonmatrix layers, tokenization, control flow, and host work. Memory capacity may make a workload possible without making every kernel faster.

### Scale-up versus scale-out

NVLink and NVSwitch provide high-bandwidth scale-up communication inside supported domains. Larger deployments still rely on external networking. Collective efficiency falls below link peak when message sizes, routing, software, or workload balance are unfavorable.

### Software maturity

TMEM, `tcgen05`, FP4, and newer cluster features require Blackwell-aware tools and kernels. A framework release may support B200 without supporting every precision or topology. Version pinning, correctness tests, and profiling are necessary when moving from Hopper or between GB100 and GB20x.

### Procurement and pricing

Data center Blackwell accelerators are normally sold through complete systems, OEMs, cloud providers, and negotiated enterprise agreements. The official sources used here do not establish one durable standalone B200 street price. Cloud hourly rates and system quotes vary by region, contract, networking, storage, and availability, so a single unsourced price would be misleading.

## Status at the research cutoff

At the end-of-day July 28, 2026 research cutoff, Blackwell was a shipping multi-market family rather than a future-only roadmap. NVIDIA had reported mass production of the original data center generation, GeForce RTX 50 and RTX PRO Blackwell products had launched, GB200 NVL72 had entered full-scale production, and Blackwell Ultra had been introduced as a separate platform evolution.[15][16][17][19][20] MLPerf and independent research provided measured evidence for B200 and GB200-class systems by that date.[22][24][26]

The cutoff does not freeze all product pages at one specification. NVIDIA documentation continued to distinguish B200 at 180 GB, GB200 at 186 GB, and RTX GB202 at compute capability 12.0.[4] Any future update should preserve those boundaries, identify the exact document revision, and avoid copying a later Ultra feature into the original B200.

## See also

- [GPU computing](https://aiwiki.ai/wiki/gpu_computing)
- [Data Center](https://aiwiki.ai/wiki/data_center)
- [Quantization](https://aiwiki.ai/wiki/quantization)
- [Model Compression](https://aiwiki.ai/wiki/model_compression)
- [Inference](https://aiwiki.ai/wiki/inference)
- [TensorRT](https://aiwiki.ai/wiki/tensorrt)

## References

1. NVIDIA Corporation. "NVIDIA Blackwell Platform Arrives to Power a New Era of Computing." March 18, 2024. https://nvidianews.nvidia.com/news/nvidia-blackwell-platform-arrives-to-power-a-new-era-of-computing
2. Bickel, Peter J. "David Blackwell, 1919-2010: An explorer in mathematics and statistics." Proceedings of the National Academy of Sciences 117, no. 48 (2020). https://statistics.berkeley.edu/sites/default/files/news/blackwellpnasremembrance.pdf
3. NVIDIA Corporation. Blackwell Tuning Guide, CUDA Toolkit 12.8 archive. January 2025. https://docs.nvidia.com/cuda/archive/12.8.0/pdf/Blackwell_Tuning_Guide.pdf
4. NVIDIA Corporation. NVIDIA Multi-Instance GPU User Guide, Release r580. November 2025. https://docs.nvidia.com/datacenter/tesla/pdf/NVIDIA_MIG_User_Guide.pdf
5. NVIDIA Corporation. NVIDIA Blackwell Architecture Datasheet. October 2025. https://nvdam.widen.net/s/wwnsxrhm2w/blackwell-datasheet-3384703
6. NVIDIA Corporation. NVIDIA DGX B200 User Guide. June 2026. https://docs.nvidia.com/dgx/dgxb200-user-guide/dgxb200-user-guide.pdf
7. Goldwasser, Ivan, Harry Petty, Pradyumna Desale, and Kirthi Devleker. "NVIDIA GB200 NVL72 Delivers Trillion-Parameter LLM Training and Real-Time Inference." NVIDIA Technical Blog, March 18, 2024. https://developer.nvidia.com/blog/nvidia-gb200-nvl72-delivers-trillion-parameter-llm-training-and-real-time-inference/
8. NVIDIA Corporation. "GB200 NVL72." https://www.nvidia.com/en-us/data-center/gb200-nvl72/
9. NVIDIA Corporation. "TensorCore 5th Generation Instructions." Parallel Thread Execution ISA 8.8, CUDA Toolkit 12.9.2 archive. https://docs.nvidia.com/cuda/archive/12.9.2/parallel-thread-execution/#tcgen05-family-instructions
10. Rouhani, Bita Darvish, et al. "Microscaling Data Formats for Deep Learning." October 2023. https://www.microsoft.com/en-us/research/publication/microscaling-data-formats-for-deep-learning/
11. Sharify, Sayeh, et al. "Post Training Quantization of Large Language Models with Microscaling Formats." Proceedings of Machine Learning Research 262 (2024): 241-258. https://proceedings.mlr.press/v262/sharify24a.html
12. NVIDIA Corporation. NVIDIA Secure AI with Blackwell and Hopper GPUs, version 1.3. January 2026. https://docs.nvidia.com/nvidia-secure-ai-with-blackwell-and-hopper-gpus-whitepaper.pdf
13. NVIDIA Corporation. "Quarterly Report for the period ended July 28, 2024." Form 10-Q. United States Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1045810/000104581024000264/nvda-20240728.htm
14. NVIDIA Corporation. "Q3 Fiscal 2025 CFO Commentary." November 2024. United States Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1045810/000104581024000315/q3fy25cfocommentary.htm
15. NVIDIA Corporation. "NVIDIA Announces Financial Results for Fourth Quarter and Fiscal 2025." February 26, 2025. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-fourth-quarter-and-fiscal-2025
16. NVIDIA Corporation. "NVIDIA Announces Financial Results for First Quarter Fiscal 2026." May 28, 2025. https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-first-quarter-fiscal-2026
17. NVIDIA Corporation. "NVIDIA Blackwell GeForce RTX 50 Series Opens New World of AI Computer Graphics." January 6, 2025. https://nvidianews.nvidia.com/news/nvidia-blackwell-geforce-rtx-50-series-opens-new-world-of-ai-computer-graphics
18. NVIDIA Corporation. NVIDIA RTX Blackwell PRO GPU Architecture, version 1.1. July 2025. https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/pdf/NVIDIA-RTX-Blackwell-PRO-GPU-Architecture-v1_1.pdf
19. NVIDIA Corporation. "NVIDIA Blackwell RTX PRO Comes to Workstations and Servers." March 18, 2025. https://nvidianews.nvidia.com/news/nvidia-blackwell-rtx-pro-workstations-servers-agentic-ai
20. NVIDIA Corporation. "NVIDIA Blackwell Ultra AI Factory Platform Paves Way for Age of AI Reasoning." March 18, 2025. https://nvidianews.nvidia.com/news/nvidia-blackwell-ultra-ai-factory-platform-paves-way-for-age-of-ai-reasoning
21. MLCommons. "MLCommons Releases New MLPerf Inference v5.1 Benchmark Results." September 9, 2025. https://mlcommons.org/2025/09/mlperf-inference-v5-1-results/
22. MLCommons. "MLPerf Inference v5.1 Results: Datacenter, Available, Closed Division." September 2025. https://docs.mlcommons.org/inference_results_v5.1/
23. Mattson, Peter, et al. "MLPerf Training Benchmark." Proceedings of Machine Learning and Systems 2 (2020): 336-349. https://proceedings.mlsys.org/paper_files/paper/2020/hash/411e39b117e885341f25efb8912945f7-Abstract.html
24. Jarmusch, Aaron, and Sunita Chandrasekaran. "Microbenchmarking NVIDIA's Blackwell Architecture: An in-depth Architectural Analysis." IEEE International Parallel and Distributed Processing Symposium, 2026. https://arxiv.org/pdf/2512.02189
25. Jarmusch, Aaron. "Publications: Microbenchmarking NVIDIA's Blackwell Architecture." https://ajarmusch.github.io/publications/
26. MLCommons. "MLPerf Training v6.0 Results: New MoE Benchmarks and Record System Diversity." June 16, 2026. https://mlcommons.org/2026/06/mlperf-training-v6-0-results/

