Current flagship (2024-2025)

NVIDIA
Blackwell

NVIDIA Blackwell is the current flagship GPU architecture, featuring HBM3e memory with up to 8.0 TB/s bandwidth, NVLink 5 interconnects (1.8 TB/s), and 2.5× faster performance than H100 for LLM workloads. Available in configurations up to 288GB VRAM.

GPU Models in this Family

Click any card to expand detailed specifications

GB300 NVL72

GB300 NVL72

13,824 GB HBM3e (192 GB per GPU × 72)8.0 TB/s per GPUNVL72 rack (72 GPUs)

NVIDIA's flagship rack-scale AI system with 72 Blackwell Ultra GPUs. Designed for trillion-parameter model training.

Trillion-parameter LLM trainingExascale HPCAI factories
Architecture
Blackwell Ultra
VRAM
13,824 GB HBM3e (192 GB per GPU × 72)
Memory Bandwidth
8.0 TB/s per GPU
Interconnect
NVLink 5 (1.8 TB/s per GPU)
FP16 Performance
1.1 ExaFLOPS (rack-level)
FP32 Performance
TBD
FP64 Performance
TBD
TDP
TBD
CUDA Cores
TBD
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
NVL72 rack (72 GPUs)
B300

B300

288 GB HBM3e8.0 TB/sSXM (dual-die)

Blackwell Ultra standalone GPU. 288GB HBM3e — 1.5× more VRAM than B200. Flagship for extreme-scale AI training.

Trillion-parameter LLM trainingExascale HPCAI factories
Architecture
Blackwell Ultra
VRAM
288 GB HBM3e
Memory Bandwidth
8.0 TB/s
Interconnect
NVLink 5 (1.8 TB/s)
FP16 Performance
2,500 TFLOPS (with sparsity)
FP32 Performance
125 TFLOPS
FP64 Performance
63 TFLOPS
TDP
1,200W
CUDA Cores
TBD
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
SXM (dual-die)
GB200 NVL72

GB200 NVL72

13,824 GB HBM3e (192 GB per GPU × 72)8.0 TB/s per GPUNVL72 rack (72 GPUs)

NVIDIA's flagship rack-scale AI system with 72 Blackwell GPUs and 36 Grace CPUs. Designed for trillion-parameter model training.

Trillion-parameter LLM trainingExascale HPCAI factories
Architecture
Blackwell
VRAM
13,824 GB HBM3e (192 GB per GPU × 72)
Memory Bandwidth
8.0 TB/s per GPU
Interconnect
NVLink 5 (1.8 TB/s per GPU)
FP16 Performance
720 PFLOPS (rack-level)
FP32 Performance
TBD
FP64 Performance
TBD
TDP
1,000W per GPU
CUDA Cores
18,432 per GPU
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
NVL72 rack (72 GPUs)
B200

B200

192 GB HBM3e8.0 TB/sSXM (HGX) / PCIe (B200 NVL)

NVIDIA B200 Blackwell flagship accelerator. Available in two form factors: SXM for highest power density (1000W, used in HGX B200 baseboards with NVLink 5), and PCIe as B200 NVL for traditional air-cooled server chassis (800W). Features 192 GB HBM3e memory with 8.0 TB/s bandwidth. 2.5× faster AI training than H100.

AI trainingLLM training (70B+ params)Multi-GPU trainingHPCInference
Architecture
Blackwell
VRAM
192 GB HBM3e
Memory Bandwidth
8.0 TB/s
Interconnect
NVLink 5 (1.8 TB/s)
FP16 Performance
2,250 TFLOPS (with sparsity)
FP32 Performance
90 TFLOPS
FP64 Performance
45 TFLOPS
TDP
1000W (SXM) / 800W (NVL)
CUDA Cores
18,432
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
SXM (HGX) / PCIe (B200 NVL)
B100

B100

192 GB HBM3e8.0 TB/sSXM (dual-die)

Entry-level Blackwell GPU. Lower power variant of B200 for mainstream AI workloads.

AI trainingInferenceHPC
Architecture
Blackwell
VRAM
192 GB HBM3e
Memory Bandwidth
8.0 TB/s
Interconnect
NVLink 5 (1.8 TB/s)
FP16 Performance
1,800 TFLOPS (with sparsity)
FP32 Performance
60 TFLOPS
FP64 Performance
30 TFLOPS
TDP
700W
CUDA Cores
TBD
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
SXM (dual-die)
RTX PRO 6000 Blackwell

RTX PRO 6000 Blackwell

96 GB GDDR71,792 GB/sPCIe Gen5 / SXM

Professional workstation/datacenter GPU based on Blackwell architecture. 96GB GDDR7 for AI, rendering, and graphics workloads.

Workstation AIRenderingGraphicsVideo processingVirtual workstations
Architecture
Blackwell
VRAM
96 GB GDDR7
Memory Bandwidth
1,792 GB/s
Interconnect
PCIe Gen5 (128 GB/s) / NVLink 5
FP16 Performance
1,250 TFLOPS (with sparsity)
FP32 Performance
125 TFLOPS
FP64 Performance
1.95 TFLOPS
TDP
600W (SXM), 300W (PCIe)
CUDA Cores
24,064
Tensor Cores
752
Process Node
TSMC 4NP
Form Factor
PCIe Gen5 / SXM
RTX PRO 6000 SE

RTX PRO 6000 SE

96 GB GDDR71,792 GB/sPCIe Gen5

Server Edition of RTX PRO 6000 Blackwell. 96GB GDDR7. Optimized for datacenter deployment with lower TDP (300W PCIe).

Datacenter AI inferenceRenderingGraphicsVideo processingVirtual workstations
Architecture
Blackwell
VRAM
96 GB GDDR7
Memory Bandwidth
1,792 GB/s
Interconnect
PCIe Gen5 (128 GB/s) — no NVLink
FP16 Performance
1,250 TFLOPS (with sparsity)
FP32 Performance
125 TFLOPS
FP64 Performance
1.95 TFLOPS
TDP
300W
CUDA Cores
24,064
Tensor Cores
752
Process Node
TSMC 4NP
Form Factor
PCIe Gen5
GB200

GB200

192 GB HBM3e8.0 TB/sSuperchip (Grace CPU + Blackwell GPU)

Grace + Blackwell superchip with unified memory. The foundation of the GB200 NVL72 rack system.

Large-scale AI trainingHPCInference at scale
Architecture
Blackwell
VRAM
192 GB HBM3e
Memory Bandwidth
8.0 TB/s
Interconnect
NVLink 5 (1.8 TB/s)
FP16 Performance
1,800 TFLOPS (with sparsity)
FP32 Performance
90 TFLOPS
FP64 Performance
45 TFLOPS
TDP
1,000W
CUDA Cores
18,432
Tensor Cores
TBD
Process Node
TSMC 4NP
Form Factor
Superchip (Grace CPU + Blackwell GPU)

Quick Comparison

GPU ModelVRAMBandwidthFP16TDPStarting Price
GB300 NVL7213,824 GB HBM3e (192 GB per GPU × 72)8.0 TB/s per GPU1.1 ExaFLOPS (rack-level)TBD
B300288 GB HBM3e8.0 TB/s2,500 TFLOPS (with sparsity)1,200W
GB200 NVL7213,824 GB HBM3e (192 GB per GPU × 72)8.0 TB/s per GPU720 PFLOPS (rack-level)1,000W per GPU
B200192 GB HBM3e8.0 TB/s2,250 TFLOPS (with sparsity)1000W (SXM) / 800W (NVL)
B100192 GB HBM3e8.0 TB/s1,800 TFLOPS (with sparsity)700W
RTX PRO 6000 Blackwell96 GB GDDR71,792 GB/s1,250 TFLOPS (with sparsity)600W (SXM), 300W (PCIe)
RTX PRO 6000 SE96 GB GDDR71,792 GB/s1,250 TFLOPS (with sparsity)300W
GB200192 GB HBM3e8.0 TB/s1,800 TFLOPS (with sparsity)1,000W

Why Choose Blackwell?

Proven Performance

Blackwell GPUs power the world's largest AI training clusters. Battle-tested in production at scale for LLM training, fine-tuning, and inference.

Cloud-Native Ready

Available across multiple cloud providers with hourly billing. Compare real-time pricing and pick the cheapest option for your workload.

Ecosystem Support

Full compatibility with PyTorch, TensorFlow, JAX, vLLM, TensorRT, and all major ML frameworks. NVLink for multi-GPU scaling.