GPU Cloud Comparison

Compare every NVIDIA GPU for AI training, inference, and HPC.

From Volta (2017) to Rubin Ultra (2026+) — 44 GPU models across 7 families. Find the right GPU by specs, use case, and cloud provider pricing.

Next-generation (2026+). HBM4 memory, NVLink 6, designed for trillion-parameter models and exascale HPC.

Blackwell10 GPUs

Current flagship (2024-2025). HBM3e memory, NVLink 5, 2.5× faster than H100 for LLM workloads. Up to 288GB VRAM.

Available2025

B300

Blackwell Ultra standalone GPU. 288GB HBM3e — 1.5× more VRAM than B200. Flagship for extreme-scale AI training.

VRAM288 GB HBM3e
Bandwidth8.0 TB/s
InterconnectNVLink 5 (1.8 TB/s)
Available2025

GB300 NVL72

NVIDIA's flagship rack-scale AI system with 72 Blackwell Ultra GPUs. Designed for trillion-parameter model training.

VRAM13,824 GB HBM3e (192 GB per GPU × 72)
Bandwidth8.0 TB/s per GPU
InterconnectNVLink 5 (1.8 TB/s per GPU)
Available2025

HGX B300

8-GPU HGX server board with Blackwell Ultra GPUs. The successor to HGX B200 with higher memory and performance.

VRAM1,536 GB HBM3e (192 GB per GPU × 8)
Bandwidth8.0 TB/s per GPU
InterconnectNVLink 5 (1.8 TB/s), NVSwitch
Available2025

RTX PRO 6000 Blackwell

Professional workstation/datacenter GPU based on Blackwell architecture. 96GB GDDR7 for AI, rendering, and graphics workloads.

VRAM96 GB GDDR7
Bandwidth1,792 GB/s
InterconnectPCIe Gen5 (128 GB/s) / NVLink 5
Available2025

RTX PRO 6000 SE

Server Edition of RTX PRO 6000 Blackwell. 96GB GDDR7. Optimized for datacenter deployment with lower TDP (300W PCIe).

VRAM96 GB GDDR7
Bandwidth1,792 GB/s
InterconnectPCIe Gen5 (128 GB/s) — no NVLink
Available2024

B100

Entry-level Blackwell GPU. Lower power variant of B200 for mainstream AI workloads.

VRAM192 GB HBM3e
Bandwidth8.0 TB/s
InterconnectNVLink 5 (1.8 TB/s)
Available2024

B200

NVIDIA's flagship Blackwell GPU for AI training and inference. 2.5× faster than H100 for LLM workloads.

VRAM192 GB HBM3e
Bandwidth8.0 TB/s
InterconnectNVLink 5 (1.8 TB/s)
Available2024

GB200

Grace + Blackwell superchip with unified memory. The foundation of the GB200 NVL72 rack system.

VRAM192 GB HBM3e
Bandwidth8.0 TB/s
InterconnectNVLink 5 (1.8 TB/s)
Available2024

GB200 NVL72

NVIDIA's flagship rack-scale AI system with 72 Blackwell GPUs and 36 Grace CPUs. Designed for trillion-parameter model training.

VRAM13,824 GB HBM3e (192 GB per GPU × 72)
Bandwidth8.0 TB/s per GPU
InterconnectNVLink 5 (1.8 TB/s per GPU)
Available2024

HGX B200

8-GPU HGX server board with Blackwell GPUs. The standard for next-generation AI training servers.

VRAM1,536 GB HBM3e (192 GB per GPU × 8)
Bandwidth8.0 TB/s per GPU
InterconnectNVLink 5 (1.8 TB/s), NVSwitch
Hopper6 GPUs

Industry standard (2022-2024). HBM3 memory, Transformer Engine, NVLink 4. The most widely-deployed AI GPU architecture.

Inference & rendering (2022-2024). GDDR6 memory, optimized for AI inference, graphics, and video processing.

Ampere12 GPUs

Mainstream (2020-2022). HBM2e memory, MIG support, NVLink 3. The most widely-deployed AI GPU in cloud today.

Available2021

A10

Mid-range AI inference and graphics GPU. 24GB GDDR6 with low power consumption.

VRAM24 GB GDDR6
Bandwidth600 GB/s
InterconnectPCIe Gen4 (64 GB/s)
Available2021

A100 NVLink

PCIe Gen4 form factor with NVLink bridge. 80GB HBM2e. NVLink 3 (600 GB/s) via bridge card — enables multi-GPU on PCIe servers.

VRAM80 GB HBM2e
Bandwidth2.0 TB/s
InterconnectNVLink 3 (600 GB/s) via NVLink bridge
Available2021

A100 PCIe

Standard PCIe Gen4 form factor (no NVLink). 80GB HBM2e. 300W TDP — no multi-GPU NVLink interconnect.

VRAM80 GB HBM2e
Bandwidth2.0 TB/s
InterconnectPCIe Gen4 (64 GB/s) — no NVLink
Available2021

A100 SXM 80GB

SXM4 form factor, 80GB HBM2e variant. Full NVLink 3 (600 GB/s). 2× VRAM of original A100 for larger models.

VRAM80 GB HBM2e
Bandwidth2.0 TB/s
InterconnectNVLink 3 (600 GB/s)
Available2021

A16

Multi-instance GPU for VDI (Virtual Desktop Infrastructure). 4 GPUs in one board for virtualization.

VRAM64 GB GDDR6 (16 GB × 4 GPUs)
BandwidthTBD
InterconnectPCIe Gen4
Available2021

A2

Entry-level AI inference GPU. 40-60W TDP — lowest power datacenter GPU. Single-slot, passive cooling.

VRAM16 GB GDDR6
Bandwidth200 GB/s
InterconnectPCIe Gen4 (64 GB/s)
Available2021

A30

Mid-range AI and HPC GPU with HBM2 memory. Lower power than A100.

VRAM24 GB HBM2
Bandwidth933 GB/s
InterconnectNVLink 3 (600 GB/s)
Available2021

RTX A4000

Entry-level Ampere workstation GPU. 16GB GDDR6, single-slot design for compact workstations.

VRAM16 GB GDDR6
Bandwidth448 GB/s
InterconnectPCIe Gen4 (64 GB/s)
Available2021

RTX A5000

Mid-range Ampere workstation GPU. 24GB GDDR6 for professional rendering and AI workloads.

VRAM24 GB GDDR6
Bandwidth768 GB/s
InterconnectPCIe Gen4 (64 GB/s)
Available2020

A100 SXM 40GB

SXM4 form factor, 40GB HBM2e variant. Full NVLink 3 (600 GB/s). The original A100 for AI training and HPC.

VRAM40 GB HBM2e
Bandwidth1.55 TB/s
InterconnectNVLink 3 (600 GB/s)
Available2020

A40

Datacenter GPU for visual computing. 48GB GDDR6 with ECC for professional workloads.

VRAM48 GB GDDR6
Bandwidth696 GB/s
InterconnectPCIe Gen4 (64 GB/s)
Available2020

RTX A6000

Flagship Ampere workstation GPU. 48GB GDDR6 with ECC for professional AI, rendering, and graphics.

VRAM48 GB GDDR6
Bandwidth768 GB/s
InterconnectPCIe Gen4 (64 GB/s) / NVLink Bridge
Turing3 GPUs

Budget inference (2018-2020). GDDR6 memory, low power. The most affordable AI GPUs for light inference workloads.

Volta1 GPU

Legacy (2017). First GPU with Tensor Cores. Mostly replaced by Ampere/Hopper but still in use.