GPU Cloud Comparison
Next-generation (2026+). Superchip (Vera CPU + Rubin GPU), HBM4 memory, NVLink 6, designed for trillion-parameter models and exascale HPC.
R100
Entry-level Rubin GPU for AI workloads. Successor to B100.
R200
High-performance Rubin GPU for AI training and HPC. Successor to B200.
Rubin Ultra
NVIDIA's next-generation flagship GPU for extreme-scale AI training and exascale HPC. Successor to Blackwell Ultra.
Vera Rubin
Vera CPU + Rubin GPU superchip — NVIDIA's next-generation integrated architecture for AI factories.
Current flagship (2024-2025). HBM3e memory, NVLink 5, 2.5× faster than H100 for LLM workloads. Up to 288GB VRAM.
B300
Blackwell Ultra standalone GPU. 288GB HBM3e — 1.5× more VRAM than B200. Flagship for extreme-scale AI training.
GB300 NVL72
NVIDIA's flagship rack-scale AI system with 72 Blackwell Ultra GPUs. Designed for trillion-parameter model training.
RTX PRO 6000 Blackwell
Professional workstation/datacenter GPU based on Blackwell architecture. 96GB GDDR7 for AI, rendering, and graphics workloads.
RTX PRO 6000 SE
Server Edition of RTX PRO 6000 Blackwell. 96GB GDDR7. Optimized for datacenter deployment with lower TDP (300W PCIe).
B100
Entry-level Blackwell GPU. Lower power variant of B200 for mainstream AI workloads.
B200
NVIDIA B200 Blackwell flagship accelerator. Available in two form factors: SXM for highest power density (1000W, used in HGX B200 baseboards with NVLink 5), and PCIe as B200 NVL for traditional air-cooled server chassis (800W). Features 192 GB HBM3e memory with 8.0 TB/s bandwidth. 2.5× faster AI training than H100.
GB200
Grace + Blackwell superchip with unified memory. The foundation of the GB200 NVL72 rack system.
GB200 NVL72
NVIDIA's flagship rack-scale AI system with 72 Blackwell GPUs and 36 Grace CPUs. Designed for trillion-parameter model training.
Industry standard (2022-2024). HBM3 memory, Transformer Engine, NVLink 4. The most widely-deployed AI GPU architecture.
H200 NVL
NVL Gen5 form factor of H200. 141GB HBM3e but no NVLink — uses NVL Gen5 (128 GB/s) instead. Lower power (350W) for environments without SXM.
H200 SXM
SXM5 form factor of H200. 141GB HBM3e with full NVLink 4 (900 GB/s). Highest bandwidth Hopper variant for large model training.
GH200
Grace + Hopper superchip with unified memory. Designed for giant-scale AI and HPC workloads.
H100 NVL
PCIe variant of H100 with 94GB HBM3 (more VRAM than SXM version). Uses NVLink bridge for multi-GPU on PCIe.
H20
China compliance variant of the H100. Meets US export controls with reduced compute performance but retains full 96GB HBM3 memory and NVLink interconnect. Designed for large-model inference in the China market.
H800 SXM
China export variant of the H100 SXM. Meets US export controls with reduced NVLink bandwidth (400 GB/s vs 900 GB/s) while retaining 80GB HBM3 and full compute density. Built for AI training in compliant regions.
GH100
The GH100 is NVIDIA's Hopper architecture GPU silicon — the foundational chip that powers all H100 variants (SXM5, PCIe, NVL). Built on TSMC's custom 4N process with 80 billion transistors across an 814 mm² die, it introduced the Transformer Engine for accelerated LLM training and inference. The GH100 die is the largest and most complex GPU NVIDIA had shipped at its 2022 launch.
H100 PCIe
Standard PCIe Gen5 form factor of H100 (no NVLink). 350W TDP — lower performance than SXM but easier to deploy. No multi-GPU NVLink interconnect.
H100 SXM
SXM5 form factor of H100. 80GB HBM3 with full NVLink 4 (900 GB/s). 700W TDP — the highest performance H100 variant for multi-GPU training.
Inference & rendering (2022-2024). GDDR6 memory, optimized for AI inference, graphics, and video processing.
RTX 2000 Ada
Entry-level Ada Lovelace workstation GPU. 16GB GDDR6 with low 70W TDP for single-slot deployments. Optimized for AI inference, CAD, and content creation workloads.
L2
China market variant of the L4. 24GB GDDR6 in a low-profile 150W form factor for AI inference and video workloads in compliant regions.
L20
China market variant of the L40. 48GB GDDR6 for AI inference and graphics. Meets US export controls while delivering high performance for data center workloads.
L40S
Datacenter GPU optimized for AI inference and graphics. Successor to A40 with 2× performance.
RTX 4000 Ada
Entry-level Ada Lovelace workstation GPU. 20GB GDDR6, single-slot design for compact workstations.
RTX 4500 Ada
Mid-range Ada Lovelace workstation GPU. 24GB GDDR6 for professional rendering and AI workloads.
RTX 5000 Ada
Mid-range Ada Lovelace workstation GPU. 32GB GDDR6 for professional rendering and AI workloads.
L4
Low-power, single-slot GPU for AI inference and video processing. 72W TDP — no extra power connector needed.
L40
Datacenter GPU for visual computing and rendering. Lower AI performance than L40S.
RTX 6000 Ada
Workstation GPU for professional AI, rendering, and graphics. 48GB GDDR6 with ECC support.
Mainstream (2020-2022). HBM2e memory, MIG support, NVLink 3. The most widely-deployed AI GPU in cloud today.
A800 SXM4
China export variant of the A100 SXM4 80GB. Meets US export controls with reduced interconnect bandwidth. Built on the Ampere architecture for AI training and HPC workloads in compliant regions.
A10
Mid-range AI inference and graphics GPU. 24GB GDDR6 with low power consumption.
A100 PCIe
Standard PCIe Gen4 form factor (no NVLink). 80GB HBM2e. 300W TDP — no multi-GPU NVLink interconnect.
A100 SXM 80GB
SXM4 form factor, 80GB HBM2e variant. Full NVLink 3 (600 GB/s). 2× VRAM of original A100 for larger models.
A16
Multi-instance GPU for VDI (Virtual Desktop Infrastructure). 4 GPUs in one board for virtualization.
A2
Entry-level AI inference GPU. 40-60W TDP — lowest power datacenter GPU. Single-slot, passive cooling.
A30
Mid-range AI and HPC GPU with HBM2 memory. Lower power than A100.
RTX A4000
Entry-level Ampere workstation GPU. 16GB GDDR6, single-slot design for compact workstations.
RTX A5000
Mid-range Ampere workstation GPU. 24GB GDDR6 for professional rendering and AI workloads.
A40
Datacenter GPU for visual computing. 48GB GDDR6 with ECC for professional workloads.
RTX A6000
Flagship Ampere workstation GPU. 48GB GDDR6 with ECC for professional AI, rendering, and graphics.
Budget inference (2018-2020). GDDR6 memory, low power. The most affordable AI GPUs for light inference workloads.
Quadro RTX 6000
Professional workstation GPU (Turing). 24GB GDDR6 with ECC for rendering, AI, and graphics workloads.
Quadro RTX 8000
High-end professional workstation GPU (Turing). 48GB GDDR6 — the largest VRAM in the Turing workstation lineup.
T4
Low-cost, low-power inference GPU. 70W TDP — no extra power connector needed. Most affordable AI GPU.
Legacy (2017). First GPU with Tensor Cores. Mostly replaced by Ampere/Hopper but still in use.