Mainstream (2020-2022)

NVIDIA
Ampere

NVIDIA Ampere is the most widely-deployed AI GPU architecture in the cloud. Featuring HBM2e memory, MIG (Multi-Instance GPU) support, and NVLink 3 interconnects, the A100 remains a workhorse for AI training and inference.

GPU Models in this Family

Click any card to expand detailed specifications

A100 SXM 80GB

A100 SXM 80GB

80 GB HBM2e2.0 TB/sSXM4

SXM4 form factor, 80GB HBM2e variant. Full NVLink 3 (600 GB/s). 2× VRAM of original A100 for larger models.

AI trainingLLM trainingHPCMIG workloads
Architecture
Ampere
VRAM
80 GB HBM2e
Memory Bandwidth
2.0 TB/s
Interconnect
NVLink 3 (600 GB/s)
FP16 Performance
624 TFLOPS (with sparsity)
FP32 Performance
19.5 TFLOPS
FP64 Performance
9.7 TFLOPS
TDP
400W
CUDA Cores
6,912
Tensor Cores
432
Process Node
TSMC 7N
Form Factor
SXM4
A100 PCIe

A100 PCIe

80 GB HBM2e2.0 TB/sPCIe Gen4 (no NVLink)

Standard PCIe Gen4 form factor (no NVLink). 80GB HBM2e. 300W TDP — no multi-GPU NVLink interconnect.

AI inferenceSingle-GPU trainingHPCBudget workloads
Architecture
Ampere
VRAM
80 GB HBM2e
Memory Bandwidth
2.0 TB/s
Interconnect
PCIe Gen4 (64 GB/s) — no NVLink
FP16 Performance
624 TFLOPS (with sparsity)
FP32 Performance
19.5 TFLOPS
FP64 Performance
9.7 TFLOPS
TDP
300W
CUDA Cores
6,912
Tensor Cores
432
Process Node
TSMC 7N
Form Factor
PCIe Gen4 (no NVLink)
A800 SXM4

A800 SXM4

80 GB HBM2e2.0 TB/sSXM4

China export variant of the A100 SXM4 80GB. Meets US export controls with reduced interconnect bandwidth. Built on the Ampere architecture for AI training and HPC workloads in compliant regions.

AI trainingHPCInference (China market)
Architecture
Ampere
VRAM
80 GB HBM2e
Memory Bandwidth
2.0 TB/s
Interconnect
NVLink 3 (400 GB/s)
FP16 Performance
624 TFLOPS (with Tensor Cores)
FP32 Performance
19.5 TFLOPS
FP64 Performance
9.7 TFLOPS
TDP
400W
CUDA Cores
6,912
Tensor Cores
432
Process Node
TSMC 7N
Form Factor
SXM4
A40

A40

48 GB GDDR6696 GB/sPCIe Gen4

Datacenter GPU for visual computing. 48GB GDDR6 with ECC for professional workloads.

GraphicsRenderingVirtual workstationsVideo processing
Architecture
Ampere
VRAM
48 GB GDDR6
Memory Bandwidth
696 GB/s
Interconnect
PCIe Gen4 (64 GB/s)
FP16 Performance
300 TFLOPS (with sparsity)
FP32 Performance
37 TFLOPS
FP64 Performance
1.15 TFLOPS
TDP
300W
CUDA Cores
10,752
Tensor Cores
336
Process Node
TSMC 7N
Form Factor
PCIe Gen4
A30

A30

24 GB HBM2933 GB/sPCIe Gen4

Mid-range AI and HPC GPU with HBM2 memory. Lower power than A100.

AI training (small models)HPCInference
Architecture
Ampere
VRAM
24 GB HBM2
Memory Bandwidth
933 GB/s
Interconnect
NVLink 3 (600 GB/s)
FP16 Performance
330 TFLOPS (with sparsity)
FP32 Performance
10.3 TFLOPS
FP64 Performance
5.2 TFLOPS
TDP
165W
CUDA Cores
3,584
Tensor Cores
224
Process Node
TSMC 7N
Form Factor
PCIe Gen4
A16

A16

64 GB GDDR6 (16 GB × 4 GPUs)TBDPCIe Gen4

Multi-instance GPU for VDI (Virtual Desktop Infrastructure). 4 GPUs in one board for virtualization.

VDIVirtual desktopsCloud gamingMulti-tenant workloads
Architecture
Ampere
VRAM
64 GB GDDR6 (16 GB × 4 GPUs)
Memory Bandwidth
TBD
Interconnect
PCIe Gen4
FP16 Performance
TBD
FP32 Performance
TBD
FP64 Performance
TBD
TDP
250W
CUDA Cores
TBD
Tensor Cores
TBD
Process Node
TSMC 7N
Form Factor
PCIe Gen4
A10

A10

24 GB GDDR6600 GB/sPCIe Gen4

Mid-range AI inference and graphics GPU. 24GB GDDR6 with low power consumption.

AI inferenceGraphicsVideo processingVDI
Architecture
Ampere
VRAM
24 GB GDDR6
Memory Bandwidth
600 GB/s
Interconnect
PCIe Gen4 (64 GB/s)
FP16 Performance
250 TFLOPS (with sparsity)
FP32 Performance
31.2 TFLOPS
FP64 Performance
0.97 TFLOPS
TDP
150W
CUDA Cores
9,216
Tensor Cores
288
Process Node
TSMC 7N
Form Factor
PCIe Gen4
A2

A2

16 GB GDDR6200 GB/sPCIe Gen4 (single-slot, passive)

Entry-level AI inference GPU. 40-60W TDP — lowest power datacenter GPU. Single-slot, passive cooling.

AI inference (light)Video processingEdge computingVDI
Architecture
Ampere
VRAM
16 GB GDDR6
Memory Bandwidth
200 GB/s
Interconnect
PCIe Gen4 (64 GB/s)
FP16 Performance
45 TFLOPS (with sparsity)
FP32 Performance
4.5 TFLOPS
FP64 Performance
0.14 TFLOPS
TDP
40-60W
CUDA Cores
1,280
Tensor Cores
40
Process Node
TSMC 7N
Form Factor
PCIe Gen4 (single-slot, passive)
RTX A6000

RTX A6000

48 GB GDDR6768 GB/sPCIe Gen4

Flagship Ampere workstation GPU. 48GB GDDR6 with ECC for professional AI, rendering, and graphics.

Workstation AIRenderingGraphicsVirtual workstationsVideo processing
Architecture
Ampere
VRAM
48 GB GDDR6
Memory Bandwidth
768 GB/s
Interconnect
PCIe Gen4 (64 GB/s) / NVLink Bridge
FP16 Performance
310 TFLOPS (with sparsity)
FP32 Performance
38.7 TFLOPS
FP64 Performance
1.21 TFLOPS
TDP
300W
CUDA Cores
10,752
Tensor Cores
336
Process Node
TSMC 7N
Form Factor
PCIe Gen4
RTX A5000

RTX A5000

24 GB GDDR6768 GB/sPCIe Gen4

Mid-range Ampere workstation GPU. 24GB GDDR6 for professional rendering and AI workloads.

Workstation AIRenderingGraphicsVideo processing
Architecture
Ampere
VRAM
24 GB GDDR6
Memory Bandwidth
768 GB/s
Interconnect
PCIe Gen4 (64 GB/s)
FP16 Performance
222 TFLOPS (with sparsity)
FP32 Performance
27.8 TFLOPS
FP64 Performance
0.87 TFLOPS
TDP
230W
CUDA Cores
8,192
Tensor Cores
256
Process Node
TSMC 7N
Form Factor
PCIe Gen4
RTX A4000

RTX A4000

16 GB GDDR6448 GB/sPCIe Gen4 (single-slot)

Entry-level Ampere workstation GPU. 16GB GDDR6, single-slot design for compact workstations.

Workstation AIRenderingGraphicsBudget professional workloads
Architecture
Ampere
VRAM
16 GB GDDR6
Memory Bandwidth
448 GB/s
Interconnect
PCIe Gen4 (64 GB/s)
FP16 Performance
154 TFLOPS (with sparsity)
FP32 Performance
19.2 TFLOPS
FP64 Performance
0.60 TFLOPS
TDP
140W
CUDA Cores
6,144
Tensor Cores
192
Process Node
TSMC 7N
Form Factor
PCIe Gen4 (single-slot)

Quick Comparison

GPU ModelVRAMBandwidthFP16TDPStarting Price
A100 SXM 80GB80 GB HBM2e2.0 TB/s624 TFLOPS (with sparsity)400W
A100 PCIe80 GB HBM2e2.0 TB/s624 TFLOPS (with sparsity)300W
A800 SXM480 GB HBM2e2.0 TB/s624 TFLOPS (with Tensor Cores)400W
A4048 GB GDDR6696 GB/s300 TFLOPS (with sparsity)300W
A3024 GB HBM2933 GB/s330 TFLOPS (with sparsity)165W
A1664 GB GDDR6 (16 GB × 4 GPUs)TBDTBD250W
A1024 GB GDDR6600 GB/s250 TFLOPS (with sparsity)150W
A216 GB GDDR6200 GB/s45 TFLOPS (with sparsity)40-60W
RTX A600048 GB GDDR6768 GB/s310 TFLOPS (with sparsity)300W
RTX A500024 GB GDDR6768 GB/s222 TFLOPS (with sparsity)230W
RTX A400016 GB GDDR6448 GB/s154 TFLOPS (with sparsity)140W

Why Choose Ampere?

Proven Performance

Ampere GPUs power the world's largest AI training clusters. Battle-tested in production at scale for LLM training, fine-tuning, and inference.

Cloud-Native Ready

Available across multiple cloud providers with hourly billing. Compare real-time pricing and pick the cheapest option for your workload.

Ecosystem Support

Full compatibility with PyTorch, TensorFlow, JAX, vLLM, TensorRT, and all major ML frameworks. NVLink for multi-GPU scaling.