
Mainstream (2020-2022)
NVIDIA
Ampere
NVIDIA Ampere is the most widely-deployed AI GPU architecture in the cloud. Featuring HBM2e memory, MIG (Multi-Instance GPU) support, and NVLink 3 interconnects, the A100 remains a workhorse for AI training and inference.
GPU Models in this Family
Click any card to expand detailed specifications

A100 SXM 80GB
80 GB HBM2e2.0 TB/sSXM4

A100 SXM 80GB
SXM4 form factor, 80GB HBM2e variant. Full NVLink 3 (600 GB/s). 2× VRAM of original A100 for larger models.

A100 PCIe
80 GB HBM2e2.0 TB/sPCIe Gen4 (no NVLink)

A100 PCIe
Standard PCIe Gen4 form factor (no NVLink). 80GB HBM2e. 300W TDP — no multi-GPU NVLink interconnect.

A800 SXM4
80 GB HBM2e2.0 TB/sSXM4

A800 SXM4
China export variant of the A100 SXM4 80GB. Meets US export controls with reduced interconnect bandwidth. Built on the Ampere architecture for AI training and HPC workloads in compliant regions.

A40
48 GB GDDR6696 GB/sPCIe Gen4

A40
Datacenter GPU for visual computing. 48GB GDDR6 with ECC for professional workloads.

A30
24 GB HBM2933 GB/sPCIe Gen4

A30
Mid-range AI and HPC GPU with HBM2 memory. Lower power than A100.

A16
64 GB GDDR6 (16 GB × 4 GPUs)TBDPCIe Gen4

A16
Multi-instance GPU for VDI (Virtual Desktop Infrastructure). 4 GPUs in one board for virtualization.

A10
24 GB GDDR6600 GB/sPCIe Gen4

A10
Mid-range AI inference and graphics GPU. 24GB GDDR6 with low power consumption.

A2
16 GB GDDR6200 GB/sPCIe Gen4 (single-slot, passive)

A2
Entry-level AI inference GPU. 40-60W TDP — lowest power datacenter GPU. Single-slot, passive cooling.

RTX A6000
48 GB GDDR6768 GB/sPCIe Gen4

RTX A6000
Flagship Ampere workstation GPU. 48GB GDDR6 with ECC for professional AI, rendering, and graphics.

RTX A5000
24 GB GDDR6768 GB/sPCIe Gen4

RTX A5000
Mid-range Ampere workstation GPU. 24GB GDDR6 for professional rendering and AI workloads.

RTX A4000
16 GB GDDR6448 GB/sPCIe Gen4 (single-slot)

RTX A4000
Entry-level Ampere workstation GPU. 16GB GDDR6, single-slot design for compact workstations.
Quick Comparison
| GPU Model | VRAM | Bandwidth | FP16 | TDP | Starting Price |
|---|---|---|---|---|---|
| A100 SXM 80GB | 80 GB HBM2e | 2.0 TB/s | 624 TFLOPS (with sparsity) | 400W | — |
| A100 PCIe | 80 GB HBM2e | 2.0 TB/s | 624 TFLOPS (with sparsity) | 300W | — |
| A800 SXM4 | 80 GB HBM2e | 2.0 TB/s | 624 TFLOPS (with Tensor Cores) | 400W | — |
| A40 | 48 GB GDDR6 | 696 GB/s | 300 TFLOPS (with sparsity) | 300W | — |
| A30 | 24 GB HBM2 | 933 GB/s | 330 TFLOPS (with sparsity) | 165W | — |
| A16 | 64 GB GDDR6 (16 GB × 4 GPUs) | TBD | TBD | 250W | — |
| A10 | 24 GB GDDR6 | 600 GB/s | 250 TFLOPS (with sparsity) | 150W | — |
| A2 | 16 GB GDDR6 | 200 GB/s | 45 TFLOPS (with sparsity) | 40-60W | — |
| RTX A6000 | 48 GB GDDR6 | 768 GB/s | 310 TFLOPS (with sparsity) | 300W | — |
| RTX A5000 | 24 GB GDDR6 | 768 GB/s | 222 TFLOPS (with sparsity) | 230W | — |
| RTX A4000 | 16 GB GDDR6 | 448 GB/s | 154 TFLOPS (with sparsity) | 140W | — |
Why Choose Ampere?
Proven Performance
Ampere GPUs power the world's largest AI training clusters. Battle-tested in production at scale for LLM training, fine-tuning, and inference.
Cloud-Native Ready
Available across multiple cloud providers with hourly billing. Compare real-time pricing and pick the cheapest option for your workload.
Ecosystem Support
Full compatibility with PyTorch, TensorFlow, JAX, vLLM, TensorRT, and all major ML frameworks. NVLink for multi-GPU scaling.