Budget inference (2018-2020)

NVIDIA
Turing

NVIDIA Turing GPUs offer the lowest cost of entry for AI inference. With GDDR6 memory and low power consumption (70W T4), they're ideal for budget inference workloads and edge computing.

GPU Models in this Family

Click any card to expand detailed specifications

T4

T4

16 GB GDDR6320 GB/sPCIe Gen3 (single-slot)

Low-cost, low-power inference GPU. 70W TDP — no extra power connector needed. Most affordable AI GPU.

AI inferenceVideo processingEdge computingBudget workloads
Architecture
Turing
VRAM
16 GB GDDR6
Memory Bandwidth
320 GB/s
Interconnect
PCIe Gen3 (32 GB/s)
FP16 Performance
65 TFLOPS (with Tensor Cores)
FP32 Performance
8.1 TFLOPS
FP64 Performance
0.25 TFLOPS
TDP
70W
CUDA Cores
2,560
Tensor Cores
320
Process Node
TSMC 12N
Form Factor
PCIe Gen3 (single-slot)
Quadro RTX 8000

Quadro RTX 8000

48 GB GDDR6672 GB/sPCIe Gen3

High-end professional workstation GPU (Turing). 48GB GDDR6 — the largest VRAM in the Turing workstation lineup.

Workstation AIRenderingGraphicsLarge datasets
Architecture
Turing
VRAM
48 GB GDDR6
Memory Bandwidth
672 GB/s
Interconnect
PCIe Gen3 (32 GB/s)
FP16 Performance
131 TFLOPS (with Tensor Cores)
FP32 Performance
16.3 TFLOPS
FP64 Performance
0.51 TFLOPS
TDP
260W
CUDA Cores
4,608
Tensor Cores
576
Process Node
TSMC 12N
Form Factor
PCIe Gen3
Quadro RTX 6000

Quadro RTX 6000

24 GB GDDR6672 GB/sPCIe Gen3

Professional workstation GPU (Turing). 24GB GDDR6 with ECC for rendering, AI, and graphics workloads.

Workstation AIRenderingGraphicsVideo processing
Architecture
Turing
VRAM
24 GB GDDR6
Memory Bandwidth
672 GB/s
Interconnect
PCIe Gen3 (32 GB/s)
FP16 Performance
131 TFLOPS (with Tensor Cores)
FP32 Performance
16.3 TFLOPS
FP64 Performance
0.51 TFLOPS
TDP
260W
CUDA Cores
4,608
Tensor Cores
576
Process Node
TSMC 12N
Form Factor
PCIe Gen3

Quick Comparison

GPU ModelVRAMBandwidthFP16TDPStarting Price
T416 GB GDDR6320 GB/s65 TFLOPS (with Tensor Cores)70W
Quadro RTX 800048 GB GDDR6672 GB/s131 TFLOPS (with Tensor Cores)260W
Quadro RTX 600024 GB GDDR6672 GB/s131 TFLOPS (with Tensor Cores)260W

Why Choose Turing?

Proven Performance

Turing GPUs power the world's largest AI training clusters. Battle-tested in production at scale for LLM training, fine-tuning, and inference.

Cloud-Native Ready

Available across multiple cloud providers with hourly billing. Compare real-time pricing and pick the cheapest option for your workload.

Ecosystem Support

Full compatibility with PyTorch, TensorFlow, JAX, vLLM, TensorRT, and all major ML frameworks. NVLink for multi-GPU scaling.