Train models with hundreds of billions of parameters. Serve 70B+ models at full precision. Run multi-node distributed training with NVLink interconnects. This tier — A100 80 GB, H100, H200, MI300X, and Blackwell-generation GPUs — uses high-bandwidth memory (HBM2e/HBM3/HBM3e) that delivers 2–5x the bandwidth of GDDR, directly impacting training throughput and large-batch inference speed.
| Provider | Price / hr |
|---|---|
$1.57/hr 1× | |
$1.59/hr12mo 1× | |
$1.73/hr 4× | |
$2.30/hr 4× | |
$2.63/hr 4× | |
$2.74/hr 1× | |
$3.75/hr 1× | |
$4.54/hr 1× | |
$6.89/hr 2× |
Showing 9 of 642 price points. Visit individual GPU pages above for full pricing.
If you're training models larger than 13B parameters at full precision, running inference on 70B+ models, or need large batch sizes for production throughput, 80 GB+ is recommended. For smaller workloads, 24–48 GB GPUs are more cost-effective.
HBM (High Bandwidth Memory) provides 2–5x the bandwidth of GDDR6X, which directly impacts training throughput and large-model inference speed. HBM is standard on datacenter GPUs (A100, H100, MI300X) while consumer GPUs use GDDR.