Best GPUs for LLM Training
Multi-GPU nodes with HBM and high-bandwidth interconnects.
What this workload needs
Recommended GPUs for LLM Training
Ordered by suitability for this workload, not by price.
B200
#1GB300
#2H200
#3H100 SXM
#4GB200
#5A100 SXM
#6MI300X
#7LLM Training GPU Pricing by Provider
| Provider | Price / hr |
|---|---|
$0.580/hr 1× | |
$0.890/hr12mo 1×4×8× | |
$0.895/hr 1×2×4×8× | |
$0.990/hr6mo 1×4×8× | |
$1.09/hr 1× | |
$1.14/hr 1× | |
$1.15/hr 8× | |
$1.19/hr 8× | |
$1.24/hr 8× | |
$1.29/hr 2×8× | |
$1.30/hr 1× | |
$1.31/hr 1× | |
$1.35/hr 2×4× | |
$1.38/hr 1×8× | |
$1.39/hr 1×2×3×4×5×6×7×8× | |
$1.39/hr36mo 1×2×4×8× | |
$1.40/hr 1× | |
$1.45/hr 1×2×4× | |
$1.47/hr 1× | |
$1.49/hr 1× | |
$1.49/hr24mo 1×2×4×8× | |
$1.50/hr 2× | |
$1.54/hr 1×8× | |
$1.59/hr 1× | |
$1.59/hr 1×2×3×4×5×6×7×8× | |
$1.59/hr12mo 2× | |
$1.60/hr 1×2×4×8× | |
$1.63/hr 1×2×4×8× | |
$1.65/hr 8× | |
$1.66/hr 1× | |
$1.69/hr6mo 2× | |
$1.79/hr 1×2×4×8× | |
$1.79/hr 1× | |
$1.79/hr 1×2×4×8× | |
$1.79/hr 8× | |
$1.81/hr 1×2×4× | |
$1.81/hr 2×4× | |
$1.89/hr 1× | |
$1.89/hr 1× | |
$1.99/hr 1×2×4×8× | |
$1.99/hr 1× | |
$1.99/hr 1×2×4× | |
$2.00/hr 1× | |
$2.00/hr 1×2×4×8× | |
$2.05/hr 1× | |
$2.06/hr 2× | |
$2.06/hr 1× | |
$2.06/hr 4× | |
$2.06/hr 8× | |
$2.10/hr 1× | |
$2.10/hr 1× | |
$2.10/hr 1× | |
$2.18/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 8× | |
$2.20/hr24mo 8× | |
$2.24/hr 4× | |
$2.30/hr 8× | |
$2.39/hr 1× | |
$2.40/hr12mo 8× | |
$2.44/hr 8× | |
$2.45/hr 1×2×4×8× | |
$2.49/hr36mo 1×8× | |
$2.50/hr 1× | |
$2.51/hr 1× | |
$2.58/hr 8× | |
$2.59/hr 1×8× | |
$2.59/hr24mo 1×8× | |
$2.59/hr 1×2×4×8× | |
$2.60/hr 1× | |
$2.60/hr 1× | |
$2.69/hr 1× | |
$2.69/hr 1× | |
$2.69/hr12mo 1×8× | |
$2.70/hr 8× | |
$2.71/hr 1× | |
$2.73/hr 1×2×4× | |
$2.74/hr 8× | |
$2.79/hr 1× | |
$2.79/hr 8× | |
$2.79/hr6mo 1×8× | |
$2.80/hr 8× | |
$2.80/hr36mo 8× | |
$2.95/hr 1× | |
$2.99/hr 4× | |
$2.99/hr 1×8× | |
$2.99/hr36mo 1×8× | |
$3.00/hr 1× | |
$3.00/hr 1× | |
$3.09/hr24mo 1×8× | |
$3.11/hr 4×8× | |
$3.12/hr 4× | |
$3.12/hr 1×2× | |
$3.12/hr 8× | |
$3.13/hr 2× | |
$3.14/hr 8× | |
$3.19/hr 1× | |
$3.19/hr12mo 1×8× | |
$3.20/hr 1× | |
$3.25/hr 1×2×4×8× | |
$3.28/hr 1× | |
$3.29/hr 1×2×3×4×5×6×7×8× | |
$3.29/hr6mo 1×8× | |
$3.31/hr 1× | |
$3.44/hr 4×8× | |
$3.46/hr 1×2× | |
$3.49/hr 1×8× | |
$3.49/hr 1× | |
$3.49/hr 1× | |
$3.50/hr 1× | |
$3.50/hr 1× | |
$3.59/hr 1×2×3×4×5× | |
$3.62/hr 1×2× | |
$3.63/hr 1×2×4×8× | |
$3.63/hr 2× | |
$3.67/hr 1×2×4× | |
$3.69/hr 1× | |
$3.69/hr 1×2×4× | |
$3.75/hr 8× | |
$3.75/hr 1× | |
$3.79/hr 1× | |
$3.79/hr36mo 8× | |
$3.79/hr 1× | |
$3.82/hr 2× | |
$3.82/hr 4×8× | |
$3.87/hr 8× | |
$3.87/hr1mo 8× | |
$3.89/hr 2×4×8× | |
$3.90/hr 1× | |
$3.93/hr 1× | |
$3.95/hr 1× | |
$3.95/hr 1× | |
$3.98/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 8× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$3.99/hr24mo 8× | |
$4.00/hr 1× | |
$4.00/hr 1× | |
$4.00/hr 1×2×4×8× | |
$4.02/hr 1× | |
$4.08/hr 8× | |
$4.09/hr 4× | |
$4.10/hr 8× | |
$4.19/hr 2× | |
$4.26/hr 8× | |
$4.27/hr 8× | |
$4.29/hr 1× | |
$4.31/hr 1×2×4× | |
$4.33/hr 4× | |
$4.33/hr 2× | |
$4.34/hr 4× | |
$4.41/hr 1×8× | |
$4.47/hr 1×8× | |
$4.49/hr12mo 8× | |
$4.50/hr 4× | |
$4.52/hr 8× | |
$4.54/hr 1× | |
$4.59/hr 2×3×4×5×6×7×8× | |
$4.63/hr 4× | |
$4.63/hr 1× | |
$4.71/hr 8× | |
$4.92/hr 1× | |
$5.10/hr 8× | |
$5.18/hr 1× | |
$5.40/hr 1×2×4×8× | |
$5.95/hr 8× | |
$5.98/hr 1× | |
$5.99/hr 1× | |
$5.99/hr 1× | |
$6.00/hr 8× | |
$6.00/hr 1× | |
$6.00/hr 1× | |
$6.16/hr 8× | |
$6.23/hr 2× | |
$6.23/hr 1× | |
$6.23/hr 4×8× | |
$6.25/hr 1× | |
$6.29/hr 1× | |
$6.30/hr 8× | |
$6.60/hr 1×2×4×8× | |
$6.69/hr 8× | |
$6.79/hr 1×2× | |
$6.79/hr 4× | |
$6.88/hr 8× | |
$6.89/hr 2× | |
$6.99/hr 1× | |
$7.91/hr 8× | |
$8.00/hr 1× | |
$8.60/hr 8× | |
$8.62/hr 1×2×4× | |
$9.00/hr 1×2×4×8× | |
$9.84/hr 1× | |
$10.00/hr 1× | |
$10.60/hr 8× | |
$11.06/hr 8× | |
$14.00/hr 1× | |
$16.00/hr 1× | |
$18.00/hr 1× |
How to choose a GPU for llm training
Training a large language model from scratch is bound by two things before it is bound by raw compute: how much of the model, its gradients and its optimizer state fit in GPU memory, and how fast gradients move between GPUs. That is why this workload lives almost exclusively on HBM-equipped datacenter parts with NVLink or Infinity Fabric interconnects, rather than on consumer cards with more attractive headline FLOPS.
A practical floor is 80 GB per GPU. Optimizer state for a model trained in mixed precision typically costs several times the parameter count in bytes, so a 7B model in a standard Adam setup already needs well over 80 GB across the node before activations are counted. Sharded training (FSDP, DeepSpeed ZeRO, tensor and pipeline parallelism) spreads that across GPUs, which makes inter-GPU bandwidth the next constraint — NVLink between GPUs in a node, and InfiniBand or equivalent between nodes.
Numeric format matters as much as capacity. Hopper introduced FP8 via the Transformer Engine and Blackwell extends it, so a training run that can use FP8 for the bulk of its matmuls sees a large step up over BF16-only hardware. Check whether your framework and model actually support the format before paying for it: an FP8-capable GPU running a BF16-only stack is just an expensive BF16 GPU.
When comparing providers, look past the hourly rate to what you are actually renting. A single-GPU instance and an 8-GPU node with NVLink are not substitutes for this workload, and interconnect quality between nodes varies widely between providers. The table below normalizes every price to per-GPU per-hour so the comparison is like-for-like; multi-GPU offerings are labelled with their GPU count.
Related Reading
Claude's March 2026 Double Usage Promotion: What AI Teams Should Know
Anthropic is doubling Claude usage limits during off-peak hours through March 27. Here's how AI and ML teams can take advantage of the promotion across Claude, Claude Code, and more.
H200 vs B200 vs GB300: Which NVIDIA GPU Should You Rent for AI?
NVIDIA now has three generations of data center GPUs available in the cloud at once. Here's how H200, B200, and the Blackwell Ultra lineup (HGX B300 and GB300) actually differ, and which one fits your workload.
5 Ways to Cut AI Training Costs with Smart GPU Choice
Explore effective strategies to significantly reduce AI training costs by optimizing GPU choices and usage without compromising performance.
Frequently Asked Questions
How much VRAM do I need to train an LLM?
For full fine-tuning or pretraining in mixed precision, budget roughly 16–20 bytes per parameter across the whole node once weights, gradients and Adam optimizer state are counted. A 7B model therefore needs well over 100 GB in total, which in practice means several 80 GB GPUs with sharding. Parameter-efficient methods like LoRA cut this dramatically — see the fine-tuning page.
Do I need NVLink for LLM training?
For single-GPU training, no. As soon as the model is sharded across GPUs, gradient all-reduce traffic becomes a bottleneck and NVLink (or AMD Infinity Fabric) makes a substantial difference to step time. PCIe-only multi-GPU nodes are usable for data-parallel training of small models but scale poorly for tensor or pipeline parallelism.
Is renting cheaper than buying for training runs?
It depends on utilization. Rental makes sense for bursty runs, for evaluating hardware before committing, and for scaling beyond what you own. Sustained, near-continuous training over a period of years is where owned hardware starts to compete, once power, cooling and operations are included. Reserved and committed-use rates sit between the two — compare current pricing above.