Best GPUs for AI Inference
Throughput per dollar, with enough VRAM to hold the weights.
What this workload needs
Recommended GPUs for AI Inference
Ordered by suitability for this workload, not by price.
L40S
#1H100 SXM
#2L4
#3A100 PCIE
#4RTX 4090
#5H200
#6MI300X
#7A10
#8AI Inference GPU Pricing by Provider
| Provider | Price / hr |
|---|---|
$0.160/hr 1× | |
$0.340/hr 1×2×3×4×5×6× | |
$0.420/hr 1× | |
$0.420/hr 8× | |
$0.439/hr 1× | |
$0.440/hr 1× | |
$0.440/hr 8× | |
$0.472/hr 1× | |
$0.490/hr 1×2×3×4×5×6×7× | |
$0.540/hr 1× | |
$0.630/hr 1×4× | |
$0.640/hr 1× | |
$0.645/hr 1×8× | |
$0.657/hr 1× | |
$0.660/hr 1×4× | |
$0.662/hr 1× | |
$0.668/hr 1× | |
$0.685/hr 1×2×4×8× | |
$0.690/hr36mo 1×2×4×8× | |
$0.720/hr 1× | |
$0.740/hr 1×2×3×4×5×6× | |
$0.740/hr 1× | |
$0.765/hr 4× | |
$0.765/hr 2× | |
$0.790/hr 1×2×3×4×5×6×7× | |
$0.790/hr24mo 1×2×4×8× | |
$0.799/hr 1× | |
$0.805/hr 1× | |
$0.840/hr 2×4×8× | |
$0.870/hr 1× | |
$0.871/hr 8× | |
$0.880/hr 1×2×4×8× | |
$0.890/hr 1× | |
$0.890/hr12mo 1×2×4×8× | |
$0.890/hr36mo 1×2×4×8× | |
$0.909/hr 1×2×4×8× | |
$0.957/hr 1× | |
$0.981/hr 1× | |
$0.990/hr 1×2×3×4×5×6× | |
$0.990/hr6mo 1×2×4×8× | |
$0.990/hr24mo 1×2×4×8× | |
$0.998/hr 1×2×4×8× | |
$1.01/hr 1× | |
$1.04/hr 1×2×4×8× | |
$1.09/hr 1×2×4×8× | |
$1.09/hr12mo 1×2×4×8× | |
$1.10/hr 1× | |
$1.15/hr 4× | |
$1.15/hr 4× | |
$1.19/hr 1× | |
$1.19/hr 1×2×3×4×5×6× | |
$1.19/hr6mo 1×2×4×8× | |
$1.20/hr 1× | |
$1.23/hr 1× | |
$1.28/hr 1× | |
$1.29/hr 1× | |
$1.29/hr 1×2×4×8× | |
$1.29/hr 1×8× | |
$1.35/hr 2× | |
$1.35/hr 1× | |
$1.37/hr 1×2×4×8× | |
$1.39/hr 1×2×3×4×5×6× | |
$1.42/hr 4× | |
$1.42/hr 1× | |
$1.45/hr 1×2×4×8×10× | |
$1.50/hr 2× | |
$1.50/hr 1× | |
$1.50/hr 1× | |
$1.55/hr 1× | |
$1.57/hr 1× | |
$1.63/hr 1×2×4×8× | |
$1.65/hr 2× | |
$1.66/hr 1× | |
$1.67/hr 8× | |
$1.70/hr 1×2×4×8× | |
$1.73/hr 1× | |
$1.79/hr 1× | |
$1.79/hr 8× | |
$1.80/hr 2× | |
$1.82/hr 4× | |
$1.86/hr 1× | |
$1.89/hr 1× | |
$1.89/hr 1× | |
$1.95/hr 1× | |
$1.99/hr 1×2×4×8× | |
$1.99/hr 1× | |
$2.00/hr 1× | |
$2.00/hr 1× | |
$2.00/hr 1×2×4×8× | |
$2.04/hr 8× | |
$2.05/hr 1× | |
$2.06/hr 2× | |
$2.06/hr 1× | |
$2.06/hr 4× | |
$2.06/hr 8× | |
$2.10/hr 1×2×4×8× | |
$2.10/hr 1× | |
$2.10/hr 1× | |
$2.18/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 8× | |
$2.20/hr24mo 8× | |
$2.24/hr 4× | |
$2.25/hr 8× | |
$2.30/hr 8× | |
$2.39/hr 1× | |
$2.40/hr12mo 8× | |
$2.44/hr 8× | |
$2.45/hr 1×2×4×8× | |
$2.49/hr36mo 1×8× | |
$2.50/hr 1× | |
$2.51/hr 1× | |
$2.58/hr 8× | |
$2.59/hr 1×8× | |
$2.59/hr24mo 1×8× | |
$2.60/hr 1× | |
$2.62/hr 4× | |
$2.69/hr 1× | |
$2.69/hr 1× | |
$2.69/hr12mo 1×8× | |
$2.71/hr 1× | |
$2.73/hr 1×2×4× | |
$2.79/hr 1× | |
$2.79/hr6mo 1×8× | |
$2.80/hr 8× | |
$2.95/hr 1× | |
$2.99/hr 4× | |
$2.99/hr 1×8× | |
$2.99/hr36mo 1×8× | |
$3.00/hr 1× | |
$3.00/hr 1× | |
$3.09/hr24mo 1×8× | |
$3.14/hr 8× | |
$3.19/hr 1× | |
$3.19/hr12mo 1×8× | |
$3.20/hr 1× | |
$3.25/hr 1×2×4×8× | |
$3.29/hr 1×2×3×4×5×6×7×8× | |
$3.29/hr6mo 1×8× | |
$3.31/hr 1× | |
$3.44/hr 4×8× | |
$3.46/hr 1×2× | |
$3.49/hr 1×8× | |
$3.50/hr 1× | |
$3.50/hr 1× | |
$3.50/hr 1× | |
$3.59/hr 1×2×3×4×5× | |
$3.62/hr 1×2× | |
$3.63/hr 1×2×4×8× | |
$3.63/hr 2× | |
$3.69/hr 1×2×4× | |
$3.77/hr 8× | |
$3.79/hr 1× | |
$3.79/hr 1× | |
$3.82/hr 2× | |
$3.82/hr 4×8× | |
$3.87/hr 8× | |
$3.87/hr1mo 8× | |
$3.95/hr 1× | |
$3.98/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$3.99/hr 8× | |
$3.99/hr 1× | |
$3.99/hr 1× | |
$4.00/hr 1×2×4×8× | |
$4.02/hr 1× | |
$4.08/hr 8× | |
$4.09/hr 4× | |
$4.19/hr 2× | |
$4.27/hr 8× | |
$4.29/hr 1× | |
$4.33/hr 4× | |
$4.33/hr 2× | |
$4.34/hr 4× | |
$4.41/hr 1×8× | |
$4.47/hr 1×8× | |
$4.50/hr 4× | |
$4.52/hr 8× | |
$4.54/hr 1× | |
$4.59/hr 2×3×4×5×6×7×8× | |
$4.63/hr 4× | |
$4.63/hr 1× | |
$4.92/hr 1× | |
$5.40/hr 1×2×4×8× | |
$5.95/hr 8× | |
$5.99/hr 1× | |
$6.00/hr 8× | |
$6.00/hr 1× | |
$6.16/hr 8× | |
$6.30/hr 8× | |
$6.60/hr 1×2×4×8× | |
$6.88/hr 8× | |
$7.91/hr 8× | |
$9.84/hr 1× | |
$10.00/hr 1× | |
$10.60/hr 8× | |
$11.06/hr 8× |
How to choose a GPU for ai inference
Serving a model is a different problem from training one. There is no optimizer state and no gradient traffic, so the memory requirement collapses to the weights plus the KV cache plus activations. The question becomes how much throughput you can extract per dollar per hour, and the answer is usually the smallest GPU your weights comfortably fit on rather than the largest one available.
Token generation is memory-bandwidth bound, not compute bound. Each generated token requires reading the entire set of active weights out of GPU memory, so bandwidth sets the ceiling on single-stream latency. This is why an accelerator with modest FLOPS but fast HBM can outperform a higher-FLOPS card with GDDR for decode, and why batching matters so much: larger batches amortize those weight reads across more tokens and move the workload back toward compute bound.
Quantization changes the hardware calculus more than any other single decision. Running weights at INT8 or FP8 roughly halves the memory footprint and the bandwidth per token versus BF16, often moving a model down a hardware tier with limited quality loss. INT4 halves it again with more measurable degradation. Decide the numeric format first, then size the GPU, not the other way around.
Do not forget the KV cache. At long context lengths and high concurrency it can rival or exceed the weights in size, and it grows linearly with both. A deployment sized only for weights will fit on paper and then fall over under real concurrency — size for your target context length and batch size together.
Related Reading
Frequently Asked Questions
What is the cheapest GPU for AI inference?
The cheapest workable GPU is the smallest one your quantized weights and KV cache fit on with headroom. For models up to about 8B parameters that is often a 24 GB card; 30B–70B models generally need 48–80 GB unless heavily quantized. Compare current hourly rates in the table above — the cheapest per-hour GPU is not always the cheapest per token.
Should I use an inference API instead of renting a GPU?
Per-token APIs avoid idle cost and operational work, which usually wins at low or spiky volume. Dedicated GPUs win at sustained high utilization, when you need a model no API serves, or when data residency rules out third-party APIs. Our inference API pricing page lists per-token rates so you can work out your own crossover point.
How much VRAM does the KV cache need?
KV cache size scales with context length, batch size, layer count and hidden size. For a 7B-class model it is on the order of half a gigabyte per thousand tokens of context per sequence, so long-context serving at high concurrency can require more memory than the weights themselves. Techniques like grouped-query attention and paged attention reduce it substantially.