Best GPUs for RAG Pipelines
An embedding model and a generator, sized separately.
What this workload needs
Recommended GPUs for RAG Pipelines
Ordered by suitability for this workload, not by price.
L4
#1L40S
#2A10
#3RTX 4090
#4A100 PCIE
#5H100 SXM
#6RTX A4000
#7RAG Pipelines GPU Pricing by Provider
| Provider | Price / hr |
|---|---|
$0.086/hr 1× | |
$0.094/hr 2× | |
$0.150/hr 1× | |
$0.160/hr 1× | |
$0.170/hr 1× | |
$0.250/hr 1× | |
$0.340/hr 1×2×3×4×5×6× | |
$0.420/hr 1× | |
$0.420/hr 8× | |
$0.439/hr 1× | |
$0.440/hr 1× | |
$0.440/hr 8× | |
$0.472/hr 1× | |
$0.490/hr 1×2×3×4×5×6×7× | |
$0.540/hr 1× | |
$0.630/hr 1×4× | |
$0.640/hr 1× | |
$0.645/hr 1×8× | |
$0.657/hr 1× | |
$0.660/hr 1×4× | |
$0.662/hr 1× | |
$0.668/hr 1× | |
$0.685/hr 1×2×4×8× | |
$0.690/hr36mo 1×2×4×8× | |
$0.720/hr 1× | |
$0.740/hr 1×2×3×4×5×6× | |
$0.740/hr 1× | |
$0.765/hr 4× | |
$0.765/hr 2× | |
$0.790/hr 1×2×3×4×5×6×7× | |
$0.790/hr24mo 1×2×4×8× | |
$0.799/hr 1× | |
$0.800/hr 1×2×4× | |
$0.805/hr 1× | |
$0.840/hr 2×4×8× | |
$0.870/hr 1× | |
$0.871/hr 8× | |
$0.880/hr 1×2×4×8× | |
$0.880/hr 2×4× | |
$0.890/hr 1× | |
$0.890/hr12mo 1×2×4×8× | |
$0.890/hr36mo 1×2×4×8× | |
$0.909/hr 1×2×4×8× | |
$0.957/hr 1× | |
$0.981/hr 1× | |
$0.990/hr 1×2×3×4×5×6× | |
$0.990/hr6mo 1×2×4×8× | |
$0.990/hr24mo 1×2×4×8× | |
$0.998/hr 1×2×4×8× | |
$1.01/hr 1× | |
$1.04/hr 1×2×4×8× | |
$1.09/hr 1×2×4×8× | |
$1.09/hr12mo 1×2×4×8× | |
$1.10/hr 1× | |
$1.15/hr 4× | |
$1.15/hr 4× | |
$1.19/hr 1× | |
$1.19/hr 1×2×3×4×5×6× | |
$1.19/hr6mo 1×2×4×8× | |
$1.20/hr 1× | |
$1.23/hr 1× | |
$1.28/hr 1× | |
$1.29/hr 1× | |
$1.29/hr 1×2×4×8× | |
$1.29/hr 1×8× | |
$1.35/hr 2× | |
$1.35/hr 1× | |
$1.37/hr 1×2×4×8× | |
$1.39/hr 1×2×3×4×5×6× | |
$1.42/hr 4× | |
$1.42/hr 1× | |
$1.45/hr 1×2×4×8×10× | |
$1.50/hr 1× | |
$1.50/hr 1× | |
$1.55/hr 1× | |
$1.57/hr 1× | |
$1.63/hr 1×2×4×8× | |
$1.65/hr 2× | |
$1.66/hr 1× | |
$1.67/hr 8× | |
$1.70/hr 1×2×4×8× | |
$1.73/hr 1× | |
$1.79/hr 1× | |
$1.79/hr 8× | |
$1.80/hr 2× | |
$1.82/hr 4× | |
$1.86/hr 1× | |
$1.89/hr 1× | |
$1.89/hr 1× | |
$1.95/hr 1× | |
$1.99/hr 1×2×4×8× | |
$2.00/hr 1× | |
$2.00/hr 1× | |
$2.04/hr 8× | |
$2.05/hr 1× | |
$2.06/hr 2× | |
$2.06/hr 1× | |
$2.06/hr 4× | |
$2.06/hr 8× | |
$2.10/hr 1×2×4×8× | |
$2.18/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 8× | |
$2.20/hr24mo 8× | |
$2.24/hr 4× | |
$2.25/hr 8× | |
$2.30/hr 8× | |
$2.40/hr12mo 8× | |
$2.44/hr 8× | |
$2.45/hr 1×2×4×8× | |
$2.49/hr36mo 1×8× | |
$2.50/hr 1× | |
$2.51/hr 1× | |
$2.59/hr24mo 1×8× | |
$2.62/hr 4× | |
$2.69/hr 1× | |
$2.69/hr 1× | |
$2.69/hr12mo 1×8× | |
$2.73/hr 1×2×4× | |
$2.79/hr 1× | |
$2.79/hr6mo 1×8× | |
$2.95/hr 1× | |
$2.99/hr 1×8× | |
$3.14/hr 8× | |
$3.19/hr 1× | |
$3.20/hr 1× | |
$3.25/hr 1×2×4×8× | |
$3.29/hr 1×2×3×4×5×6×7×8× | |
$3.31/hr 1× | |
$3.46/hr 1×2× | |
$3.50/hr 1× | |
$3.50/hr 1× | |
$3.63/hr 1×2×4×8× | |
$3.63/hr 2× | |
$3.77/hr 8× | |
$3.82/hr 2× | |
$3.87/hr 8× | |
$3.95/hr 1× | |
$3.98/hr 1× | |
$3.99/hr 8× | |
$4.08/hr 8× | |
$4.09/hr 4× | |
$4.19/hr 2× | |
$4.29/hr 1× | |
$4.41/hr 1×8× | |
$4.50/hr 4× | |
$5.40/hr 1×2×4×8× | |
$5.95/hr 8× | |
$5.99/hr 1× | |
$6.16/hr 8× | |
$6.88/hr 8× | |
$10.00/hr 1× | |
$11.06/hr 8× |
How to choose a GPU for rag pipelines
A retrieval-augmented generation pipeline is not one model but at least two, and they have different hardware profiles. The embedding model is small, runs over large batches during ingestion, and is throughput bound. The generator is the large language model that consumes retrieved context, and it is latency and memory bound. Sizing one GPU for both usually overprovisions for one job and underprovisions for the other.
Embedding is the easier half. Common embedding models are a few hundred million parameters and fit in a handful of gigabytes, so an entry or mid-tier GPU saturates on batch size long before it runs out of memory. Ingestion is also a batch job rather than a serving path, which makes it a good candidate for spot capacity and for scaling to zero between corpus updates.
Generation is where the cost sits, and RAG makes it harder than plain chat in one specific way: retrieved context makes prompts long. Prefill cost scales with prompt length, and the KV cache grows with it too, so a RAG deployment needs materially more memory headroom per concurrent request than the same model serving short prompts. Size for your retrieval budget — top-k times chunk size — not for the question alone.
Reranking sits between the two and is often forgotten. A cross-encoder reranker runs over every retrieved candidate rather than once per query, so its throughput requirement scales with your candidate count. If your pipeline reranks, budget for it as a third model rather than as a rounding error on the generator.
Related Reading
Frequently Asked Questions
Do I need a GPU for embeddings?
For one-off ingestion of a small corpus, CPU is often adequate. For large corpora or frequent re-indexing, a GPU is typically an order of magnitude faster per document and cheaper overall despite the higher hourly rate. Entry and mid-tier GPUs are usually sufficient, since embedding models are small.
Should the embedding model and LLM share a GPU?
They can, and co-locating avoids a second instance. But the two have different scaling profiles — embedding is throughput bound and bursty, generation is latency bound and continuous — so at production volume separate GPUs usually give better utilization and more predictable serving latency.
How does long context affect GPU requirements for RAG?
Retrieved context lengthens every prompt, which increases both prefill compute and KV cache size per request. A RAG deployment therefore needs more VRAM headroom per concurrent user than the same model handling short prompts. Size the KV cache for top-k times chunk size, plus the answer.