Skip to main content
Use case

Best GPUs for RAG Pipelines

An embedding model and a generator, sized separately.

GPUs 7
Providers 44
From $0.086/hr

What this workload needs

Practical VRAM floor
16 GB
What to compare on
VRAM capacity · Memory bandwidth · Batch throughput · Cost per hour

Recommended GPUs for RAG Pipelines

Ordered by suitability for this workload, not by price.

RAG Pipelines GPU Pricing by Provider

ProviderPrice / hr
$0.086/hr
1×
$0.094/hr
2×
$0.150/hr
1×
$0.160/hr
1×
Runpod logo
RunpodCommunity Cloud
$0.170/hr
1×
Runpod logo
RunpodSecure Cloud
$0.250/hr
1×
Runpod logo
RunpodCommunity Cloud
$0.340/hr
1×2×3×4×5×6×
$0.420/hr
1×
$0.420/hr
8×
$0.439/hr
1×
$0.440/hr
1×
$0.440/hr
8×
$0.472/hr
1×
Runpod logo
RunpodSecure Cloud
$0.490/hr
1×2×3×4×5×6×7×
$0.540/hr
1×
$0.630/hr
1×4×
$0.640/hr
1×
$0.645/hr
1×8×
$0.657/hr
1×
$0.660/hr
1×4×
$0.662/hr
1×
$0.668/hr
1×
$0.685/hr
1×2×4×8×
$0.690/hr36mo
1×2×4×8×
$0.720/hr
1×
Runpod logo
RunpodSecure Cloud
$0.740/hr
1×2×3×4×5×6×
$0.740/hr
1×
$0.765/hr
4×
$0.765/hr
2×
Runpod logo
RunpodCommunity Cloud
$0.790/hr
1×2×3×4×5×6×7×
$0.790/hr24mo
1×2×4×8×
$0.799/hr
1×
$0.800/hr
1×2×4×
Amazon AWS logo
Amazon AWSus-east-1
$0.805/hr
1×
$0.840/hr
2×4×8×
$0.870/hr
1×
$0.871/hr
8×
$0.880/hr
1×2×4×8×
$0.880/hr
2×4×
$0.890/hr
1×
$0.890/hr12mo
1×2×4×8×
$0.890/hr36mo
1×2×4×8×
$0.909/hr
1×2×4×8×
$0.957/hr
1×
$0.981/hr
1×
Runpod logo
RunpodSecure Cloud
$0.990/hr
1×2×3×4×5×6×
$0.990/hr6mo
1×2×4×8×
$0.990/hr24mo
1×2×4×8×
$0.998/hr
1×2×4×8×
Amazon AWS logo
Amazon AWSus-east-1
$1.01/hr
1×
$1.04/hr
1×2×4×8×
$1.09/hr
1×2×4×8×
$1.09/hr12mo
1×2×4×8×
$1.10/hr
1×
$1.15/hr
4×
Amazon AWS logo
Amazon AWSus-east-1
$1.15/hr
4×
$1.19/hr
1×
Runpod logo
RunpodCommunity Cloud
$1.19/hr
1×2×3×4×5×6×
$1.19/hr6mo
1×2×4×8×
$1.20/hr
1×
$1.23/hr
1×
$1.28/hr
1×
$1.29/hr
1×
$1.29/hr
1×2×4×8×
$1.29/hr
1×8×
$1.35/hr
2×
$1.35/hr
1×
$1.37/hr
1×2×4×8×
Runpod logo
RunpodSecure Cloud
$1.39/hr
1×2×3×4×5×6×
Amazon AWS logo
Amazon AWSus-east-1
$1.42/hr
4×
$1.42/hr
1×
$1.45/hr
1×2×4×8×10×
$1.50/hr
1×
$1.50/hr
1×
$1.55/hr
1×
$1.57/hr
1×
$1.63/hr
1×2×4×8×
$1.65/hr
2×
$1.66/hr
1×
Amazon AWS logo
Amazon AWSus-east-1
$1.67/hr
8×
$1.70/hr
1×2×4×8×
$1.73/hr
1×
$1.79/hr
1×
$1.79/hr
8×
$1.80/hr
2×
$1.82/hr
4×
Amazon AWS logo
Amazon AWSus-east-1
$1.86/hr
1×
$1.89/hr
1×
$1.89/hr
1×
$1.95/hr
1×
$1.99/hr
1×2×4×8×
$2.00/hr
1×
$2.00/hr
1×
Amazon AWS logo
Amazon AWSus-east-1
$2.04/hr
8×
$2.05/hr
1×
$2.06/hr
2×
$2.06/hr
1×
$2.06/hr
4×
$2.06/hr
8×
$2.10/hr
1×2×4×8×
$2.18/hr
1×
$2.19/hr
1×
$2.19/hr
1×
$2.19/hr
8×
$2.20/hr24mo
8×
$2.24/hr
4×
CoreWeave logo
CoreWeaveEUROPE
$2.25/hr
8×
$2.30/hr
8×
$2.40/hr12mo
8×
CoreWeave logo
CoreWeaveEUROPE
$2.44/hr
8×
$2.45/hr
1×2×4×8×
$2.49/hr36mo
1×8×
$2.50/hr
1×
$2.51/hr
1×
$2.59/hr24mo
1×8×
Amazon AWS logo
Amazon AWSus-east-1
$2.62/hr
4×
Runpod logo
RunpodCommunity Cloud
$2.69/hr
1×
$2.69/hr
1×
$2.69/hr12mo
1×8×
$2.73/hr
1×2×4×
$2.79/hr
1×
$2.79/hr6mo
1×8×
$2.95/hr
1×
$2.99/hr
1×8×
$3.14/hr
8×
$3.19/hr
1×
$3.20/hr
1×
$3.25/hr
1×2×4×8×
Runpod logo
RunpodSecure Cloud
$3.29/hr
1×2×3×4×5×6×7×8×
$3.31/hr
1×
$3.46/hr
1×2×
$3.50/hr
1×
$3.50/hr
1×
$3.63/hr
1×2×4×8×
$3.63/hr
2×
Amazon AWS logo
Amazon AWSus-east-1
$3.77/hr
8×
$3.82/hr
2×
$3.87/hr
8×
$3.95/hr
1×
$3.98/hr
1×
$3.99/hr
8×
$4.08/hr
8×
$4.09/hr
4×
$4.19/hr
2×
$4.29/hr
1×
$4.41/hr
1×8×
$4.50/hr
4×
$5.40/hr
1×2×4×8×
$5.95/hr
8×
$5.99/hr
1×
CoreWeave logo
CoreWeaveEUROPE
$6.16/hr
8×
Amazon AWS logo
Amazon AWSus-east-1
$6.88/hr
8×
$10.00/hr
1×
$11.06/hr
8×
Direct from providerVia marketplace

How to choose a GPU for rag pipelines

A retrieval-augmented generation pipeline is not one model but at least two, and they have different hardware profiles. The embedding model is small, runs over large batches during ingestion, and is throughput bound. The generator is the large language model that consumes retrieved context, and it is latency and memory bound. Sizing one GPU for both usually overprovisions for one job and underprovisions for the other.

Embedding is the easier half. Common embedding models are a few hundred million parameters and fit in a handful of gigabytes, so an entry or mid-tier GPU saturates on batch size long before it runs out of memory. Ingestion is also a batch job rather than a serving path, which makes it a good candidate for spot capacity and for scaling to zero between corpus updates.

Generation is where the cost sits, and RAG makes it harder than plain chat in one specific way: retrieved context makes prompts long. Prefill cost scales with prompt length, and the KV cache grows with it too, so a RAG deployment needs materially more memory headroom per concurrent request than the same model serving short prompts. Size for your retrieval budget — top-k times chunk size — not for the question alone.

Reranking sits between the two and is often forgotten. A cross-encoder reranker runs over every retrieved candidate rather than once per query, so its throughput requirement scales with your candidate count. If your pipeline reranks, budget for it as a third model rather than as a rounding error on the generator.

Related Reading

Frequently Asked Questions

Do I need a GPU for embeddings?

For one-off ingestion of a small corpus, CPU is often adequate. For large corpora or frequent re-indexing, a GPU is typically an order of magnitude faster per document and cheaper overall despite the higher hourly rate. Entry and mid-tier GPUs are usually sufficient, since embedding models are small.

Should the embedding model and LLM share a GPU?

They can, and co-locating avoids a second instance. But the two have different scaling profiles — embedding is throughput bound and bursty, generation is latency bound and continuous — so at production volume separate GPUs usually give better utilization and more predictable serving latency.

How does long context affect GPU requirements for RAG?

Retrieved context lengthens every prompt, which increases both prefill compute and KV cache size per request. A RAG deployment therefore needs more VRAM headroom per concurrent user than the same model handling short prompts. Size the KV cache for top-k times chunk size, plus the answer.

Related Use Cases

Browse Related Categories