An embedding model and a generator, sized separately.
Ordered by suitability for this workload, not by price.
| Provider | Price / hr |
|---|---|
$0.060/hr 1× | |
$0.070/hr 2× | |
$0.110/hr 1× | |
$0.130/hr 1× | |
$0.150/hr 1× | |
$0.152/hr 1× | |
$0.160/hr 1× | |
$0.170/hr 1× | |
$0.200/hr 1× | |
$0.230/hr12mo 4× | |
$0.230/hr24mo 4× | |
$0.250/hr 1×2× | |
$0.252/hr 1× | |
$0.260/hr 4× | |
$0.300/hr 1×2×8× | |
$0.318/hr36mo 1× | |
$0.335/hr 1× | |
$0.340/hr 1×2×3×4×5×6× | |
$0.353/hr 1× | |
$0.380/hr 1×2×4× | |
$0.380/hr12mo 8× | |
$0.380/hr24mo 8× | |
$0.400/hr 8× | |
$0.420/hr 8× | |
$0.430/hr 8× | |
$0.440/hr 1× | |
$0.440/hr 8× | |
$0.440/hr 1×8× | |
$0.442/hr 1× | |
$0.442/hr12mo 1× | |
$0.450/hr 1× | |
$0.450/hr36mo 2×4×8× | |
$0.490/hr 1×2×3×4×5× | |
$0.497/hr 1× | |
$0.500/hr 1× | |
$0.530/hr 1× | |
$0.540/hr 1× | |
$0.540/hr 1× | |
$0.550/hr 1× | |
$0.600/hr 1×4× | |
$0.600/hr 4×8× | |
$0.620/hr12mo 8× | |
$0.620/hr24mo 8× | |
$0.630/hr 1×4× | |
$0.630/hr12mo 2×4×8× | |
$0.640/hr 1× | |
$0.645/hr 1×8× | |
$0.650/hr 1× | |
$0.660/hr 1×4× | |
$0.662/hr 1× | |
$0.664/hr 1× | |
$0.670/hr 1× | |
$0.685/hr 1×2×4×8× | |
$0.690/hr 1× | |
$0.690/hr36mo 1×2×4×8× | |
$0.700/hr 8× | |
$0.740/hr 1×2×3×4×5×6×7× | |
$0.760/hr 1× | |
$0.765/hr 4× | |
$0.765/hr 2× | |
$0.778/hr 0.5× | |
$0.790/hr 1×2×3×4×5× | |
$0.790/hr24mo 1×2×4×8× | |
$0.799/hr 1× | |
$0.800/hr 1×2×4× | |
$0.805/hr 1× | |
$0.840/hr 4×8× | |
$0.870/hr 1× | |
$0.880/hr 1×2×4×8× | |
$0.880/hr 2×4× | |
$0.880/hr 1× | |
$0.890/hr 1× | |
$0.890/hr12mo 1×2×4×8× | |
$0.890/hr36mo 1×2×4×8× | |
$0.915/hr 1×2×4×8× | |
$0.925/hr 2×4× | |
$0.930/hr 1× | |
$0.950/hr 1× | |
$0.950/hr 1×2×4×8× | |
$0.960/hr 1× | |
$0.966/hr 1× | |
$0.970/hr 1×2×4×8× | |
$0.970/hr 1× | |
$0.985/hr 8× | |
$0.988/hr 1× | |
$0.990/hr6mo 1×2×4×8× | |
$0.990/hr24mo 1×2×4×8× | |
$1.00/hr 2×4×8× | |
$1.01/hr 1× | |
$1.04/hr 1×2×4×8× | |
$1.05/hr 8× | |
$1.05/hr 1× | |
$1.06/hr 4× | |
$1.07/hr 2× | |
$1.08/hr 1× | |
$1.09/hr 1×2×3×4×5×6×7× | |
$1.09/hr 1×2×4×8× | |
$1.09/hr12mo 1×2×4×8× | |
$1.10/hr 1× | |
$1.11/hr 1× | |
$1.15/hr 4× | |
$1.15/hr 1× | |
$1.15/hr 4× | |
$1.19/hr 1×2×3×4×5× | |
$1.19/hr6mo 1×2×4×8× | |
$1.20/hr 1× | |
$1.20/hr 1× | |
$1.23/hr 1× | |
$1.26/hr 1×2×4× | |
$1.27/hr 8× | |
$1.29/hr 1× | |
$1.29/hr 1× | |
$1.29/hr 1×2×4×8× | |
$1.29/hr 1×8× | |
$1.35/hr 1× | |
$1.37/hr 2× | |
$1.37/hr 1×2×4×8× | |
$1.42/hr 4× | |
$1.42/hr 1× | |
$1.42/hr 1× | |
$1.45/hr 2× | |
$1.47/hr 1× | |
$1.47/hr 8× | |
$1.50/hr 1× | |
$1.50/hr 1× | |
$1.50/hr 1× | |
$1.54/hr 1× | |
$1.57/hr 1× | |
$1.59/hr 1×2×3×4× | |
$1.60/hr 8× | |
$1.60/hr 1×2×4×8× | |
$1.63/hr 1×2×4×8× | |
$1.65/hr 2× | |
$1.65/hr 1× | |
$1.67/hr 0.5× | |
$1.67/hr 8× | |
$1.67/hr 4× | |
$1.67/hr 0.25× | |
$1.68/hr 1× | |
$1.69/hr 1× | |
$1.70/hr 1×2× | |
$1.70/hr 1× | |
$1.71/hr 1×2×4×8× | |
$1.75/hr 2× | |
$1.78/hr 1×2×4×8× | |
$1.79/hr 1× | |
$1.86/hr 1× | |
$1.93/hr 1× | |
$1.95/hr 1× | |
$1.95/hr 2× | |
$1.95/hr 1× | |
$1.96/hr 1× | |
$1.99/hr 1×2×4×8× | |
$1.99/hr 1× | |
$2.00/hr 1× | |
$2.00/hr 1× | |
$2.00/hr 1× | |
$2.04/hr 8× | |
$2.04/hr 1× | |
$2.07/hr 1× | |
$2.08/hr 1× | |
$2.09/hr 1× | |
$2.10/hr 1× | |
$2.10/hr 1×2×4×8× | |
$2.15/hr 1× | |
$2.15/hr 1× | |
$2.19/hr 1× | |
$2.19/hr 8× | |
$2.19/hr 1× | |
$2.20/hr 1× | |
$2.20/hr 1× | |
$2.20/hr12mo 8× | |
$2.20/hr24mo 8× | |
$2.25/hr 8× | |
$2.30/hr 8× | |
$2.40/hr 8× | |
$2.45/hr 8× | |
$2.45/hr 1×2×4×8× | |
$2.46/hr 8× | |
$2.49/hr36mo 1×8× | |
$2.50/hr 1×2×4×8× | |
$2.50/hr 1×2× | |
$2.51/hr 1× | |
$2.59/hr24mo 1×8× | |
$2.60/hr 2×4× | |
$2.62/hr 4× | |
$2.63/hr 2×4× | |
$2.64/hr 8× | |
$2.69/hr 1× | |
$2.69/hr12mo 1×8× | |
$2.72/hr 1× | |
$2.73/hr 4× | |
$2.73/hr 1×2× | |
$2.73/hr 1× | |
$2.73/hr 2×4× | |
$2.74/hr 1× | |
$2.75/hr 4× | |
$2.75/hr 8× | |
$2.75/hr 2× | |
$2.79/hr6mo 1×8× | |
$2.89/hr 1× | |
$2.95/hr 1× | |
$2.98/hr 1× | |
$2.99/hr 1×8× | |
$3.00/hr 1×2×4×8× | |
$3.00/hr 2× | |
$3.14/hr 8× | |
$3.19/hr 1× | |
$3.19/hr 1× | |
$3.20/hr 1× | |
$3.20/hr 8× | |
$3.25/hr 1×2×4×8× | |
$3.29/hr 1× | |
$3.30/hr 1× | |
$3.33/hr 1×2× | |
$3.36/hr 8× | |
$3.39/hr 1× | |
$3.48/hr 8× | |
$3.49/hr 1×2×3×4×5×6×7×8× | |
$3.50/hr 1× | |
$3.50/hr 1× | |
$3.59/hr 1× | |
$3.63/hr 1×2×4×8× | |
$3.77/hr 8× | |
$3.80/hr 1× | |
$3.85/hr 1× | |
$3.90/hr 1× | |
$3.95/hr 1× | |
$3.98/hr 1× | |
$4.09/hr 4× | |
$4.09/hr 4× | |
$4.19/hr 2× | |
$4.19/hr 2× | |
$4.29/hr 1× | |
$4.39/hr 8× | |
$4.41/hr 1×8× | |
$4.50/hr 4× | |
$4.86/hr36mo 8× | |
$5.40/hr 1×2×4×8× | |
$5.95/hr 1×8× | |
$5.99/hr 1× | |
$6.16/hr 8× | |
$6.64/hr 8× | |
$6.88/hr 8× | |
$7.23/hr 1× | |
$7.67/hr12mo 8× | |
$10.00/hr 1× | |
$11.06/hr 8× | |
$11.06/hr 8× | |
$0.550/hr 1× | |
$2.69/hr 1× | |
$3.99/hr 8× |
A retrieval-augmented generation pipeline is not one model but at least two, and they have different hardware profiles. The embedding model is small, runs over large batches during ingestion, and is throughput bound. The generator is the large language model that consumes retrieved context, and it is latency and memory bound. Sizing one GPU for both usually overprovisions for one job and underprovisions for the other.
Embedding is the easier half. Common embedding models are a few hundred million parameters and fit in a handful of gigabytes, so an entry or mid-tier GPU saturates on batch size long before it runs out of memory. Ingestion is also a batch job rather than a serving path, which makes it a good candidate for spot capacity and for scaling to zero between corpus updates.
Generation is where the cost sits, and RAG makes it harder than plain chat in one specific way: retrieved context makes prompts long. Prefill cost scales with prompt length, and the KV cache grows with it too, so a RAG deployment needs materially more memory headroom per concurrent request than the same model serving short prompts. Size for your retrieval budget — top-k times chunk size — not for the question alone.
Reranking sits between the two and is often forgotten. A cross-encoder reranker runs over every retrieved candidate rather than once per query, so its throughput requirement scales with your candidate count. If your pipeline reranks, budget for it as a third model rather than as a rounding error on the generator.
For one-off ingestion of a small corpus, CPU is often adequate. For large corpora or frequent re-indexing, a GPU is typically an order of magnitude faster per document and cheaper overall despite the higher hourly rate. Entry and mid-tier GPUs are usually sufficient, since embedding models are small.
They can, and co-locating avoids a second instance. But the two have different scaling profiles — embedding is throughput bound and bursty, generation is latency bound and continuous — so at production volume separate GPUs usually give better utilization and more predictable serving latency.
Retrieved context lengthens every prompt, which increases both prefill compute and KV cache size per request. A RAG deployment therefore needs more VRAM headroom per concurrent user than the same model handling short prompts. Size the KV cache for top-k times chunk size, plus the answer.