Diffusion models fit in 24 GB; batch size buys throughput.
Ordered by suitability for this workload, not by price.
| Provider | Price / hr |
|---|---|
$0.080/hr 1× | |
$0.090/hr 1× | |
$0.130/hr 4× | |
$0.130/hr 1× | |
$0.151/hr 1× | |
$0.160/hr 1× | |
$0.160/hr 8× | |
$0.180/hr 1×3× | |
$0.200/hr 1× | |
$0.200/hr 1× | |
$0.220/hr 1×2×3×4× | |
$0.252/hr 1× | |
$0.280/hr 1× | |
$0.300/hr 1×2×8× | |
$0.305/hr 1×2×4×8× | |
$0.311/hr 1× | |
$0.318/hr36mo 1× | |
$0.335/hr 1× | |
$0.340/hr 1×2×3×4×5×6× | |
$0.350/hr 1× | |
$0.350/hr 1× | |
$0.353/hr 1× | |
$0.380/hr 1×2×4× | |
$0.380/hr12mo 8× | |
$0.380/hr24mo 8× | |
$0.390/hr 1× | |
$0.400/hr 1× | |
$0.400/hr 8× | |
$0.410/hr 1× | |
$0.420/hr 2×4× | |
$0.420/hr 8× | |
$0.430/hr 8× | |
$0.440/hr 1× | |
$0.440/hr 8× | |
$0.440/hr 1×8× | |
$0.442/hr 1× | |
$0.442/hr12mo 1× | |
$0.450/hr 1× | |
$0.450/hr36mo 2×4×8× | |
$0.490/hr 1×2×3×4×5× | |
$0.493/hr 1×2×4×8× | |
$0.497/hr 1× | |
$0.500/hr 1×2× | |
$0.500/hr 1× | |
$0.500/hr 1× | |
$0.500/hr 1×2×4×8× | |
$0.520/hr 1×2×4×8× | |
$0.530/hr 1× | |
$0.530/hr 1×2×3×4×5×6× | |
$0.540/hr 1× | |
$0.540/hr 1× | |
$0.550/hr 1× | |
$0.550/hr 1× | |
$0.550/hr 1× | |
$0.550/hr 1× | |
$0.570/hr 1×2×4×8× | |
$0.570/hr 1×2×4×8× | |
$0.600/hr 1× | |
$0.600/hr 1×4× | |
$0.600/hr 4×8× | |
$0.610/hr 1×2×4×8× | |
$0.620/hr12mo 8× | |
$0.620/hr24mo 8× | |
$0.627/hr 2×4×8× | |
$0.630/hr 1×4× | |
$0.630/hr12mo 2×4×8× | |
$0.640/hr 1× | |
$0.660/hr12mo 8× | |
$0.660/hr24mo 8× | |
$0.660/hr 1×4× | |
$0.662/hr 1× | |
$0.664/hr 1× | |
$0.670/hr 1× | |
$0.680/hr 1× | |
$0.680/hr 8× | |
$0.682/hr 1×2×4×8× | |
$0.685/hr 1×2×4×8× | |
$0.690/hr 1× | |
$0.690/hr 2×8× | |
$0.700/hr 8× | |
$0.740/hr 1×2×3×4×5×6×7× | |
$0.740/hr 1×2×3× | |
$0.750/hr 1× | |
$0.765/hr 4× | |
$0.765/hr 2× | |
$0.778/hr 0.5× | |
$0.790/hr 1×2×3×4×5× | |
$0.790/hr 1×2×4×8× | |
$0.799/hr 1× | |
$0.805/hr 1× | |
$0.840/hr 1× | |
$0.840/hr 4×8× | |
$0.860/hr 1× | |
$0.870/hr 1× | |
$0.880/hr 1×2×4×8× | |
$0.890/hr36mo 1×2×4×8× | |
$0.915/hr 1×2×4×8× | |
$0.925/hr 2×4× | |
$0.927/hr 1× | |
$0.930/hr 1× | |
$0.945/hr 8× | |
$0.950/hr 1×2×4×8× | |
$0.960/hr 1× | |
$0.966/hr 1× | |
$0.970/hr 1×2×4× | |
$0.970/hr 1×2×4×8× | |
$0.970/hr 1× | |
$0.985/hr 8× | |
$0.988/hr 1× | |
$0.990/hr24mo 1×2×4×8× | |
$1.00/hr 2×4×8× | |
$1.01/hr 1× | |
$1.02/hr 1× | |
$1.04/hr 1×2×4×8× | |
$1.04/hr 1×2×4×8× | |
$1.05/hr 8× | |
$1.05/hr 1× | |
$1.06/hr 4× | |
$1.07/hr 1×2×4× | |
$1.07/hr 1× | |
$1.07/hr 2× | |
$1.09/hr 2× | |
$1.09/hr 1×2×3×4×5×6×7× | |
$1.09/hr 1×4× | |
$1.09/hr 1×4× | |
$1.09/hr12mo 1×2×4×8× | |
$1.10/hr 1× | |
$1.11/hr 1× | |
$1.14/hr 1× | |
$1.15/hr 4× | |
$1.19/hr6mo 1×2×4×8× | |
$1.20/hr 1× | |
$1.23/hr 1× | |
$1.24/hr 1× | |
$1.25/hr 1× | |
$1.26/hr 1×2×4× | |
$1.27/hr 8× | |
$1.29/hr 1× | |
$1.29/hr 1× | |
$1.29/hr 1× | |
$1.29/hr 1×2×4×8× | |
$1.30/hr 2× | |
$1.37/hr 2× | |
$1.37/hr 1×2×4×8× | |
$1.42/hr 4× | |
$1.42/hr 1× | |
$1.42/hr 1× | |
$1.45/hr 2× | |
$1.49/hr 1×2×4×8× | |
$1.50/hr 1× | |
$1.57/hr 1× | |
$1.65/hr 2× | |
$1.65/hr 1× | |
$1.67/hr 0.5× | |
$1.67/hr 8× | |
$1.67/hr 4× | |
$1.67/hr 0.25× | |
$1.69/hr 1× | |
$1.70/hr 1×2× | |
$1.71/hr 1×2×4×8× | |
$1.86/hr 1× | |
$1.89/hr 1× | |
$1.93/hr 1×2×4× | |
$1.95/hr 2× | |
$1.95/hr 1× | |
$1.96/hr 1× | |
$2.00/hr 1× | |
$2.03/hr 4× | |
$2.03/hr 2× | |
$2.04/hr 8× | |
$2.10/hr 1×2×4×8× | |
$2.25/hr 8× | |
$2.62/hr 4× | |
$2.73/hr 1× | |
$3.19/hr 1× | |
$3.50/hr 1× | |
$3.77/hr 8× | |
$0.330/hr 1× | |
$0.550/hr 1× | |
$1.09/hr 2× |
Diffusion models are small compared with language models. Stable Diffusion XL and FLUX-class models fit comfortably within 24 GB even at full precision, and considerably less with the offloading and quantization options built into common inference stacks. That makes image generation one of the few serious AI workloads where consumer-class GPUs are genuinely competitive with datacenter parts.
The workload is compute bound rather than bandwidth bound, because each denoising step runs the full UNet or transformer over a comparatively small latent. Raw FP16 throughput and sustained clocks therefore predict images per second better than memory bandwidth does. This is the opposite of LLM token generation, and it is why a card like the RTX 4090 punches above its position in the datacenter hierarchy here.
VRAM sets the ceiling on resolution and batch size rather than on whether the model runs at all. Higher resolutions grow the latent quadratically, and video or multi-frame generation multiplies that again, so 24 GB that is ample for single 1024px images gets tight quickly for upscaling pipelines, ControlNet stacks, or video models. If your pipeline chains several models, count them all.
For production serving, throughput per dollar usually favours several mid-tier GPUs over one large one, since requests parallelize cleanly across devices and no single request needs more memory than a mid-tier card provides. Datacenter cards still earn their place where you need ECC memory, sustained duty cycles, or the provider simply does not offer consumer hardware.
A 12–16 GB card runs SDXL-class models with standard optimizations, and 24 GB gives comfortable headroom for higher resolutions, larger batches, ControlNet stacks and LoRA training. Below 12 GB you can still generate images with offloading, at a noticeable speed penalty.
It is well matched to the workload: image generation is compute bound rather than memory-bandwidth bound, and 24 GB of GDDR6X covers the models in common use. It lacks ECC memory and NVLink, which matters for training at scale but not for generation. Availability varies by provider — see the table above.
LoRA training for a diffusion model fits within 16–24 GB. Full fine-tuning of an SDXL-class model wants 40 GB or more, and training from scratch is a multi-GPU exercise. Video models raise every one of these figures substantially.