The Nvidia L40S provides multi-workload acceleration for large language model (LLM) inference and training, graphics, and video applications. Based on the latest Ada Lovelace architecture.

| Provider | Price / hr |
|---|---|
$0.380/hr 1×2×4× | |
$0.530/hr 7× | |
$0.600/hr 1× | |
$0.620/hr24mo 8× | |
$0.700/hr12mo 8× | |
$0.706/hr 1×2×4×8× | |
$0.750/hr 8× | |
$0.770/hr 8× | |
$0.790/hr 1×2×3×4×5×6×7×8× | |
$0.790/hr 1× | |
$0.800/hr 1× | |
$0.800/hr 4× | |
$0.801/hr 2× | |
$0.801/hr 1× | |
$0.840/hr 4×8× | |
$0.870/hr 1× | |
$0.880/hr 1×2×4×8× | |
$0.890/hr36mo 1×2×4×8× | |
$0.930/hr 1×2×4×8× | |
$0.950/hr 1× | |
$0.960/hr 1× | |
$0.970/hr 1×2×4×8× | |
$0.970/hr 1× | |
$0.982/hr 1× | |
$0.985/hr 8× | |
$0.990/hr24mo 1×2×4×8× | |
$1.00/hr 1× | |
$1.09/hr 1×2×3×4×5×6×7× | |
$1.09/hr12mo 1×2×4×8× | |
$1.19/hr 1× | |
$1.19/hr6mo 1×2×4×8× | |
$1.20/hr 1× | |
$1.25/hr 1×2×4× | |
$1.26/hr 8× | |
$1.28/hr 1× | |
$1.29/hr 1×2×4×8× | |
$1.41/hr 1× | |
$1.41/hr 1×2×4× | |
$1.41/hr 8× | |
$1.45/hr 2×4×8×10× | |
$1.50/hr 1× | |
$1.50/hr 1× | |
$1.52/hr 10× | |
$1.57/hr 1× | |
$1.59/hr 10× | |
$1.65/hr 2× | |
$1.67/hr 0.5× | |
$1.67/hr 4× | |
$1.67/hr 0.25× | |
$1.70/hr 1×2×4×8× | |
$1.70/hr 1× | |
$1.78/hr 2× | |
$1.86/hr 1× | |
$1.94/hr 2× | |
$1.95/hr 1× | |
$1.96/hr 1× | |
$2.00/hr 1×2× | |
$2.10/hr 1×2×4×8× | |
$2.25/hr 8× | |
$2.62/hr 4× | |
$2.73/hr 4×8× | |
$3.19/hr 1× | |
$3.21/hr 3× | |
$3.50/hr 1× | |
$3.77/hr 8× | |
$0.550/hr 1× |
Prices updated daily. Last check: Sep 15, 2026
Every configuration, price rank, and alternative for one provider at a time.
The L40S is well-suited for organizations requiring combined AI and graphics capabilities in cloud environments. Its 48GB memory capacity and Transformer Engine make it effective for large language model inference, generative AI applications, and medium-scale training workloads. The inclusion of RT Cores and DLSS 3 support enables professional rendering, architectural visualization, and content creation workflows. The GPU's 24/7 data center design makes it appropriate for production AI inference services, while its dual-purpose nature serves environments running NVIDIA Omniverse for collaborative 3D workflows alongside AI applications.
L40S pricing varies by provider, region, and commitment level. Check the pricing table above for current rates across all providers.
The L40S excels at combined AI inference and graphics workloads, particularly large language model inference with its 48GB memory and Transformer Engine, generative AI applications, professional rendering with RT Core acceleration, and mixed enterprise workloads requiring both compute and visualization capabilities.
The H100 offers superior AI training performance with HBM3 memory and higher tensor throughput, while the L40S provides a balance of AI inference capabilities and graphics rendering with its RT Cores and DLSS 3 support. The L40S's 48GB GDDR6 memory is sufficient for most inference tasks, while the H100's 80GB HBM3 better serves large-scale training workloads.