AI cloud for inference and compute
Last reviewed Aug 18, 2026
Runcrate provides a cloud-based platform for inference and compute services, focused on scaling AI workflows.
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 5, 2026
Configurations, price rank, and alternatives for one GPU at a time.
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 4, 2026
| Model | Input/1M | Output/1M |
|---|---|---|
| $0.037 | $0.170 | |
| $0.050 | $0.200 | |
| $0.090 | $0.550 | |
| $0.140 | $0.280 | |
| $0.150 | $1.15 | |
| $0.200 | $0.800 | |
| $0.300 | $1.00 | |
| $0.540 | $0.540 | |
| $0.600 | $2.08 | |
| $0.750 | $3.50 |
Input, output, and batch rates, plus alternatives, for one model at a time.
Getting started guide coming soon.
Runcrate offers various GPU types including RTX A5000, A100 PCIE, RTX A6000, A100 SXM, H100 SXM, GH200, H200, L40, A10, RTX 4000 Ada, RTX 4090, RTX 4090, A40, RTX 6000 Ada, RTX A4000, L40S, RTX 5090, Tesla V100, A16, Gaudi 2, L4, RTX PRO 6000. Check the pricing table above for current availability and pricing.
Visit Runcrate's website to create an account and start using their GPU services.
Find the best prices for the same GPUs and models from other providers