Serverless GPUs with per-second billing and global deployment
Last reviewed May 8, 2026
Koyeb is a serverless platform offering on-demand NVIDIA GPUs from RTX-class up through B200 with per-second billing, scale-to-zero and global edge deployment. Koyeb is joining Mistral AI to build the future of AI infrastructure.
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 16, 2026
Configurations, price rank, and alternatives for one GPU at a time.
Per-second billing with scale-to-zero across the GPU catalog
From RTX-4000-SFF-ADA and L4 through L40S, A100, H100, H200 and B200
Deploy applications close to users across multiple regions from a single push
| Option | Details |
|---|---|
| Per-Second GPU Billing | Charged per second of GPU runtime, with scale-to-zero when idle |
| Multi-GPU Configurations | Pre-configured 2x, 4x and 8x GPU instances for larger workloads |
Global edge presence across North America, Europe and Asia
Documentation, community forum and paid support tiers
Sign up via the Koyeb console using email or GitHub
Select an instance type from the GPU catalog and configure your service
Push from a GitHub repo or Docker image and Koyeb handles the rest
Koyeb offers various GPU types including A100 PCIE, A100 PCIE, A100 PCIE, A100 PCIE, RTX A6000, A100 SXM, A100 SXM, A100 SXM, A100 SXM, H100 SXM, H100 SXM, H100 SXM, H100 SXM, L40S. Check the pricing table above for current availability and pricing.
Create an account, Pick a GPU instance, Deploy from Git or container
Koyeb's main advantages include: Per-second billing across the full GPU range, Multi-GPU configurations (2x, 4x, 8x) available on-demand, Serverless model removes idle-instance costs.
Koyeb's main limitations include: Smaller per-cluster scale than dedicated training neoclouds, Some GPU SKUs require requesting access, Long-running fixed-instance training is not the primary use case.
Find the best prices for the same GPUs and models from other providers