Skip to main content
Modal logo

Modal

Serverless GPUs for AI workloads with per-second billing

Inference specialist🇺🇸 US

Last reviewed May 8, 2026

Modal is a serverless GPU platform that lets developers run Python functions, jobs and inference endpoints on NVIDIA GPUs with per-second billing and scale-to-zero.

GPU Models
10
From / hour
$0.59

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: August 7, 2026

GPU Model
Price / hr
A10
$1.10/hr
1×
A100 SXM
$2.10/hr
1×
B200
$6.25/hr
1×
H100 SXM
$3.95/hr
1×
H200
$4.54/hr
1×
HGX B300
$7.10/hr
1×
L4
$0.799/hr
1×
L40S
$1.95/hr
1×
RTX PRO 6000
$3.03/hr
1×
Tesla T4
$0.590/hr
1×

Configurations, price rank, and alternatives for one GPU at a time.

Pros & Cons

Advantages

  • Serverless model removes idle-instance costs
  • Per-second billing across the full GPU range
  • Strong fit for inference, batch jobs and ML pipelines

Limitations

  • Long-running, fixed-instance training is not the primary use case
  • Cold starts and storage limits require some application design
  • No bare-metal access; workloads run inside Modal's runtime

Key Features

Serverless GPUs

Run Python functions on NVIDIA GPUs without provisioning instances; cold starts in seconds

Per-Second Billing

Pay for actual GPU runtime at sub-minute granularity, with scale-to-zero by default

Container-Native

Define environments in code, with automatic image building and caching

Wide GPU Catalog

From T4 and L4 through A100, L40S, H100, H200 and B200

Pricing Options

OptionDetails
Per-Second GPU BillingCharged per second of GPU runtime, with scale-to-zero when idle
Free TierMonthly free credits for experimentation and personal projects
Team and Enterprise PlansVolume commitments and enterprise support for production deployments

Availability & Support

Regions

Multi-region availability across North America and Europe

Support

Documentation, community forum, and enterprise support for paid plans

Getting Started

  1. 1

    Install the SDK

    Run `pip install modal` and authenticate via the CLI

  2. 2

    Define a function

    Decorate a Python function with the desired GPU and image specification

  3. 3

    Run or deploy

    Invoke locally or deploy as a long-lived endpoint or scheduled job

Compare Providers

Find the best prices for the same GPUs and models from other providers

IO.NET logo

IO.NET

8 shared GPUs with Modal

Compare Prices
Oracle Cloud logo

Oracle Cloud

8 shared GPUs with Modal

Compare Prices
Runpod logo

Runpod

8 shared GPUs with Modal

Compare Prices