Skip to main content
Modal logo

Modal

Serverless GPUs for AI workloads with per-second billing

Inference specialist🇺🇸 US

Last reviewed May 8, 2026

Modal is a serverless GPU platform that lets developers run Python functions, jobs and inference endpoints on NVIDIA GPUs with per-second billing and scale-to-zero.

GPU Models
11
From / hour
$0.59

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: September 24, 2026

A10
$1.10/hr
1×
A100 PCIE
$2.10/hr
1×
A100 SXM
$2.50/hr
1×
B200
$6.25/hr
1×
H100 SXM
$3.95/hr
1×
H200
$4.54/hr
1×
HGX B300
$7.10/hr
1×
L4
$0.799/hr
1×
L40S
$1.95/hr
1×
RTX PRO 6000
$3.03/hr
1×

Configurations, price rank, and alternatives for one GPU at a time.

Pros & Cons

Advantages

  • Serverless model removes idle-instance costs
  • Per-second billing across the full GPU range
  • Strong fit for inference, batch jobs and ML pipelines

Limitations

  • Long-running, fixed-instance training is not the primary use case
  • Cold starts and storage limits require some application design
  • No bare-metal access; workloads run inside Modal's runtime

Key Features

Serverless GPUs

Run Python functions on NVIDIA GPUs without provisioning instances; cold starts in seconds

Per-Second Billing

Pay for actual GPU runtime at sub-minute granularity, with scale-to-zero by default

Container-Native

Define environments in code, with automatic image building and caching

Wide GPU Catalog

From T4 and L4 through A100, L40S, H100, H200 and B200

Pricing Options

OptionDetails
Per-Second GPU BillingCharged per second of GPU runtime, with scale-to-zero when idle
Free TierMonthly free credits for experimentation and personal projects
Team and Enterprise PlansVolume commitments and enterprise support for production deployments

Availability & Support

Regions

Multi-region availability across North America and Europe

Support

Documentation, community forum, and enterprise support for paid plans

Getting Started

  1. 1

    Install the SDK

    Run `pip install modal` and authenticate via the CLI

  2. 2

    Define a function

    Decorate a Python function with the desired GPU and image specification

  3. 3

    Run or deploy

    Invoke locally or deploy as a long-lived endpoint or scheduled job

Frequently Asked Questions

What GPU types does Modal offer?

Modal offers various GPU types including A100 PCIE, Tesla T4, A100 SXM, H100 SXM, H200, A10, HGX B300, L40S, B200, L4, RTX PRO 6000. Check the pricing table above for current availability and pricing.

How do I get started with Modal?

Install the SDK, Define a function, Run or deploy

What are Modal's main advantages?

Modal's main advantages include: Serverless model removes idle-instance costs, Per-second billing across the full GPU range, Strong fit for inference, batch jobs and ML pipelines.

What are Modal's limitations?

Modal's main limitations include: Long-running, fixed-instance training is not the primary use case, Cold starts and storage limits require some application design, No bare-metal access; workloads run inside Modal's runtime.

Compare Providers

Find the best prices for the same GPUs and models from other providers

GPU Outlet logo

GPU Outlet

11 shared GPUs with Modal

Compare Prices
Vast.ai logo

Vast.ai

10 shared GPUs with Modal

Compare Prices
AtmosCompute logo

AtmosCompute

9 shared GPUs with Modal

Compare Prices