Skip to main content
fal.ai logo

fal.ai

Serverless inference platform optimized for generative media

Inference specialist🇺🇸 US

Last reviewed May 8, 2026

fal.ai is a serverless inference platform focused on generative media (image, video, audio) with a hosted model catalog and on-demand GPU runtimes for custom deployments.

GPU Models
3
From / hour
$2.99
LLM Models
1
From / 1M input

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: September 23, 2026

B200
$6.25/hr
1×
HGX B300
$8.50/hr
1×
RTX PRO 6000
$2.99/hr
1×

fal.ai pricing by GPU

Configurations, price rank, and alternatives for one GPU at a time.

LLM API Pricing

Pay-per-token pricing. Prices shown per 1M tokens.

Prices last updated: August 25, 2026

ModelInput/1MOutput/1M
$0.020/MP-

Pros & Cons

Advantages

  • Strong catalog of generative media models behind a single API
  • Per-second billing for serverless GPU deployments
  • Specialized inference optimizations for diffusion and audio

Limitations

  • Less suited to long-running fixed-instance training
  • Custom GPU deployments require contacting support to get started
  • Lower-tier consumer GPUs are not part of the catalog

Key Features

Hosted Model Catalog

Production endpoints for image, video and audio models billed by output unit (per image, per megapixel, per second or per video)

Custom GPU Deployments

Run private models on dedicated NVIDIA GPUs with autoscaling and scale-to-zero

Optimized Runtimes

Inference engines tuned for diffusion and audio workloads

Pricing Options

OptionDetails
Output-Based Model API PricingHosted model endpoints billed per generated output unit — per image, per megapixel, per second of video or per video
Per-Second GPU PricingCustom deployments billed per second of GPU runtime, with scale-to-zero
Enterprise ContractsVolume commitments and dedicated capacity for high-throughput customers, with discounted GPU rates below the published list price

Availability & Support

Regions

Multi-region serverless infrastructure

Support

Documentation, community channels and enterprise support for paid customers

Getting Started

  1. 1

    Create an account

    Sign up and generate an API key

  2. 2

    Pick a hosted model or upload your own

    Choose from the catalog or define a custom GPU-backed deployment

  3. 3

    Call the API

    Invoke endpoints from any language using the REST or SDK clients

Frequently Asked Questions

What GPU types does fal.ai offer?

fal.ai offers various GPU types including HGX B300, B200, RTX PRO 6000. Check the pricing table above for current availability and pricing.

How do I get started with fal.ai?

Create an account, Pick a hosted model or upload your own, Call the API

What are fal.ai's main advantages?

fal.ai's main advantages include: Strong catalog of generative media models behind a single API, Per-second billing for serverless GPU deployments, Specialized inference optimizations for diffusion and audio.

What are fal.ai's limitations?

fal.ai's main limitations include: Less suited to long-running fixed-instance training, Custom GPU deployments require contacting support to get started, Lower-tier consumer GPUs are not part of the catalog.

Compare Providers

Find the best prices for the same GPUs and models from other providers

1Legion logo

1Legion

3 shared GPUs with fal.ai

Compare Prices
AtmosCompute logo

AtmosCompute

3 shared GPUs with fal.ai

Compare Prices
CoreWeave logo

CoreWeave

3 shared GPUs with fal.ai

Compare Prices

All fal.ai Comparisons

fal.ai vs 1Legionfal.ai vs AtmosComputefal.ai vs CoreWeavefal.ai vs GPU Outletfal.ai vs gpu.aifal.ai vs Hyperstackfal.ai vs Lium.iofal.ai vs Massed Computefal.ai vs Modalfal.ai vs Oracle Cloudfal.ai vs Runpodfal.ai vs Spheronfal.ai vs UpCloudfal.ai vs Vast.aifal.ai vs Verdafal.ai vs VoltageGPUfal.ai vs Civofal.ai vs Daytonafal.ai vs Deep Infrafal.ai vs Google Cloudfal.ai vs Hexgrid Cloudfal.ai vs Latitude.shfal.ai vs Lyceumfal.ai vs Nebiusfal.ai vs TheAI Cloudfal.ai vs AceCloudfal.ai vs Akamai Cloudfal.ai vs Aquanodefal.ai vs Bentausfal.ai vs Beyond.plfal.ai vs Compute Cheapfal.ai vs DigitalOceanfal.ai vs EcoHashfal.ai vs Gcorefal.ai vs GMI Cloudfal.ai vs Hinodefal.ai vs Hyperbolicfal.ai vs IO.NETfal.ai vs Jarvis Labsfal.ai vs Lambda Labsfal.ai vs Omega Gradientfal.ai vs Packet AIfal.ai vs Runcratefal.ai vs Scalewayfal.ai vs Seewebfal.ai vs Sestercefal.ai vs Shadeformfal.ai vs Together AIfal.ai vs Vultrfal.ai vs Amazon AWSfal.ai vs Atlas Cloudfal.ai vs Beamfal.ai vs Crusoefal.ai vs Cudo Computefal.ai vs Denvr Dataworksfal.ai vs Hot Aislefal.ai vs Koyebfal.ai vs Microsoft Azurefal.ai vs Novita AIfal.ai vs Oblivusfal.ai vs Paperspacefal.ai vs QuickPodfal.ai vs Salad Cloudfal.ai vs SwissGPUfal.ai vs Theta EdgeCloudfal.ai vs Thunder Computefal.ai vs Voltage Parkfal.ai vs Zettabytefal.ai vs OpenRouterfal.ai vs Anthropicfal.ai vs Cerebrasfal.ai vs Coherefal.ai vs Fireworks AIfal.ai vs Geoddfal.ai vs Groqfal.ai vs Mistral AIfal.ai vs OpenAIfal.ai vs Perplexityfal.ai vs Replicatefal.ai vs Velokey