fal.ai
Serverless inference platform optimized for generative media
Last reviewed May 8, 2026
fal.ai is a serverless inference platform focused on generative media (image, video, audio) with a hosted model catalog and on-demand GPU runtimes for custom deployments.
Available GPUs
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 23, 2026
| B200 | $6.25/hr 1× | 1× |
| HGX B300 | $8.50/hr 1× | 1× |
| RTX PRO 6000 | $2.99/hr 1× | 1× |
fal.ai pricing by GPU
Configurations, price rank, and alternatives for one GPU at a time.
LLM API Pricing
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: August 25, 2026
| Model | Input/1M | Output/1M |
|---|---|---|
| $0.020/MP | - |
Pros & Cons
Advantages
- Strong catalog of generative media models behind a single API
- Per-second billing for serverless GPU deployments
- Specialized inference optimizations for diffusion and audio
Limitations
- Less suited to long-running fixed-instance training
- Custom GPU deployments require contacting support to get started
- Lower-tier consumer GPUs are not part of the catalog
Key Features
Hosted Model Catalog
Production endpoints for image, video and audio models billed by output unit (per image, per megapixel, per second or per video)
Custom GPU Deployments
Run private models on dedicated NVIDIA GPUs with autoscaling and scale-to-zero
Optimized Runtimes
Inference engines tuned for diffusion and audio workloads
Pricing Options
| Option | Details |
|---|---|
| Output-Based Model API Pricing | Hosted model endpoints billed per generated output unit — per image, per megapixel, per second of video or per video |
| Per-Second GPU Pricing | Custom deployments billed per second of GPU runtime, with scale-to-zero |
| Enterprise Contracts | Volume commitments and dedicated capacity for high-throughput customers, with discounted GPU rates below the published list price |
Availability & Support
Regions
Multi-region serverless infrastructure
Support
Documentation, community channels and enterprise support for paid customers
Getting Started
- 1
Create an account
Sign up and generate an API key
- 2
Pick a hosted model or upload your own
Choose from the catalog or define a custom GPU-backed deployment
- 3
Call the API
Invoke endpoints from any language using the REST or SDK clients
Frequently Asked Questions
What GPU types does fal.ai offer?
fal.ai offers various GPU types including HGX B300, B200, RTX PRO 6000. Check the pricing table above for current availability and pricing.
How do I get started with fal.ai?
Create an account, Pick a hosted model or upload your own, Call the API
What are fal.ai's main advantages?
fal.ai's main advantages include: Strong catalog of generative media models behind a single API, Per-second billing for serverless GPU deployments, Specialized inference optimizations for diffusion and audio.
What are fal.ai's limitations?
fal.ai's main limitations include: Less suited to long-running fixed-instance training, Custom GPU deployments require contacting support to get started, Lower-tier consumer GPUs are not part of the catalog.
Compare Providers
Find the best prices for the same GPUs and models from other providers