Run open-source models at scale
Last reviewed Mar 14, 2026
Replicate is a platform for running machine learning models in the cloud, offering thousands of open-source models with simple API access and pay-per-use pricing.
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 11, 2026
| Model | Input/1M | Output/1M |
|---|---|---|
| $0.025/img | - | |
| $0.040/img | - | |
| $0.040/img | - | |
| $0.090/sec | - | |
| $0.090/img | - | |
| $0.250/sec | - |
Access thousands of open-source models including LLMs, image generators, and more
Consistent REST API across all models with webhooks for async processing
Deploy your own models using Cog containerization
Automatic scaling with cold-start optimization
| Option | Details |
|---|---|
| Pay-per-prediction | Charged per model run based on compute time and hardware |
| Free tier | Limited free predictions for new users |
US-based infrastructure with global CDN
Documentation, Discord community, email support
Sign up at replicate.com with GitHub or email
Copy your API token from account settings
Use the API or Python client to run any model
Check the pricing table above for the models Replicate currently serves and their per-token rates.
Create an account, Get API token, Run a prediction
Replicate's main advantages include: Largest selection of open-source models on one platform, Simple pay-per-prediction pricing with no minimum, Easy deployment of custom models via Cog, Active community contributing new models daily.
Replicate's main limitations include: Cold start latency for less popular models, Pricing can be unpredictable for high-volume use, Less optimized than specialized inference providers.
Find the best prices for the same GPUs and models from other providers