Hugging Face
Model hub with dedicated and routed inference
Last reviewed Sep 28, 2026
Hugging Face runs the Hub for open models and datasets. Inference Endpoints deploys a model on dedicated GPUs on AWS or Google Cloud, billed per hour of instance time, and Inference Providers routes API requests to partner inference providers at those providers' own rates.
We're actively tracking prices for Hugging Face. Check back soon, or browse other providers with current pricing.
Pros & Cons
Advantages
- Deploy any compatible Hub model to dedicated GPUs without managing infrastructure
- Choice of cloud and region per endpoint
- Inference Providers adds no markup over partner rates
- Monthly inference credits on free and PRO accounts
Limitations
- Endpoint GPU rates are above raw cloud rental for the same card
- GPU availability differs by cloud and region
- Inference Providers rates are the partners' rates, so costs vary with the provider a request is routed to
Key Features
Inference Endpoints
Dedicated deployments of Hub models on NVIDIA GPUs (T4 to H200 and RTX PRO 6000) on AWS or Google Cloud, with a choice of region and 1 to 8 GPUs
Inference Providers
One API and token for models served by partner providers such as Together, Fireworks, Groq, Cerebras, Novita and DeepInfra, billed at each provider's rate with no Hugging Face markup
Model Hub
Hosting for open models, datasets, and Spaces demo apps
Autoscaling
Endpoints scale replicas with traffic and can scale to zero
Pricing Options
| Option | Details |
|---|---|
| Inference Endpoints | Per hour of GPU instance time, billed by the minute |
| Inference Providers | Pass-through of the partner provider's per-token rate |
| Enterprise | Team and Enterprise Hub plans with managed billing |
Availability & Support
Support
Documentation, community forums, and paid support on Enterprise plans
Company
Published by Hugging Face to identify the company behind this listing.
- Headquarters
- 🇺🇸 New York, United States
Getting Started
- 1
Create an account
Sign up at huggingface.co and add a payment method
- 2
Deploy an endpoint
Pick a model, cloud, region, and GPU instance in the Inference Endpoints console
- 3
Or use Inference Providers
Call a model through the router with a Hugging Face access token
Frequently Asked Questions
What GPU types does Hugging Face offer?
Check the pricing table above for Hugging Face's current GPU availability and pricing.
How do I get started with Hugging Face?
Create an account, Deploy an endpoint, Or use Inference Providers
What are Hugging Face's main advantages?
Hugging Face's main advantages include: Deploy any compatible Hub model to dedicated GPUs without managing infrastructure, Choice of cloud and region per endpoint, Inference Providers adds no markup over partner rates, Monthly inference credits on free and PRO accounts.
What are Hugging Face's limitations?
Hugging Face's main limitations include: Endpoint GPU rates are above raw cloud rental for the same card, GPU availability differs by cloud and region, Inference Providers rates are the partners' rates, so costs vary with the provider a request is routed to.