Skip to main content
Hugging Face logo

Hugging Face

Model hub with dedicated and routed inference

Cloud marketplace🇺🇸 USInference EndpointsInference ProvidersModel HubOpen Source

Last reviewed Sep 28, 2026

Hugging Face runs the Hub for open models and datasets. Inference Endpoints deploys a model on dedicated GPUs on AWS or Google Cloud, billed per hour of instance time, and Inference Providers routes API requests to partner inference providers at those providers' own rates.

We're actively tracking prices for Hugging Face. Check back soon, or browse other providers with current pricing.

Pros & Cons

Advantages

  • Deploy any compatible Hub model to dedicated GPUs without managing infrastructure
  • Choice of cloud and region per endpoint
  • Inference Providers adds no markup over partner rates
  • Monthly inference credits on free and PRO accounts

Limitations

  • Endpoint GPU rates are above raw cloud rental for the same card
  • GPU availability differs by cloud and region
  • Inference Providers rates are the partners' rates, so costs vary with the provider a request is routed to

Key Features

Inference Endpoints

Dedicated deployments of Hub models on NVIDIA GPUs (T4 to H200 and RTX PRO 6000) on AWS or Google Cloud, with a choice of region and 1 to 8 GPUs

Inference Providers

One API and token for models served by partner providers such as Together, Fireworks, Groq, Cerebras, Novita and DeepInfra, billed at each provider's rate with no Hugging Face markup

Model Hub

Hosting for open models, datasets, and Spaces demo apps

Autoscaling

Endpoints scale replicas with traffic and can scale to zero

Pricing Options

OptionDetails
Inference EndpointsPer hour of GPU instance time, billed by the minute
Inference ProvidersPass-through of the partner provider's per-token rate
EnterpriseTeam and Enterprise Hub plans with managed billing

Availability & Support

Support

Documentation, community forums, and paid support on Enterprise plans

Company

Published by Hugging Face to identify the company behind this listing.

Headquarters
🇺🇸 New York, United States

Getting Started

  1. 1

    Create an account

    Sign up at huggingface.co and add a payment method

  2. 2

    Deploy an endpoint

    Pick a model, cloud, region, and GPU instance in the Inference Endpoints console

  3. 3

    Or use Inference Providers

    Call a model through the router with a Hugging Face access token

Frequently Asked Questions

What GPU types does Hugging Face offer?

Check the pricing table above for Hugging Face's current GPU availability and pricing.

How do I get started with Hugging Face?

Create an account, Deploy an endpoint, Or use Inference Providers

What are Hugging Face's main advantages?

Hugging Face's main advantages include: Deploy any compatible Hub model to dedicated GPUs without managing infrastructure, Choice of cloud and region per endpoint, Inference Providers adds no markup over partner rates, Monthly inference credits on free and PRO accounts.

What are Hugging Face's limitations?

Hugging Face's main limitations include: Endpoint GPU rates are above raw cloud rental for the same card, GPU availability differs by cloud and region, Inference Providers rates are the partners' rates, so costs vary with the provider a request is routed to.