Skip to main content
Baseten logo

Baseten

Inference platform for open and custom models

Inference specialist🇺🇸 USInference APIDedicated InferenceGPU CloudTraining

Last reviewed Sep 28, 2026

Baseten is an inference platform that serves popular open models through per-token Model APIs and runs custom or fine-tuned models on dedicated GPU deployments billed per minute, with managed training jobs on the same hardware.

We're actively tracking prices for Baseten. Check back soon, or browse other providers with current pricing.

Pros & Cons

Advantages

  • Per-minute billing on dedicated GPUs
  • Per-token Model APIs and dedicated deployments from one account
  • SOC 2 Type II certified and HIPAA compliant
  • Multi-GPU instance sizes published with vCPU and memory

Limitations

  • Dedicated GPU rates are higher than raw GPU rental because they include the serving platform
  • Model API catalog is a curated set of models
  • Volume discounts and custom regions require a sales conversation

Key Features

Model APIs

OpenAI- and Anthropic-compatible endpoints for hosted open models, billed per input, cached input, and output token

Dedicated deployments

Custom, fine-tuned, or open-source models on dedicated GPUs from T4 to B200 and RTX PRO 6000, in 1 to 8-GPU instances, with autoscaling

Training

Managed training jobs on the same GPU instance types, with checkpoints deployable to inference

Truss

Open-source packaging standard for serving models built in any framework

Self-hosted and hybrid

Enterprise deployments in the customer's VPC or split across Baseten's cloud and the customer's

Pricing Options

OptionDetails
Model APIsPer 1M input, cached input, and output tokens
Dedicated deploymentsPer minute of GPU instance time, scaled to zero when idle
Pro and EnterpriseVolume discounts, priority GPU access, and self-hosted options

Availability & Support

Support

Email and in-app chat on Basic; dedicated Slack and Zoom support on Pro and Enterprise

Company

Published by Baseten to identify the company behind this listing.

Headquarters
🇺🇸 San Francisco, United States

Getting Started

  1. 1

    Create an account

    Sign up; new accounts include credits

  2. 2

    Call a Model API

    Use an OpenAI-compatible client against a hosted model

  3. 3

    Deploy a model

    Package a model with Truss or pick one from the model library and choose a GPU instance type

Frequently Asked Questions

What GPU types does Baseten offer?

Check the pricing table above for Baseten's current GPU availability and pricing.

How do I get started with Baseten?

Create an account, Call a Model API, Deploy a model

What are Baseten's main advantages?

Baseten's main advantages include: Per-minute billing on dedicated GPUs, Per-token Model APIs and dedicated deployments from one account, SOC 2 Type II certified and HIPAA compliant, Multi-GPU instance sizes published with vCPU and memory.

What are Baseten's limitations?

Baseten's main limitations include: Dedicated GPU rates are higher than raw GPU rental because they include the serving platform, Model API catalog is a curated set of models, Volume discounts and custom regions require a sales conversation.