Skip to main content
Together AI logo

Together AI

The AI Native Cloud

Inference specialist🇺🇸 USinferenceopen-sourcetraining

Last reviewed Mar 14, 2026

Together AI is the AI Native Cloud platform engineered for developers building with open-source and frontier AI models. They provide serverless inference, fine-tuning, and GPU clusters with industry-leading performance optimizations.

GPU Models
6
From / hour
$1.49
LLM Models
52
From / 1M input
$0.01

Available GPUs

Hourly on-demand pricing. Click column headers to sort.

Prices last updated: July 24, 2026

GPU Model
Price / hr
A100 SXM
$2.59/hr
1×2×4×8×
B200
$9.00/hr
1×2×4×8×
H100 SXM
$5.40/hr
1×2×4×8×
H200
$6.60/hr
1×2×4×8×
L40
$1.49/hr
1×2×4×8×
L40S
$2.10/hr
1×2×4×8×

LLM API Pricing

Pay-per-token pricing. Prices shown per 1M tokens.

Prices last updated: July 24, 2026

ModelInput/1MOutput/1M
$0.0080$0.0080
$0.020$0.020
$0.050$0.200
$0.060$0.250
$0.060$0.060
$0.060$0.120
$0.060$0.060
$0.100$0.300
$0.150$1.50
$0.150$0.600

Pros & Cons

Advantages

  • 2x faster inference and 90% faster pre-training with cutting-edge research optimizations
  • Competitive pricing with 50% batch API discount
  • Wide selection of 100+ open-source models
  • OpenAI-compatible APIs for easy migration
  • Research leadership with FlashAttention contributions
  • Global data center network across 25+ cities

Limitations

  • Primarily focused on open-source models
  • GPU cluster pricing requires custom quotes for reserved capacity
  • Smaller ecosystem compared to major cloud providers

Key Features

100+ Open-Source Models

Access to Llama, DeepSeek, Qwen, and other leading open-source models

Serverless Inference

Pay-per-token API with OpenAI-compatible endpoints

Fine-Tuning Platform

LoRA and full fine-tuning with proprietary optimizations

GPU Clusters

Instant self-service or reserved dedicated clusters with H100, H200, B200, GB200, GB300 access

Batch API

50% cost reduction for non-urgent inference workloads

Provisioned Throughput

Reserve dedicated capacity in throughput units (PTUs) with SLAs

Code Interpreter

Execute LLM-generated code in sandboxed environments

AI Factory

Custom infrastructure at frontier scale

Sandbox

Build development environments for AI

Managed Storage

Store model weights & data securely

Dedicated Inference

Deploy models on custom hardware with guaranteed performance

Evaluations

Measure model quality

Pricing Options

OptionDetails
Serverless pay-per-tokenPer-token pricing scales based on model size, from small open-source models to 405B parameter frontier models
Batch API50% discount for non-urgent inference workloads
Provisioned ThroughputReserve dedicated capacity priced in throughput units (PTUs) with guaranteed SLAs
Fine-tuningPer-token pricing for LoRA and full fine-tuning based on model size and dataset
GPU Clusters - On-demandHourly GPU pricing for instant self-service clusters
GPU Clusters - ReservedCustom pricing for reserved capacity with significant discounts for longer commitments
Dedicated InferenceSingle-tenant GPU instances with guaranteed performance

Availability & Support

Regions

Global data center network across 25+ cities with frontier hardware including GB300, GB200, B200, H200, H100

Support

Documentation, community Discord, email support, and expert support for reserved cluster customers

Getting Started

  1. 1

    Create an account

    Sign up at together.ai

  2. 2

    Get API key

    Generate an API key from your dashboard

  3. 3

    Choose a model

    Browse 100+ models for chat, code, images, video, and audio

  4. 4

    Make API calls

    Use OpenAI-compatible endpoints or Together SDK

Compare Providers

Find the best prices for the same GPUs and models from other providers

AtmosCompute logo

AtmosCompute

6 shared GPUs with Together AI

CoreWeave logo

CoreWeave

6 shared GPUs with Together AI

Runpod logo

Runpod

6 shared GPUs with Together AI