The AI Native Cloud
Last reviewed Mar 14, 2026
Together AI is the AI Native Cloud platform engineered for developers building with open-source and frontier AI models. They provide serverless inference, fine-tuning, and GPU clusters with industry-leading performance optimizations.
Hourly on-demand pricing. Click column headers to sort.
Prices last updated: September 8, 2026
Configurations, price rank, and alternatives for one GPU at a time.
Pay-per-token pricing. Prices shown per 1M tokens.
Prices last updated: September 8, 2026
| Model | Input/1M | Output/1M | |||
|---|---|---|---|---|---|
| $0.050 | $0.200 | ||||
| $0.060 | $0.250 | ||||
| $0.060 | $0.060 | ||||
| $0.060 | $0.060 | ||||
| $0.140 | $0.280 | ||||
| $0.150 | $0.470 | ||||
| $0.150 | $1.50 | ||||
| $0.150 | $0.500 | ||||
| $0.150 | $0.600 | ||||
| $0.150 | $1.50 | ||||
Input, output, and batch rates, plus alternatives, for one model at a time.
Access to Llama, DeepSeek, Qwen, and other leading open-source models
Pay-per-token API with OpenAI-compatible endpoints
LoRA and full fine-tuning with proprietary optimizations
Instant self-service or reserved dedicated clusters with H100, H200, B200, GB200, GB300 access
50% cost reduction for non-urgent inference workloads
Reserve dedicated capacity in throughput units (PTUs) with SLAs
Execute LLM-generated code in sandboxed environments
Custom infrastructure at frontier scale
Build development environments for AI
Store model weights & data securely
Deploy models on custom hardware with guaranteed performance
Measure model quality
| Option | Details |
|---|---|
| Serverless pay-per-token | Per-token pricing scales based on model size, from small open-source models to 405B parameter frontier models |
| Batch API | 50% discount for non-urgent inference workloads |
| Provisioned Throughput | Reserve dedicated capacity priced in throughput units (PTUs) with guaranteed SLAs |
| Fine-tuning | Per-token pricing for LoRA and full fine-tuning based on model size and dataset |
| GPU Clusters - On-demand | Hourly GPU pricing for instant self-service clusters |
| GPU Clusters - Reserved | Custom pricing for reserved capacity with significant discounts for longer commitments |
| Dedicated Inference | Single-tenant GPU instances with guaranteed performance |
Global data center network across 25+ cities with frontier hardware including GB300, GB200, B200, H200, H100
Documentation, community Discord, email support, and expert support for reserved cluster customers
Sign up at together.ai
Generate an API key from your dashboard
Browse 100+ models for chat, code, images, video, and audio
Use OpenAI-compatible endpoints or Together SDK
Together AI offers various GPU types including A100 PCIE, A100 PCIE, A100 PCIE, A100 PCIE, A100 SXM, A100 SXM, A100 SXM, A100 SXM, H100 SXM, H100 SXM, H100 SXM, H100 SXM, H200, H200, H200, H200, L40, L40, L40, L40, RTX 6000 Ada, RTX 6000 Ada, RTX 6000 Ada, RTX 6000 Ada, L40S, L40S, L40S, L40S, B200, B200, B200, B200. Check the pricing table above for current availability and pricing.
Create an account, Get API key, Choose a model, Make API calls
Together AI's main advantages include: 2x faster inference and 90% faster pre-training with cutting-edge research optimizations, Competitive pricing with 50% batch API discount, Wide selection of 100+ open-source models, OpenAI-compatible APIs for easy migration, Research leadership with FlashAttention contributions, Global data center network across 25+ cities.
Together AI's main limitations include: Primarily focused on open-source models, GPU cluster pricing requires custom quotes for reserved capacity, Smaller ecosystem compared to major cloud providers.
Find the best prices for the same GPUs and models from other providers