Compare per-token rates across 34 providers and 264 LLM models
LLM APIs charge separately for input tokens (your prompts) and output tokens (model responses). Output tokens are usually 2-5x more expensive than input tokens.
The context window determines how much text a model can process at once. Larger context windows allow for longer conversations and document analysis.
Open-source models (Llama, Mistral) are often cheaper but may require more tuning. Proprietary models (GPT-4, Claude) typically offer better out-of-box performance.
Many providers offer 50% discounts for batch/async API usage. Consider batch APIs for non-time-sensitive workloads to reduce costs.
LLM inference pricing refers to the cost of using Large Language Model APIs. Providers charge per token, with separate rates for input tokens (your prompts) and output tokens (model responses). Output tokens typically cost 2-5x more than input tokens.
LLM API costs are calculated by multiplying your input token count by the input price per million tokens, multiplying your output token count by the output price per million tokens, and adding the two together. Because output tokens usually cost several times more than input tokens, workloads that generate long responses are more sensitive to the output rate — see the live comparison table above for current per-model rates.
The cheapest LLM provider depends on the model you need. For open-source models like Llama, providers like Together AI, Fireworks AI, and DeepInfra offer competitive rates. For proprietary models like GPT-4 or Claude, you'll typically need to use OpenAI or Anthropic directly. Check the live comparison table above for current rates across providers.
A context window is the maximum number of tokens an LLM can process in a single request. Larger context windows (128K-1M+ tokens) allow for longer conversations and document analysis but may cost more. Most frontier models now support at least 128K tokens.