Skip to main content
Anthropic

Claude Sonnet 5.5

Claude Sonnet 5.5 is a chat-oriented large language model from Anthropic, part of the Claude Sonnet line that sits between the smaller Haiku and larger Opus tiers.

Input from
$1.00 / 1M tokens
across 2 providers

API Pricing

Cheapest on OpenRouter — 40% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$1.00$5.00$0.100
$2.00$10.00$0.200
$2.00$10.00$0.200

Prices updated daily. Last check: Sep 29, 2026

Claude Sonnet 5.5 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
56.0 / 100
Output Speed
145 t/s
Latency (TTFT)
304.8s

Reasoning & Knowledge

  • Humanity's Last Exam55.0%

Coding

  • SciCode61.0%

Instruction & Long Context

  • Long-Context Reasoning82.7%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Anthropic
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Part of Anthropic's Claude Sonnet tier, which targets a balance between capability and serving cost rather than maximum model size
  • Measured output throughput of approximately 89.3 tokens per second in Artificial Analysis testing, suitable for streaming chat interfaces
  • Time to first token of roughly 860 ms in the same testing, giving reasonably prompt response starts for interactive use
  • Belongs to a model family with an established API surface and documentation from Anthropic, easing migration from earlier Claude versions
  • Mid-tier positioning makes it a common default choice for teams that find Haiku-class models too limited and Opus-class models more than they need
  • Available through multiple hosting providers, so buyers can compare terms in the pricing table on this page

Limitations

  • Context window size is not confirmed in our database, so plan long-document workloads against Anthropic's published specifications
  • We do not track standardized quality benchmark scores for this model, making direct quality comparisons with peers harder from this page alone
  • Reported throughput of roughly 89.3 tokens per second is below what some smaller, lighter models achieve, which matters for very high-volume batch generation
  • Approximately 860 ms time to first token adds noticeable delay in latency-sensitive voice or real-time pipelines
  • Mid-tier models generally trade some headroom on the hardest reasoning tasks compared with larger models in the same family

Key Features

•Chat and instruction-following interface in the Claude message format
•Streaming token output, measured at roughly 89.3 tokens per second by Artificial Analysis
•Time to first token of approximately 860 ms in third-party latency testing
•Sonnet-tier positioning within Anthropic's Haiku / Sonnet / Opus lineup
•Accessible via API from multiple hosting providers listed in the pricing table
•Compatible with existing Claude API integration patterns and SDKs

About Claude Sonnet 5.5

Claude Sonnet 5.5 is a conversational large language model developed by Anthropic and released as part of the Claude family. Within Anthropic's naming scheme, Sonnet models occupy the middle tier of the lineup, positioned between the compact Haiku models and the larger Opus models, and are generally aimed at workloads that need stronger reasoning than a small model provides without committing to the largest available option. On measured serving performance, third-party testing from Artificial Analysis reports Claude Sonnet 5.5 generating roughly 89.3 output tokens per second with a time to first token of about 860 ms. Those figures describe responsiveness during streaming generation and initial latency respectively, and they are useful for sizing interactive applications such as chat interfaces, coding assistants, and agent loops where perceived latency matters. Note that both numbers are provider- and load-dependent and can differ from what you observe against a specific endpoint. We do not currently track a confirmed context window, modality list, or standardized quality benchmark set for Claude Sonnet 5.5 in our database, so this page focuses on what has been verified. Readers evaluating the model against siblings in the Claude family, or against mid-tier models from other providers, should check Anthropic's own documentation for capability details and use the pricing table on this page to compare what each provider charges for access.

Common Use Cases

Claude Sonnet 5.5 fits the workloads that typically land on a mid-tier chat model: customer-facing assistants, internal knowledge chatbots, drafting and editing text, code explanation and generation inside developer tools, and multi-step agent workflows where each step needs solid reasoning but the pipeline runs often enough that the largest model tier becomes impractical. The measured throughput of around 89.3 output tokens per second supports streaming responses that feel continuous to a reader, while the roughly 860 ms time to first token is acceptable for text chat though less ideal for real-time voice interfaces. For very high-volume, low-complexity tasks such as bulk classification or short-form tagging, a smaller Haiku-class model is usually the more economical fit; for the hardest analytical or long-horizon agentic work, teams often escalate to an Opus-tier model and keep Sonnet as the default for everything in between.

Frequently Asked Questions

How much does Claude Sonnet 5.5 cost?

Pricing depends on which provider hosts the model and on the pricing type — input versus output tokens, on-demand versus committed capacity, and any batch or caching discounts a provider offers. Rates also change over time. Check the pricing table on this page for current per-provider figures rather than relying on a quoted number.

What is Claude Sonnet 5.5 best used for?

It is best suited to general-purpose chat and assistant workloads: conversational agents, document drafting and revision, coding assistance, and agent workflows that need dependable reasoning at mid-tier serving cost. Its measured streaming speed makes it workable for interactive interfaces where users read output as it is generated.

How does Claude Sonnet 5.5 compare to Haiku and Opus models from Anthropic?

Anthropic organizes the Claude family into Haiku (smallest and fastest), Sonnet (middle tier), and Opus (largest). Sonnet 5.5 sits in that middle position, which generally means more capability than a Haiku model at higher cost and latency, and less headroom than an Opus model at lower cost. The right choice depends on whether your task is bottlenecked by reasoning quality or by throughput and budget.

How fast is Claude Sonnet 5.5?

Artificial Analysis measured roughly 89.3 output tokens per second with a time to first token of about 860 ms. Real-world numbers vary by provider, region, prompt length, and current load, so treat these as a reference point rather than a guarantee.

What context window does Claude Sonnet 5.5 support?

We do not have a confirmed context window recorded for this model in our database. Consult Anthropic's official model documentation or your chosen provider's API reference for the supported input length before designing long-document or large-codebase workflows.