Skip to main content
Anthropic

Claude Haiku 5.5

Claude Haiku 5.5 is a chat model from Anthropic in the Claude Haiku line, the company's small, speed-oriented tier of Claude models.

Input from
$0.050 / 1M tokens
across 1 provider

API Pricing

Cheapest on OpenRouter — 33% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$0.050$0.250$0.0050
$0.100$0.500$0.010

Prices updated daily. Last check: Oct 8, 2026

Claude Haiku 5.5 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
43.4 / 100
Output Speed
239 t/s
Latency (TTFT)
292.4s

Reasoning & Knowledge

  • Humanity's Last Exam44.4%

Coding

  • SciCode55.0%

Instruction & Long Context

  • Long-Context Reasoning82.7%

Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Anthropic
Modalities
Text

Capabilities

Strengths & Limitations

Strengths

  • Part of Anthropic's Haiku tier, which is designed around low latency and high request volume rather than maximum reasoning depth
  • Measured output speed of roughly 137.8 tokens per second (Artificial Analysis), suitable for streaming responses to end users
  • Positioned below Sonnet and Opus in the Claude lineup, giving teams a cheaper tier to route simple requests to within the same model family
  • Shares the Claude prompting conventions and API surface, so prompts and integrations can often be reused across Claude tiers
  • Available through multiple serving providers, allowing price, region, and rate-limit comparison
  • Well suited to being the high-throughput worker model in a multi-model pipeline

Limitations

  • Measured time to first token of about 4,697 ms is high relative to the model's generation speed, which affects short interactive turns
  • As a Haiku-tier model, it is not the tier Anthropic targets at the hardest reasoning, long-horizon agentic, or complex coding workloads
  • Context window, modality support, and tool-calling details are not tracked in our database for this entry — verify against Anthropic's documentation before building against them
  • Benchmark accuracy scores for this model are not available in our dataset, so quality comparisons against peers must come from your own evaluations
  • Throughput and latency figures vary by serving provider, so the measured numbers may not match what you observe on a specific endpoint

Key Features

•Anthropic Claude Haiku tier — the compact, speed-oriented branch of the Claude family
•Text chat and instruction-following interface
•Measured output throughput of approximately 137.8 tokens/second (Artificial Analysis)
•Measured time to first token of approximately 4,697 ms (Artificial Analysis)
•Streaming token output for incremental response rendering
•Compatible with the Claude API message format and prompting conventions
•Multi-provider availability for price and region comparison

About Claude Haiku 5.5

Claude Haiku 5.5 is a text chat model built by Anthropic. It belongs to the Haiku branch of the Claude family, which Anthropic positions as its compact, latency-sensitive tier — sitting below the Sonnet and Opus lines that are aimed at heavier reasoning and long-horizon agentic work. Haiku models are generally selected when a workload needs many requests handled quickly and consistently rather than the deepest possible reasoning on each one. On throughput, independent measurements from Artificial Analysis put Claude Haiku 5.5 at roughly 137.8 output tokens per second, with a measured time to first token of about 4,697 ms. That combination is worth reading carefully: once generation starts, tokens arrive at a brisk rate, but the measured initial latency is substantial, which matters more for short interactive turns than for long generations where the streaming rate dominates total completion time. Other specifications for this entry — context window, image input, and tool-calling details — are not tracked in our database, so we do not list them here; check Anthropic's model documentation for the authoritative specification. In practice, Haiku-tier Claude models are used for high-volume, cost-sensitive text work: classification, extraction, summarization, routing, content tagging, and as the worker model inside larger pipelines where a Sonnet- or Opus-tier model handles planning. Because Claude Haiku 5.5 is served through multiple providers, the practical decision usually comes down to the price, rate limits, and region offered by each one — compare those in the pricing table on this page.

Common Use Cases

Claude Haiku 5.5 fits workloads where request volume is high and each individual call is relatively bounded: classifying support tickets, extracting structured fields from documents, summarizing conversation logs, tagging or moderating user-generated text, drafting short responses, and acting as the routing or triage layer in front of a larger model. Its measured generation rate of around 138 tokens per second makes it reasonable for streaming chat responses where the user watches text appear, though the measured initial latency means very short, snappy exchanges may feel less immediate than the raw token rate suggests. Teams already standardized on Claude often use a Haiku-tier model as the default and escalate only the hard requests to Sonnet- or Opus-tier models, which keeps per-request cost down without changing prompt format. For tasks that require deep multi-step reasoning, large-scale code refactoring, or long autonomous agent runs, a higher Claude tier is the more typical choice.

Frequently Asked Questions

How much does Claude Haiku 5.5 cost?

Pricing depends on the provider you use and the pricing model — input versus output tokens, batch versus real-time, and any caching or committed-use discounts. Rates also change over time. See the pricing table on this page for current per-provider figures.

What is Claude Haiku 5.5 best used for?

It is best suited to high-volume, latency-sensitive text tasks: classification, extraction, summarization, tagging, routing, and short response generation. It is commonly used as the worker or triage model in pipelines where a larger Claude tier handles the harder cases.

How fast is Claude Haiku 5.5?

Artificial Analysis measured roughly 137.8 output tokens per second with a time to first token of about 4,697 ms. The high generation rate suits long streaming outputs; the initial latency is more noticeable on very short exchanges. Actual numbers vary by serving provider and load.

How does Claude Haiku 5.5 compare to Anthropic's Sonnet and Opus models?

Haiku is Anthropic's compact tier, intended for speed and volume, while Sonnet and Opus tiers are aimed at more demanding reasoning, coding, and agentic workloads. If your task needs multi-step reasoning or long autonomous runs, a higher tier is usually the better fit; if it is a repetitive, well-scoped task run many times, Haiku is the typical choice.

What context window does Claude Haiku 5.5 support?

We do not currently track a confirmed context window for this model entry, so we do not list one. Check Anthropic's official model documentation or your serving provider's model card for the authoritative limit.