Claude Haiku 5.5
Claude Haiku 5.5 is a chat model from Anthropic in the Claude Haiku line, the company's small, speed-oriented tier of Claude models.
API Pricing
Cheapest on OpenRouter — 33% below avg| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $0.050 | $0.250 | $0.0050 | |
| $0.100 | $0.500 | $0.010 |
Prices updated daily. Last check: Oct 8, 2026
Claude Haiku 5.5 pricing by provider
Input, output, and batch rates, plus alternatives, for one provider at a time.
Performance & Benchmarks
Source: Artificial Analysis →Reasoning & Knowledge
- Humanity's Last Exam44.4%
Coding
- SciCode55.0%
Instruction & Long Context
- Long-Context Reasoning82.7%
Benchmarks measured Oct 2026. Scores are independent evaluations, not vendor-reported.
Model Details
General
- Creator
- Anthropic
- Modalities
- Text
Capabilities
Strengths & Limitations
Strengths
- Part of Anthropic's Haiku tier, which is designed around low latency and high request volume rather than maximum reasoning depth
- Measured output speed of roughly 137.8 tokens per second (Artificial Analysis), suitable for streaming responses to end users
- Positioned below Sonnet and Opus in the Claude lineup, giving teams a cheaper tier to route simple requests to within the same model family
- Shares the Claude prompting conventions and API surface, so prompts and integrations can often be reused across Claude tiers
- Available through multiple serving providers, allowing price, region, and rate-limit comparison
- Well suited to being the high-throughput worker model in a multi-model pipeline
Limitations
- Measured time to first token of about 4,697 ms is high relative to the model's generation speed, which affects short interactive turns
- As a Haiku-tier model, it is not the tier Anthropic targets at the hardest reasoning, long-horizon agentic, or complex coding workloads
- Context window, modality support, and tool-calling details are not tracked in our database for this entry — verify against Anthropic's documentation before building against them
- Benchmark accuracy scores for this model are not available in our dataset, so quality comparisons against peers must come from your own evaluations
- Throughput and latency figures vary by serving provider, so the measured numbers may not match what you observe on a specific endpoint
Key Features
About Claude Haiku 5.5
Common Use Cases
Claude Haiku 5.5 fits workloads where request volume is high and each individual call is relatively bounded: classifying support tickets, extracting structured fields from documents, summarizing conversation logs, tagging or moderating user-generated text, drafting short responses, and acting as the routing or triage layer in front of a larger model. Its measured generation rate of around 138 tokens per second makes it reasonable for streaming chat responses where the user watches text appear, though the measured initial latency means very short, snappy exchanges may feel less immediate than the raw token rate suggests. Teams already standardized on Claude often use a Haiku-tier model as the default and escalate only the hard requests to Sonnet- or Opus-tier models, which keeps per-request cost down without changing prompt format. For tasks that require deep multi-step reasoning, large-scale code refactoring, or long autonomous agent runs, a higher Claude tier is the more typical choice.
Frequently Asked Questions
How much does Claude Haiku 5.5 cost?
Pricing depends on the provider you use and the pricing model — input versus output tokens, batch versus real-time, and any caching or committed-use discounts. Rates also change over time. See the pricing table on this page for current per-provider figures.
What is Claude Haiku 5.5 best used for?
It is best suited to high-volume, latency-sensitive text tasks: classification, extraction, summarization, tagging, routing, and short response generation. It is commonly used as the worker or triage model in pipelines where a larger Claude tier handles the harder cases.
How fast is Claude Haiku 5.5?
Artificial Analysis measured roughly 137.8 output tokens per second with a time to first token of about 4,697 ms. The high generation rate suits long streaming outputs; the initial latency is more noticeable on very short exchanges. Actual numbers vary by serving provider and load.
How does Claude Haiku 5.5 compare to Anthropic's Sonnet and Opus models?
Haiku is Anthropic's compact tier, intended for speed and volume, while Sonnet and Opus tiers are aimed at more demanding reasoning, coding, and agentic workloads. If your task needs multi-step reasoning or long autonomous runs, a higher tier is usually the better fit; if it is a repetitive, well-scoped task run many times, Haiku is the typical choice.
What context window does Claude Haiku 5.5 support?
We do not currently track a confirmed context window for this model entry, so we do not list one. Check Anthropic's official model documentation or your serving provider's model card for the authoritative limit.