Claude Opus 5 is a chat model from Anthropic in the Claude Opus line, the company's naming tier for its largest, highest-capability models.
| Provider | Input / 1M | Output / 1M | Cached / 1M |
|---|---|---|---|
| $2.50 | $12.50 | $0.250 | |
| $2.50 | $12.50 | $0.250 | |
| $5.00 | $25.00 | - | |
| $5.00 | $25.00 | $0.500 | |
| $5.00 | $25.00 | $0.500 | |
| $5.00 | $25.00 | $0.500 |
Prices updated daily. Last check: Sep 8, 2026
Input, output, and batch rates, plus alternatives, for one provider at a time.
Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.
Claude Opus 5 fits the workloads teams reserve for their top capability tier: multi-file code changes and reviews that require holding several constraints at once, analysis over long or dense source material, drafting and critiquing technical or legal text, and agent runs where an early mistake compounds across later steps. Its measured profile — around 53 output tokens per second with a roughly 9.6 second time to first token — argues against using it for autocomplete, live voice, or any UI where the user watches an empty box, and in favor of asynchronous and batch-style jobs: overnight document processing, evaluation and grading pipelines, CI-triggered code analysis, or a router's escalation path when a Sonnet- or Haiku-class model returns a low-confidence answer. For high-volume classification, extraction, or routing, a lighter model in the Claude family is usually the more economical choice, with Opus 5 handling only the residual hard cases.
Pricing depends on which provider is serving the model and which pricing mode you use — on-demand per-token rates, batch discounts, and cached-input rates can all differ, and they change over time. See the pricing table on this page for the current figures across providers.
It is aimed at the hardest slice of a workload: complex coding and refactoring, long-document analysis, and multi-step agentic tasks where a wrong intermediate step is costly. Its high time to first token (~9.6 seconds measured) makes asynchronous, batch, and escalation-path usage a better fit than latency-critical interactive features.
Opus is Anthropic's naming tier for its largest models, so Claude Opus 5 is positioned above Sonnet-class (balanced cost and speed) and Haiku-class (lightweight, high-volume) siblings in the same family. The common pattern is to route most traffic to a smaller tier and reserve Opus 5 for requests that the smaller models handle unreliably.
Third-party measurements put it at roughly 53 output tokens per second with a time to first token of about 9.6 seconds. Once streaming begins the pace is workable for reading, but the initial delay is long by interactive-chat standards, so most real-time products either show interim UI state or route those requests to a faster model.
Anthropic delivers Claude models through hosted APIs and partner clouds rather than as downloadable weights, so access is via an API provider. The pricing table on this page lists the providers we track for this model.
We do not have a confirmed context window recorded for this entry, so we would rather not guess. Check Anthropic's official model documentation, or your provider's model card, for the exact input and output limits.