Skip to main content
Anthropic

Claude Fable 5.1

Claude Fable 5.1 is a chat-oriented large language model from Anthropic, tracked on this page with measured output throughput of roughly 61 tokens per second.

Input from
$5.00 / 1M tokens
across 3 providers

API Pricing

Cheapest on Anthropic 38% below avg
ProviderInput / 1MOutput / 1MCached / 1M
$5.00$25.00$0.125
$5.00$25.00$0.125
$10.00$50.00$0.250
$10.00$50.00$0.250
$10.00$50.00$0.250

Prices updated daily. Last check: Sep 2, 2026

Claude Fable 5.1 pricing by provider

Input, output, and batch rates, plus alternatives, for one provider at a time.

Performance & Benchmarks

Source: Artificial Analysis →
Intelligence
65.7 / 100
Coding
81.6 / 100
Output Speed
67.6 t/s
Latency (TTFT)
244.4s

Reasoning & Knowledge

  • GPQA Diamond93.7%
  • Humanity's Last Exam59.1%

Coding

  • SciCode62.0%

Agentic & Tool Use

  • Terminal-Bench v2.191.4%
  • τ-bench Banking47.2%

Instruction & Long Context

  • Long-Context Reasoning80.0%

Benchmarks measured Sep 2026. Scores are independent evaluations, not vendor-reported.

Model Details

General

Creator
Anthropic
Modalities
Text

Capabilities

Tool Calling
No
Open Source
No

Strengths & Limitations

Strengths

  • Measured sustained output of about 61.4 tokens per second, adequate for long-form drafting and batch generation
  • Part of Anthropic's Claude line, so it slots into existing Anthropic-oriented tooling and prompting conventions
  • Independent third-party serving benchmarks (Artificial Analysis) are available rather than only vendor-reported numbers
  • Chat-formatted interface suits multi-turn assistants, summarization, and rewriting workloads
  • Listed across the providers shown in the pricing table on this page, allowing side-by-side cost comparison

Limitations

  • Time to first token measured at roughly 16.1 seconds, which is poor for interactive or streaming-first user interfaces
  • We do not currently track a confirmed context window for this model, so long-document limits must be checked with the provider
  • No reasoning, coding, or knowledge benchmark scores are recorded in our database for this entry
  • Modality support beyond text chat is not confirmed in our data
  • Output throughput of ~61 tokens/sec trails faster small models used for high-volume, low-latency tasks

Key Features

Chat-completion (multi-turn conversational) interface
Anthropic Claude model family lineage
Measured output throughput of ~61.4 tokens/second
Measured time to first token of ~16.1 seconds
Third-party performance data sourced from Artificial Analysis
Multi-provider availability tracked in the pricing table on this page

About Claude Fable 5.1

Claude Fable 5.1 is a text chat model created by Anthropic and listed in the Claude line of models. Our catalog entry for it is intentionally narrow: we record it as a conversational (chat) model from Anthropic, and we carry independent serving measurements for it, but we do not currently track a confirmed context window, modality list, or parameter count for this entry. Where those fields are blank here, treat them as unknown rather than absent — Anthropic's own model documentation is the authoritative source for the full specification sheet. On the performance side, the figures we do have come from Artificial Analysis. Claude Fable 5.1 has been measured at approximately 61.4 output tokens per second, with a time to first token of roughly 16.1 seconds. That combination — a moderate steady-state generation rate paired with a long initial delay before the first token appears — is the profile typically seen from models that perform substantial work before emitting visible output, or from endpoints serving long prompts. For interactive interfaces, the first-token latency is the number that will dominate perceived responsiveness; for batch or background generation, the sustained tokens-per-second figure matters more. In practice, Claude Fable 5.1 is used the way other Anthropic chat models are: multi-turn conversation, drafting and rewriting text, summarization, and question answering against supplied material. Because throughput here is mid-range and time to first token is high, it fits asynchronous or long-form work more comfortably than latency-sensitive chat widgets, autocomplete, or voice front-ends. Availability and cost differ by provider — see the pricing table on this page for the providers currently serving it.

Common Use Cases

Claude Fable 5.1 suits text workloads where total completion quality matters more than the delay before the first character appears: long-form drafting, document summarization, rewriting and editing passes, question answering over supplied context, and background content generation in queues or cron jobs. The measured ~16 second time to first token makes it a weak fit for typeahead, live chat widgets, voice agents, or any UI where a user is watching a cursor blink; if you do serve it interactively, plan for a loading state rather than token-by-token streaming from the outset. Because our entry does not record confirmed context-window or vision support, validate those requirements directly with your provider before committing to document-heavy or image-input pipelines. For teams already standardized on Anthropic prompting patterns, it is worth evaluating against other Claude entries in this catalog on a per-task basis.

Frequently Asked Questions

How much does Claude Fable 5.1 cost?

Pricing varies by provider and by pricing model — some bill separately for input and output tokens, others offer cached-input rates or dedicated capacity. Because rates change frequently, we do not quote figures in this write-up. Check the pricing table on this page for the current per-provider rates.

What is Claude Fable 5.1 best used for?

It is a chat model, so it fits multi-turn conversation, summarization, drafting, rewriting, and question answering. Its measured sustained throughput of roughly 61 tokens per second works well for long-form and batch generation, while its high time to first token makes it less suited to latency-sensitive interactive interfaces.

How fast is Claude Fable 5.1?

Artificial Analysis measurements put it at about 61.4 output tokens per second with a time to first token of roughly 16.1 seconds. The generation rate is mid-range; the first-token latency is on the slow side, so perceived responsiveness in a live chat UI will be limited by that initial wait.

What context window does Claude Fable 5.1 support?

We do not currently have a confirmed context window recorded for this model. That means unknown rather than small — check Anthropic's documentation or your chosen provider's model page for the supported input length before building around long documents.

Does Claude Fable 5.1 accept images or other non-text input?

Our database records it as a chat model and does not confirm additional input modalities. Absence of that field in our catalog is not evidence that the capability is missing, so verify multimodal support with the provider you plan to use.