Claude Fable 5.1, released September 1, 2026, costs the same as Fable 5 on paper: $10 per million input tokens, $50 per million output. Most coverage stopped there. The actual pricing change is one line down: cache reads dropped 75%, from $1.00 to $0.25 per million tokens. For agentic workloads that repeatedly re-read the same context, that's the number that moves your bill, not the headline rate.
01What Actually Changed
Every other line on Fable 5.1's rate card is identical to Fable 5:
| Rate | |
|---|---|
| Input | $10.00 / MTok |
| Output | $50.00 / MTok |
| Cache write (5-min) | $12.50 / MTok |
| Cache write (1-hour) | $20.00 / MTok |
| Cache read | $0.25 / MTok (was $1.00) |
| Batch (input / output) | $5.00 / $25.00 |
Update, Sep 23: the batch rate above was correct at original publication but had not yet been independently verified against Anthropic's own pricing page at the time — it was the standard 50% batch discount carried over from Fable 5's confirmed rate. It's now directly confirmed on Anthropic's official pricing docs rather than inferred.
Most current Claude models price a cache read at 10% of their base input rate. Fable 5.1 broke that pattern first: its cache read sits at 2.5% of base input. Update, Sep 24: Fable 5.1 is no longer the only exception. Claude Opus 5.5, launched September 22, prices its own cache read at 5% of input, half the standard ratio but still not as steep as Fable 5.1's 2.5%. That makes three different cache-read ratios now live across the current Claude lineup (2.5%, 5%, and the standard 10%), rather than one deliberate exception against an otherwise-uniform rule.
One side effect worth knowing: Fable 5.1's cache reads ($0.25) are cheaper than Claude Opus 5's ($0.50), even though Fable 5.1's base input price is double Opus 5's. Update, Sep 24: that comparison no longer holds against Opus 5.5, though. At $0.20 per million tokens, Opus 5.5's cache read actually undercuts Fable 5.1's $0.25 in absolute dollar terms, despite Fable 5.1 having the deeper discount as a percentage of its own (much higher) input rate. Fable 5.1's distinction is discount depth, not the lowest sticker price on the lineup. A second detail not widely reported: US-only inference, for workloads that need to stay in-region, runs at a 1.1x uplift on input and output tokens, though the cache-read rate itself isn't affected by that surcharge.
Run your own cache-heavy workload through the calculator
Model your real cache-read share, not a rule of thumb02Why This Matters More For Some Workloads Than Others
Cache reads only cost you money if you're actually caching. A one-off question with no reused context sees zero benefit from this change, since there's nothing being re-read.
Agentic work is the opposite case. A coding agent that keeps a system prompt, repo structure, and tool definitions cached across dozens of tool calls in one session is re-reading that same prefix over and over. The more of your token spend that's cache reads rather than fresh input or output, the more this cut is worth to you.
A worked example: an agent session with a 100,000-token cached prefix (system prompt, project context, tool specs), reused across 20 tool calls, each adding 1,000 tokens of new input and generating 800 tokens of output.
| Fable 5 | Fable 5.1 | |
|---|---|---|
| Cache read (100K × 20 calls) | $2.00 | $0.50 |
| New input (1K × 20 calls) | $0.20 | $0.20 |
| Output (800 × 20 calls) | $0.80 | $0.80 |
| Total | $3.00 | $1.50 |
In this example, cache reads made up two-thirds of the original bill, so cutting that line 75% cuts the total in half. A session with a smaller cached prefix, or fewer repeated calls, would see a smaller effect, since less of the bill was ever coming from cache reads to begin with. Anthropic's own figures, drawn from internal usage data, put the range at roughly 25% savings on typical workloads and up to roughly 45% on the most cache-heavy agentic ones. Where your workload lands in that range depends entirely on what share of your token spend is currently cache reads.
03The Complication Most Coverage Left Out
Independent benchmarking adds a wrinkleThe worked example above assumes output volume stays constant between Fable 5 and Fable 5.1. Independent testing from Artificial Analysis found that isn't guaranteed: they measured Fable 5.1 using roughly 1.7x the output tokens of Fable 5 to reach comparable results on the same task. Output is billed at $50/MTok regardless of cache savings, so a task that genuinely needs 1.7x the output tokens could cost more overall on Fable 5.1 than Fable 5, even with cache reads 75% cheaper. This doesn't cancel the cache-read discount, a well-cached, low-output task still gets cheaper, but it means the 25-45% figures are best case, not guaranteed, and the only way to know where your specific workload actually lands is to measure it.
04What to Check in Your Own Setup
If you're already running agentic workloads on Fable 5 or Fable 5.1, the number worth pulling from your own usage logs is: what percentage of your billed tokens are cache reads versus fresh input and output? That percentage, not the headline rate, tells you how much this change is actually worth to you. Worth checking output token counts too, given the Artificial Analysis finding above, rather than assuming they'll match what you saw on Fable 5.
If your cache-read share is low, it's worth asking why. A large, stable system prompt or tool-definition block that isn't currently cached is a common miss, and caching it is free upside regardless of which model you're on.
Estimate your actual savings with your real token counts
Compare Fable 5.1 against Opus 5.5 and Sonnet 5 on your workload