TokenRateCalc ← Back to LLM cost calculator
◆ Explainer

Claude Fable 5.1's Real Price Cut Isn't the Headline Rate, It's Cache Reads

Input and output pricing is unchanged from Fable 5. Cache reads dropped 75%, and that's the line that actually moves your bill, if you're caching at all.

Published Sep 11, 2026 (ET) · Updated Sep 24, 2026 (ET) · 6 min read · Confirmed via anthropic.com and platform.claude.com/docs
Sample bill composition — same task, before and after
Total bill shrinks from $3.00 to $1.50; the cache-read share does the work
Fable 5$3.00
Fable 5.1$1.50
Cache reads (the line that changed)
Input + output (unchanged, $1.00 both times)

Claude Fable 5.1, released September 1, 2026, costs the same as Fable 5 on paper: $10 per million input tokens, $50 per million output. Most coverage stopped there. The actual pricing change is one line down: cache reads dropped 75%, from $1.00 to $0.25 per million tokens. For agentic workloads that repeatedly re-read the same context, that's the number that moves your bill, not the headline rate.

01What Actually Changed

Every other line on Fable 5.1's rate card is identical to Fable 5:

Rate
Input$10.00 / MTok
Output$50.00 / MTok
Cache write (5-min)$12.50 / MTok
Cache write (1-hour)$20.00 / MTok
Cache read$0.25 / MTok (was $1.00)
Batch (input / output)$5.00 / $25.00

Update, Sep 23: the batch rate above was correct at original publication but had not yet been independently verified against Anthropic's own pricing page at the time — it was the standard 50% batch discount carried over from Fable 5's confirmed rate. It's now directly confirmed on Anthropic's official pricing docs rather than inferred.

Most current Claude models price a cache read at 10% of their base input rate. Fable 5.1 broke that pattern first: its cache read sits at 2.5% of base input. Update, Sep 24: Fable 5.1 is no longer the only exception. Claude Opus 5.5, launched September 22, prices its own cache read at 5% of input, half the standard ratio but still not as steep as Fable 5.1's 2.5%. That makes three different cache-read ratios now live across the current Claude lineup (2.5%, 5%, and the standard 10%), rather than one deliberate exception against an otherwise-uniform rule.

One side effect worth knowing: Fable 5.1's cache reads ($0.25) are cheaper than Claude Opus 5's ($0.50), even though Fable 5.1's base input price is double Opus 5's. Update, Sep 24: that comparison no longer holds against Opus 5.5, though. At $0.20 per million tokens, Opus 5.5's cache read actually undercuts Fable 5.1's $0.25 in absolute dollar terms, despite Fable 5.1 having the deeper discount as a percentage of its own (much higher) input rate. Fable 5.1's distinction is discount depth, not the lowest sticker price on the lineup. A second detail not widely reported: US-only inference, for workloads that need to stay in-region, runs at a 1.1x uplift on input and output tokens, though the cache-read rate itself isn't affected by that surcharge.

Run your own cache-heavy workload through the calculator

Model your real cache-read share, not a rule of thumb
Estimate your Fable 5.1 costs →

02Why This Matters More For Some Workloads Than Others

Cache reads only cost you money if you're actually caching. A one-off question with no reused context sees zero benefit from this change, since there's nothing being re-read.

Agentic work is the opposite case. A coding agent that keeps a system prompt, repo structure, and tool definitions cached across dozens of tool calls in one session is re-reading that same prefix over and over. The more of your token spend that's cache reads rather than fresh input or output, the more this cut is worth to you.

A worked example: an agent session with a 100,000-token cached prefix (system prompt, project context, tool specs), reused across 20 tool calls, each adding 1,000 tokens of new input and generating 800 tokens of output.

Fable 5Fable 5.1
Cache read (100K × 20 calls)$2.00$0.50
New input (1K × 20 calls)$0.20$0.20
Output (800 × 20 calls)$0.80$0.80
Total$3.00$1.50

In this example, cache reads made up two-thirds of the original bill, so cutting that line 75% cuts the total in half. A session with a smaller cached prefix, or fewer repeated calls, would see a smaller effect, since less of the bill was ever coming from cache reads to begin with. Anthropic's own figures, drawn from internal usage data, put the range at roughly 25% savings on typical workloads and up to roughly 45% on the most cache-heavy agentic ones. Where your workload lands in that range depends entirely on what share of your token spend is currently cache reads.

03The Complication Most Coverage Left Out

Independent benchmarking adds a wrinkle

The worked example above assumes output volume stays constant between Fable 5 and Fable 5.1. Independent testing from Artificial Analysis found that isn't guaranteed: they measured Fable 5.1 using roughly 1.7x the output tokens of Fable 5 to reach comparable results on the same task. Output is billed at $50/MTok regardless of cache savings, so a task that genuinely needs 1.7x the output tokens could cost more overall on Fable 5.1 than Fable 5, even with cache reads 75% cheaper. This doesn't cancel the cache-read discount, a well-cached, low-output task still gets cheaper, but it means the 25-45% figures are best case, not guaranteed, and the only way to know where your specific workload actually lands is to measure it.

04What to Check in Your Own Setup

If you're already running agentic workloads on Fable 5 or Fable 5.1, the number worth pulling from your own usage logs is: what percentage of your billed tokens are cache reads versus fresh input and output? That percentage, not the headline rate, tells you how much this change is actually worth to you. Worth checking output token counts too, given the Artificial Analysis finding above, rather than assuming they'll match what you saw on Fable 5.

If your cache-read share is low, it's worth asking why. A large, stable system prompt or tool-definition block that isn't currently cached is a common miss, and caching it is free upside regardless of which model you're on.

Estimate your actual savings with your real token counts

Compare Fable 5.1 against Opus 5.5 and Sonnet 5 on your workload
Open the calculator →

Frequently Asked Questions

Did Claude Fable 5.1's base pricing change?

No. Input and output remain $10 and $50 per million tokens, identical to Fable 5. Only the cache-read rate changed.

How much will I actually save?

Depends on what share of your current bill is cache reads. Anthropic reports roughly 25% savings on typical workloads and up to roughly 45% on highly agentic ones. A workload with no caching sees no savings from this change, and independent benchmarking suggests Fable 5.1 may use more output tokens per task than Fable 5, which works against the cache savings.

Is Fable 5.1's cache-read price lower than other Claude models?

It's cheaper than Claude Opus 5's $0.50, and it still has the steepest cache-read discount of any current Claude model as a percentage of input (2.5%, versus the usual 10%). But as of Claude Opus 5.5's September 22, 2026 launch, it's no longer the lowest absolute cache-read price on Claude: Opus 5.5 undercuts it at $0.20 per million tokens (a 5% ratio on a lower $4 input rate), matching Sonnet 5's existing $0.20 figure. Fable 5.1's real distinction is discount depth relative to its own input price, not the lowest sticker price outright.

What's the minimum prompt size to use caching at all?

512 tokens, unchanged from before. Anything shorter than that can't be cached regardless of model.

Cache-read pricing ($0.25, down from $1.00) and the 25%/45% savings figures confirmed directly via Anthropic's own product page (anthropic.com/claude/fable) and Claude's official account, cross-checked against 8 independent sources as of September 10, 2026. The Artificial Analysis output-token finding is independent third-party benchmarking, not an Anthropic figure, cited via secondary reporting rather than the original benchmark directly. The worked example table is calculated independently for this article rather than reused from any source's math. Updated Sep 23, 2026: the Batch API rate ($5.00/$25.00) is now independently confirmed directly against Anthropic's official pricing documentation (platform.claude.com/docs/en/about-claude/pricing) — it was accurate at original publication but had been carried over from the standard 50% discount rather than separately verified at that time. Updated Sep 24, 2026: Claude Opus 5.5 launched September 22 at a $0.20 cache-read rate, confirmed against platform.claude.com/docs/en/about-claude/pricing — this piece's original claims that Fable 5.1 was "the only" 10%-pattern exception and had the lowest absolute Claude cache-read price were corrected to account for it. See the live calculator for current Fable 5.1 pricing.