Frontier model comparison · Anthropic · two announcements, two system cards, one rate card, six outside evaluators and 23 dated user reports
Claude Opus 5.5 vs Claude Fable 5.1: level at their default settings, and Fable costs 2.9 times as much per task
At each model's default effort, Artificial Analysis's Intelligence Index v4.3.2 scores Claude Opus 5.5 at 51.24 for $1.34 per task and Claude Fable 5.1 at 51.15 for $3.91, and Anthropic's own docs say to start with Opus 5.5. Fable 5.1 still ranks first on two agent leaderboards that have not yet listed Opus 5.5. Both models' Anthropic documents, one rate card, six outside evaluators and 23 dated user reports, read on 23 September 2026.
Our verdict
Run coding, knowledge work and most agent jobs on Claude Opus 5.5, which Artificial Analysis scores above Claude Fable 5.1 at every effort setting but low for less per task, and move a job to Fable 5.1 only when Opus 5.5 at high or xhigh effort fails your own evaluations on long-running agent work, the one area where outside leaderboards rank Fable first and have not yet measured Opus 5.5.
Claude Opus 5.5 vs Claude Fable 5.1 specs
- Released
- Claude Opus 5.5: 22 September 2026, per Anthropic's announcement. Claude Fable 5.1: 1 September 2026, per Anthropic's news index and the system card cover
- API model ID
- claude-opus-5-5 / claude-fable-5-1, per Anthropic's models overview
- Context window
- Both 1M tokens, per the models overview
- Maximum output
- Both 128K tokens, per the models overview
- Knowledge cutoff
- Both June 2026, per the models overview and both system cards
- Effort settings and default
- Both: low, medium, high, xhigh and max. API default: Opus 5.5 medium, Fable 5.1 high. Fable 5.1 defaults to high in Claude Code and to medium in Claude Cowork and Claude.ai, per its announcement; no product default is stated for Opus 5.5 in what we read
- Thinking
- Both adaptive and always on, per the models overview; the Opus 5.5 system card says thinking cannot be disabled in the API for either
- Comparative latency
- Opus 5.5 Moderate, Fable 5.1 Slower, per the models overview
- Safeguards
- Both: cyber classifiers fall back to Claude Opus 4.8, biology and frontier-AI development classifiers to Claude Opus 5; on the API fallback must be set up by the developer on both. Anthropic calls Opus 5.5's the same class as Fable 5.1's
- Data retention
- Opus 5.5: available with zero data retention, per its announcement. Fable 5.1: 30-day retention for safety monitoring by default, zero retention for eligible enterprise customers, per its product page
- Availability
- Both on the Claude API, Amazon Web Services, Google Cloud and Microsoft Azure, per each announcement. Fable 5.1 on Pro, Max, Team and Enterprise plans, per its product page
Claude Opus 5.5 vs Claude Fable 5.1 pricing
- Input, per 1M tokens
- Claude Opus 5.5 $4.00 / Claude Fable 5.1 $10.00. Fable is 2.5x (our division)
- Output, per 1M tokens
- Claude Opus 5.5 $20.00 / Claude Fable 5.1 $50.00. Fable is 2.5x (our division)
- Cache read, per 1M tokens
- Claude Opus 5.5 $0.20 / Claude Fable 5.1 $0.25. Fable is 1.25x (our division)
- Cache write, 5-minute TTL, per 1M tokens
- Claude Opus 5.5 $5.00 / Claude Fable 5.1 $12.50. Fable is 2.5x (our division)
- US-only inference
- Both 1.1x on input and output: Claude Opus 5.5 $4.40 and $22.00, Claude Fable 5.1 $11.00 and $55.00 (our arithmetic)
- Batch
- The pricing page states a 50% saving with batch processing and names no model. If it applies as it reads: Claude Opus 5.5 $2.00 and $10.00, Claude Fable 5.1 $5.00 and $25.00 (our arithmetic)
- Fast mode
- Claude Opus 5.5 $8.00 and $40.00, 2x standard. Claude Fable 5.1: the pricing page we read gives no fast-mode rate
- Requests rerouted by safeguards
- Claude Fable 5.1: not billed at Fable rates, per its product page. Claude Opus 5.5: no matching line in what we read
- Cost per index task at each model's API default
- Claude Opus 5.5 at medium $1.34 / Claude Fable 5.1 at high $3.91, on Artificial Analysis Intelligence Index v4.3.2. The pairing of the two defaults is ours
Rate card read at the vendor’s own documentation on Sep 23, 2026.
Claude Opus 5.5 and Claude Fable 5.1 score almost the same at their default settings, and Opus 5.5 gets there for about a third of the cost per task. On Artificial Analysis's Intelligence Index v4.3.2, Opus 5.5 at medium effort, its default, scores 51.24 for $1.34 per index task; Fable 5.1 at high, its default, scores 51.15 for $3.91. Anthropic's own advice is to "start with Claude Opus 5.5 for most workloads".
Fable 5.1 still ranks first where Opus 5.5 has not been measured. Artificial Analysis's Coding Agent Index v1.5 and LMArena's Agent board both put Fable 5.1 at the top, and neither had an Opus 5.5 entry on 23 September 2026, so those first places have no Opus 5.5 figure to be set against.
We ran neither model. Both are Anthropic's, so every vendor figure is one lab describing its own products, from documents read on 23 September 2026: the two announcements and system cards, Fable's product page, the models overview and the pricing page. Six outside evaluators were read the same day, and the user reports below date from 1 to 23 September. Every number names its publisher, benchmark version and effort setting.
Is Claude Opus 5.5 better than Claude Fable 5.1?
On Artificial Analysis's Intelligence Index v4.3.2, Claude Opus 5.5 scores higher at four of the five effort settings, by 2.3 to 4.3 points at max, xhigh, high and medium, and costs less per index task at each. Claude Fable 5.1 scores higher only at low effort, 46.82 against 42.31.
| Effort | Opus 5.5 score | Opus 5.5 cost per index task | Fable 5.1 score | Fable 5.1 cost per index task |
|---|---|---|---|---|
| max | 57.62 | $5.98 | 53.35 | $7.63 |
| xhigh | 55.99 | $3.46 | 53.20 | $5.98 |
| high | 53.58 | $1.82 | 51.15 | $3.91 |
| medium | 51.24 | $1.34 | 48.92 | $2.98 |
| low | 42.31 | $0.55 | 46.82 | $2.37 |
Artificial Analysis Intelligence Index v4.3.2, five model pages per model, read 23 September 2026. Unrounded scores are from the data behind the charts; the pages show 58, 56, 54, 51 and 42 for Opus 5.5 and 53, 53, 51, 49 and 47 for Fable 5.1. Both models ran with Anthropic's default fallback enabled. Medium is Opus 5.5's default, high is Fable 5.1's.
Every Fable 5.1 row has an Opus 5.5 row that scores higher for less. Opus 5.5 at high (53.58 for $1.82) edges Fable 5.1 at max (53.35 for $7.63) at 24% of the cost, and Opus 5.5 at medium beats Fable's one winning row, low, for $1.34 against $2.37. At the same setting Fable 5.1 costs 1.3 to 4.3 times as much per index task, our division.
The index combines ten evaluations, and one of them can matter more to a job than the total. Outside it the rows below mostly point the same way, with Fable 5.1 ahead on Anthropic's own OfficeQA runs by small margins. METR, which evaluated Opus 5.5 before release, called it "an incremental improvement above Fable 5.1 on our quantitative evaluations" and published no numbers.
What does Anthropic recommend, Opus 5.5 or Fable 5.1?
Anthropic recommends Claude Opus 5.5 first and Claude Fable 5.1 as the next step. Its models overview, read 23 September 2026, says: "If you're unsure which model to use, start with Claude Opus 5.5 for most workloads. Use Claude Fable 5.1 for demanding reasoning and long-horizon agentic work, or when your evals on Claude Opus 5.5 at higher effort still fall short."
So the lab tells developers to try the cheaper of its two models first. The Opus 5.5 launch post says it "performs at the level of Claude Fable 5.1 on most work", then qualifies its own table, where Opus 5.5 leads Fable 5.1 on all nine rows: "In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest." Fable's product page, read the same day, still called Fable 5.1 "our most capable generally available model".
Is Claude Fable 5.1 worth 2.5 times the price of Opus 5.5?
Claude Fable 5.1 lists at $10 per million input tokens and $50 per million output, 2.5 times Claude Opus 5.5's $4 and $20, and Artificial Analysis's index scores it higher only at low effort. Cache writes keep the 2.5 times ratio ($12.50 against $5); cache reads do not, at $0.25 against $0.20.
| Per million tokens | Claude Opus 5.5 | Claude Fable 5.1 | Fable as a multiple |
|---|---|---|---|
| Input | $4.00 | $10.00 | 2.5x |
| Output | $20.00 | $50.00 | 2.5x |
| Cache write, 5-minute TTL | $5.00 | $12.50 | 2.5x |
| Cache read | $0.20 | $0.25 | 1.25x |
claude.com/pricing, read 23 September 2026 in US dollars. Multiples are our division.
The narrow cache-read gap dates from Fable 5.1's launch on 1 September, when Anthropic cut it from Fable 5's $1: "Cache reads now cost 75% less, or $0.25 per million tokens."
US-only inference costs 1.1 times on both ($4.40 and $22.00 against $11.00 and $55.00, our arithmetic), and the page's 50% batch saving names no model. Fast mode is priced for Opus 5.5 only, at $8 and $40. Fable's product page adds that users are not "charged Fable prices for rerouted requests", the flagged queries another model answers; we found no such line for Opus 5.5.
What effort does each model default to, and what does it do to the bill?
On the API, Claude Opus 5.5 defaults to medium effort and Claude Fable 5.1 to high, per Anthropic's models overview. At those defaults Artificial Analysis's index v4.3.2 has them level, 51.24 against 51.15, and Fable 5.1 costs $3.91 per index task against $1.34, 2.9 times as much, our pairing and division.
Effort sets how long the model reasons before answering, and so how many output tokens a task bills. Both models take the same five settings, and both think adaptively with no off switch. A developer who swaps the model ID and leaves effort unset therefore changes two things: a rate 2.5 times higher and one more notch of reasoning.
Fable's defaults also vary by product. Its launch post says Fable 5.1 "defaults to High effort in Claude Code, and to Medium in Claude Cowork and on Claude.ai". We found no statement of Opus 5.5's default inside those products.
One user's harness shows a wider gap. gwd, 22 September 2026, running a patch-review harness over 12 patches with 14 issues between them, reported: "Opus 5.5: Found 8/14 issues. Total cost: $15.40. Fable 5.1: Found 7/14 issues. Total cost: $66.34". That is 4.3 times the cost for one fewer issue, on the user's figures, with no effort setting given.
Does prompt caching narrow the price gap between Opus 5.5 and Fable 5.1?
For the same tokens, yes: the more of a request is read from cache, the closer Fable 5.1's bill comes to 1.25 times Opus 5.5's, since cache reads are the one line not priced at 2.5 times. Three requests of our own land at 2.5, 2.0 and 1.4 times.
| Request, same tokens on both | Claude Opus 5.5 | Claude Fable 5.1 | Fable as a multiple |
|---|---|---|---|
| 50,000 new input, 5,000 output, nothing cached | $0.30 | $0.75 | 2.5x |
| 400,000 cached, 10,000 new input, 4,000 output | $0.20 | $0.40 | 2.0x |
| 900,000 cached, 2,000 new input, 1,000 output | $0.208 | $0.295 | 1.4x |
Our arithmetic from the list rates above, assuming a live cache written by an earlier call and leaving out cache writes.
Anthropic's own workload mix lands in the middle. Fable 5.1's launch post charts a typical and a highly agentic workload in index units; at Fable 5.1 prices, cache reads are 17 of 55 units in the agentic mix and 11 of 75 in the typical one. Repriced at Opus 5.5's rates with the same tokens, Fable costs about 1.9 times as much on the agentic mix and 2.2 times on the typical one, our arithmetic.
Same tokens is the assumption that breaks. At max effort Artificial Analysis counted about 119,000 output tokens per index task for Opus 5.5 against about 78,000 for Fable 5.1, so the per-task gap at max is only 1.3 times. At high the order flips (62 million output tokens across the index for Fable 5.1, 53 million for Opus 5.5) and the gap passes 2 times.
How do Claude Opus 5.5 and Fable 5.1 compare on benchmarks?
Where one publisher ran both on the same benchmark version, Claude Opus 5.5 scores higher on almost every row, including 1846 against 1735 Elo on GDPval-AA v2.1 and 59.6% against 52.0% on Terminal-Bench 4.0, both Artificial Analysis runs at max effort. Claude Fable 5.1's leads are narrow except at low effort, and are listed below.
| Benchmark and version | Setting | Who ran it | Claude Opus 5.5 | Claude Fable 5.1 |
|---|---|---|---|---|
| GDPval-AA v2.1 | max, both | Artificial Analysis | 1846 Elo | 1735 Elo |
| Terminal-Bench 4.0, mini-swe-agent harness | max, both | Artificial Analysis | 59.6% | 52.0% |
| Terminal-Bench 4.0, Claude Code harness | Opus 5.5 xhigh, Fable 5.1 max | Anthropic | 66.4% | 55.8% |
| Humanity's Last Exam | max, both | Artificial Analysis | 61.4% | 59.1% |
| CursorBench 4.0 | max, both | Cursor, reported by Anthropic | 57.8% | 51.8% |
| FrontierSWE v2 | max, both | Proximal, reported by Anthropic | 62.3% | 56.3% |
| AutomationBench v1.0.6 | max, both | Zapier | 40.0% | 31.4% |
| ARC-AGI-2, Semi-Private | each at its best: Opus 5.5 high, Fable 5.1 max | ARC Prize | 93.3% | 90.0% |
| OSWorld 2.0, partial / strict | max, both | Anthropic | 81.8% / 48.7% | 80.7% / 42.8% |
Artificial Analysis, Zapier and ARC Prize rows read at source on 23 September 2026, some from page data. Other rows from Anthropic's Opus 5.5 system card (pp. 174 to 212); we have not read Cursor's or Proximal's own pages.
Two rows need a caution. On Terminal-Bench, Anthropic's table shows Opus 5.5 at xhigh, its best result; at max it scored 64.8% (p. 178). On AutomationBench, Claude Opus 5 handled about 40% of Fable 5.1's tasks (260 of 657) and they count in the 31.4%, while Anthropic says Opus 5.5 ran without fallback. No Fable 5.1 cell in Anthropic's table is a rival lab's self-report: each is Anthropic's own run or an outside operator's.
Why do Fable 5.1's GDPval and CursorBench scores differ between sites?
Because they are on different versions of each benchmark. Claude Fable 5.1 scores 1853 on GDPval-AA v2 and 1735 on v2.1, and 73.4% on CursorBench 3.2.0 and 51.8% on 4.0. Neither pair is a change in the model, and only the v2.1 and 4.0 figures can be set beside Claude Opus 5.5.
| Benchmark and version | Claude Opus 5.5 | Claude Fable 5.1 | Where the figures appear |
|---|---|---|---|
| GDPval-AA v2 | not published | 1853 Elo | Fable 5.1 launch post and system card; Artificial Analysis, 1 September |
| GDPval-AA v2.1 | 1846 Elo | 1735 Elo | Opus 5.5 launch post; Artificial Analysis data |
| CursorBench 3.2.0 | not published | 73.4% | Fable 5.1 launch post and system card |
| CursorBench 4.0 | 57.8% | 51.8% | Opus 5.5 launch post and system card |
| Intelligence Index, a version before 4.3 | not listed | 66 | Artificial Analysis article, 1 September |
| Intelligence Index v4.3.2 | 57.62 | 53.35 | Artificial Analysis model pages |
All figures at max effort.
On GDPval-AA the scale moved. Artificial Analysis's version history says that for v2.1 "the Elo scale is now anchored to DeepSeek V4.1 Flash (max) at 1600". Set Fable's 1853 on v2 beside Opus 5.5's 1846 on v2.1 and Fable looks 7 points ahead; on the same v2.1 scale Opus 5.5 leads by 111.
CursorBench 3.2.0 and 4.0 are different task sets at different costs: Fable 5.1 at max costs $9.64 per task on 3.2.0, in the data behind Anthropic's chart, and $17.28 on 4.0. Both system cards warn that "previous system cards reported older versions of CursorBench, so the scores are not comparable".
The launch-day 66 is the same case. Artificial Analysis's 1 September article cites Terminal-Bench v2.1, which the publisher dropped in index version 4.3, so the 66 sits on an earlier index; that dating is our inference. On v4.3.2 the same setting scores 53.35.
Where does Claude Fable 5.1 beat Opus 5.5?
In a few places: by under two points at matched settings from medium to max, and by wider margins at low effort. In Anthropic's Opus 5.5 system card, Claude Fable 5.1 scores 80.2% on OfficeQA and 69.0% on OfficeQA Pro against 78.9% and 67.7% for Claude Opus 5.5, both at max effort over five runs (p. 208), and the two tie on Toolathlon at 77.8% (p. 211).
On Artificial Analysis's v4.3.2 data at matched settings, Fable 5.1 leads AA-LCR, the publisher's long-context reasoning test, at four of five settings (85.3% against 84.7% at max), and Humanity's Last Exam at xhigh, high and low (58.7% against 57.5% at xhigh), each by under two points. At low effort the gaps widen: 40.4% against 31.3% on Terminal-Bench 4.0, 1450 against 1224 Elo on GDPval-AA v2.1 and 27.7% against 17.7% on CritPt.
Fable's larger results sit on boards with no Opus 5.5 row. On the Coding Agent Index v1.5, Claude Code running Fable 5.1 at max with fallback scores 62.2, first of the eight entries in its chart, at a mean $12.39 per task; Opus 5.5 is in the page's model list with no agent entry. On LMArena's Agent Arena, board dated 15 September, claude-fable-5.1-max is first with a net improvement of 13.71% (plus or minus 1.72) over 13,320 sessions; Opus 5.5 is on no LMArena board. Epoch AI's Capability Index has Fable 5.1 second overall at 165.0 and no Opus 5.5 score. A first place without Opus 5.5 on the board says Fable beats the models listed, and no more.
Is Claude Opus 5.5 or Fable 5.1 better for coding?
On the coding benchmarks both have been run on, Claude Opus 5.5 scores higher at max effort: 57.8% against 51.8% on CursorBench 4.0 at max, 59.6% against 52.0% on Artificial Analysis's Terminal-Bench 4.0 run at max, and 62.3% against 56.3% on Proximal's FrontierSWE v2. Claude Fable 5.1 is first on the Coding Agent Index v1.5, which has no Opus 5.5 entry yet.
Cost per coding task widens the gap. Cursor's figure for Fable 5.1 at max on CursorBench 4.0 is $17.28 per task, and Anthropic's card puts Opus 5.5 at medium, its default, at 52.5% for about $3, above Fable's max score. On Artificial Analysis's Terminal-Bench run, Fable 5.1 does best at xhigh, 55.1%, above its own 52.0% at max; Opus 5.5 scores 59.6% at xhigh. Terminal-Bench is one of the three benchmarks inside the Coding Agent Index, but that index runs full coding agents such as Claude Code, so Opus 5.5's lead here may not carry over.
Anthropic's launch page adds a HAProxy rewrite of its own, in which Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1 at 51% lower cost, with both versions passing nearly all of HAProxy's regression tests. Anthropic ran and reported that task, and we cannot check it.
Which is better for long-running agents, Opus 5.5 or Fable 5.1?
No outside measurement settles it as of 23 September 2026. Anthropic points long-horizon agentic work to Claude Fable 5.1, and both agent leaderboards that list Fable 5.1 rank it first, but neither lists Claude Opus 5.5 yet. The models overview describes Opus 5.5 as "For long-running agentic coding and knowledge work" and Fable 5.1 as "For demanding reasoning and long-horizon agentic work".
Vendor documents give each model a claim; the one outside comparison, METR's, is qualitative and on AI R&D tasks, and says Opus 5.5 likely represents a modest improvement on Fable 5.1. Fable's product page calls it "best for ambitious, long-running, asynchronous work", and Opus 5.5's system card reports gains over Claude Opus 5 in "long-horizon professional work". METR's text in that card says Opus 5.5 "still has qualitative weaknesses that an expert human is unlikely to exhibit when solving hard, long-horizon tasks" (p. 43); METR has published no standalone study of Fable 5.1.
Long runs cost Fable 5.1 time and money. On the Coding Agent Index, Claude Code with Fable 5.1 at max averages $12.39 and about 35 minutes of agent time per task (2,089.9 seconds). In Anthropic's ProgramBench long-context test, 72% of Fable 5.1's episodes had at least one turn answered by the fallback model, Claude Opus 5, though that touched under 1% of turns (Fable 5.1 system card, p. 183).
Users disagree directly. waruyamaZero, 22 September 2026, gave one feature request to Opus 5.5 at xhigh and Fable 5.1 at high: "Opus 5.5 made some very questionable architectural decisions and agreed that they were not great. Fable 5.1 just worked like a charm, just as I expected." 9to5grinder, 22 September 2026, running a subagent workflow on Opus 5.5: "I'm not seeing any regression in long horizon agentic work. In fact quite the opposite, it follows instructions just as well as Fable 5.1 or better at 800k+ tokens in context." In an earlier comment the same user described Fable 5.1 taking up a worktree-splitting workflow for a day and then dropping it, a workflow Opus 5.5 now follows; the user suspects the Fable 5.1 that followed it that day was Opus 5.5 served under Fable's name. Two direct comparisons pointing opposite ways is where the evidence stands.
Is Claude Fable 5.1 slower than Opus 5.5?
On Artificial Analysis's measurements, yes: Claude Fable 5.1 generates 55.0 to 65.7 output tokens per second across its five settings, and Claude Opus 5.5 75.0 to 92.7 at the four settings the publisher has measured. At high effort the figures are 55.9 against 90.4.
Anthropic's models overview labels Fable 5.1's comparative latency "Slower" and Opus 5.5's "Moderate". Speed per token is not time per task: Opus 5.5 writes more tokens per task at max, where the publisher shows no Opus speed, so no max-against-max timing exists. One user reported the reverse on launch day. KasianFranks, 22 September 2026: "Back to Fable 5.1 - Opus 5.5 is now taking 10x longer just as Opus 5." The comment names no setting or task.
Does Claude Fable 5.1 refuse or fall back more often than Opus 5.5?
At similar rates on Gray Swan's prompt-injection benchmark, as Anthropic's system cards report it: the two hand requests to Claude Opus 4.8 in 18% of Opus 5.5's rollouts against 23% of Fable 5.1's (Opus 5.5 card p. 86). The Opus 5.5 system card also says Opus 5.5 blocks defensive vulnerability-discovery requests less often than earlier Opus and Fable models, in a chart with no values in the text (p. 57).
The safeguards are now a similar class. The Opus 5.5 launch post calls it "the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently." On both, flagged cyber requests go to Claude Opus 4.8 and flagged biology and frontier-AI development requests to Claude Opus 5. On the API neither falls back unless the developer sets it up: Opus 5.5's card says "the developer must opt in to automatic fallbacks" (p. 48), and Fable's product page says "API customers must configure their settings with our new Fallback API." For cyber work the two end in the same place; Fable's own card says its "performance on cyber tasks is nearly identical to that of Opus 4.8" (p. 46). How verified security teams get past those blocks is covered in the safeguards section of our Opus 5.5 review.
Data retention is the clearest difference. Fable's product page: "Using Fable requires 30-day data retention for safety monitoring by default." Zero retention is open only to eligible enterprise customers until Anthropic's customer-hosted storage arrives. The Opus 5.5 launch post: "Like previous Opus models, Opus 5.5 is available with zero data retention."
Users have filed refusal reports against both. On Claude Code's GitHub tracker, Mooooseeee, 4 September 2026, wrote "Fable 5.1 keeps reverting to Opus 4.8." mannewalis, 12 September 2026, after benign messages were flagged: "There is no configuration in which this work proceeds on Fable 5.1." lmcoelho, 22 September 2026, a Cyber Verification Program participant: "I have not been able to complete a single task on Fable 5.1." Opus 5.5 drew the same kind of report on its first day: John-Lussier, 22 September 2026, "New default Opus 5.5 refuses to do any cyber work." Fable has had three weeks to collect complaints and Opus 5.5 one day, so the count says nothing about rates.
When should I use Claude Fable 5.1 instead of Opus 5.5?
Use Claude Fable 5.1 when Claude Opus 5.5 at high or xhigh effort still fails your own evaluations, most plausibly on long-running agent work, the case Anthropic's docs name. On Artificial Analysis's Intelligence Index v4.3.2, every Fable 5.1 setting has an Opus 5.5 setting that scores higher for less, so raising Opus 5.5's effort is the cheaper first step.
The two steps are not priced alike. Moving Opus 5.5 from high ($1.82 per index task) to xhigh ($3.46) adds 2.41 points. Moving from Opus 5.5 at high to Fable 5.1 at its default, high, costs $3.91 and loses 2.43 points.
The published data points to a few candidates: office-document questions, where Fable 5.1 leads Anthropic's OfficeQA runs by 1.3 points; long-context reasoning, where it edges Opus 5.5 on AA-LCR at most settings; and long agent runs, where it holds two leaderboards that have not measured Opus 5.5. Other facts point away from it: cyber work ends at Claude Opus 4.8 on both, Fable 5.1 needs 30-day data retention by default, and at the two defaults it costs 2.9 times as much per index task.
Some users split the work. termmonkey, 23 September 2026, had compared a Fable 5.1 orchestrator over an Opus 5 builder with an all-Opus setup, and after the launch moved everything to Opus 5.5 at medium, keeping complex plans at xhigh. macaronianddeeez, 22 September 2026, had used Opus as "a Fable driven agent" to hold Fable usage down, and found Opus 5.5 did "about as well as Fable 5.1" on a project spec and website rebuild.
What do users report about Opus 5.5 vs Fable 5.1?
Of eight accounts that describe running both models on their own work, six report Claude Opus 5.5 matching or beating Claude Fable 5.1, and two preferred Fable 5.1, one on a feature's design and one on speed. The one report with dollar figures for both, gwd's above, has Opus 5.5 at under a quarter of Fable's cost. All eight posted on 22 or 23 September 2026, and self-selected launch-day comments show where the argument sits without settling it.
Five of the eight are quoted above: waruyamaZero, 9to5grinder, gwd, KasianFranks and macaronianddeeez. The others: roflc0ptic, 23 September 2026, back-testing against pull-request comments: "opus 5.5 matched fable 5.1. About $24 for 10 PRs." senko, 22 September 2026: "Opus 5.5 is neck-and-neck with Fable 5.1 and Astra 6 in my vibe-coding tests - maybe even better than Fable 5.1". thefourthchime, 23 September 2026: "I have to say so far Opus 5.5 gives me much better results than Fable 5.1 did. I can actually understand what it's talking about." On style, senderista, 23 September 2026: "Fable 5.1 is still intolerable."
Others argued the benchmarks miss where Fable is strong. En-tro-py, 22 September 2026, wrote that "the benchmarks only tell part of the capability differences". Alt_Restorer, 22 September 2026, offered a theory that "the larger a model is, the better its out of distribution performance"; none of the documents we read gives either model's size. VexObserver, 22 September 2026, called Fable's CursorBench cost "not a small efficiency gap" yet allowed that "Fable could still be better on certain long-horizon/agentic tasks that Cursorbench doesn't capture well."
Before Opus 5.5 existed, quota was a recurring Fable 5.1 complaint in the reports we read, mostly on Claude Code's GitHub tracker. xthakila, 1 September 2026: "Fable 5.1 medium burns through 100% of claude code limit in under 20 mins!!!" Fr33lumby, 4 September 2026, comparing two sessions on one plan: "Identical token cost within 1%. Quota consumption differs ~6-12x." luke14free, 8 September 2026, with auto-resume on: "even asking a simple question causes immediate burn of ~half of the quota." tleonhardt, 12 September 2026, on a code review at high effort, said it "rapidly blows through my entire session budget without actually finishing the review". voidfreud, 20 September 2026, had set high effort everywhere, saw the model report a reasoning effort of 10, and wrote: "I spent over 14 hours in Claude Code on Fable today and completed zero work." skunk_of_thunder, 21 September 2026, after a week at low on a Max plan, read the meters as meaning "it's impossible to use 100% of your fable-specific usage before you run out of all-models usage." These describe subscription limits in one client; none is an API bill, and we cannot check them.
What has changed since Claude Fable 5.1 launched?
Claude Fable 5.1's list price has not changed since its release on 1 September 2026. What changed is the model beside it and the scale its outside scores use.
- 1 September 2026: Anthropic releases Fable 5.1 at $10 and $50 per million tokens with cache reads cut to $0.25. Artificial Analysis's launch article gives it 66 on an earlier index version and 1853 on GDPval-AA v2.
- Up to 13 September 2026: a Claude Code banner quoted in GitHub issue 92576 reads "weekly Claude Code limit raised 50% until Sep 13".
- September 2026: Artificial Analysis moves to Intelligence Index v4.3.2 and GDPval-AA v2.1 on a re-anchored Elo scale; Fable 5.1 at max becomes 53.35 and 1735.
- 22 September 2026: Anthropic releases Claude Opus 5.5 at $4 and $20, and its models overview recommends Opus 5.5 first. METR and Artificial Analysis publish on it the same day.
- 23 September 2026: Fable's product page still calls it Anthropic's most capable generally available model, and Opus 5.5 has no entry on the Coding Agent Index, LMArena or Epoch AI.
Still open on the day of publication: Opus 5.5 results on those three, which would turn Fable's first places into comparisons; Opus 5.5's default effort inside Claude Code and Claude.ai; and how often either model falls back in real traffic. We will re-read the pricing page, the models overview and each evaluator at least monthly, sooner when any of them publishes on either model, and date every change on this page.
Claude Opus 5.5 vs Claude Fable 5.1: independent evaluations
Six organisations outside Anthropic were checked on 23 September 2026, and three publish figures for both models on the same instrument. Artificial Analysis, on Intelligence Index v4.3.2 with every entry for both models run under Anthropic's default fallback, scores Claude Opus 5.5 at 57.62, 55.99, 53.58, 51.24 and 42.31 at max, xhigh, high, medium and low, for $5.98, $3.46, $1.82, $1.34 and $0.55 per index task, and Claude Fable 5.1 at 53.35, 53.20, 51.15, 48.92 and 46.82 for $7.63, $5.98, $3.91, $2.98 and $2.37. On its own Terminal-Bench 4.0 run (mini-swe-agent harness) it has Opus 5.5 at 59.6% and Fable 5.1 at 52.0% at max, and on GDPval-AA v2.1 1846 against 1735 Elo at max. ARC Prize verifies Opus 5.5 at 98.5% on ARC-AGI-1 and 93.3% on ARC-AGI-2 Semi-Private at high, and Fable 5.1 at 97.5% and 90.0% at max. Zapier's AutomationBench v1.0.6 has Opus 5.5 at 40.0% at max, run during Anthropic's early access, and Fable 5.1 with Claude Opus 5 fallback at 31.4% at max, with Opus 5 handling 260 of 657 tasks. Artificial Analysis's Coding Agent Index v1.5 ranks Claude Code with Fable 5.1 at max first at 62.2 and has no Opus 5.5 entry; LMArena's Agent Arena, board dated 15 September, ranks claude-fable-5.1-max first at a net improvement of 13.71% and lists Opus 5.5 on no board; Epoch AI gives Fable 5.1 a Capability Index of 165.0 and Opus 5.5 no score. METR published a qualitative Opus 5.5 summary calling it an incremental improvement over Fable 5.1, with no numbers, and has no standalone Fable 5.1 publication.
What we read
- Anthropic, Introducing Claude Opus 5.5, the announcement page: the benchmark table and its footnotes, the Fable 5.1 comparison lines, the HAProxy example, the safeguards, preserved-thinking and data-retention sections, saved as textThe lab's own document · read Sep 23, 2026
- Anthropic, Claude Opus 5.5 System Card: the capability table and section 8 pages 174 to 212 (including OfficeQA and Toolathlon), the safeguards and fallback pages 12, 13, 48 and 57, the prompt-injection fallback rates on pages 86 and 88, METR's text on pages 41 to 44 and the self-preference result on page 127The lab's own document · read Sep 23, 2026
- Anthropic, Introducing Claude Fable 5.1 and Claude Mythos 5.1, the announcement page: the benchmark table and caption, the data behind its per-effort charts, the effort-default note, the safeguards section and the cost section with its workload chartThe lab's own document · read Sep 23, 2026
- Anthropic, System Card: Claude Fable 5.1 and Claude Mythos 5.1, 212 pages: the capability table and settings on pages 167 to 173, the ProgramBench and Toolathlon results, the safeguards and fallback pages 45 to 58, the prompt-injection fallback pages 83 to 85 and the alignment findings on pages 91 to 96 and 124 to 125The lab's own document · read Sep 23, 2026
- Anthropic, the Claude Fable product page: price, US-only inference, rerouted-request billing, 30-day data retention, plan availability and the Fallback API lineThe lab's own document · read Sep 23, 2026
- Anthropic, models overview in the Claude Platform docs: the model-choice guidance and the Opus 5.5 and Fable 5.1 columns for latency, price, API ID, thinking, default effort, context window, maximum output and knowledge cutoffThe lab's own document · read Sep 23, 2026
- Anthropic, claude.com pricing page, the API rows for Opus 5.5 and Fable 5.1 with the batch, fast mode, US-only and cache TTL lines, served in US dollars from our locationThe lab's own document · read Sep 23, 2026
- Anthropic, news index, for the dates of the Fable 5.1 announcement and the alignment report it calls yesterdayThe lab's own document · read Sep 23, 2026
- Artificial Analysis, the five Claude Fable 5.1 model pages on Intelligence Index v4.3.2, one per effort setting, including per-evaluation values and output speed in the page dataIndependent · read Sep 23, 2026
- Artificial Analysis, the five Claude Opus 5.5 model pages on Intelligence Index v4.3.2, one per effort setting, including per-evaluation values and output speedIndependent · read Sep 23, 2026
- Artificial Analysis, Claude Fable 5.1 launch article dated 1 September 2026, on an index version before 4.3Independent · read Sep 23, 2026
- Artificial Analysis, Claude Opus 5.5 launch article dated 22 September 2026, including its output-token counts per index taskIndependent · read Sep 23, 2026
- Artificial Analysis, intelligence-benchmarking methodology and version history, for the GDPval-AA v2.1 re-anchoring and the version 4.3 benchmark changesIndependent · read Sep 23, 2026
- Artificial Analysis, Coding Agent Index v1.5, with a Fable 5.1 entry and no Opus 5.5 entryIndependent · read Sep 23, 2026
- LMArena (arena.ai), the Text, Code/WebDev, Vision, Document, Agent and Search boards, with claude-fable-5.1-max listed and no Opus 5.5 entryIndependent · read Sep 23, 2026
- ARC Prize, results page for Claude Fable 5.1, verified scores at five effort settingsIndependent · read Sep 23, 2026
- ARC Prize, results page for Claude Opus 5.5, verified scores at five effort settingsIndependent · read Sep 23, 2026
- Zapier, AutomationBench leaderboard version 1.0.6, the Opus 5.5 rows and the Fable 5.1 with Opus 5 fallback row and its footnoteIndependent · read Sep 23, 2026
- Epoch AI, Capability Index and benchmark data files, with a Fable 5.1 score and no Opus 5.5 scoreIndependent · read Sep 23, 2026
- METR, Summary of METR's predeployment evaluation of Claude Opus 5.5, and its blog index and time-horizons page, checked for any Fable 5.1 publication and finding noneIndependent · read Sep 23, 2026
- Reddit comments comparing Opus 5.5 and Fable 5.1 in r/ClaudeCode, r/ClaudeAI and r/cursor, read through a public archive searchIndependent · read Sep 23, 2026
- Hacker News comments naming Fable 5.1 since 1 September 2026, read through the site's search APIIndependent · read Sep 23, 2026
- GitHub, public issues on the anthropics/claude-code tracker naming Fable 5.1 or Opus 5.5Independent · read Sep 23, 2026
What we did not read
- Anthropic's Fable 5.1 model page in the platform docs. Fable 5.1's API facts on this page come from the models overview, which we read, and not from that page.
- Anthropic's Fallback API help article and its pages on Enterprise Frontier Safeguards, the customer-hosted storage that is to replace the zero-retention exception for Fable 5.1. The data-retention lines here are from the Fable product page and the Opus 5.5 announcement only.
- Large parts of the Fable 5.1 system card: the welfare section (pages 139 to 166), most of the harmlessness and malicious-use sections (pages 59 to 81), the life-sciences, multilingual and healthcare sections (pages 198 to 205), the software-engineering sections beyond the lines quoted and the multi-agent section (pages 179 to 182). Pages 30 to 39, 46 to 51, 97 to 119 and 126 to 138 were searched but not read line by line. Values the card shows only as chart images were not transcribed.
- The Opus 5.5 system card beyond the pages named in the source list. The full record of what we read of it is in our Opus 5.5 review.
- The customer quotes on both launch pages. They are chosen by Anthropic, and none is reproduced or counted here.
- Cursor's CursorBench leaderboard, Proximal's FrontierSWE results, Cognition's FrontierCode results and the Terminal-Bench public leaderboards. Every figure from those operators on this page comes through Anthropic's documents.
- Artificial Analysis's Devin Fusion entry on the Coding Agent Index, which includes Fable 5.1 but was not extracted, and time-to-first-token values for either model, which are drawn as charts with no value printed.
- LMArena's changelog on the Fable pass, and its Image-to-WebDev board. Epoch AI's hub pages were not re-read; its figures here come from its data files.
- Vals.ai, Scale's SEAL leaderboards, LiveBench and the official SWE-bench leaderboard. They were out of scope and were not checked for either model.
- Any source that explains why Anthropic's Opus 5.5 launch table prints Fable 5.1 at 65.6% on Humanity's Last Exam with tools while Fable's own launch post prints 65.0%. We found none, so the page uses Artificial Analysis's same-harness figures for that benchmark.
- Cloud marketplace rate cards for either model on Amazon Bedrock, Google Cloud Vertex AI or Microsoft Azure, and any page stating which consumer plans include Opus 5.5 or its default effort inside Claude Code and Claude.ai.
- The third-party robot-safety benchmark one Reddit post relays for Fable 5.1, and Reddit, Hacker News and GitHub comments beyond what the archive search, the search API and the issue search returned on 23 September 2026. G2, Trustpilot and Capterra carry no dated reviews of either model.
- Both models. We ran no prompt through Claude Opus 5.5 or Claude Fable 5.1 for this page and report no result of our own.
- Anything published after 23 September 2026.
Disclosure. Our writing workflow runs on Anthropic models, including Claude Opus 5.5, one of the two models compared; both models on this page are Anthropic's. The Opus 5.5 system card measures a small but statistically significant bias toward itself when reminded that it is Claude (0.07 points out of 10, p. 127). So every comparison here rests on published figures, each named with its publisher and setting.
We run no hands-on tests. This comparison is built from the lab’s own published documents, independent evaluations by other organisations, and dated user reports, each named above with the date we read it. How we investigate →