Anthropic · shipped September 1, 2026

Fable 5.1

The frontier model, one decimal bigger. This is a screen-recorded walkthrough of what actually changed — the benchmarks, the price, what the builders are saying, and one thing worth building next.

The 30-second version

Four numbers that tell the whole story

0
Artificial Analysis Intelligence Index at max effort — the highest score they've ever measured
−75%
cache-read price vs Fable 5 — $1.00 → $0.25 per million tokens
2.1×
Fable 5 on agentic science (Terminal-Bench-Science: 24.7% → 52.6%)
~1.7×
the output tokens Fable 5 burns — the catch we'll come back to
Sources: Artificial Analysis — "Claude Fable 5.1 tops the Intelligence Index"; Anthropic launch post; Artificial Analysis token-use data (cited by Matthew Berman).

What you're actually looking at

Fable 5.1 — and its twin, Mythos 5.1

Same model, two safety tiers. Fable 5.1 is generally available. Mythos 5.1 is the same weights behind lighter guardrails — invite-only, through Project Glasswing, for vetted cybersecurity and life-sciences work.

1M token context128K max output Adaptive thinking, always onKnowledge cutoff June 2026 $10 / $50 per M in / outAPI id claude-fable-5-1
Source: Claude Platform docs — Fable 5.1 overview; Anthropic launch post.

Fable 5 → Fable 5.1 · compare and contrast

The gains pile up where the work is long

Fable 5 Fable 5.1

Terminal-Bench-Science · Terminal-Bench 4.0 · AutomationBench · browser-agent tasks · FrontierFinance · CursorBench 3.2 · Humanity's Last Exam (tools). On short, simple prompts the two are nearly identical — you're paying for headroom on long runs.

Source: Anthropic launch post & model card; Artificial Analysis benchmark tables.

Side by side

Fable 5 vs Fable 5.1

Fable 5Fable 5.1
Releasedearlier in 2026Sept 1, 2026
Cache read / M tokens$1.00$0.25
Input / output / M$10 / $50$10 / $50 — unchanged
Knowledge cutoffearlierJune 2026
Terminal-Bench 4.042.0%55.8%
Terminal-Bench-Science24.7%52.6%
Intelligence Index (max)6266
Output tokens per taskbaseline~1.7× more
Effort controlfixed per conversationper-message (beta)
Vulnerability discoveryblockedallowed (defensive only)
Sources: Claude Platform docs; Anthropic launch post; Artificial Analysis; The New Stack.

The case for it

What Fable 5.1 does better

  • +The agentic leap is real. Browser tasks 57% → 82%. AutomationBench 17% → 31%. Long, tool-heavy runs is where it separates.
  • +Cheaper to run on big context. Cache reads down 75% → ~25% lower cost on typical work, ~45% on highly agentic work.
  • +Fewer bogus refusals. ~60% fewer cybersecurity false positives, ~85% fewer on biology safeguards.
  • +New controls. Per-message effort, turn-scoped system messages, readable progress updates between tool calls, content provenance watermark.
  • +#1 on the Intelligence Index at 66 — ahead of Opus 5 (63), GPT-5.6 Sol (61), Grok 4.6 (61).
Sources: Anthropic launch post; Artificial Analysis; TechCrunch.

The case against it

Where Fable 5.1 will bite you

  • Base price didn't move. Still $10 / $50 — twice Opus 5, and far above GPT-5.6 Sol ($4 / $20) or Gemini 3.7 Flash ($0.75 / $3.75).
  • It talks more. ~1.7× the output tokens of Fable 5 — per finished task it can cost more, even with cheaper cache. (Berman / Artificial Analysis)
  • Launch-week meltdown. Max-plan users torched 5-hour quotas in minutes; a single prompt ate a 20× Max window in 52 minutes. Suspected Claude Code cache bug.
  • Three breaking API changes from Fable 5: forced tool use errors out, thinking blocks are model-locked, editing old turns invalidates them.
  • Overkill for most work. Simple prompts ≈ Fable 5. The FT reported Fable 5 was only ~11% of Anthropic spend across ~70,000 companies.
Sources: X trending — "usage limits in minutes"; Claude Platform docs (migration); Financial Times via VentureBeat; 4sysops.

Real data · one benchmark

Artificial Analysis Intelligence Index — max effort

0
Claude Fable 5.1
0
Claude Opus 5
0
Claude Fable 5
0
GPT-5.6 Sol
0
Grok 4.6 (high)

Highest score Artificial Analysis has recorded. Margin over the field is real but narrow — 3–5 points.

Source: artificialanalysis.ai/articles/claude-fable-5-1 (Sept 2026).

Real data · coding & agents

Where the gap is actually wide

80.0
58.6
54.2
SWE-bench Pro
Fable 5.1 · GPT-5.6 Sol · Gemini 3.1 Pro
52.6
29.0
24.7
22.4
Terminal-Bench-Science
Fable 5.1 · Opus 5 · Fable 5 · GPT-5.6 Sol
Fable 5.1everyone else
Sources: Anthropic launch post; Artificial Analysis; morphllm Claude benchmark roundup.

Real data · what it costs

Cheaper cache, hungrier tokens

Cache read — $ / million tokens

$1.00
Fable 5
$0.25
Fable 5.1

Input — $ / million tokens

$10
Fable 5.1
$4
GPT-5.6 Sol
$0.75
Gemini 3.7 Flash
$0.43
DeepSeek V4 Pro

Cost per solved SWE-bench task: DeepSeek V4 Pro ~$0.31 vs Fable ~$0.81. Output: Fable 5.1 $50 / M vs GPT-5.6 Sol $20 / M. The cheaper cache only wins back money if your prompts are big and reused.

Sources: VentureBeat; Together.ai & Fireworks.ai DeepSeek-vs-Fable cost analyses; Claude Platform pricing.

The room's reaction

What the builders are saying

"No company has released a model as good as Mythos / Fable… I was clearly wrong about Anthropic."

Elon Musk — xAI, on X (via TechCrunch / Yahoo Finance)

"Fable 5.1 is our best model yet for coding, data analysis, computer use, design, presentations, and long-running agentic work."

Boris Cherny — Anthropic, creator of Claude Code, on X

"Raises the ceiling on what Fable 5 could do — particularly for coding. The most powerful AI model ever released."

Alex Finn — @AlexFinn on X

"Fable 5.1 FINALLY kills AI website slop." Layered scroll, touch-friendly nav, mobile-first grids that don't read as AI-generated.

Nate Herk — AI Automation Society (YouTube / Skool)

The /goal command: "Claude keeps working, checking its own progress, until the OUTCOME is met" — built for long agentic runs, not single replies.

Sabrina Ramonov — sabrina.dev, "6 INSANE Projects to Learn Claude Fable"

"Mythos (Fable) is AGI." Rebuilt his Lovable-style app-builder in 5 prompts.

Riley Brown — @rileybrown on X

Independent testing: Fable 5.1 uses ~1.7× the output tokens of Fable 5 — it can cost more per completed task despite the promised savings.

Matthew Berman — on X, citing Artificial Analysis

Asked to optimize a system, Fable delivered a 17.7× speedup on one benchmark and ~22% average across the suite.

Wes Roth — @WesRoth on X

"A slightly better benchmark hack with yet another price hike."

4sysops — editorial take
All quotes attributed to their authors. Full source links in the deck's SOURCES.md and the video description.

Named customers, on the record

What early access actually did with it

  • Millennium — identified a rare crash that had gone unsolved for 4–5 years, by disassembling a vendor library and matching a core dump.
  • Ramp — ran an unattended 38-hour ML job: six experiments, findings written up.
  • Jane Street — "solves more coding problems than Fable 5 or Opus 5"; state-of-the-art on their trading-intuition eval.
  • MongoDB — complex prototype in three days, running unattended for hours with its own verification loops.
  • SpaceX AI — 73.4% on CursorBench 3.2, "especially skilled at verifying its own work on hard coding tasks."
Source: Anthropic launch post (customer testimonials); VentureBeat.

The competitive picture

How it stacks up against everyone else

OpenAI — GPT-5.6 Sol

Close on knowledge, behind on agentic coding (SWE-bench Pro 58.6 vs 80). But ~2.5× cheaper in and out. "OpenAI returns to the ring" on price.

Google — Gemini 3.1 Pro / 3.7 Flash

Wins on raw price ($0.75 / $3.75 Flash), context and multimodal. Trails badly on agentic coding (SWE-bench Pro 54.2). Plays the volume game.

DeepSeek — V4 Pro

20–90× cheaper per solved task. Tops some SWE cost-efficiency leaderboards. The pick when compute cost is the hard constraint.

xAI — Grok 4.6 / 4.7

Grok 4.6 (high) ties GPT at 61. Musk is aiming 4.7 straight at Fable — "Opus-class, faster, cheaper" — while conceding Fable is "definitely better."

Anthropic's moat isn't the benchmark points — it's long-horizon reliability: self-verification, root-cause debugging, hours-to-days unattended, the Claude Code ecosystem, and enterprise zero-data-retention.

Sources: VentureBeat; The New Stack ("Grok 4.5 Opus-killer"); morphllm benchmark roundup; Artificial Analysis.

Prove it on camera · next video

Build the "Overnight Prospect Teardown" agent

One /goal. It runs while you sleep.

Point Fable 5.1 at 10 luxury-remodeler competitors. It crawls each site, audits the funnel and offer, drives a browser to test what it finds, then produces a comparison spreadsheet, a one-page "here's your gap" deck per prospect, and drafted outreach emails — all verified against its own checklist before it stops.

Why it proves 5.1 specifically: long-horizon agentic run (Ramp's 38 hours), self-verification (SpaceX's "checks its own work"), native slide + spreadsheet output, and cheap cache reads over a big repeated context.

Backups: self-grading content factory (7-day calendar auto-scored to 8+/10, 24/7 A/B loop on DM hooks — Sabrina's projects) · podcast-episode-in-a-box agent (research → script → metadata → thumbnail → digest draft, unattended) · "one sentence, came back from lunch to a working lead-magnet app."

Grounded in: Anthropic customer testimonials; Sabrina Ramonov's Fable project set; Riley Brown / MongoDB unattended-build demos.

Bottom line

Should you switch?

If you do hours-long agentic coding or research — Fable 5.1 is the best there is right now. Switch.
If you do chat and short tasks — stay on Opus 5 or Sonnet 5 and pocket the difference.
Watch the bill either way. Cheaper cache, hungrier tokens — measure cost per finished task, not per token.

Fable 5.1 · Sept 2026

Thanks for watching

Every number on screen is sourced — full list in the description. Like, subscribe, and tell me what you want me to build with it next.

5.1
The Business Algorithm
01 / 16
Navigation
/ Space next  ·  back  ·  Home restart
F fullscreen  ·  ? close this