Anthropic · shipped September 1, 2026
Fable 5.1
The frontier model, one decimal bigger. This is a screen-recorded walkthrough of what actually changed — the benchmarks, the price, what the builders are saying, and one thing worth building next.
Anthropic · shipped September 1, 2026
The frontier model, one decimal bigger. This is a screen-recorded walkthrough of what actually changed — the benchmarks, the price, what the builders are saying, and one thing worth building next.
The 30-second version
What you're actually looking at
Same model, two safety tiers. Fable 5.1 is generally available. Mythos 5.1 is the same weights behind lighter guardrails — invite-only, through Project Glasswing, for vetted cybersecurity and life-sciences work.
Fable 5 → Fable 5.1 · compare and contrast
Terminal-Bench-Science · Terminal-Bench 4.0 · AutomationBench · browser-agent tasks · FrontierFinance · CursorBench 3.2 · Humanity's Last Exam (tools). On short, simple prompts the two are nearly identical — you're paying for headroom on long runs.
Side by side
| Fable 5 | Fable 5.1 | |
|---|---|---|
| Released | earlier in 2026 | Sept 1, 2026 |
| Cache read / M tokens | $1.00 | $0.25 |
| Input / output / M | $10 / $50 | $10 / $50 — unchanged |
| Knowledge cutoff | earlier | June 2026 |
| Terminal-Bench 4.0 | 42.0% | 55.8% |
| Terminal-Bench-Science | 24.7% | 52.6% |
| Intelligence Index (max) | 62 | 66 |
| Output tokens per task | baseline | ~1.7× more |
| Effort control | fixed per conversation | per-message (beta) |
| Vulnerability discovery | blocked | allowed (defensive only) |
The case for it
The case against it
Real data · one benchmark
Highest score Artificial Analysis has recorded. Margin over the field is real but narrow — 3–5 points.
Real data · coding & agents
Real data · what it costs
Cache read — $ / million tokens
Input — $ / million tokens
Cost per solved SWE-bench task: DeepSeek V4 Pro ~$0.31 vs Fable ~$0.81. Output: Fable 5.1 $50 / M vs GPT-5.6 Sol $20 / M. The cheaper cache only wins back money if your prompts are big and reused.
The room's reaction
"No company has released a model as good as Mythos / Fable… I was clearly wrong about Anthropic."
"Fable 5.1 is our best model yet for coding, data analysis, computer use, design, presentations, and long-running agentic work."
"Raises the ceiling on what Fable 5 could do — particularly for coding. The most powerful AI model ever released."
"Fable 5.1 FINALLY kills AI website slop." Layered scroll, touch-friendly nav, mobile-first grids that don't read as AI-generated.
The /goal command: "Claude keeps working, checking its own progress, until the OUTCOME is met" — built for long agentic runs, not single replies.
"Mythos (Fable) is AGI." Rebuilt his Lovable-style app-builder in 5 prompts.
Independent testing: Fable 5.1 uses ~1.7× the output tokens of Fable 5 — it can cost more per completed task despite the promised savings.
Asked to optimize a system, Fable delivered a 17.7× speedup on one benchmark and ~22% average across the suite.
"A slightly better benchmark hack with yet another price hike."
Named customers, on the record
The competitive picture
Close on knowledge, behind on agentic coding (SWE-bench Pro 58.6 vs 80). But ~2.5× cheaper in and out. "OpenAI returns to the ring" on price.
Wins on raw price ($0.75 / $3.75 Flash), context and multimodal. Trails badly on agentic coding (SWE-bench Pro 54.2). Plays the volume game.
20–90× cheaper per solved task. Tops some SWE cost-efficiency leaderboards. The pick when compute cost is the hard constraint.
Grok 4.6 (high) ties GPT at 61. Musk is aiming 4.7 straight at Fable — "Opus-class, faster, cheaper" — while conceding Fable is "definitely better."
Anthropic's moat isn't the benchmark points — it's long-horizon reliability: self-verification, root-cause debugging, hours-to-days unattended, the Claude Code ecosystem, and enterprise zero-data-retention.
Prove it on camera · next video
/goal. It runs while you sleep.Point Fable 5.1 at 10 luxury-remodeler competitors. It crawls each site, audits the funnel and offer, drives a browser to test what it finds, then produces a comparison spreadsheet, a one-page "here's your gap" deck per prospect, and drafted outreach emails — all verified against its own checklist before it stops.
Why it proves 5.1 specifically: long-horizon agentic run (Ramp's 38 hours), self-verification (SpaceX's "checks its own work"), native slide + spreadsheet output, and cheap cache reads over a big repeated context.
Backups: self-grading content factory (7-day calendar auto-scored to 8+/10, 24/7 A/B loop on DM hooks — Sabrina's projects) · podcast-episode-in-a-box agent (research → script → metadata → thumbnail → digest draft, unattended) · "one sentence, came back from lunch to a working lead-magnet app."
Bottom line
Fable 5.1 · Sept 2026
Every number on screen is sourced — full list in the description. Like, subscribe, and tell me what you want me to build with it next.