← Back to Hello, AI
2 min

Claude Opus 5 Just Beat Anthropic's Own Flagship

Anthropic's $5/$25 Opus 5 beats its own $10/$50 Fable 5 on Frontier-Bench and ARC-AGI-3, raising an uncomfortable question about who really leads the Claude lineup.

Anthropic's cheapest flagship is now its best one. Claude Opus 5 landed on July 24, 2026 at $5/$25 per million tokens — the same rate card as Opus 4.8 — and it beats Fable 5, the $10/$50 premium tier, on Frontier-Bench and ARC-AGI-3. On Frontier-Bench v0.1 the spread is not subtle: 43.3% for Opus 5 against 33.7% for Fable 5 and 18.7% for Opus 4.8. A lab does not usually ship a model that makes its own top SKU look overpriced.

The price comparison is the part that should make anyone with a budget line pay attention. Opus 5 costs exactly half of Fable 5 on both input and output. Fable 5's position has also been eroding independently — since its June 30 relaunch it has been credits-only, no longer bundled into standard Claude plans, and it ships without Zero Data Retention support. For a team evaluating both today, the expensive model is harder to buy, harder to compliance-approve, and now behind on the published numbers.

Early third-party signal points the same direction. Lovable's internal eval found that Opus 5 "isn't just better on our hardest agentic coding tasks, up 22% over Opus 4.7, it's steadier, with far less variance run to run." Variance is the metric that matters for agentic workloads and the one benchmarks rarely capture. A model that scores well on average but fails one run in five forces you to build retry scaffolding around it; a steadier model at half the price changes the shape of the system you build, not just the invoice.

The caveats are real. The Frontier-Bench figures are a launch-day Anthropic claim, reported by Vellum and tech-ish but not independently rerun. Benchmarks published by the vendor on release day have a way of narrowing under outside scrutiny. Fable 5 may still hold edges Anthropic has not chosen to publish — long-horizon agent runs, safety headroom, behavior under adversarial load are all places where a premium tier could justify itself without ever appearing on a leaderboard. Declaring it dead 24 hours in would be its own kind of hype.

That leaves an uncomfortable question for this site. helloai's models.json has already swapped Opus 4.8 for Opus 5, but Fable 5 remains separately tracked at 1507 Elo, still listed as leader in both Overall Preference and Coding & Engineering. Those two facts are now in tension, and the honest answer is that our admission rules do not yet say what to do when a lab undercuts its own flagship — Elo and published benchmarks are disagreeing, and we have not decided which wins. If same-lab cannibalization becomes the pattern rather than the exception, every directory tracking premium tiers as a separate rank will need a rule for how long a flagship keeps its crown after a cheaper sibling passes it.

← More from Hello, AI