Book a Free Strategy Call
Skip the read: talk to Walid in 30 min.
Free strategy call. We map your AI engineering team, you keep the notes.
Claude Opus 5 ranks #1 of 170 models on the Artificial Analysis Intelligence Index, scoring 61 in max reasoning mode against Claude Fable 5's 60. Anthropic released the model on July 24, 2026, and backed it with wins across four other named benchmarks, plus one honest loss.
That loss matters as much as the wins. Anthropic's own launch disclosure places Opus 5 behind Mythos 5 on cybersecurity benchmarks, the single category where the company did not claim the top spot. Here is the full rundown, benchmark by benchmark, with the source for every number.
Claude Opus 5 benchmark results at a glance
| Benchmark | Result | Source |
|---|---|---|
| Artificial Analysis Intelligence Index | 61 (max), #1 of 170 models; Fable 5 scores 60 | Artificial Analysis |
| Frontier-Bench | Surpasses all models tested, more than 2x Opus 4.8 | Anthropic |
| ARC-AGI 3 | 3x the next-best model | Anthropic |
| OSWorld 2.0 (computer use) | Beats Fable 5 at roughly one third the cost | Anthropic |
| Zapier AutomationBench | Roughly 1.5x the next-best model | Anthropic |
| Cybersecurity benchmarks | Behind Mythos 5 | Anthropic |
Related Reads
Artificial Analysis Intelligence Index: Opus 5 takes the top spot
The Artificial Analysis Intelligence Index aggregates results across a broad benchmark suite to produce a single comparative score for every major model it tracks, currently 170 of them. Opus 5 in its max reasoning configuration scores 61, edging out Fable 5's 60 and putting Opus 5 in first place on the entire leaderboard.
A one-point gap is not a landslide. What makes it notable is the price attached to it: Opus 5 costs half of what Fable 5 costs per token, so the model taking the top intelligence score is also the cheaper of the two. That combination, best score plus lower price, is the core argument in Anthropic's launch messaging and is covered in more depth in our Claude Opus 5 pricing breakdown.
Free weekly brief
Steal our production automations
The exact n8n flows, Claude Code setups, and prompts we ship for clients, broken down step by step. No spam, unsubscribe anytime.
Frontier-Bench: more than double Opus 4.8
Anthropic says Opus 5 surpasses every model it tested on Frontier-Bench, and specifically more than doubles the score of its own prior flagship, Opus 4.8. Anthropic has not published the full underlying task list for Frontier-Bench in the same detail as third-party suites like Artificial Analysis, so this figure should be read as a self-reported comparison against Anthropic's own prior model, a useful data point, not an independently verified leaderboard result.
The scale of the jump (more than 2x a model released earlier the same year) is one of the larger single-generation gains Anthropic has disclosed for this benchmark.
ARC-AGI 3: 3x the next-best model
ARC-AGI is designed to test reasoning that resists memorization: each task uses novel visual patterns a model has not seen in training. Anthropic reports Opus 5 at 3x the score of the next-best model on the newest version of the test, ARC-AGI 3.
This is the widest claimed margin in the entire release. A 3x gap on a benchmark specifically built to resist pattern-matching from training data is a stronger signal than a similar margin on a benchmark closer to typical training distributions, though as with Frontier-Bench, this is Anthropic's own reported comparison rather than an independently published leaderboard.
OSWorld 2.0: computer use at a third of the cost
OSWorld 2.0 measures how well a model can operate a real computer environment: navigating applications, completing multi-step UI tasks, and recovering from errors without step-by-step human instructions. Anthropic reports that Opus 5 beats Fable 5 on this benchmark while costing roughly a third as much to run.
This is the clearest bridge between the benchmark numbers and the pricing story. Computer-use tasks are agentic by nature, meaning the model takes many sequential actions per task, so a cost claim here reflects total tokens to completion, not just the sticker price per million tokens. For the token-price side of that comparison, see our Opus 5 pricing guide.
Zapier AutomationBench: 1.5x the next-best model
Zapier's AutomationBench tests models on real workflow-automation tasks, the kind of multi-step, tool-calling work that runs inside no-code and low-code automation platforms. Anthropic reports Opus 5 at roughly 1.5x the next-best model's score here.
Of the five benchmarks in this release, AutomationBench is the most directly relevant to teams building or running AI agents inside automation tools, since it is scored on tasks that resemble production automation work rather than academic reasoning problems.
Where Opus 5 does not win: cybersecurity
Anthropic's disclosure is specific about this one: Opus 5 trails Mythos 5 on cybersecurity benchmarks. Anthropic did not soften or omit this result in its own materials, which is worth noting given how much of the rest of the release is framed around wins.
If your workload is security research, vulnerability analysis, or red-team tooling, this is the one category in the entire launch where the newest model is not the strongest choice inside Anthropic's own lineup. Mythos 5 remains the model to evaluate for that specific use case.
How Opus 5 compares to Fable 5 overall
Putting the results together: on the one benchmark where Fable 5 is named directly (the Artificial Analysis Intelligence Index), Opus 5 comes out ahead, and on the one benchmark with a direct Fable 5 comparison in Anthropic's own materials (OSWorld 2.0), Opus 5 wins on both capability and cost. None of Anthropic's disclosed results show Opus 5 losing to Fable 5 specifically, cybersecurity is a loss against Mythos 5, a different model. And Opus 5 does all of this at half Fable 5's token price.
For the full side-by-side of the two models, capability and price together, see our Claude Opus 5 vs Claude Fable 5 comparison. If you are also weighing non-Anthropic options, our Claude Fable 5 alternatives roundup covers models outside this family, and our models page has current specs and pricing for the full Anthropic lineup.
A note on how to read self-reported benchmarks
Three of the five numbers in this release (Frontier-Bench, ARC-AGI 3, and Zapier AutomationBench) come from Anthropic's own launch materials rather than an independently maintained public leaderboard. That does not make them false, but it means they have not been reproduced by a third party the way the Artificial Analysis Intelligence Index has. The Intelligence Index score (61, #1 of 170) is the one figure in this rundown you can check yourself, updated on Artificial Analysis's own leaderboard, independent of Anthropic's launch messaging.
FAQ
What is Claude Opus 5's Intelligence Index score?
Claude Opus 5 scores 61 in max reasoning mode on the Artificial Analysis Intelligence Index, ranking #1 out of 170 models tracked. Claude Fable 5 scores 60, one point behind.
Is Claude Opus 5 better than Claude Fable 5?
On the Artificial Analysis Intelligence Index, Opus 5 scores 61 against Fable 5's 60, a narrow but confirmed win. Anthropic's other disclosed results (Frontier-Bench, ARC-AGI 3, Zapier AutomationBench, and a direct OSWorld 2.0 win) are framed against "the next-best model" or "all models tested" rather than Fable 5 by name, and cybersecurity is the one area where Anthropic places Opus 5 behind Mythos 5, not Fable 5.
Does Claude Opus 5 beat Mythos 5?
Not across the board. Anthropic's own disclosure places Opus 5 behind Mythos 5 specifically on cybersecurity benchmarks, the one category in the July 24, 2026 release where Opus 5 is not the top Anthropic model.
What is ARC-AGI 3 and why does Opus 5's score matter?
ARC-AGI 3 is a reasoning benchmark built from novel visual puzzles a model cannot have memorized from training data. Anthropic reports Opus 5 scoring 3x the next-best model on it, the widest margin in this release, which is a stronger signal of genuine reasoning gains than similar margins on more familiar benchmark formats.
How good is Claude Opus 5 at computer use?
Anthropic reports Opus 5 beating Fable 5 on OSWorld 2.0, the standard computer-use benchmark, while costing roughly a third as much per task to run. See our Opus 5 pricing guide for the token-price side of that comparison.
Are Anthropic's Opus 5 benchmark claims independently verified?
Partially. The Artificial Analysis Intelligence Index score (61, #1 of 170) comes from a third-party leaderboard you can check directly. Frontier-Bench, ARC-AGI 3, and Zapier's AutomationBench results are Anthropic's own self-reported figures from its launch materials, not yet reproduced on an independent public leaderboard.
Is Claude Opus 5 good for AI agents and automation?
Anthropic reports Opus 5 at roughly 1.5x the next-best model on Zapier's AutomationBench, a benchmark built specifically around workflow-automation tasks, plus a computer-use win on OSWorld 2.0 at lower cost. Both point toward Opus 5 as a strong fit for agentic and automation workloads specifically, not just general chat or writing tasks.
Was Claude Opus 5 the only model Anthropic released recently?
No. Anthropic describes Opus 5 as its fourth model launch in two months, following a rapid release cadence that also included Fable 5 and Mythos 5 earlier in the same window.
Sources: Anthropic, Introducing Claude Opus 5, Artificial Analysis Intelligence Index leaderboard, CNBC, Anthropic's Claude Opus 5 launch, TechCrunch, Anthropic launches Opus 5
Continue Reading
Claude Opus 5 Pricing: $5 and $25 per Million Tokens
Claude Opus 5 costs $5 per million input tokens and $25 per million output, half of Fable 5's rate. What it means for cost per task and who should switch.
Claude Opus 5 vs Fable 5: Which Should You Actually Use in 2026?
Claude Opus 5 costs half of Fable 5 per token and ranks #1 on the Artificial Analysis Intelligence Index. Here is the honest comparison and which one to use.
7 Best AI Agent Security Tools in 2026 (Verified, Compared)
What the Hugging Face breach showed. On July 20, 2026, Axios reported that Hugging Face had disclosed a breach of part of its production infrastructure that it described, in [its own incident write…
Book a Free Strategy Call
Building this in production?
Walid runs a 30-min call to map your AI engineering team. Free, no slides.
Free weekly brief
Steal our production automations
The exact n8n flows, Claude Code setups, and prompts we ship for clients, broken down step by step. No spam, unsubscribe anytime.

Adel keeps the engine running at AY Automate. He owns internal processes, team coordination, and the operational excellence that lets us ship fast for clients.



