Claude Opus 5: Same Price, a Generational Leap

Anthropic has shipped four models in under two months. The latest arrived on Friday, July 24: Claude Opus 5.
This time the surprise is not in the price. Opus 5 costs exactly the same as its predecessor, Opus 4.8: $5 per million input tokens and $25 per million output tokens. The surprise is the size of the leap you get for that unchanged price.
Anthropic positions the model as one that approaches Fable 5’s frontier intelligence at half the cost. Yet the published numbers show Opus 5 doing more than approaching: it overtakes the twice-as-expensive Fable 5 on most coding and knowledge-work evaluations.
Claude Opus 5 in 30 seconds
• Release date: July 24, 2026
• Pricing: $5 input / $25 output per million tokens — identical to Opus 4.8
• Context window: 1 million tokens
• Knowledge cutoff: May 2026 — the most recent among current Claude models
• Standout benchmarks: Frontier-Bench, OSWorld 2.0, GDPval-AA v2, ARC-AGI 3, AutomationBench
• Where it trails: SWE-bench Pro by 0.8 points; cyber exploitation, where Mythos 5 leads
What has happened over the past two months
A short look back is needed to see where Opus 5 fits. The model lineup has changed almost entirely since early June.
• May 28 — Opus 4.8. Held the top-of-frontier position for a while.
• June 9 — the Mythos tier. Anthropic opened a new tier above Opus. The first two models, Claude Mythos 5 and Claude Fable 5, launched on the same day. They share the same underlying model; Fable 5 ships with additional safeguards in biology, cybersecurity and LLM R&D.
• June 12 — access suspended. Anthropic halted access to both models to comply with U.S. Department of Commerce export controls. For roughly three weeks the top tier was effectively unavailable.
• June 30 — controls lifted. Anthropic restored access on July 1. Sonnet 5 shipped the same day: a model that pulled the mid tier into Opus territory, offered at a promotional $2 / $10 through August 31.
• July 24 — Opus 5. Reshuffled the table once more, this time with the leap showing up purely in performance rather than in price.
That pace creates a practical problem: a model choice made two months ago is very likely out of date today. Existing pipelines may still be pinned to Opus 4.8 or something older. The difference Opus 5 brings is large enough to justify revisiting those decisions.
Claude Opus 5 specifications and pricing
Specification | Value |
Context window | 1 million tokens (both default and maximum) |
Maximum output | 128,000 tokens |
Thinking | On by default |
Knowledge cutoff | May 2026 (the most recent among current Claude models) |
Pricing | $5 input / $25 output per million tokens — identical to Opus 4.8 |
Data retention | No data-retention requirement for general access |
It is now the default model on the Claude Max plan, and the most capable model available on Claude Pro.
The model also exposes effort levels: low, medium, high (default), xhigh and max. Some of the scores below shift with this setting, which is worth keeping in mind when reading the benchmark figures.
Why the Opus tier is different this time
The familiar story with model tiers is that the lower tier closes in on the one above it. With Opus 5 the opposite happened: the Opus tier not only catches up with the Fable tier above it, it passes it in several places — at half the price.
Anthropic’s own emphasis is not on any single benchmark but on the shape of the curve. Nearly every published chart plots score against dollars spent, and the Opus 5 curve sits above and to the left of Fable 5’s. The claim is not “pay more, get more” but “pay less, get more.”
Claude Opus 5 benchmark results
The figures below come from Anthropic’s launch notes and system card; some have been cross-checked against independent write-ups of the same material.
Two caveats are worth stating up front:
• A significant share of the scores shift noticeably with the effort setting.
• In the Frontier-Bench run, calls refused by the safety classifier were routed down to Opus 4.8. That result is therefore not a pure model measurement but a system measurement.
Coding benchmarks
On Anthropic’s new agentic coding evaluation, Frontier-Bench v0.1, Opus 5 scores 43.3 percent. Opus 4.8 managed 18.7 percent; Fable 5 reached 33.7 percent and GPT-5.6 Sol 34.4 percent. That is more than double the previous generation, at a lower cost per task.
On SWE-bench Pro the field is tighter: Opus 5 scores 79.2 percent, while Fable 5 (80.0 percent) and Mythos 5 (80.3 percent) stay narrowly ahead. But with Opus 4.8’s 69.2 percent as the reference point, the distance closed in a single generation is clear. Paying double for the remaining 0.8 points is not a compelling trade for most teams.
On SWE-bench Verified, a score of 96.0 percent suggests the benchmark is effectively saturated.
On CursorBench 3.2, Opus 5 lands within 0.5 points of Fable 5’s top score at roughly half the cost per task.
Computer use and automation
On OSWorld 2.0, Opus 5 scores 70.6 percent against Fable 5’s 66.1 percent and Opus 4.8’s 55.7 percent. In Anthropic’s framing, Opus 5 beats Fable 5’s best result at roughly a third of the cost.
Zapier AutomationBench — which measures whether business tasks are completed end to end — gives Opus 5 26.0 percent, against 17.0 percent for Opus 4.8 and 17.4 percent for Fable 5. Even at its lowest effort setting it clears more tasks than the other models.
Zapier CEO Wade Foster reports an end-to-end workflow that starts from a raw account-health table, flags at-risk accounts, notifies the right owner and produces a summary — a flow no previous model completed, and which Opus 5 finished 100 percent of the time.
Knowledge work and reasoning
On GDPval-AA v2, Opus 5 takes the lead with an Elo of 1,861, against 1,747 for Fable 5 and 1,736 for GPT-5.6 Sol. One note is needed here: Elo tables are periodically rebaselined, so placing scores published on different dates side by side is misleading. What matters is the ranking within the same run.
On ARC-AGI 3 — which measures performance on problem types the model has not seen before — Opus 5 scores 30.2 percent while the nearest competitor sits around 8 percent, a gap of roughly four times. A margin that large is unusual and deserves a degree of healthy scepticism; rather than trusting a single run, it is worth verifying against your own workload.
Science and visual output
Opus 5 outperforms Opus 4.8 across the whole of Anthropic’s life-sciences evaluation suite. The clearest gains are in organic chemistry:
• a 10.2-point gap in deriving molecular structure from spectroscopy data
• a 7.7-point gap in predicting how sequence variation in a protein affects its function
There is a marked step up in visual output as well; at launch, Anthropic showcased interactive simulations produced by the model.
What it changes in practice
Anthropic’s narrative converges on a single behaviour: the model verifies its own work and iterates until it succeeds.
The launch examples point that way. In one Frontier-Bench task, the model is asked to turn an engineering drawing of a machine part into a 3D FreeCAD model, with direct access to the drawing deliberately withheld. Opus 5 writes its own computer-vision pipeline, recovers the geometry from raw pixels and rebuilds the part. No other model solved it in five attempts under the same setup.
On a real bug in a widely used package manager, it found the edge case the community patch had missed and fixed the root cause. A competing model addressed only the surface symptom and declared the bug resolved.
Early access feedback
The common thread is, notably, efficiency:
• Harvey: reports performance comparable to Opus 4.8’s maximum reasoning mode while generating on average 26 percent fewer tokens.
• Fundamental Research Lab: reports on average 9 points higher accuracy on hard financial-modelling tasks, using a third fewer turns and tool calls, in 60 percent less time.
• Box: reports an 8 percent overall improvement over Opus 4.8, 11 percent on data analysis and 17 percent on due-diligence workflows.
That these are all efficiency numbers matters. They do not say “smarter” so much as “the same work with fewer tokens and fewer turns” — and that is the side that shows up directly on the bill.
Does anything change in how you write prompts?
Instructions you may have grown used to adding for older models “check your work, then review it again” can now trigger something Opus 5 already does, burning tokens for nothing.
When porting existing prompts to Opus 5, it is worth revisiting these self-verification instructions. The model exhibits the behaviour by default.
Safety and alignment
In Anthropic’s automated behavioural audit, Opus 5 stands out as its most aligned model to date. Its overall misalignment score of 2.3 is the lowest among recent models.
It adheres to Claude’s constitution more closely than Opus 4.8, Sonnet 5 and Fable 5; it has the lowest rate of deceptive behaviour and is the most resistant to being manipulated into misuse. It also produces the safest result on avoiding reckless actions with hard-to-reverse side effects.
In other words, capability and safety moved in the same direction with Opus 5. The “more capable but riskier” trade-off common across tiers does not apply here.
The line drawn on the cyber side
Opus 5 was deliberately not trained on cyber tasks. Even so, the general capability increase brought it close to Mythos 5 on vulnerability discovery. It remains clearly behind on turning a vulnerability into an actual threat — that is, on exploitation.
The classifiers permit vulnerability discovery in source code while blocking binary-based scanning, penetration testing and exploit generation. Roughly 85 percent fewer interventions are expected compared with Fable 5.
On the biology side, Opus 5 is now the most capable scientific research model available in general access.
Opus 5 vs Fable 5 vs Sonnet 5 vs GPT-5.6 Sol
Opus 5 is best read not in isolation but alongside the other options on the table at the same time:
Model | Price (input / output, per million tokens) | Position relative to Opus 5 |
Opus 5 | $5 / $25 | The reference point. Leads on Frontier-Bench, OSWorld 2.0, GDPval-AA v2, ARC-AGI 3 and AutomationBench. No data-retention requirement. |
Fable 5 | $10 / $50 | Behind on most shared benchmarks, at twice the price. What it retains: 0.8 points on SWE-bench Pro and a few decimals on CursorBench. It also carries a data-retention requirement. |
Mythos 5 | Not in general access | Still ahead on cyber exploitation and autonomous biology research. In use with a small number of organisations under Project Glasswing. |
Sonnet 5 | $2 / $10 (through August 31), then $3 / $15 | The difference is no longer intelligence but unit cost. Still an overwhelming advantage for high-volume, repetitive work. |
GPT-5.6 Sol | Same as Opus 5 on input, roughly 17% more expensive on output | Opus 5 leads on Frontier-Bench, ARC-AGI 3, GDPval-AA v2, OSWorld 2.0, CursorBench and AutomationBench. Sol leads on DeepSWE; the coding-agent index is close to a tie. |
The summary of that table: Fable 5 is no longer the default top step, while Sonnet 5 holds its ground at the bottom. Opus 5 has widened the band between them considerably.
Who is Claude Opus 5 for?
• Teams building long-running, unsupervised automation: the verify-and-iterate behaviour pays off most directly in day-to-day operational automation.
• Computer use and interface automation: a result on OSWorld 2.0 that is both higher and cheaper closes the argument.
• Knowledge-work-heavy teams: for producing reports, analyses, spreadsheets and presentations, the GDPval-AA v2 lead is the benchmark that maps most closely onto real work.
• Current Fable 5 users: for most workloads, moving to Opus 5 is both cheaper and better on most benchmarks. Keeping Fable 5 makes more sense only where you have verified it makes a real difference on a task you actually measure.
• High-volume, simple work: Opus 5 is the wrong tool. Sonnet 5 still wins on unit cost.
Conclusion
Opus 5 is a release that shows a generational jump does not have to show up on the price tag. For the same $5 / $25, you get a model that more than doubles Opus 4.8 on Frontier-Bench and beats the twice-as-expensive Fable 5 on computer use and knowledge work.
Where it trails is narrow and well defined: under a point on SWE-bench Pro, and a deliberate gap against Mythos 5 on cyber exploitation.
After Opus 5, model selection has actually got a little simpler: Sonnet 5 for work where volume and cost dominate, Opus 5 for almost everything else. Fable 5 is no longer the default, but the upper step you reach for once you have a verified need.



