claude opus 5 · released 24 july 2026
Opus 5, no hype.
5 rules nobody is talking about.
Everything here comes from Anthropic's own charts and docs. Including the parts that make them look bad.
drag to spin · built by opus 5
Never use max effort.
It costs 14% more and gives you a worse answer. The peak is one setting below.
Coding goes to xhigh. Everything else starts at medium.
The line goes up, then turns down at the last dot
tap the chart to open it full sizenotes
The numbers. Last two dots on the red Opus 5 line: 44.4% at $14.50, then 43.3% at $16.50. More money, lower score. Each dot is one effort level: low, medium, high, xhigh, max.
Also true. The same dip appears on two more benchmarks. The AA Coding Agent Index loses 2.8% for 16% more money. Humanity's Last Exam with tools costs 38% more to score the same. Anthropic's summary table prints the max number, so their headline understates their own model.
Delete "verify your work" from your prompt.
Opus 5 already checks itself. Telling it to check makes it check twice and bill you for both.
Open your prompt file. Delete the red. Paste the blue. One minute, works forever.
old way
Written for Opus 4.8
- Always verify your work before responding.
- Double-check your answer.
- Use a subagent to verify the result.
- Delegate aggressively to subagents.
new way
Written for Opus 5
- Keep responses focused, brief, and concise.
- Delegate to a subagent only for large tasks that are genuinely independent and parallelizable.
- Do not use subagents to verify your own work.
notes
Anthropic's own wording. Verification instructions "cause over-verification on Claude Opus 5, and removing them reduces wasted tokens with no loss in quality." Every line in the blue card is lifted straight from their prompting guide.
The conciseness line is a separate fix. Effort controls how much the model thinks, not how much it says. Turning effort down will not shorten the answer, so you have to ask for it.
Skip Fable 5.
Twice the price of Opus 5. Loses on 9 of the 14 benchmarks Anthropic published.
Use Opus 5. Touch Fable only after xhigh has failed the same problem twice.
Opus 5 wins the five that matter. Fable 5 wins four by a rounding error.
| benchmark | opus 5 $5 / $25 | fable 5 $10 / $50 |
|---|---|---|
| Frontier-Bench | 43.3% | 33.7% |
| AutomationBench | 26.0% | 17.4% |
| OSWorld 2.0 | 70.6% | 66.1% |
| BrowseComp | 90.8% | 87.4% |
| ARC-AGI-3 | 30.2% | no score |
| Legal Agent | 11.7% | 13.3% |
| DeepSWE | 68.8% | 69.7% |
| HLE, no tools | 56.3% | 56.5% |
| FrontierCode | 53.4% | 53.5% |
Frontier-Bench
opus 5
43.3%
fable 5
33.7%
AutomationBench
opus 5
26.0%
fable 5
17.4%
OSWorld 2.0
opus 5
70.6%
fable 5
66.1%
BrowseComp
opus 5
90.8%
fable 5
87.4%
ARC-AGI-3
opus 5
30.2%
fable 5
no score
Legal Agent
opus 5
11.7%
fable 5
13.3%
DeepSWE
opus 5
68.8%
fable 5
69.7%
HLE, no tools
opus 5
56.3%
fable 5
56.5%
FrontierCode
opus 5
53.4%
fable 5
53.5%
price per million tokens, input then output. opus 5 $5 / $25, fable 5 $10 / $50.
notes
The rounding errors. Fable 5's last three wins are 0.9, 0.2 and 0.1 points. Its one real win is legal work, where the best score on the entire board is 13.3%, so every model fails 87% of the time there.
Never shown on a chart. Fable 5 requires 30 day data retention, cannot be used under zero data retention, and its safety classifiers interrupt roughly seven times more often than Opus 5's.
Match the model to the task.
The gap between cheapest and priciest is ten times. Most people run one model for everything.
Haiku for boring repeat work. Sonnet for writing and answers. Opus for long jobs it finishes on its own.
What people actually use Claude for, and what to run it on
Work
43%
Personal
40%
Coursework
17%
The cheapest model is 10 times cheaper than the priciest
notes
Price is per million tokens, input then output. Output always costs five times input, on every model without exception. Bar length is the input price.
Real tasks. Haiku: clean 200 messy addresses, 50 subject line variants, tag a folder of receipts. Sonnet: draft the email, summarize the 40 page PDF, write one function, explain a contract. Opus: refactor 20 files, run an agent unattended for an hour, automate a workflow end to end.
The overspend. Hobbies and lifestyle is 9.3% of all usage and the single biggest place people burn Opus money for zero gain.
It finds security bugs. Do not believe it cannot exploit them.
Anthropic says the gap between finding and weaponizing is what makes Opus 5 safe to ship.
Finding the hole: Opus 5 is already at the frontier
Building the weapon: Opus 5 is far behind
Finding the flaw needs intelligence. Writing the exploit needs time. They gated the cheap half.
notes
What the charts measure. Top: percentage of OSS-Fuzz challenges where the model correctly identified the vulnerability, higher is better. Bottom: number of challenges where it produced a working exploit, out of the same set. Mythos 5 is Anthropic's restricted top model, so Opus 5 is 0.6 points off the frontier at finding and three times behind at weaponizing.
The tell. Anthropic runs a Cyber Verification Program that hands approved companies a build of Opus 5 with the safeguards removed. The classifiers block binary scanning, penetration testing and exploit writing, while explicitly allowing source code vulnerability hunting.
The fair counter. Turning a memory bug into a working exploit really does need separate skills, like defeating memory layout randomization. This is opinion, not a finding.
the bottom line
57 days. Same price. Double the score.
The last Opus came out in May. This one costs exactly the same and scores twice as high on agentic coding.
28 may 2026
Claude Opus 4.8
21.1%
$5 in / $25 out
24 july 2026 · 57 days later
Claude Opus 5
43.3%
$5 in / $25 out
notes
The score. Frontier-Bench v0.1, agentic terminal coding, at max effort for both models. Opus 4.8 scored 21.1%, Opus 5 scored 43.3%, and the price per token did not move between the two releases.
The point. Every cost curve in this release shifted sideways toward cheaper, not upward toward a new ceiling. So the move is to spend less for the same result, not more for a better one. All five rules are versions of that.