oleg melnikovyoutube

claude opus 5 · released 24 july 2026

Opus 5, no hype.
5 rules nobody is talking about.

Everything here comes from Anthropic's own charts and docs. Including the parts that make them look bad.

drag to spin · built by opus 5

01

Never use max effort.

It costs 14% more and gives you a worse answer. The peak is one setting below.

do this

Coding goes to xhigh. Everything else starts at medium.

The line goes up, then turns down at the last dot

Anthropic's Frontier-Bench chart plotting score against cost per attempt. The Opus 5 line rises to 44.4 percent at about 14.50 dollars and then falls to 43.3 percent at about 16.50 dollars.tap the chart to open it full size

notes

The numbers. Last two dots on the red Opus 5 line: 44.4% at $14.50, then 43.3% at $16.50. More money, lower score. Each dot is one effort level: low, medium, high, xhigh, max.

Also true. The same dip appears on two more benchmarks. The AA Coding Agent Index loses 2.8% for 16% more money. Humanity's Last Exam with tools costs 38% more to score the same. Anthropic's summary table prints the max number, so their headline understates their own model.

02

Delete "verify your work" from your prompt.

Opus 5 already checks itself. Telling it to check makes it check twice and bill you for both.

do this

Open your prompt file. Delete the red. Paste the blue. One minute, works forever.

old way

Written for Opus 4.8

  • Always verify your work before responding.
  • Double-check your answer.
  • Use a subagent to verify the result.
  • Delegate aggressively to subagents.

new way

Written for Opus 5

  • Keep responses focused, brief, and concise.
  • Delegate to a subagent only for large tasks that are genuinely independent and parallelizable.
  • Do not use subagents to verify your own work.

notes

Anthropic's own wording. Verification instructions "cause over-verification on Claude Opus 5, and removing them reduces wasted tokens with no loss in quality." Every line in the blue card is lifted straight from their prompting guide.

The conciseness line is a separate fix. Effort controls how much the model thinks, not how much it says. Turning effort down will not shorten the answer, so you have to ask for it.

03

Skip Fable 5.

Twice the price of Opus 5. Loses on 9 of the 14 benchmarks Anthropic published.

do this

Use Opus 5. Touch Fable only after xhigh has failed the same problem twice.

Opus 5 wins the five that matter. Fable 5 wins four by a rounding error.

  • Frontier-Bench

    opus 5

    43.3%

    fable 5

    33.7%

  • AutomationBench

    opus 5

    26.0%

    fable 5

    17.4%

  • OSWorld 2.0

    opus 5

    70.6%

    fable 5

    66.1%

  • BrowseComp

    opus 5

    90.8%

    fable 5

    87.4%

  • ARC-AGI-3

    opus 5

    30.2%

    fable 5

    no score

  • Legal Agent

    opus 5

    11.7%

    fable 5

    13.3%

  • DeepSWE

    opus 5

    68.8%

    fable 5

    69.7%

  • HLE, no tools

    opus 5

    56.3%

    fable 5

    56.5%

  • FrontierCode

    opus 5

    53.4%

    fable 5

    53.5%

price per million tokens, input then output. opus 5 $5 / $25, fable 5 $10 / $50.

notes

The rounding errors. Fable 5's last three wins are 0.9, 0.2 and 0.1 points. Its one real win is legal work, where the best score on the entire board is 13.3%, so every model fails 87% of the time there.

Never shown on a chart. Fable 5 requires 30 day data retention, cannot be used under zero data retention, and its safety classifiers interrupt roughly seven times more often than Opus 5's.

04

Match the model to the task.

The gap between cheapest and priciest is ten times. Most people run one model for everything.

do this

Haiku for boring repeat work. Sonnet for writing and answers. Opus for long jobs it finishes on its own.

What people actually use Claude for, and what to run it on

Work

43%

Personal

40%

Coursework

17%

Content creation22.9%Sonnet 5
Education & learning12.8%Sonnet 5
Software development11.7%Opus 5
Research10.9%Opus 5
Hobbies & lifestyle9.3%Haiku 4.5
Business ops4.6%Opus 5
Document processing4.1%Haiku 4.5
Data analysis3.5%Sonnet 5

The cheapest model is 10 times cheaper than the priciest

Haiku 4.5$1 / $5
Sonnet 5$3 / $15
Opus 5$5 / $25
Fable 5$10 / $50

notes

Price is per million tokens, input then output. Output always costs five times input, on every model without exception. Bar length is the input price.

Real tasks. Haiku: clean 200 messy addresses, 50 subject line variants, tag a folder of receipts. Sonnet: draft the email, summarize the 40 page PDF, write one function, explain a contract. Opus: refactor 20 files, run an agent unattended for an hour, automate a workflow end to end.

The overspend. Hobbies and lifestyle is 9.3% of all usage and the single biggest place people burn Opus money for zero gain.

05

It finds security bugs. Do not believe it cannot exploit them.

Anthropic says the gap between finding and weaponizing is what makes Opus 5 safe to ship.

Finding the hole: Opus 5 is already at the frontier

Opus 4.861.5%
Opus 579.4%
Mythos 580.0%

Building the weapon: Opus 5 is far behind

Opus 4.80
Opus 54
Mythos 513
my take

Finding the flaw needs intelligence. Writing the exploit needs time. They gated the cheap half.

notes

What the charts measure. Top: percentage of OSS-Fuzz challenges where the model correctly identified the vulnerability, higher is better. Bottom: number of challenges where it produced a working exploit, out of the same set. Mythos 5 is Anthropic's restricted top model, so Opus 5 is 0.6 points off the frontier at finding and three times behind at weaponizing.

The tell. Anthropic runs a Cyber Verification Program that hands approved companies a build of Opus 5 with the safeguards removed. The classifiers block binary scanning, penetration testing and exploit writing, while explicitly allowing source code vulnerability hunting.

The fair counter. Turning a memory bug into a working exploit really does need separate skills, like defeating memory layout randomization. This is opinion, not a finding.

the bottom line

57 days. Same price. Double the score.

The last Opus came out in May. This one costs exactly the same and scores twice as high on agentic coding.

28 may 2026

Claude Opus 4.8

21.1%

$5 in / $25 out

24 july 2026 · 57 days later

Claude Opus 5

43.3%

$5 in / $25 out

notes

The score. Frontier-Bench v0.1, agentic terminal coding, at max effort for both models. Opus 4.8 scored 21.1%, Opus 5 scored 43.3%, and the price per token did not move between the two releases.

The point. Every cost curve in this release shifted sideways toward cheaper, not upward toward a new ceiling. So the move is to spend less for the same result, not more for a better one. All five rules are versions of that.