buildfastwithaibuildfastwithai
AI WorkshopsAll blogsAgentic AI Launchpad
Agentic AI Launchpad
Download Unrot App
Free AI Workshop
Mentorship

Agentic AI Launchpad

Go from user to builder in 6 weeks.

Explore Program
Claude Mastery Course
Share
Back to blogs
Reviews

Claude Opus 5 Review: Benchmarks, Price & Honest Take

July 25, 2026
15 min read
Share:
Claude Opus 5 Review: Benchmarks, Price & Honest Take
Share:

Claude Opus 5 Review: Fable 5 Power at Half the Price?

Anthropic just released a model that beats its own flagship on some benchmarks and costs half as much. Claude Opus 5 landed on July 24, 2026, priced at $5 per million input tokens against Claude Fable 5's $10, and it posts a Frontier-Bench score of 43.3% that tops Fable 5's own 33.7%. On paper, that makes the cheaper model look like the better buy.

On paper is doing a lot of work in that sentence. Opus 5 is genuinely impressive, it more than doubles Opus 4.8 on hard coding, and it also lost four published benchmarks, ran its own evaluations rather than independent ones, and quietly leaned on Opus 4.8 as a fallback on the very benchmark it leads. This review covers all of it: the wins, the losses, the caveat most coverage skipped, and whether Opus 5 should replace Opus 4.8 in your stack.

The one-line verdict: Claude Opus 5 is the best value in Anthropic's lineup and a clear upgrade over Opus 4.8, but it is a near-frontier model, not the frontier. Fable 5 is still the flagship, the benchmark table has an asterisk, and the effort dial is the feature that actually matters.

What Is Claude Opus 5?

Claude Opus 5 is Anthropic's new near-frontier model, released July 24, 2026, that reaches roughly Claude Fable 5-level intelligence at half the price, with a low/medium/high effort toggle and a 1M-token context window. It is available across Claude.ai, the API as claude-opus-5, Claude Code and Claude Cowork, and it is the default model on the Claude Max plan.

The positioning is the thing to get right, because it is easy to misread. Opus 5 is not Anthropic's most capable model. Claude Fable 5 remains the frontier flagship, and Mythos 5 still leads on restricted cybersecurity tasks. Opus 5 is the value tier: most of Fable 5's capability, at Opus 4.8's price, with a dial that lets you spend more only when a task needs it. Anthropic is competing with itself on cost, which is a good problem for buyers.

Table 1: Claude Opus 5 at a glance

Screenshot 2026-07-25 154855

Fable 5 stays the frontier flagship at $10 / $50. Opus 5 targets the same quality at half the input price.

If you are upgrading from the previous generation, our Claude Opus 4.8 review is the right baseline to read alongside this one, since Opus 5 keeps its price and roughly doubles its hardest coding scores.

Claude Opus 5 Pricing and the Effort Dial

Claude Opus 5 costs $5 per million input tokens and $25 per million output, identical to Opus 4.8 and exactly half of Fable 5's $10 and $50. The headline is not a price cut, it is that you now get near-Fable capability at the Opus price you were already paying.

The genuinely new feature is the effort toggle. You set low, medium or high effort per request, trading cost and latency for capability. A routine call runs cheap and fast on low; a hard reasoning problem gets the full model on high. It is the same idea appearing across the industry this month, from Inkling's 0.2-to-0.99 dial to GPT-5.6's ultra mode, and it is quickly becoming a standard control rather than a differentiator.

The strategic read is that Anthropic put a dial on your AI bill. For agent workflows, where token consumption runs high and most steps are routine, defaulting to low or medium effort and reserving high for the hard steps can cut costs substantially without hurting output quality. That routing discipline is where the real savings live, more than the sticker price.

A COMPLIANCE DETAIL WORTH FLAGGING

Opus 5 carries no 30-day data-retention requirement, unlike Fable 5. For teams under strict data-protection rules, that difference can matter as much as any benchmark, because it changes what you are allowed to send the model in the first place.

Benchmarks: Where Opus 5 Wins

Opus 5 sets new highs on agentic coding, computer use, business automation and abstract reasoning, in several cases beating Fable 5 outright. The standout is ARC-AGI-3, where it scores roughly three times the next-best model and twenty times Opus 4.8.

Table 2: Where Opus 5 leads

claude opus 5

Opus 5 more than doubles Opus 4.8 on Frontier-Bench. On ARC-AGI-3, Opus 4.8 scored just 1.5%.

A few of these deserve a second look. Beating Fable 5 on OSWorld 2.0 and GDPval, at half the cost, is the clearest evidence for the value pitch, because those are computer-use and knowledge-work tasks where the frontier flagship was supposed to lead. On science, Opus 5 also improves 10.2 points over Opus 4.8 on organic chemistry and 7.7 points on protein prediction, which points at real gains in structured reasoning rather than just coding.

The ARC-AGI-3 number is the one to remember: 30.2% against the next model's 7.8% is not an incremental win, it is a different tier of abstract reasoning. If it holds up under independent testing, that gap is the most impressive thing about this release.

Benchmarks: Where Opus 5 Loses

Opus 5 lost four published benchmarks, and to Anthropic's credit those losses appear in the same materials as the wins. A launch that only shows victories is an advertisement; one that shows defeats is closer to a report.

Table 3: Where Opus 5 trails

claude opus 5

The Humanity's Last Exam gap is 0.2 points, effectively a tie. The HealthBench gap to Mythos 5 is the widest.

Read the pattern, not just the rows. Opus 5 trails GPT-5.6 Sol on one coding benchmark while beating it on others, which tells you these two are genuinely close on software work and the winner depends on the specific test. It trails Fable 5 by a whisker on reasoning and by more on law, and it sits well behind Mythos 5 on professional health tasks, which is expected since Mythos is the restricted, safety-lifted tier. None of these are embarrassing, and publishing them buys credibility for the wins.

My take: the honest loss table is the reason to trust the win table more, not less. A model that only ever wins is a model whose benchmarks you should discount. Opus 5 loses in believable places, which makes its victories more believable too. That said, believable is not the same as verified, and the next section is why.

🚀 Cohort Waitlist Open
Go From AI User to AI Builder

Don't just use ChatGPT. Learn to build custom LLM agents, RAG pipelines, and full-stack Agentic AI apps in our intensive 6-week program.

6 Weeks Live Mentorship
Deploy 5+ Real-world Apps
Weekly App Templates & Code
No Coding Experience Required
Explore Program
Join 1,000+ graduates•Free Registration

The Safety-Classifier Caveat You Should Know

On Frontier-Bench, the benchmark Opus 5 leads, Anthropic used Opus 4.8 as a fallback whenever a safety classifier refused a request, and it did not disclose how often that happened. That single footnote is the most important thing in the entire launch, and it changes how you should read the headline score.

Here is why it matters. If a safety classifier blocks Opus 5 on some fraction of Frontier-Bench tasks, and the older Opus 4.8 quietly answers those instead, then the reported 43.3% is a blend of two models, not a clean measurement of Opus 5 alone. Without the substitution rate, there is no way to know whether the number reflects Opus 5's real capability or a spliced result. The magnitude is undisclosed, so the effect could be trivial or significant, and we cannot tell which.

WHY THIS IS THE STORY, NOT A FOOTNOTE

A benchmark where a different model stands in for refused requests is not a benchmark of one model. Anthropic disclosed the substitution, which is to its credit, but withholding the rate means the flagship coding score cannot be taken at face value. Until that number is published or independent testers reproduce the result, treat 43.3% on Frontier-Bench as provisional.

The broader caveat compounds it: several of these benchmarks were built by outside organisations, but the runs were Anthropic's own rather than independent. That is normal for a launch and it is still a limitation. Self-run numbers are a claim, not a verified fact, and the safety-classifier fallback is a concrete reason to wait for third-party confirmation before betting a migration on the coding lead specifically.

My contrarian point: this is the rare case where a model's honesty about its own testing is also the reason to be cautious about its top score. Both things are true. Respect the transparency, and still verify on your own workload before you trust the headline.

This is not the first time an Opus benchmark story needed a closer look. Our Claude Opus 4.7 regression analysis is a reminder that launch numbers and lived experience do not always match.

What Actually Changed: Agentic Reasoning

The real upgrade in Opus 5 is agentic reasoning, not raw intelligence. Its strength is checking its own work, iterating when it hits a blocker, and building internal tools like test harnesses and computer-vision pipelines to solve a task, rather than a step change in single-shot answers.

That distinction shapes where it shines and where it costs you. On a long-chain agentic task, self-verification and autonomous iteration mean Opus 5 keeps making progress where a lesser model stalls or hands back a wrong answer confidently. Users report it completes autonomous tasks without needing external validation, which is exactly the behaviour that makes an agent trustworthy enough to leave running. The trade-off is token consumption: all that self-checking and tool-building burns tokens, so the effort dial and disciplined routing are how you keep the bill sane.

Quotable version: Opus 5 is not much smarter than Opus 4.8 at answering, it is much better at not giving up. For agents, that matters more than a few benchmark points.

If you want to build reliable agent loops around a model like this, the self-verification pattern is the whole game, and it pairs directly with the multi-agent approach we see across the July 2026 model field.

Opus 5 vs Opus 4.8 vs Fable 5

Opus 5 is a clear upgrade over Opus 4.8 at the same price, and a near-match for Fable 5 at half the price, with Fable 5 still ahead on the very hardest reasoning. Here is the three-way picture that decides which Claude you should actually run.

Table 4: The Claude lineup compared

claude opus 5

Opus 4.8's Frontier-Bench figure is implied by Anthropic's claim that Opus 5 more than doubles it.

The upgrade case from Opus 4.8 is straightforward: same price, meaningfully better on nearly everything, plus the effort dial and the retention advantage. I would move to Opus 5 for most Opus 4.8 workloads without hesitation. The Fable 5 question is subtler. If you need the absolute best reasoning and can afford double the price, Fable 5 still wins the hardest tasks. If you want most of that quality at half the cost, Opus 5 is now the smart default, with the safety-classifier caveat noted.

For how these Claude models stack against the open and OpenAI competition on coding specifically, our GLM-5.2 vs Claude vs GPT-5.6 vs Kimi comparison is the companion read.

Fast Mode: Is It Worth 2x the Price?

Opus 5 Fast mode runs at roughly 2.5x the default speed for double the price, which is a better ratio than it first sounds. You pay 2x and get 2.5x the speed, so the cost per unit of time actually improves for latency-sensitive work.

Whether it is worth it comes down to a simple question: is a human or a pipeline waiting on the output? For an interactive coding session, a live agent, or anything where a developer sits watching a spinner, the speed is worth the premium because engineer time costs far more than tokens. For batch jobs, overnight runs, or anything asynchronous, the extra cost buys you nothing you can use, and default mode is the right call. This is the same calculus Anthropic has offered since the Opus 4.6 and 4.7 Fast modes, and the answer has not changed: pay for speed only when speed is the constraint.

My rule: Fast mode for anything a person watches, default mode for anything a queue processes. The 2.5x-for-2x ratio makes it an easy yes for interactive work and an easy no for background jobs.

We ran the full cost-benefit on the previous generation in our Claude Opus 4.7 Fast Mode guide, and the framework there applies cleanly to Opus 5.

🚀 Cohort Program Open
Claude Mastery: Cowork & Code

The only comprehensive program designed to take you from basic prompting to building interactive Artifacts, custom integrations, and deploying production-ready code with Claude Code.

No coding experience needed
Build interactive Artifacts & Agents
Deploy apps with Claude Code
Cohort-based learning & mentorship
Explore Program
Cohort-based training•Register Now

Who Should Use Claude Opus 5

Use Claude Opus 5 as your new default if you run agentic coding, computer-use automation, or knowledge work and want near-Fable quality at half the price. Stay on Fable 5 only if you need the absolute best on the hardest reasoning, law, or professional health tasks and can pay double.

Table 5: Fit by use case

claude opus 5claude opus 5

Opus 5 is the default recommendation for most teams. Reserve Fable 5 and Mythos 5 for the specific tasks where they still lead.

Verdict, 9 out of 10: Claude Opus 5 is the best value in AI coding right now and an easy upgrade from Opus 4.8. I dock one point for the self-run benchmarks and the undisclosed safety-classifier fallback on its headline score, which mean you should verify the coding lead on your own repository before betting a migration on it. Everything else about this release earns the hype, and the effort dial is the quiet feature that will save teams the most money.

For the full month's context and where Opus 5 lands against every rival, see our best AI models July 2026 ranking and the GPT-5.6 review it competes with most directly.

Frequently Asked Questions

Q: What is Claude Opus 5?

Claude Opus 5 is Anthropic's near-frontier AI model, released July 24, 2026, that reaches roughly Claude Fable 5-level intelligence at half the price. It has a low/medium/high effort toggle, a 1M-token context window, and is available on Claude.ai, the API as claude-opus-5, Claude Code and Claude Cowork. It is the default model on Claude Max.

Q: How much does Claude Opus 5 cost?

Claude Opus 5 costs $5 per million input tokens and $25 per million output tokens, the same as Opus 4.8 and half of Fable 5's $10 and $50. A Fast mode runs at about 2.5x speed for double the price, roughly $10 input and $50 output per million tokens.

Q: Is Claude Opus 5 better than Fable 5?

On some benchmarks yes, overall not quite. Opus 5 beats Fable 5 on Frontier-Bench (43.3% vs 33.7%), OSWorld 2.0 and GDPval, at half the price. Fable 5 still leads the hardest reasoning, such as Humanity's Last Exam, and remains Anthropic's frontier flagship. Opus 5 is the value pick, not the outright best.

Q: Is Claude Opus 5 better than Opus 4.8?

Yes, clearly, at the same price. Opus 5 more than doubles Opus 4.8 on Frontier-Bench, jumps from 1.5% to 30.2% on ARC-AGI-3, improves 10.2 points on organic chemistry, and adds an effort toggle and no 30-day data-retention requirement. For most Opus 4.8 workloads, Opus 5 is a straightforward upgrade.

Q: What is the effort toggle in Claude Opus 5?

The effort toggle lets you set low, medium or high effort per request, trading cost and latency for capability. Routine calls run cheap on low, while hard reasoning tasks get the full model on high. Defaulting to low or medium and reserving high for difficult steps can cut agent costs substantially.

Q: What is Claude Opus 5 Fast mode?

Fast mode runs Opus 5 at roughly 2.5x the default speed for double the price, about $10 input and $50 output per million tokens. Because you get 2.5x speed for 2x cost, it is worth it for interactive work where a person waits on output, and not worth it for asynchronous batch jobs.

Q: Is Claude Opus 5 good for coding?

Yes. Opus 5 leads Frontier-Bench at 43.3% and comes within 0.5% of Fable 5 on CursorBench 3.2 at half the cost, with strong self-verification for agentic coding. It trails GPT-5.6 Sol on DeepSWE v1.1, so the two are close on software work. Note that its headline Frontier-Bench score used Opus 4.8 as a fallback on refused requests, so verify on your own code.

Q: Should I switch to Claude Opus 5?

If you use Opus 4.8, yes, since Opus 5 costs the same and performs meaningfully better. If you use Fable 5 and can accept near-frontier rather than absolute-best quality, switching halves your cost. Before migrating coding workloads, verify the benchmark lead on your own repository, given the self-run tests and the safety-classifier caveat.

Recommended Reads

  • Claude Opus 4.8 review
  • Claude Opus 4.7 Fast Mode guide
  • GLM-5.2 vs Claude vs GPT-5.6 vs Kimi
  • Best AI models July 2026 ranking
  • GPT-5.6 Sol Terra Luna review

Half the price for near-flagship quality is the kind of release that resets a whole stack. Test Opus 5 on a real task this week, and follow Build Fast with AI for honest coverage of every major model launch.

References

  • Anthropic, Introducing Claude Opus 5
  • VentureBeat, Anthropic launches Opus 5
  • Tech-ish, Opus 5 launch benchmarks and price
  • Quartz, Opus 5 at half of Fable 5 price
  • Vals AI, Claude Opus 5 evaluations
  • Finout, Claude Opus 5 pricing 2026
Enjoyed this article? Share it →
Share:
    You Might Also Like
    Qwen3.6-27B: 27B Model Beats 397B on Coding (2026)
    Reviews
    Qwen3.6-27B: 27B Model Beats 397B on Coding (2026)

    Qwen3.6-27B scores 77.2% on SWE-bench Verified, beats a 397B MoE, runs on 18GB VRAM, and matches Claude 4.5 Opus on Terminal-Bench. Full review inside.

    Qwen 3.6 Plus Preview: 1M Context, Speed & Benchmarks 2026
    Reviews
    Qwen 3.6 Plus Preview: 1M Context, Speed & Benchmarks 2026

    Qwen 3.6 Plus Preview drops on OpenRouter with a 1M token context, free access, and up to 3x faster speed vs Claude Opus 4.6. Here's the full breakdown.