Anthropic on Thursday introduced Claude Opus 5, a new flagship artificial-intelligence model that the company says comes close to the performance of its more expensive frontier system, Claude Fable 5, while adding a potentially important advance in one of the field’s most stubborn security problems: prompt injection.
The release arrives at a moment when major A.I. companies are no longer competing only on raw benchmark scores. Increasingly, they are also trying to prove that their models can be trusted in more autonomous roles — browsing the web, using software tools and carrying out multistep tasks without being easily manipulated by malicious instructions hidden in websites or documents.
Opus 5 appears aimed squarely at that shift.
A cheaper model near the top tier
Anthropic said Opus 5, released July 24, matches or exceeds top-tier performance on coding and knowledge-work evaluations, including Frontier-Bench and GDPval-AA, while costing less than Fable 5. The company priced the model at $5 per million input tokens and $25 per million output tokens, the same rates as its earlier Opus 4.8 model.
Independent benchmark trackers have also placed the model near the top of the current field. Artificial Analysis ranked Opus 5 at 61 on its Intelligence Index, effectively tied with Fable 5 and ahead of OpenAI’s GPT-5.6 Sol. On ARC-AGI-3, a closely watched test intended to measure novel problem-solving rather than memorized knowledge, ARC Prize listed Opus 5 at 30.2 percent as of July 24, the highest published score for any model.
That score drew particular attention because ARC-style tests have become a proxy, however imperfect, for whether systems can reason through unfamiliar tasks. Reports from evaluators said Opus 5 displayed unusually strong logical behavior, including deriving reflection equations on its own in at least one case.
Anthropic has framed the model less as a moonshot than as a practical high-end option: a system that approaches frontier intelligence without frontier pricing. In a market where the best models are often expensive enough to limit broad deployment, that positioning could matter as much as the leaderboard wins.
A security claim that may matter even more
Yet some of the strongest reaction to Opus 5 has focused not on benchmark performance, but on security.
In its system card, Anthropic said Opus 5 was its “least prompt injectable model yet,” a notable claim in a field where prompt injection has become a central obstacle to building safe A.I. agents. In these attacks, malicious text embedded in a webpage, email or file tries to override a model’s original instructions, potentially causing it to leak information, ignore safeguards or take unintended actions.
For browser-based agents, that problem has been especially severe. A model sent to navigate the open web can encounter hostile instructions almost anywhere.
According to results cited from Anthropic’s system card, Opus 5 used with the company’s added “Auto Mode” defenses recorded a 0 percent prompt-injection success rate across 129 browser-agent test scenarios. Without those extra product-layer protections, the success rate was reported at 3.7 percent. In a separate benchmark from Gray Swan, the reported prompt-injection success rate after 15 attempts fell to 2.0 percent for Opus 5, down from 5.5 percent for Opus 4.8.
Those figures, if they hold up outside company-run testing, would represent a significant improvement. Prompt injection has long been treated by researchers as a nearly endemic weakness of large language models, especially once they are connected to tools, memory and the web. A model that is materially harder to manipulate could make agentic products far more viable in enterprise settings, where one bad interaction with a hostile webpage can become a serious security risk.
Still, the claim comes with caveats. The 0 percent figure depends not only on the model but also on Anthropic’s surrounding defenses. The open question is how well that resilience carries over into real-world deployments, where developers may use different tool chains, looser controls or fewer guardrails than Anthropic’s own stack.
Anthropic’s fast-moving lineup
The launch also underscores how quickly Anthropic has been reshaping its product lineup this year. Fable 5 and Mythos 5 arrived in June, followed by Sonnet 5 at the end of that month. Opus 5 now appears to fill a different role: a generally available premium model positioned as a more economical alternative to Fable for everyday production use.
That distinction is important. Anthropic has said Opus 5 was not intentionally trained on cybersecurity tasks, even though it has become better at identifying vulnerabilities as its overall capabilities improved. The company has also emphasized that the model still trails Mythos 5 on offensive cyber abilities, especially exploiting vulnerabilities rather than merely finding them.
In effect, Anthropic is trying to offer a powerful system without simply pushing its most capable or highest-risk model into every use case. The strategy reflects a broader tension in the industry, where companies want to expand adoption of increasingly capable systems while persuading regulators, corporate buyers and the public that they are doing so responsibly.
Anthropic said pre-deployment testing found Opus 5 to be its “most aligned model to date,” and that it posted the company’s lowest recent score for misaligned behavior in an automated behavioral audit.
Why this release stands out
Benchmark leads in A.I. are often fleeting, and they can be highly sensitive to test conditions, reasoning settings and which rival models have been published or independently evaluated. Opus 5’s edge over Fable 5 and GPT-5.6 Sol may prove narrow, temporary or difficult to generalize across all workloads.
But the broader significance of the release may lie elsewhere.
For much of the past two years, the A.I. race has centered on a simple question: which model is smartest? Opus 5 suggests that a different question is starting to matter just as much: which model is capable enough, cheap enough and secure enough to be trusted with real work?
If Anthropic’s claims hold up under wider testing, Opus 5 could strengthen the case that frontier-level systems are becoming not only more powerful, but also more deployable — a shift that may prove more consequential than another incremental win on a leaderboard.
Sources
Further reading and reporting used to add context:
- https://artificialanalysis.ai/articles/opus-5
- https://artificialanalysis.ai/changelog
- https://artificialanalysis.ai/models/claude-opus-5-high
- https://artificialanalysis.ai/articles/claude-opus-5-leader-agentic-knowledge-work
- https://artificialanalysis.ai/models/claude-opus-5-xhigh
- https://artificialanalysis.ai/models/comparisons/claude-opus-5-high-vs-claude-opus-5
- https://www.anthropic.com/news/redeploying-fable-5
- https://artificialanalysis.ai/articles/four-frontier-launches-in-eight-days-six-labs-now-field-a-model-above-50-on-the-artificial-analysis-intelligence-index
- https://artificialanalysis.ai/articles
- https://artificialanalysis.ai/articles/claude-fable-5-mythos-intelligence-index
- https://artificialanalysis.ai/evaluations/aa-briefcase
- https://artificialanalysis.ai/articles/claude-sonnet-5-agentic-cost
- https://www-cdn.anthropic.com/73ad94ca3c0502e75e46637cc62c8bd9532a7f2c/Claude%20Sonnet%205%20System%20Card.pdf
- https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf
- https://www-cdn.anthropic.com/78073f739564e986ff3e28522761a7a0b4484f84.pdf
- https://www-cdn.anthropic.com/6a5fa276ac68b9aeb0c8b6af5fa36326e0e166dd.pdf
- https://www.anthropic.com/news/claude-sonnet-5
- https://www.anthropic.com/news?type=product
- https://www.anthropic.com/news?via=p2p
- https://www.anthropic.com/news?type=company
- https://www.anthropic.com/news?_x_tr_hist=true&_x_tr_pto=tc&_x_tr_sl=en&_x_tr_tl=es
- https://www.anthropic.com/claude/fable
- https://www.anthropic.com/claude/mythos
- https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench?s=08
- https://www.anthropic.com/responsible-scaling-policy
- https://aws.amazon.com/blogs/machine-learning/introducing-claude-opus-5-on-aws-anthropics-most-capable-opus-model/
- https://www.reddit.com/r/Anthropic/comments/1v5h6r8/introducing_claude_opus_5/
- https://www.reddit.com/r/ClaudeAI/comments/1v5h6o9/introducing_claude_opus_5/
- https://qz.com/anthropic-claude-opus-5-fable-5-price-072426
- https://t.co/uI4Wu5Fj7t
- https://www.reddit.com/r/Anthropic/comments/1v5h84p/introducing_claude_opus_5/
- https://www.reddit.com/r/ClaudeCode/comments/1v5h6pr/introducing_claude_opus_5/
- https://platform.claude.com/docs/de/resources/overview
- https://www.techtimes.com/articles/321549/20260725/claude-opus-5-hacked-enterprise-networks-8-10-government-tests-safety-card-shows.htm
- https://platform.claude.com/docs/en/about-claude/model-deprecations
- https://www.reddit.com/r/ClaudeCode/comments/1v5h61g/official_opus_5_release_benchmarks_inside/
- https://cyberlab.team/blog/claude-opus-5-prompting-and-limits/
- https://www.reddit.com/r/Anthropic/comments/1v5h2w8/opus_5_has_been_released/
- https://tech.yahoo.com/ai/claude/articles/anthropic-releases-claude-opus-5-133000235.html
- https://www.reddit.com/r/Anthropic/comments/1v5q1ju/opus_5_first_impressions_vs_fable/
- https://www.reddit.com/r/Anthropic/comments/1v5o21h/anthropic_dropped_opus_5_but_wheres_the_usage/
- https://www.reddit.com/r/ClaudeCodeTLDR/comments/1v5ixis/tldr_introducing_claude_opus_5/
- https://thenextweb.com/news/anthropic-claude-opus-5-launch-frontier-bench-coding
- https://www.reddit.com/r/Anthropic/comments/1v5khv0/opus_5_omg_its_alive/
- https://www.reddit.com/r/Anthropic/comments/1v5j7f4/agi_confirmed_opus_5/
- https://www.reddit.com/r/ClaudeCode/comments/1v5h089/opus_5_incoming/
- https://podcasts.apple.com/hn/podcast/claude-opus-5-the-system-card-by-zvi/id1698192712?i=1000778343154&l=en-GB
- https://9to5mac.com/2026/07/24/anthropic-upgrades-claude-with-new-opus-5-model-details-here/
- https://arxiv.org/abs/2602.19467
- https://www.funcas.es/wp-content/uploads/2026/06/Inteligencia-artificial-y-estabilidad-financiera_esp.pdf
- https://www.funcas.es/wp-content/uploads/2026/06/Artificial-intelligence-and-financial-stability.pdf
- https://arxiv.org/abs/2510.01670
- https://www.frontiersin.org/journals/cell-and-developmental-biology/articles/10.3389/fcell.2025.1704762/pdf
- https://platform.claude.com/docs/en/resources/overview
- https://platform.claude.com/docs/id/resources/overview
- https://platform.claude.com/docs/pt-BR/resources/overview
- https://simonwillison.net/?trk=public_post-text
- Opus 5 may have solved browser-based prompt injection, the biggest security flaw haunting AI agents
- https://feeds.simonwillison.net/2025/Aug/7/gpt-5/
- https://simonwillison.net/2026/apr/18/opus-system-prompt/
- https://simonwillison.net/2024/Oct/23/model-card/
- https://simonwillison.net/2025/Jan/26/chatgpt-operator-system-prompt/
- https://simonwillison.net/2025/May/25/claude-4-system-card/
- https://static.simonwillison.net/static/2024/chrome-headless-page.pdf
- https://arcprize.org/results/anthropic-claude-opus-5
- https://claude5.ai/news/claude-opus-5-arc-agi-3-record
- https://arcprize.org/competitions/2026/arc-agi-3
- https://explainx.ai/blog/arc-agi-3-opus-5-leaderboard-july-2026
- https://wikidocs.net/blog/%40openwiki/26310/
- https://rizz.dev/blog/meta-analysis/opus-5-novel-reasoning-benchmark
- https://www.reddit.com/r/ClaudeAI/comments/1v5heie/opus_5_302_on_arcagi_3/
- https://www.techmeme.com/260724/p23
- https://decrypt.co/374305/claude-opus-5-outscores-fable-5-most-benchmarks-half-price
- https://www.vellum.ai/blog/claude-opus-5-benchmarks-explained
- https://www.digitalapplied.com/blog/arc-prize-verifies-claude-opus-5-arc-agi-3-record
- https://www.explainx.ai/blog/arc-agi-3-opus-5-leaderboard-july-2026
- https://chipos.io/news/brief/anthropic-s-opus-5-blows-past-fable-5-and-gpt-5-6-sol-on-the-benchmark-designed-to-7b2e102e8e56dcda
- https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf
- Introducing Claude Opus 5 \ Anthropic














Leave a Reply