|

Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing

Today, Anthropic released Claude Opus 5. It replaces Claude Opus 4.8 as the Opus-tier flagship. Pricing is unchanged at $5 per million enter tokens and $25 per million output tokens.

The Anthropic staff positions Opus 5 as approaching the intelligence of Claude Fable 5 at half the value. It is now the default mannequin on Claude Max and the strongest mannequin on Claude Pro.

What truly modified at the API degree

Three adjustments are fairly vital earlier than any benchmark does:

  1. Thinking is on by default. On Opus 4.8, requests ran with out pondering except you set pondering: {"kind": "adaptive"}. On Opus 5 the similar request thinks, and the effort parameter controls depth. Because max_tokens caps pondering plus response textual content, current values want assessment.
  2. There is a breaking change. Setting pondering: {"kind": "disabled"} with effort xhigh or max now returns a 400 error. The restriction is enforced per request. You both cap effort at excessive or drop the pondering subject.
  3. Anthropic tells builders to delete their verification prompts. Instructions like “embody a remaining verification step” now trigger over-verification, as a result of the mannequin already verifies its personal work. The Opus 5 prompting guide covers the tuning patterns.

The mannequin ID is claude-opus-5. Context is 1M tokens as each default and most, with no smaller variant. Maximum output is 128k tokens on the synchronous Messages API. The Message Batches API reaches 300k with the output-300k-2026-03-24 beta header. The minimal cacheable immediate drops to 512 tokens, down from 1,024.


Coding and agentic outcomes

On FrontierBench v0.1, a 74-task successor to Terminal-Bench 2.1, Opus 5 scored 43.3% at max effort. Opus 4.8 scored 18.7%. Fable 5 reached 33.7% and GPT-5.6 Sol reached 37.5%. At xhigh effort Opus 5 reaches 44.4% imply reward, its finest end result.

One element from that run is value noting. Opus 5 security classifiers flagged and refused 5% of API calls, throughout 4% of trials. Fable 5 classifiers flagged 42% of calls throughout 26% of trials.

Opus 5 scored 96.0% on SWE-bench Verified and 79.2% on SWE-bench Pro. Fable 5 nonetheless edges it on Pro at 80.0%. On SWE-bench Multimodal the soar is bigger, from 38.4% to 59.4%.

The agentic numbers are the clearest wins. Opus 5 reached 70.57% on OSWorld 2.0 in opposition to 55.7% for Opus 4.8. On Zapier AutomationBench it scored 26.0%, in opposition to 17.0% for Opus 4.8 and 17.4% for Fable 5. At medium effort it nonetheless scores 24% at $0.89 per activity.

On Artificial Analysis GDPval-AA v2, Opus 5 takes the prime two leaderboard spots at ELO 1861 and 1827. The xhigh setting beats each different mannequin whereas utilizing 25% fewer output tokens than max.

Reasoning, and the ARC-AGI-3 end result

Anthropic prompted Opus 5 on all six IMO 2026 issues with out instruments or an agent harness. A 3-model decide panel scored all 24 generated options appropriate. Human consultants independently graded one pre-specified resolution per downside at 7/7. The remaining 42/42 is gold-medal degree, above the 29/42 cutoff.

The ARC Prize Foundation reviews a verified 30.16% on ARC-AGI-3 at excessive effort. That is roughly 4 occasions the finest beforehand reported leaderboard rating. GPT-5.6 Sol reached 7.78% and Opus 4.8 reached 1.52%. Opus 5 outcomes at max effort have been unavailable at launch.

On Humanity's Last Exam, Opus 5 scored 56.3% with out instruments and 64.7% with them.

Tools beat pondering on multimodal work

The multimodal part carries a sensible lesson. Agentic device use scales test-time compute extra cost-effectively than adaptive pondering alone.

On Chartography, Opus 5 scored 29.6% with out instruments and 83.0% with a container and an image-cropping device. On BenchCAD Vision2Code, voxel IoU strikes from 0.366 to 0.821. With instruments, that beats Claude Mythos 5 at 0.678 by a large margin.


Cyber functionality rose, and safeguards have been relaxed in a single place

Anthropic didn't practice Opus 5 on cybersecurity duties. Capability rose anyway, as a byproduct of common functionality beneficial properties.

On ExploitBench, Opus 5 captured 10.14 imply functionality flags in the AutoNudge arm. It produced 99 full arbitrary-code-execution exploits. Mythos 5 produced 132. On OSS-Fuzz, Opus 5 scored non-zero on 79.4% of targets. Mythos 5 reached roughly 80%. But Opus 5 accomplished 4 full exploits to Mythos 5's 13.

That hole defines the safeguard design. Opus 5 is almost as robust as Mythos 5 at discovering vulnerabilities, and considerably behind at exploiting them. So Anthropic unblocked vulnerability discovering in supply code. Binary-based scanning, penetration testing and exploit technology keep blocked. Classifiers are anticipated to intervene round 85% much less typically than on Fable 5. Defenders can apply to the Cyber Verification Program.

UK AISI examined early checkpoints on three cyber ranges at 100M tokens per try. Opus 5 solved 'The Last Ones' end-to-end in 8 of 10 makes an attempt. It didn't remedy the tougher 'Doing Life' vary. It reached step 22 of 23, additional than any mannequin examined.

Under the RSP, Anthropic treats Opus 5 as having CB-1 capabilities however not CB-2. It applies the similar ASL-3 protections used for Opus 4.8. The AI R&D threshold just isn't crossed.

Prompt injection is the standout security quantity

On the Gray Swan oblique immediate injection benchmark, attacker success inside 15 makes an attempt fell from 5.5% on Opus 4.8 to 2.0%. Mythos 5 sits at 2.6% and GPT-5.6 Sol at 20.0%.

In browser environments run via Claude Cowork, assault success dropped from 31.5% on Opus 4.8 to three.70%. That determine is with no safeguards utilized. With auto mode enabled, it reached 0% throughout all 129 environments.

Community Sentiment Analysis


Key Takeaways

  • Opus 5 ships at unchanged $5/$25 pricing with a 1M-token context window as each default and most.
  • Thinking is now on by default, and disabling it above excessive effort returns a 400 error.
  • Agentic evaluations are the clearest wins: OSWorld 2.0 at 70.57%, AutomationBench at 26.0%, ARC-AGI-3 at 30.16%.
  • Cyber safeguards calm down just for source-code vulnerability discovering; exploitation paths keep blocked.
  • Anthropic publishes its personal negatives, together with barely increased factual hallucination than Opus 4.8.


Sources: Anthropic launch post, Claude Opus 5 System Card, What's new in Claude Opus 5, TechCrunch, and CodeRabbit independent review

The submit Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing appeared first on MarkTechPost.

Similar Posts