Anthropic Releases Claude Opus 5: Frontier Coding and Agentic Performance at Half the Cost of Claude Fable 5
AI

Anthropic Releases Claude Opus 5: Frontier Coding and Agentic Performance at Half the Cost of Claude Fable 5

Claude Opus 5 launched July 24 with state-of-the-art scores on coding and knowledge-work benchmarks, prices starting at $5 per million input tokens, and early-access results showing large gains on financial modeling, due diligence, and agentic software engineering tasks.

5 min readJuly 28, 2026
Back to News
Photo by Pavel Danilyuk on Pexels
TL;DR
  • -Anthropic released Claude Opus 5 on July 24, 2026, priced at $5 per million input tokens and $25 per million output tokens, matching the Opus 4.8 rate, with state-of-the-art results on Frontier-Bench v0.1 and GDPval-AA claimed by Anthropic.
  • -Early-access customers report gains including 9 percentage points higher accuracy on some of their hardest financial-modeling tasks compared to Opus 4.8, a 17% improvement in due diligence workflows, and 22% gains on the hardest agentic coding tasks; Anthropic states the model still trails Mythos 5 on cybersecurity evaluations.
  • -Enterprise AI leaders should run a bounded pilot on their three highest-value workload types, gate autonomous actions behind human-review checkpoints until internal evals confirm the model's self-verification claims, and confirm data-handling and US-only inference options with Anthropic before production deployment.

Anthropic Releases Claude Opus 5

Anthropic released Claude Opus 5 on July 24, 2026. The model is available today through the Claude API, Amazon Web Services, Google Cloud, and Microsoft Foundry.

Pricing starts at $5 per million input tokens and $25 per million output tokens. That matches the rate for Opus 4.8. Anthropic says prompt caching can cut input costs by up to 90%. Batch processing offers 50% savings. For workloads that must stay in US infrastructure, US-only inference is available at 1.1× the standard token price.

The model becomes the new default on Claude Max and the strongest model on Claude Pro.

What the Benchmarks Show

Anthropic positions Opus 5 as near-frontier performance at a lower price than Claude Fable 5. Anthropic describes it as state-of-the-art on Frontier-Bench and GDPval-AA. On Frontier-Bench v0.1, Opus 5 surpasses all other models. It more than doubles Opus 4.8's score at a lower cost per task.

On CursorBench 3.2 at max effort, Opus 5 scores within 0.5% of Fable 5 but at half the cost per task.

The gap widens on reasoning and automation tasks. On ARC-AGI 3, Opus 5's score is three times that of the next-best model. On Zapier AutomationBench, Opus 5's pass rate is roughly 1.5× the next-best model at the same cost per task. On OSWorld 2.0, Opus 5 surpasses Fable 5's best result at just over a third of the cost.

Anthropic also reports gains in life sciences. Opus 5 scores 10.2 percentage points higher than Opus 4.8 on organic chemistry tasks, such as inferring molecular structures from spectroscopy data. It scores 7.7 percentage points higher on protein-sequence function prediction.

One area where Opus 5 does not lead: Anthropic states the model remains behind Mythos 5 on cybersecurity tasks.

What Early-Access Customers Report

The following are direct reports from named early-access deployments. These are the claims of those customers, not independently verified outcomes.

Box reports Opus 5 outperforms Opus 4.8 by 8% overall, with an 11% improvement in data analysis and 17% in due diligence workflows across technology, healthcare, and public sector use cases.

One financial modeling team reports that on some of their hardest financial-modeling tasks, Opus 5 averaged 9 percentage points higher accuracy than Opus 4.8. The team also reports a third fewer turns and tool calls, and 60% less time on hard financial-domain tasks.

Lovable reports Opus 5 is up 22% over Opus 4.7 on its hardest agentic coding tasks, with far less variance run to run. Within Devin, the model shows particular strength on difficult debugging and root-cause analysis.

Zapier reports Opus 5 completed a full churn-prevention sequence end to end, flagging at-risk accounts, alerting owners, and summarizing for retention ops. Previous models did not pass.

Anthropic's own testing includes a notable agentic example. On one Frontier-Bench task, Opus 5 was given a drawing of a machine part with no direct way to view it. It wrote its own computer vision pipeline to extract geometry from raw pixels and reconstructed the full 3D FreeCAD model. No competing model with the same setup could solve it after five attempts.

The Effort Setting Changes Your Cost Model

Opus 5 includes an effort-control system. Users can choose how much effort Claude puts into a task, trading more token spend for better response quality at higher settings. On higher effort settings, Claude thinks more frequently and more deeply. On lower effort settings, Claude responds faster and uses rate limits more slowly. Performance and cost charts in the announcement show results across effort levels, letting buyers optimize by task type.

This matters for enterprise budgeting. The announcement does not state default effort levels for API access. Enterprise teams should confirm this before estimating production spend.

Details the Sources Do Not State

Several details relevant to enterprise adoption are absent from the sources.

Whether enterprise-grade audit logging or role-based access controls are available for Opus 5 through the API or cloud platforms is not stated.

Extensive testing and evaluation ensures the release of Opus 5 meets Anthropic's standards for safety, security, and reliability, and the accompanying model card covers safety results in depth. The model card content itself is not available in these sources.

The Claude Opus product page references a 1M context window for the Opus 4.8 tier. No equivalent confirmation for Opus 5 appears in those sources.

Enterprise teams should verify these specifics directly with Anthropic or through their cloud provider before committing production workloads.

What Enterprise AI Leaders Should Do Now

The benchmark results and early-access reports make a controlled evaluation reasonable. The decision before enterprise AI, application, and platform leaders is which workloads to test first and what gates to set.

Pick three workloads with measurable outcomes. The clearest early signals come from tasks with an existing quality baseline. Financial modeling accuracy, due diligence review completion rates, and agentic coding task pass rates are good starting points. Use the same evaluation criteria you applied to Opus 4.8 so the comparison is clean.

Test across effort levels, not just at max. Cost scales with effort. The business case depends on finding the right operating point per workload. A task that passes at low effort eliminates the cost premium entirely. Testing only at max effort overstates both quality and spend.

Set human-review gates before autonomous deployment. Opus 5's agentic capability is the headline claim. The model's self-verification behavior is reported by customers and Anthropic. It is not an independently audited guarantee. For any workflow where the model takes external actions, such as sending alerts, modifying code, or updating records, require human approval until your internal evals confirm error rates you can accept.

Confirm security and data controls before production. US-only inference is available at 1.1× pricing, which matters for regulated industries. Specific isolation, logging, and retention controls are not described in the sources. Get written confirmation from Anthropic or your cloud provider before moving sensitive data through the API.

Metrotechs analysis: The cost-to-performance ratio is the most important variable for enterprise adoption decisions. If Opus 5 delivers Fable-level results on your highest-value tasks at half the cost, the ROI calculation changes quickly. A bounded pilot on three workloads gives you that data inside four to six weeks without committing production infrastructure.

Sources and supporting resources
Previous
AWS Shield Advanced Adopts the Anti-DDoS Managed Rule Group: What Changes and When
Next
Lenovo Feeds Device Telemetry Into ServiceNow's AI Control Tower to Fix IT Problems Early

Get ERP, Cloud, Data, and AI Updates

News, insights, and practical guidance across ERP, Cloud, Data, AI, digital transformation, and technology projects.

No spam. Unsubscribe anytime.