Anthropic Releases Claude Opus 5: Frontier Coding and Agentic Performance at Half the Cost Per Task on CursorBench
AI

Anthropic Releases Claude Opus 5: Frontier Coding and Agentic Performance at Half the Cost Per Task on CursorBench

Claude Opus 5 launched July 24 with state-of-the-art scores on coding and knowledge-work benchmarks, prices starting at $5 per million input tokens, and early-access results including 9 percentage points higher accuracy on some of the hardest financial-modeling tasks versus Opus 4.8, a 17% due diligence improvement reported by Box, and a 22% gain on agentic coding tasks versus Opus 4.7.

5 min read
Back to News
TL;DR
  • -Anthropic released Claude Opus 5 on July 24, 2026, priced at $5 per million input tokens and $25 per million output tokens, matching the Opus 4.8 rate, with state-of-the-art results on Frontier-Bench v0.1 and GDPval-AA claimed by Anthropic.
  • -Early-access customers report gains including 9 percentage points higher accuracy on some of the hardest financial-modeling tasks compared to Opus 4.8, a 17% improvement in due diligence workflows reported by Box, and 22% gains on the hardest agentic coding tasks versus Opus 4.7; Anthropic states the model still trails Mythos 5 on cybersecurity evaluations.
  • -Enterprise AI leaders should run a bounded pilot on their three highest-value workload types, gate autonomous actions behind human-review checkpoints until internal evals confirm the model's self-verification claims, and confirm data-handling and US-only inference options with Anthropic before production deployment.

Anthropic Releases Claude Opus 5

Anthropic released Claude Opus 5 on July 24, 2026, available immediately through the Claude API, Amazon Web Services, Google Cloud, and Microsoft Foundry.

Pricing starts at $5 per million input tokens and $25 per million output tokens, matching the Opus 4.8 rate. Prompt caching can cut input costs by up to 90%. Batch processing offers 50% savings. For workloads that must stay in US infrastructure, US-only inference is available at 1.1× the standard token price.

Opus 5 becomes the new default model on Claude Max and the strongest model on Claude Pro.

What the Benchmarks Show

Anthropic positions Opus 5 as near-frontier performance at a lower price than Claude Fable 5. The company describes it as state-of-the-art on Frontier-Bench and GDPval-AA, though it remains behind Mythos 5 on cybersecurity tasks.

On Frontier-Bench v0.1, Opus 5 surpasses all other models and more than doubles Opus 4.8's score at a lower cost per task. On CursorBench 3.2 at max effort, Opus 5 scores within 0.5% of Fable 5 but at half the cost per task.

The gap widens on reasoning and automation. On ARC-AGI 3, Opus 5's score is three times that of the next-best model. On Zapier AutomationBench, Opus 5's pass rate is roughly 1.5× the next-best model at the same cost per task. On OSWorld 2.0, Opus 5 surpasses Fable 5's best result at just over a third of the cost.

Anthropic also reports life sciences gains. Opus 5 scores 10.2 percentage points higher than Opus 4.8 on organic chemistry tasks, such as inferring molecular structures from spectroscopy data. It scores 7.7 percentage points higher on protein-sequence function prediction.

What Early-Access Customers Report

The following figures come from named early-access deployments and reflect those customers' own claims.

Box reports Opus 5 outperforms Opus 4.8 by 8% overall, with an 11% improvement in data analysis and 17% in due diligence workflows across technology, healthcare, and public sector use cases.

One financial modeling team reports that on some of their hardest financial-modeling tasks, Opus 5 averaged 9 percentage points higher accuracy than Opus 4.8, with a third fewer turns and tool calls and 60% less time on hard financial-domain tasks.

Lovable reports Opus 5 is up 22% over Opus 4.7 on its hardest agentic coding tasks, with far less variance run to run.

A separate early-access customer reports that within Devin, Opus 5 shows particular strength on difficult debugging and root-cause analysis tasks.

Zapier reports Opus 5 completed a full churn-prevention sequence end to end, flagging at-risk accounts, alerting owners, and summarizing for retention ops. Previous models did not pass.

Anthropic's own testing includes a notable agentic example. On one Frontier-Bench task, Opus 5 was given a drawing of a machine part with no direct way to view it. It wrote its own computer vision pipeline to extract geometry from raw pixels and reconstructed the full 3D FreeCAD model. No competing model with the same setup could solve it after five attempts.

The Effort Setting Changes Your Cost Model

Opus 5 includes an effort-control system. Users can control how hard Opus 5 works, trading more token spend for better response quality at higher settings. Anthropic expanded this control to all plans with the Opus 4.8 launch: on higher effort settings, Claude thinks more frequently and more deeply; on lower effort settings, Claude responds faster and uses rate limits more slowly.

Performance and cost charts in the announcement show results across effort levels, letting buyers optimize by task type. The announcement does not state a default effort level for API access. Enterprise teams should confirm this before estimating production spend.

What Remains Unconfirmed Before Production

Several details relevant to enterprise adoption are absent from Anthropic's announcement.

Anthropic's announcement does not state whether enterprise-grade audit logging or role-based access controls are available for Opus 5 through the API or cloud platforms.

Anthropic states that extensive testing and evaluation ensures the release of Opus 5 meets its standards for safety, security, and reliability, and that the accompanying model card covers safety results in depth. The model card content itself is not available in these sources.

The Claude Opus product page references a 1M context window for the Opus 4.8 tier. No equivalent confirmation for Opus 5 appears in the current sources.

Enterprise teams should verify these specifics directly with Anthropic or through their cloud provider before committing production workloads.

What Enterprise AI Leaders Should Do Now

The benchmark results and early-access reports make a controlled evaluation reasonable. The decision before enterprise AI, application, and platform leaders is which workloads to test first and what gates to set.

Pick three workloads with measurable outcomes. The clearest early signals come from tasks with an existing quality baseline. Based on the early-access results, financial modeling accuracy, due diligence review completion rates, and agentic coding task pass rates are reasonable starting points. Use the same evaluation criteria you applied to Opus 4.8 so the comparison is clean.

Test across effort levels, not just at max. Cost scales with effort. The business case depends on finding the right operating point per workload. A task that passes at low effort eliminates the cost premium entirely. Testing only at max effort overstates both quality and spend.

Set human-review gates before autonomous deployment. Agentic capability is the headline claim. The model's self-verification behavior is reported by customers and Anthropic. For any workflow where the model takes external actions, such as sending alerts, modifying code, or updating records, require human approval until your internal evals confirm error rates you can accept.

Confirm security and data controls before production. US-only inference is available at 1.1× pricing, which matters for regulated industries. Specific isolation, logging, and retention controls are not described in the announcement. Get written confirmation from Anthropic or your cloud provider before moving sensitive data through the API.

If Opus 5 delivers Fable-level results on your highest-value tasks at half the cost per task on CursorBench, the ROI calculation changes quickly. A bounded pilot on three workloads can produce that comparison data before committing production infrastructure.

Sources and supporting resources
Previous
AWS Releases CSA Compliance Guide Mapping 207 CCM v4.1 Controls to Services and Evidence
Next
Albertsons CEO Susan Morris Builds AI Into the Operating Model, Not a Side Project

Get Manufacturing Technology Updates

Problem-led guidance on manufacturing operations, integration, portals, analytics, automation, custom software, trusted records, and fit-for-purpose engineering.

No spam. Unsubscribe anytime.