Anthropic released Claude Opus 5 on July 24, its latest flagship model targeting software engineering, business automation, and autonomous agent workloads. The company claims Opus 5 delivers near-frontier intelligence comparable to its top-tier Fable 5 model at roughly half the cost per task.

Benchmark Results

The headline numbers center on agentic and coding evaluations. Anthropic said Opus 5 outperforms every other model on OSWorld 2.0, a benchmark measuring AI agents’ ability to interact with computers and complete autonomous tasks. On that benchmark, Opus 5 surpassed Fable 5’s best result at roughly one-third the cost.

On Zapier AutomationBench, which tests whether models can execute end-to-end business workflows, Opus 5 achieved a pass rate approximately 1.5 times higher than the next-best model at the same cost. Zapier CEO Wade Foster confirmed to Anthropic that Opus 5 “topped Zapier’s AutomationBench leaderboard without spending more tokens than prior Claude models,” completing a full churn-prevention sequence that previous models failed.

In coding, Opus 5 more than doubled Opus 4.8’s performance on Frontier-Bench v0.1, according to Anthropic. On CursorBench 3.2, the model scored within 0.5% of Fable 5’s peak at half the cost per task. Cursor co-founder Sualeh Asif told Anthropic that Opus 5 delivers “near Fable 5 intelligence at Opus speed and cost.”

On ARC-AGI 3, a novel problem-solving evaluation, Opus 5’s score was three times higher than the next-best model, per Anthropic’s announcement.

Real-World Agent Behavior

Anthropic highlighted several early-access examples that demonstrate agentic capability rather than raw benchmark performance. In one Frontier-Bench task, the model was given a drawing of a machine part and asked to rebuild it as a 3D FreeCAD model, but was intentionally given no way to view the drawing directly. Opus 5 wrote its own computer vision pipeline to extract geometry from the raw pixels, then reconstructed the part. Anthropic reported that no competing model succeeded after five attempts.

A trading firm engineer used Opus 5 to build a market data feed for a new exchange in a single session, according to Benzinga. When no live data source was available to validate against, the model built its own test harness to verify its code parsed the exchange’s data correctly.

Pricing and Availability

Opus 5 is now the default model on Claude Max and the strongest model available on Claude Pro. Anthropic said the model costs the same as its predecessor, Opus 4.8, while delivering substantially higher performance across evaluations. The company also introduced a “Fast mode” option for users who want increased speed at a higher cost.

Safety Constraints

Anthropic said Opus 5 showed lower rates of problematic behavior during internal evaluations but remains behind its Mythos 5 model on offensive cybersecurity and advanced biology research tasks. Benzinga reported that Anthropic introduced additional safeguards for cyber-related tasks, including restrictions on penetration testing and exploit generation.

The Agentic Benchmark Arms Race

The release signals that frontier model competition has shifted decisively toward agentic performance metrics. OSWorld, AutomationBench, and Frontier-Bench all measure whether models can complete multi-step tasks autonomously, not just answer questions correctly. With OpenAI’s Codex and ChatGPT Work crossing 10 million users this month, and Google pushing Gemini’s agent capabilities, Opus 5 positions Anthropic as a direct competitor for enterprise teams evaluating which foundation model to build autonomous workflows on.