Skip links
Claude Opus 5 by Anthropic - AI model cover image

Claude Opus 5: Anthropic’s New Flagship Model for Deep Reasoning and Long-Running Agents

On July 24, 2026, Anthropic released Claude Opus 5, the newest model in its Opus lineup and one of the most capable systems the company has shipped to date. According to Anthropic’s official announcement, Opus 5 is designed to come close to the reasoning depth of the company’s higher-tier “Fable 5” system while costing roughly half as much to run. It is built less for quick one-off questions and more for long, multi-step professional work: coding sessions that span hours, research tasks that require dozens of steps, and business workflows that need to be carried through to completion without constant supervision. Opus 5 is now the default model inside Claude Max, the strongest model available on Claude Pro, and accessible to developers through the Claude API.

What Sets Claude Opus 5 Apart

Earlier Claude models were sometimes criticized for stopping too early, settling for a shallow first attempt, or losing track of long instructions during extended sessions. Anthropic built Opus 5 specifically to address this pattern. The company reports that the model is considerably better at verifying its own work and iterating until a task is genuinely finished, rather than merely appearing finished on the surface. In practice, this shows up as Opus 5 testing its own code before handing it back, double-checking a spreadsheet or contract for inconsistencies before calling the job complete, and holding on to a coherent plan across dozens of tool calls without drifting off track. Anthropic frames this as a combination of “agency and thoroughness,” and it is the thread that connects most of the improvements described in the model’s official release notes.

Benchmark Performance

On Frontier-Bench v0.1, an internal coding benchmark, Anthropic reports that Opus 5 more than doubles the score of its predecessor, Opus 4.8, while costing less per completed task. On CursorBench, which is built around real developer workflows inside the Cursor editor, Opus 5 running at maximum effort lands within about half a percentage point of Fable 5’s best score, at roughly half the price per task. The improvements go beyond software engineering. On ARC-AGI 3, an evaluation designed to test genuinely novel problem-solving rather than recall from training data, Opus 5 is reported to score around three times higher than the next-best model tested. On Zapier’s AutomationBench, a benchmark that measures whether a model can complete realistic, multi-step business tasks end to end, Opus 5 reportedly reached a 100 percent pass rate on a churn-prevention workflow that earlier models could not complete at all. Anthropic also cites strong, cost-efficient results on OSWorld 2.0, a computer-use benchmark, along with GDPval-AA, HLE, and DeepSearchQA, three evaluations focused on professional and research-style reasoning.

Beyond Coding: Scientific Research and Visual Output

The model’s gains are not limited to programming tasks. Anthropic reports measurable improvements across its internal life-sciences evaluations, spanning structural biology, organic chemistry, and bioinformatics. The largest single jump, over ten percentage points, appears on tasks that involve inferring a molecule’s structure from spectroscopy data, while protein-related tasks, such as predicting how a change in a protein’s sequence affects its function, also show clear gains. On the creative side, Anthropic showcases Opus 5 producing more sophisticated visual and generative outputs, including a simulated wind tunnel that models airflow over both aerodynamic and irregular objects. This suggests the model’s underlying reasoning improvements extend into visual and scientific domains, not only text-based work.

Safety and Alignment

Anthropic states that, during pre-deployment testing, its automated behavioral audit found Opus 5 to be the company’s most aligned model released so far, with lower rates of deceptive behavior and reduced susceptibility to misuse compared with Opus 4.8, Sonnet 5, and Fable 5. On sensitive dual-use categories such as cybersecurity and biology, Anthropic reports that Opus 5 does not advance the frontier of dangerous capability, and remains behind another internal model, referred to in the announcement as Mythos 5, on both offensive cybersecurity and biological research risk. The company also notes that it deliberately avoided training Opus 5 on offensive cyber tasks, and that the model’s safety classifiers now intervene roughly 85 percent less often than they did for Fable 5, while continuing to block higher-risk actions such as binary-based vulnerability scanning and exploit generation. Readers interested in the full technical detail can find it in Anthropic’s official System Card, linked from the announcement page.

Pricing and Availability

Claude Opus 5 is priced the same as its predecessor: five dollars per million input tokens and twenty-five dollars per million output tokens through the Claude API, with an optional Fast mode that runs roughly two and a half times faster for double the base price. The model is already available across Claude.ai, the Claude API, Claude Code, and Claude Cowork. Alongside the release, Anthropic introduced two beta features: the ability to change which tools a model can use mid-conversation without losing prompt caching, and an automatic fallback option that reroutes flagged API requests to another available model instead of blocking them outright.

  • Input tokens: $5 per million
  • Output tokens: $25 per million
  • Fast mode: about 2.5x speed at double the base price
  • Default model on Claude Max; strongest model available on Claude Pro

Is Claude Opus 5 Right for Your Workflow?

For teams working on agentic coding, long-running research, financial modeling, or detailed document review, Opus 5’s main advantage is consistency over long sessions rather than raw speed. Several companies quoted in Anthropic’s official announcement, including engineering leaders from Cursor, Devin-maker Cognition, and Zapier, reported that the model completed tasks their previous tools could not finish at all. For shorter, simpler requests, Anthropic continues to recommend its Sonnet line as a better balance of speed and cost. Developers can begin testing Opus 5 through the Claude developer documentation, while non-technical users can try it directly inside Claude.ai.

Conclusion

Claude Opus 5 does not claim to be Anthropic’s single most powerful model overall — that position still appears to belong to the company’s higher-tier system — but it narrows the gap substantially while cutting cost roughly in half, which matters more for most real-world use cases than topping a leaderboard. As AI tools shift from answering isolated questions to completing entire workflows with less supervision, this emphasis on verification and follow-through looks set to matter more than raw benchmark scores alone.

Leave a comment

Explore
Drag