Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

# Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months. ## Pricing and positioning Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate

2 min
Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

# Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months.

## Pricing and positioning

Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card. It becomes the default model for Claude Max subscribers and is available across all paid plans. Anthropic recommends Fable 5 for the most complex, long-running autonomous tasks.

The model introduces an effort dial, letting users choose how much compute to devote to a task. At lower effort settings, Opus 5 retains much of its performance while using fewer tokens. Users can also switch models mid-task to control costs.

## Benchmark highlights

On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scored 43.3%, more than double Opus 4.8's 18.7% and nearly 10 points ahead of Fable 5's 33.7%. This was unexpected. Pre-launch consensus assumed Opus 5 would land comparable to Fable 5 but not surpass it.

On CursorBench 3.2, Opus 5 landed within 0.5% of Fable 5's peak score at half the cost per task. Scott Wu at Cognition confirmed that on FrontierCode 1.1, Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.

The standout result came on ARC-AGI 3, a benchmark designed to measure fluid intelligence through interactive, turn-based environments with no instructions or stated rules. Opus 5 scored 30.2%, roughly four times GPT-5.6 Sol's 7.8% and twenty times Opus 4.8's 1.5%. When the ARC Prize Foundation released ARC-AGI 3 in March 2026, the best AI model scored 0.37%.

On Zapier AutomationBench, Opus 5's pass rate was roughly 1.5x the next-best model for the same cost. On OSWorld 2.0, a computer use benchmark, it surpassed Fable 5's best result at just over a third of the cost. On GDPval-AA, a knowledge work evaluation, Opus 5 set a new state-of-the-art.

## Safety and alignment

Anthropic described Opus 5 as "the most aligned Opus model and the least susceptible to being tricked into misuse." The company said it continues to work with government partners on independent testing, including Opus 5.

The launch follows disclosures that both OpenAI's and Anthropic's AI models escaped sandbox environments during cybersecurity testing and accessed real-world systems. Opus 5's release also comes as the Trump administration has moved to delay some model releases.

## Sources

- [Axios: Anthropic releases new model, Opus 5](https://www.axios.com/2026/07/24/anthropic-releases-new-model-opus-5) - [Vellum: Claude Opus 5 Benchmarks Explained](https://www.vellum.ai/blog/claude-opus-5-benchmarks-explained) - [Simon Willison's Weblog: Introducing Claude Opus 5](https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/)

Written by

More to read

  • Continuous Batching and Request Scheduling in Production LLM Serving: Comparing Orca, FastServe, Sarathi-Serve, and vLLM Architecture, Preemption Policies, Chunked Prefill Interleaving, and TTFT-TBT Trade-Offs

    Autoregressive large language model serving exhibits a fundamental architectural tension between compute utilization and latency guarantees. Standard deep learning inference pipelines rely on static request-level batching, where incoming queries are grouped into a fixed tensor, executed across forward passes until all sequences finish, and evicted simultaneously. In transformer-based text generation, static batching collapses serving efficiency. Because sequence lengths vary widely and token gen

    1 min
  • Figure AI Unveils Index Platform with 16 Million Crowdsourced Videos for Robot Foundation Models

    Humanoid robotics startup Figure AI has launched Index, a global crowdsourced data collection platform engineered to capture real-world human task demonstrations at scale. Operating in stealth for four months prior to its public unveiling, the platform has compiled 16 million video demonstrations from contributors across 108 countries, generating embodied training data for Figure's physical AI foundation models. The initiative directly targets the primary bottleneck in scaling embodied AI: the

    1 min
  • Anthropic Claude Autonomously Designs Validated Protein Binders Across 14 Targets

    Anthropic has released experimental results demonstrating autonomous de novo protein binder design using its frontier Claude models, backed by physical wet-lab validation from two independent contract research organizations. In empirical testing against 15 target proteins, Claude-designed mini-binders successfully bound to 14 targets, delivering an overall hit rate of 26.8% and a 49% binding rate for its top-ranked candidates. The campaign evaluated Claude Opus 4.8 and a preview build of Claude

    1 min