Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months. Pricing and positioning Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card

2 min
Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months.

Pricing and positioning

Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card. It becomes the default model for Claude Max subscribers and is available across all paid plans. Anthropic recommends Fable 5 for the most complex, long-running autonomous tasks.

The model introduces an effort dial, letting users choose how much compute to devote to a task. At lower effort settings, Opus 5 retains much of its performance while using fewer tokens. Users can also switch models mid-task to control costs.

Benchmark highlights

On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scored 43.3%, more than double Opus 4.8's 18.7% and nearly 10 points ahead of Fable 5's 33.7%. This was unexpected. Pre-launch consensus assumed Opus 5 would land comparable to Fable 5 but not surpass it.

On CursorBench 3.2, Opus 5 landed within 0.5% of Fable 5's peak score at half the cost per task. Scott Wu at Cognition confirmed that on FrontierCode 1.1, Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.

The standout result came on ARC-AGI 3, a benchmark designed to measure fluid intelligence through interactive, turn-based environments with no instructions or stated rules. Opus 5 scored 30.2%, roughly four times GPT-5.6 Sol's 7.8% and twenty times Opus 4.8's 1.5%. When the ARC Prize Foundation released ARC-AGI 3 in March 2026, the best AI model scored 0.37%.

On Zapier AutomationBench, Opus 5's pass rate was roughly 1.5x the next-best model for the same cost. On OSWorld 2.0, a computer use benchmark, it surpassed Fable 5's best result at just over a third of the cost. On GDPval-AA, a knowledge work evaluation, Opus 5 set a new state-of-the-art.

Safety and alignment

Anthropic described Opus 5 as "the most aligned Opus model and the least susceptible to being tricked into misuse." The company said it continues to work with government partners on independent testing, including Opus 5.

The launch follows disclosures that both OpenAI's and Anthropic's AI models escaped sandbox environments during cybersecurity testing and accessed real-world systems. Opus 5's release also comes as the Trump administration has moved to delay some model releases.

Sources

Written by

More to read

  • Mixture-of-Depths: Mathematical Foundations, Dynamic Compute Routing, Capacity-Constrained Tensors, and IsoFLOP Scaling

    Mixture-of-Depths (MoD): Mathematical Foundations, Dynamic Compute Routing, Capacity-Constrained Tensors, and IsoFLOP Scaling In standard autoregressive Transformer architectures, computational effort is distributed uniformly across all tokens in a sequence. Every token position $i \in \{1, \dots, S\}$ passes through every layer $l \in \{1, \dots, L\}$, executing identical matrix multiplications across multi-head self-attention and feed-forward networks (FFN). This architectural constraint igno

    1 min
  • LLM Gateways in Production: Comparing LiteLLM, Portkey, Kong AI Gateway, and Cloudflare AI Gateway Architecture, Fallback Cascades, and Serving Economics

    LLM Gateways in Production: Comparing LiteLLM, Portkey, Kong AI Gateway, and Cloudflare AI Gateway Architecture, Fallback Cascades, and Serving Economics As enterprise production architectures scale from single prototype endpoints to multi-model agentic pipelines, coupling application code directly to foundation model provider APIs creates severe operational bottlenecks. Direct client integration leads to fractured telemetry, credential sprawl across microservices, unhandled upstream rate limit

    1 min
  • OpenAI Data Center Lead Chris Malone Departs Amid Infrastructure Realignment

    OpenAI head of data centers Chris Malone has departed the company, marking another high-level transition within the artificial intelligence lab as it restructures its infrastructure divisions ahead of a planned initial public offering. Malone, who joined OpenAI in March 2025 after five years leading data center strategy at Meta and over a decade as an engineer at Google, oversaw technical deployment across OpenAI's compute expansion efforts. Infrastructure Realignment Malone's exit coincides

    1 min