Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months. Pricing and positioning Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card

2 min
Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months.

Pricing and positioning

Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card. It becomes the default model for Claude Max subscribers and is available across all paid plans. Anthropic recommends Fable 5 for the most complex, long-running autonomous tasks.

The model introduces an effort dial, letting users choose how much compute to devote to a task. At lower effort settings, Opus 5 retains much of its performance while using fewer tokens. Users can also switch models mid-task to control costs.

Benchmark highlights

On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scored 43.3%, more than double Opus 4.8's 18.7% and nearly 10 points ahead of Fable 5's 33.7%. This was unexpected. Pre-launch consensus assumed Opus 5 would land comparable to Fable 5 but not surpass it.

On CursorBench 3.2, Opus 5 landed within 0.5% of Fable 5's peak score at half the cost per task. Scott Wu at Cognition confirmed that on FrontierCode 1.1, Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.

The standout result came on ARC-AGI 3, a benchmark designed to measure fluid intelligence through interactive, turn-based environments with no instructions or stated rules. Opus 5 scored 30.2%, roughly four times GPT-5.6 Sol's 7.8% and twenty times Opus 4.8's 1.5%. When the ARC Prize Foundation released ARC-AGI 3 in March 2026, the best AI model scored 0.37%.

On Zapier AutomationBench, Opus 5's pass rate was roughly 1.5x the next-best model for the same cost. On OSWorld 2.0, a computer use benchmark, it surpassed Fable 5's best result at just over a third of the cost. On GDPval-AA, a knowledge work evaluation, Opus 5 set a new state-of-the-art.

Safety and alignment

Anthropic described Opus 5 as "the most aligned Opus model and the least susceptible to being tricked into misuse." The company said it continues to work with government partners on independent testing, including Opus 5.

The launch follows disclosures that both OpenAI's and Anthropic's AI models escaped sandbox environments during cybersecurity testing and accessed real-world systems. Opus 5's release also comes as the Trump administration has moved to delay some model releases.

Sources

Written by

More to read

  • Amazon Data Center Could Be Powered by One of the Nation's Most Polluting Power Plants

    Amazon is investing in a new natural-gas power plant in Pecos County, Texas, to supply a West Texas data center, and the project holds a permit that would allow it to emit more carbon dioxide than any coal plant in the country, according to The Verge and the New York Times. The plant, tracked as GW Ranch by Cleanview, a service that monitors data center power projects, would deploy 35 natural-gas turbines generating about 7.65 gigawatts. At least initially, the plant would not connect to

    1 min
  • Claude Code Defaults to Auto Mode. The Classifier Catches More Than Humans.

    Claude Code Defaults to Auto Mode. The Classifier Catches More Than Humans. Claude Code will ship with Auto Mode enabled by default starting August 14 for Pro, Max, and Team subscribers, shifting the developer role further from active coding toward reviewing AI-generated output. Only Enterprise customers will need to opt in. Auto Mode lets the agent execute steps without waiting for manual approval at each one. A classifier intercepts actions the model judges dangerous or irreversible and paus

    1 min