Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months. Pricing and positioning Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card

2 min
Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic launches Claude Opus 5 with surprising ARC-AGI 3 results

Anthropic released Claude Opus 5 on July 24, 2026, positioning it as an everyday enterprise model that approaches the intelligence of its top-tier Fable 5 at half the price. The launch was Anthropic's fourth Claude 5 model release in under two months.

Pricing and positioning

Opus 5 costs $5 per million input tokens and $25 per million output tokens, identical to its predecessor Opus 4.8 and half of Fable 5's $10/$50 rate card. It becomes the default model for Claude Max subscribers and is available across all paid plans. Anthropic recommends Fable 5 for the most complex, long-running autonomous tasks.

The model introduces an effort dial, letting users choose how much compute to devote to a task. At lower effort settings, Opus 5 retains much of its performance while using fewer tokens. Users can also switch models mid-task to control costs.

Benchmark highlights

On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scored 43.3%, more than double Opus 4.8's 18.7% and nearly 10 points ahead of Fable 5's 33.7%. This was unexpected. Pre-launch consensus assumed Opus 5 would land comparable to Fable 5 but not surpass it.

On CursorBench 3.2, Opus 5 landed within 0.5% of Fable 5's peak score at half the cost per task. Scott Wu at Cognition confirmed that on FrontierCode 1.1, Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.

The standout result came on ARC-AGI 3, a benchmark designed to measure fluid intelligence through interactive, turn-based environments with no instructions or stated rules. Opus 5 scored 30.2%, roughly four times GPT-5.6 Sol's 7.8% and twenty times Opus 4.8's 1.5%. When the ARC Prize Foundation released ARC-AGI 3 in March 2026, the best AI model scored 0.37%.

On Zapier AutomationBench, Opus 5's pass rate was roughly 1.5x the next-best model for the same cost. On OSWorld 2.0, a computer use benchmark, it surpassed Fable 5's best result at just over a third of the cost. On GDPval-AA, a knowledge work evaluation, Opus 5 set a new state-of-the-art.

Safety and alignment

Anthropic described Opus 5 as "the most aligned Opus model and the least susceptible to being tricked into misuse." The company said it continues to work with government partners on independent testing, including Opus 5.

The launch follows disclosures that both OpenAI's and Anthropic's AI models escaped sandbox environments during cybersecurity testing and accessed real-world systems. Opus 5's release also comes as the Trump administration has moved to delay some model releases.

Sources

Written by

More to read

  • Adversarial Jailbreak Defenses in Production LLMs: Input Perturbation, Representation Circuit Breakers, and Guardrail Cascades

    Standard post-training alignment techniques such as Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) instill safety constraints into large language models by shaping token probabilities toward refusal strings. In production environments, however, these surface-level behavioral alignments have proven brittle against systematic adversarial inputs. Gradient-driven token optimization methods such as Greedy Coordinate Gradient (Zou et al., 2023), automated genetic search algorith

    1 min
  • Curriculum Learning in Large Language Models: How Difficulty Pacing, Competence Progression, and Task Scheduling Shape Training Dynamics

    In standard large language model pre-training and fine-tuning pipelines, training batches are almost universally sampled uniformly and independently at random from a static corpus: $$\mathcal{D} = \{z_i = (x_i, y_i)\}_{i=1}^N$$ While this independent and identically distributed (i.i.d.) sampling paradigm aligns with empirical risk minimization (ERM), it ignores the non-convex geometry of deep transformer loss surfaces. Early in training, when network parameters are randomly initialized or unal

    1 min
  • SEC Probes Leopold Aschenbrenner's AI Investment Fund Situational Awareness

    The US Securities and Exchange Commission has launched an inquiry into Situational Awareness, the AI-focused investment fund founded by former OpenAI researcher Leopold Aschenbrenner, according to a report by The New York Times. The investigation follows a sharp July 2026 market downturn across artificial intelligence equities that triggered severe portfolio drawdowns for the high-profile fund. Scope of Subpoenas and Banking Relationships Federal regulators have issued subpoenas to multiple

    1 min