AI Safety8 articles

AI Safety

Articles

  • Meta Joins OpenAI and Anthropic as Third AI Lab Whose Model Hacked External Systems During Testing

    Meta has become the third major AI company in as many weeks to disclose that one of its models breached external systems during cybersecurity testing, following similar incidents at OpenAI and Anthropic. Meta's Muse Spark model exploited a security vulnerability in another company's systems during an evaluation conducted by Irregular, an independent testing firm, a Meta spokesperson confirmed Wednesday. The breach occurred due to a misconfiguration by Irregular that inadvertently gave the model

    1 min
  • Claude breached three real organizations during Anthropic's cybersecurity tests

    Anthropic has disclosed that three of its Claude models gained unauthorized access to the production systems of three separate organizations during cybersecurity evaluations, after a misconfiguration left test environments connected to the open internet. The company began a retrospective review of 141,006 evaluation runs on July 23, following OpenAI's July 21 disclosure that its own models had escaped a sandboxed environment and accessed Hugging Face infrastructure via a zero-day exploit. Anthr

    1 min
  • Claude breached three real organizations during Anthropic's cybersecurity tests

    Anthropic has disclosed that three of its Claude models gained unauthorized access to the production systems of three separate organizations during cybersecurity evaluations, after a misconfiguration left test environments connected to the open internet. The company began a retrospective review of 141,006 evaluation runs on July 23, following OpenAI's July 21 disclosure that its own models had escaped a sandboxed environment and accessed Hugging Face infrastructure via a zero-day exploit. Anthr

    1 min
  • UK Safety Institute Finds Frontier Models Deceived Humans During Cyber Testing

    The UK's AI Security Institute (AISI) disclosed Tuesday that AI agents powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol engaged in what it described as "sustained, potentially harmful activity directed at real people and organisations" during routine cybersecurity evaluations. Across 122 test repetitions, evaluators flagged 19 concerning actions. Seventeen came from a single model, Mythos 5. Two involved GPT-5.6 Sol with cyber classifiers enabled. The institute said it was "the first ti

    1 min
  • Claude overruled a simulated Anthropic CEO in agentic misalignment test

    Anthropic researchers found that Claude disobeyed a fictional version of CEO Dario Amodei during a simulation designed to test how AI agents behave when they believe their employer is concealing a safety risk. The findings, published by Anthropic's alignment team, were reported by The Bureau of Investigative Journalism on July 20. Claude Opus 4.5, deployed under the test name "Atlas," was placed inside a fictional Anthropic AI safety team with access to staff messages, calendars, and research f

    1 min
  • OpenAI paused an experimental model that kept escaping its sandbox

    OpenAI revealed it was forced to halt internal deployment of an experimental AI model after the system repeatedly found ways to circumvent its containment restrictions, a disclosure that adds to a growing list of control failures at the company. In a blog post about safety alignment in long-horizon models, OpenAI described an autonomous system designed to operate for hours or days without human intervention. Unlike previous models that would stop and return to the user when hitting environmenta

    1 min
  • Claude finds mathematical flaws in post-quantum cryptography and AES

    Anthropic researchers using Claude Mythos Preview have discovered improved attacks against two widely studied cryptographic algorithms, demonstrating that frontier AI models can find mathematical weaknesses in encryption schemes, not just bugs in their implementations. The first result is an improved key recovery attack against HAWK, a post-quantum digital signature scheme currently under consideration by NIST. The second identifies a new approach to attacking round-reduced versions of AES, the

    1 min
  • Claude breached three real organizations during Anthropic's cybersecurity tests

    Anthropic has disclosed that three of its Claude models gained unauthorized access to the production systems of three separate organizations during cybersecurity evaluations, after a misconfiguration left test environments connected to the open internet. The company began a retrospective review of 141,006 evaluation runs on July 23, following OpenAI's July 21 disclosure that its own models had escaped a sandboxed environment and accessed Hugging Face infrastructure via a zero-day exploit. Anthr

    1 min