Safety7 articles

Safety

Articles

  • AI Guardrails in Production: Multi-Stage Filtering, Classification Models, and the Latency Tax

    AI Guardrails in Production: Multi-Stage Filtering, Classification Models, and the Latency Tax Production deployments of large language models cannot rely solely on system prompt instructions to maintain safety, prevent prompt injection, or restrict domain scope. System prompt alignment is inherently susceptible to adversarial bypasses, context dilution, and non-deterministic instruction following. To enforce strict security, compliance, and topic boundaries, engineering teams increasingly depl

    1 min
  • OpenAI Disbands Preparedness Team in Continued Safety Restructuring

    OpenAI has disbanded its dedicated Preparedness team, redistributing safety and frontier-risk evaluation responsibilities across individual functional units, according to reporting by the Financial Times and The Verge. The Preparedness team was established in late 2023 to evaluate and mitigate catastrophic risks associated with frontier AI models, focusing on cybersecurity exploits, chemical, biological, radiological, and nuclear (CBRN) threats, and autonomous model behavior. Under the new orga

    1 min
  • Researchers find that changing text color can hijack a vision-language model's reasoning

    A new research collaboration has found that the color, contrast, and brightness of text can quietly steer the outputs of vision-language models (VLMs), causing them to misread meaning and reach different conclusions without any change to the words themselves. The authors say their experiments provide a systematic analysis of how low-level visual styling of text distorts the semantic representations inside a VLM's vision encoder, and how those shifts show up as behavioral changes across both sub

    1 min
  • Mistral launches Shieldstral, a 3B open-weights policy-adaptive safety classifier

    Mistral AI released Shieldstral on Monday, a 3-billion-parameter open-weights safety classifier that matches or outperforms guard models up to seven times its size. The model is available under Apache 2.0 and runs on a single 16 GB GPU. What sets Shieldstral apart from typical guardrail models is its approach to content moderation. Rather than baking a fixed taxonomy of harm categories into the model weights — which forces developers to retrain whenever their safety requirements change — Shield

    1 min
  • Fable 5 comes back on a shorter leash

    Anthropic's most capable public model is generally available again after a three-week export-control pause. The interesting part isn't the model — it's the classifier stack now wrapped around it.

    1 min