AI security guardrails fail across non-English languages

AI models can process dozens of languages, but the safety layers meant to keep them in check predominantly speak English. A Dark Reading investigation published last month highlights a growing concern: organizations operating across multiple languages face uneven protection against jailbreaks, prompt injection, and unsafe model outputs. The report centers on findings from AI security vendor DeepKeep, which published a blog post on July 22 titled "Your AI Speaks 100 Languages. Your AI Security L

2 min
AI security guardrails fail across non-English languages

AI models can process dozens of languages, but the safety layers meant to keep them in check predominantly speak English. A Dark Reading investigation published last month highlights a growing concern: organizations operating across multiple languages face uneven protection against jailbreaks, prompt injection, and unsafe model outputs.

The report centers on findings from AI security vendor DeepKeep, which published a blog post on July 22 titled "Your AI Speaks 100 Languages. Your AI Security Layer Doesn't." The core problem is that leading models from OpenAI, Google, and Anthropic demonstrate strong performance in roughly 30 to 40 languages, but safety tuning, behavior alignment, and reinforcement learning rely heavily on English-speaking annotators.

English benefits from disproportionate training data and tokenization schemes that represent it more efficiently than other languages. Academic research shows LLMs perform logic, reasoning, coding, and math tasks best when prompted in English. Less-supported languages like Welsh and Swahili can handle only basic prompts and may produce grammatical errors.

The security gap emerges because guardrails built on top of models inherit these disparities. A jailbreak attempt phrased in a less-supported language may bypass filters that would catch the same attempt in English. Security intent can shift during translation, making malicious prompts less explicit and harder to detect.

European organizations face particular exposure. The EU's regulatory and economic bloc encompasses dozens of languages, and cross-border operations are routine. A model that safely handles French or German may not apply the same constraints when prompted in Hungarian or Maltese.

Chinese models like Qwen and DeepSeek outperform Western models on Chinese text and cultural context, representing a partial exception to the English-centric pattern. Some initiatives are underway to improve language support in Europe and the Middle East, but the gap remains wide.

Runtime AI firewalls and policy enforcement tools offer a partial mitigation, blocking suspicious tool calls, preventing data exfiltration, and logging prompt injection attempts. These tools add a separate enforcement layer independent of the model's own safety training, but their effectiveness across all languages is not yet well documented.

Sources

[Europe's Multilingual Reality Exposes AI Security Gaps](https://www.darkreading.com/cybersecurity-operations/europes-multilingual-reality-exposes-ai-security-gaps) - Dark Reading, July 24, 2026

[Your AI Speaks 100 Languages. Your AI Security Layer Doesn't](https://www.deepkeep.ai/blog/your-ai-speaks-100-languages-your-ai-security-layer-doesnt) - DeepKeep, July 22, 2026

Written by

More to read

  • OpenAI Flags Astra Model as Potentially Reaching Critical Cybersecurity Risk Level

    # OpenAI Flags Astra Model as Potentially Reaching "Critical" Cybersecurity Risk Level OpenAI has paused parts of development on its upcoming Astra model after internal evaluations indicated it could reach the highest risk tier — "Critical" — in the company's Preparedness Framework for cybersecurity capabilities. This is the first time OpenAI has flagged one of its own models as potentially reaching this level. ## Key Points - Internal tests of Astra showed "significant advancements in agenti

    1 min
  • ByteDance Trains 10 Trillion-Parameter AI Model to Rival Anthropic's Mythos

    ByteDance is pretraining a large model with up to 10 trillion parameters, a scale the Financial Times reports could put it in the same class as Anthropic's most advanced systems. The model, still in early pretraining, would be more than three times the size of Moonshot AI's Kimi K3, currently the largest Chinese model at 2.8 trillion parameters. Three people familiar with the project told the FT the model is in pretraining, a phase that typically lasts three to six months before full training a

    1 min