A new report by Telus Digital has exposed significant AI safety gaps across major generative artificial intelligence models. The findings reveal that every tested model can be manipulated into unsafe behavior under specific conditions.
The company’s second GenAI Safety Model Benchmark conducted over 620,000 adversarial tests across 34 models from 10 global providers. These providers included Anthropic, OpenAI, Google, Meta, Alibaba, Baidu, ByteDance, Zhipu AI, 01.AI, and Mistral.
Analyzing the AI Safety Gaps in Modern Models
The testing revealed that vulnerability rates ranged from 1.3% to 93%, with lower percentages indicating safer models. Notably, Anthropic’s Claude models secured five of the ten lowest vulnerability scores, including the lowest overall rate. However, the report noted that even single-digit failure rates present risks in high-stakes enterprise environments.
The research identified model size, reasoning capability, and the developer’s safety approach as the strongest predictors of resilience. Specifically, reasoning models that deliberate before responding showed a 19.9% vulnerability rate, compared to 55.1% for standard models. Meanwhile, smaller models remained highly susceptible to attacks regardless of whether they were open-source or proprietary.
Risk Categories and Model Behaviors
The benchmark highlighted that these AI safety gaps cluster heavily around privacy exploitation, fraud, and cybersecurity threats. Furthermore, researchers identified a pattern termed “refuse-but-engage,” where a model officially declines a harmful request but still provides related information that could cause harm.
The study also noted that open-source models are not inherently less safe than closed ones. For instance, the GLM 4.7 model from China’s Zhipu AI outperformed several proprietary alternatives during the testing process.
Disproportionate Spending on Security
The report highlighted a massive disparity between general artificial intelligence deployment and economy-wide security investments. While global AI spending is projected to reach $2.52 trillion in 2026, spending on AI trust, risk, and security management is projected at just $3.4 billion. This represents approximately $1 spent on security for every $735 spent on AI capabilities.
Consequently, 86% of organizations surveyed report that they have already experienced an AI-related security incident. This mismatch in funding leaves many enterprises underprepared to defend against emerging threats.
Recommendations for Enterprise Defense
To address these AI safety gaps, Telus Digital recommends that enterprises transition to continuous, automated adversarial testing. This testing should be embedded directly into development workflows and combined with human oversight and clean data practices.




