AI security guardrails hinder vulnerability research and drive shift to foreign models
Strict usage restrictions on frontier artificial intelligence models are obstructing legitimate cybersecurity research, prompting experts to warn that defenders are being pushed toward foreign-owned open-source systems.
Strict usage restrictions on frontier artificial intelligence models are increasingly obstructing legitimate cybersecurity research. While designed to prevent malicious hackers from exploiting these systems, the guardrails are now hindering the network defenders tasked with finding and fixing vulnerabilities first.
In June, the United States government imposed export control restrictions on Anthropic’s Mythos and Fable models following reports that their safeguards could be bypassed. Although controls on Fable 5 were lifted on 1 July and Mythos 5 was reintroduced to vetted US organisations, the broader climate of gatekeeping remains.
Tech giants like Anthropic and OpenAI have introduced vetted initiatives, such as the Cyber Verification Program and Trusted Access for Cyber, to grant approved researchers access to less restricted models. However, security professionals argue these measures are arbitrary and counterproductive.
Mark Dowd, a prominent security researcher who sells previously unknown software flaws to Western governments, criticised the approach. He stated that "it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not."
The dual nature of these tools complicates the issue for European and global security firms. Chris Anley, chief scientist at NCC Group, noted that asking an AI to exploit a bug is essential for confirming a vulnerability is worth fixing.
Anley explained that asking an AI to "fix this code" serves as both an essential mechanism for defence and a roadmap for finding critical flaws. He compared the technology to a hammer, noting "it’s definitely a tool but it’s also irreducibly a weapon as well."
This friction has direct implications for the cybersecurity market and corporate defence strategies. When frontier models refuse to process security-related queries, firms are forced to seek alternatives, slowing down vulnerability patching and increasing corporate risk exposure.
Paolo Stagno, chief technology officer at CrowdFense, argued that AI companies "essentially treat customers like children who need babysitting." Consequently, his firm avoids using cloud-based frontier models for exploit development to prevent sensitive data leakage, relying instead on local open-source alternatives.
Not all researchers face the same barriers. Giuseppe Cali, a security researcher, stated that guardrails do not impede his work because he uses AI only for initial reverse engineering and tool building, preferring to handle bug discovery himself.
Conversely, an anonymous researcher at a smartphone-component manufacturer reported that strict guardrails render the tools unusable for security work. Chris Thompson, chief executive of RemoteThreat, added that inconsistent model behaviour forces researchers to spend time negotiating with the AI rather than analysing vulnerabilities.
The most significant security consequence is the migration of analytical talent. Thompson warned that "responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems," such as Chinese open-source models like GLM, which can be run locally without restrictions.
Thompson urged frontier AI laboratories to provide responsible access and hold abusers accountable rather than tightening restrictions. He warned that "there’s this big wave of attacks that are going to happen at speed and scale like never before," leaving legitimate security consultants stifled in their efforts to prepare.