ai guardrails AI companies have spent the past year tightening access to their most capable models, adding vetted researcher programs and cybersecurity guardrails meant to keep the tools out of the hands of malicious hackers. But as those controls get stricter, some offensive security researchers say they are also making legitimate vulnerability discovery harder, slower, and in some cases less effective.
ai guardrails
Guardrails meant for attackers are affecting defenders too
The tension has become more visible in the wake of U.S. export control restrictions placed in June on Anthropic’s AI models Mythos and Fable. The government move was prompted at least in part by a report that said the models’ guardrails could be bypassed to help build and execute cyberattacks. Anthropic has long positioned Mythos as a highly controlled system available only to carefully vetted users, with strict restrictions layered on top.
Those controls have since changed. According to the material, export controls on Fable 5 and Mythos 5 were later lifted. Fable 5 returned to general access on July 1, while Mythos 5 was reintroduced only to vetted U.S. organizations as part of the government’s review process.
Anthropic and OpenAI both offer programs intended to give cybersecurity researchers more access with fewer restrictions: OpenAI’s Trusted Access for Cyber program and Anthropic’s Cyber Verification Program. But researchers who work offensively — meaning they actively probe systems for weaknesses, exploitability, and proof-of-concept behavior — say the restrictions can get in the way of useful work.
“A hammer” that can also be a weapon
During a recent cybersecurity podcast appearance, researcher Mark Dowd said he was uncomfortable with “these random large companies” making “arbitrary decisions about what is safe in security and what’s not.” Dowd acknowledged that his own background — he has spent decades finding and selling “zero-days” to Western governments rather than reporting them to vendors — may shape his view. Governments pay a premium for those vulnerabilities because they remain unpatched and usable for intelligence operations.
Other researchers described a more practical problem: AI models are often useful precisely when they are pushed to reason through an exploit path, but guardrails may refuse to engage once the conversation becomes too close to offensive analysis.
Chris Anley, chief scientist at NCC Group, said asking an AI model to try to exploit a bug is a key part of determining whether a flaw is real and worth fixing. If the model refuses, that can hinder defenders as much as attackers.
Anley described the issue as the overlap between offensive and defensive work. A prompt such as “fix this code,” he said, can serve both as a defensive exercise and as “a roadmap for finding critical vulnerabilities in the code base.” In his view, the same tool can be both useful and dangerous, and those uses are hard to separate cleanly.
He compared AI to “a hammer”: necessary for building, but also “irreducibly a weapon.”
Some researchers shift to local or open-source models
When guardrails block a task, Anley said he and his colleagues sometimes turn to open source AI models with no restrictions.
Paolo Stagno, chief technology officer at Crowdfense, said he and his team use frontier models mainly for reverse engineering, not for finding vulnerabilities or building exploits. He argued that sending that work to a cloud-based service risks exposing sensitive data or allowing it to be absorbed into future training runs. For more sensitive tasks, Stagno said, his team prefers open source models run locally, where data does not leave the environment.
Stagno also criticized the broader access model from AI vendors, saying they “essentially treat customers like children who need babysitting” through their vetted programs and guardrails.
Not every researcher sees guardrails as a blocker
Some offensive security researchers say their use of AI is limited enough that the restrictions are less of a problem. Giuseppe Cali, who finds zero-days and develops exploits, said guardrails do not impede his work because he does not rely on AI for offensive operations. Instead, he uses it for initial reverse engineering, to understand code and build supporting tools.
For that kind of work, Cali said, AI can speed up analysis and free him to focus on discovering vulnerabilities. But he made clear that he wants to keep the core work in human hands.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” Cali said. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
Strict limits can make models nearly unusable
One researcher at a smartphone-component manufacturer, who spoke anonymously because he was not authorized to speak to the press, said his employer is not part of Anthropic’s CVP program. As a result, he said, the company’s tools are “barely useful” for finding vulnerabilities because the guardrails are so strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” he said.
Chris Thompson, chief executive of cybersecurity firm RemoteThreat and founder of Offensive AI Con, said he has seen similar inconsistency. Even inside the looser limits of vetted access programs, he said, guardrails can behave differently from one day to the next.
That unpredictability, he argued, wastes time that should be spent on analysis.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” Thompson said. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
Why some researchers are turning elsewhere
Thompson said the result is that researchers are sometimes pushed toward Chinese open source models such as GLM, which can be downloaded freely and run locally without vetting or usage restrictions. In his view, that creates an unintended consequence: responsible researchers may migrate away from U.S.-governed systems toward foreign-owned ones.
“I think it’s more harmful than good to have these guardrails in place,” he said.
Rather than adding more restrictions, Thompson called on frontier AI labs to broaden access, provide responsible use paths, and hold abusive users accountable. His warning is that defenders need these tools now because the next wave of cyberattacks will arrive faster and at greater scale.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” Thompson said. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
Explore more: Blog Our Services Contact Us
Source: Original report
Was this helpful?
Last Modified: July 24, 2026 at 6:37 pm
4 views
