Tech

How AI guardrails are impeding the work of offensive cybersecurity researchers

Noozly Editorial Desk ·
How AI guardrails are impeding the work of offensive cybersecurity researchers

Offensive security researchers who hunt for unpatched software flaws and build tools to demonstrate how they could be exploited say the very safeguards meant to keep artificial intelligence out of criminal hands are now getting in their way. In conversations about the guardrails built into models from OpenAI and Anthropic, several of these researchers described systems that were designed to stop bad actors but that end up snagging legitimate defensive work along with it.

The tension stems from a broader industry push over the past year to keep powerful AI systems from being weaponized. Both companies have layered on vetting requirements and usage restrictions aimed at preventing their chatbots and coding assistants from being turned into engines for real-world hacking campaigns. For people whose job is to probe systems for weaknesses before criminals find them, though, those same restrictions can make the tools reluctant to help with tasks that look identical to an attack, even when the intent is defensive.

The friction became a matter of government policy in June, when U.S. officials imposed export control restrictions on two of Anthropic's flagship models, known as Mythos and Fable. Regulators acted at least partly in response to findings suggesting the models' built-in protections against misuse for cyberattacks could be circumvented, raising alarm that the systems might be coaxed into helping build or launch malicious code.

Independent of whether that jailbreak concern was the true driver behind the restrictions, Anthropic itself has long promoted Mythos as an extraordinarily capable system in the offensive-security realm, describing it in terms that cast it as a uniquely dangerous cyber tool that should only reach carefully screened users under tight controls. That framing, researchers say, has shaped how conservatively the model behaves even for accounts with legitimate security work to do.

The export limits did not stay in place for long. Restrictions on both models have since been lifted, though not on equal terms. Fable 5 returned to unrestricted public availability on July 1, while Mythos 5 has been made available again only to vetted organizations within the United States, as part of an ongoing government review process rather than a full return to general access.

The episode echoes a long-running debate in the security industry over so-called dual-use tools — software that can just as easily map a network's defenses as break through them. Penetration testers and vulnerability researchers have historically relied on unrestricted access to scripting and exploitation frameworks; as AI models increasingly sit alongside those tools, restrictions calibrated to stop criminals risk also throttling the researchers hired to find flaws before attackers do.

Not everyone in the field views the guardrails as pure overreach. Some practitioners acknowledge that a model capable of automating attack code at scale poses genuine risks if released without any friction, and that vetting programs, however imperfect, are preferable to no controls at all. The disagreement is less about whether limits should exist than about where the line should sit and how quickly access can be restored once a model proves itself safe for professional use.

For now, the uneven reinstatement of access — full availability for one model, conditional vetting for the other — suggests AI developers and regulators are still calibrating that balance in real time, with offensive researchers caught in the middle of a policy still being written.

Source: TechCrunch

techtechcrunch
Original source
TechCrunch →

Related articles