AI companies tighten guardrails for cyber misuse
TechCrunch reports offensive security researchers say refusals block exploit verification, open-source models fill the gap when cloud tools say no
AI guardrails slow offensive security work, TechCrunch reports vetted access programs still block exploit verification, researchers shift to open models when cloud tools refuse
Cybersecurity researchers who test software by breaking it are running into a new friction point: large AI systems that refuse to help. TechCrunch reports that companies building frontier models have tightened “guardrails” and introduced vetted-access programs meant to limit criminal misuse, but the same controls also interfere with legitimate work such as confirming whether a vulnerability can be exploited.
The clash is partly about who gets to decide what “too dangerous” looks like, and how that decision is enforced. TechCrunch notes that Anthropic and OpenAI offer programs—Anthropic’s Cyber Verification Program and OpenAI’s Trusted Access for Cyber—that provide vetted researchers with fewer restrictions. Even so, researchers cited by TechCrunch say models still decline requests that are central to their workflow, especially when they need to turn a suspected bug into a working exploit to prove it is real and to measure its impact.
The policy debate has already spilled into government action. TechCrunch points to U.S. export controls imposed in June 2026 on Anthropic’s models Mythos and Fable after a report claimed it was possible to bypass guardrails to build and execute malicious cyberattacks. Those export controls were later lifted, with Fable returning to general access on July 1, while Mythos was reintroduced only to vetted U.S. organizations as part of a government review process.
Researchers quoted in the piece describe the practical result: when a cloud model refuses to engage, work does not stop—it routes around the refusal. TechCrunch reports that offensive security professionals sometimes switch to open-source models that lack comparable guardrails. Others narrow how they use frontier systems to reduce the chance that sensitive material is exposed to a provider, using them for tasks like reverse engineering while avoiding vulnerability discovery or exploit building.
Mark Dowd, a security researcher cited by TechCrunch, said he is uncomfortable with large companies making what he views as arbitrary decisions about security “safety.” Chris Anley, chief scientist at NCC Group, told TechCrunch that asking an AI system to exploit a bug is often the step that confirms a vulnerability is real—meaning refusals can slow both attackers and defenders, while changing who can afford to keep going.
The controls are marketed as a brake on misuse, but the article’s reporting suggests they also act as a sorting mechanism: those with access to alternative models, private tooling, or open systems keep their pace, while everyone else waits for a refusal message.