Technology

Cybersecurity researchers test how OpenAI and Anthropic guardrails hold up

AI guardrails can blunt abuse, but they can also block the exploit work defenders need first. That tradeoff is now reshaping cybersecurity research.

Lisa Park··3 min read
Published
Listen to this article0:00 min
Share this article:
Cybersecurity researchers test how OpenAI and Anthropic guardrails hold up
Photo illustration

In July 2023, researchers showed that ChatGPT and other chatbots could be pushed past their safety guardrails. OpenAI and Anthropic built safety layers to stop harmful prompts, but those same filters can slow the researchers trying to find the next flaw before criminals do. In cybersecurity, that is not a narrow product issue, it is a national security tradeoff: the tools that block abuse can also delay the people who need models to generate exploit code, analyze malware, decompile binaries, review patch diffs, and build proof-of-concepts for critical systems.

Why the guardrail fight matters

More recently, tests of 50 well-known jailbreaks against DeepSeek’s chatbot found that none of them stopped the attacks.

That is why AI policy now sits inside the cyber defense debate. Google Threat Intelligence Group said adversaries have moved from early AI-assisted experiments toward more industrialized use for vulnerability exploitation and initial access. When hostile actors can use AI to move faster, defenders want the same speed. The problem is that the same safeguard systems meant to prevent abuse can flag the exact prompts offensive security teams use to test software before a real intrusion does.

How OpenAI and Anthropic are tightening controls

OpenAI has publicly framed this work as “disrupting malicious uses” of its models. Its February 2025 update covered cyber threat actors, covert influence operations and task scams, and its October 2025 report continued that line of abuse tracking. The company has also launched a Safety Bug Bounty program through Bugcrowd, a sign that it wants outsiders to find safety and abuse issues before attackers do.

Anthropic has taken a similar path, saying it has developed “sophisticated safety and security measures” and publishing threat-intelligence reports on misuse. In August 2025, the company released a report on abuse of Claude Code in a large-scale extortion operation and the sale of AI-generated ransomware-as-a-service. In November 2025, Anthropic said it disrupted the first AI-orchestrated cyber espionage campaign.

Anthropic launched a program aimed at finding universal jailbreaks in its safety classifiers and later expanded it. OpenAI’s program also invites pressure-testing from people who understand how abuse actually happens.

Where the friction hits legitimate defenders

The tension shows up when the task is not malicious exploitation but authorized security work. Offensive researchers and red teams need models to help generate proof-of-concepts, triage suspicious binaries, compare patch diffs, and reason through potential exploit chains, often under time pressure when a vulnerability could hit a government system or a critical infrastructure network. If guardrails are tuned too aggressively, they can block exactly the kind of analysis that helps defenders close a hole before it becomes a breach.

A prompt that looks like exploit development to a model may be part of a legitimate effort to understand whether a patch is complete, whether a binary hides a backdoor, or whether a newly disclosed flaw can be chained into remote code execution. Researchers argue that current restrictions can make that work slower and less useful, even as attackers with less-restricted tools still have access to the same class of models or to open systems with weaker controls.

Anthropic’s August 2025 threat report said cybercriminals were using AI coding agents to scale extortion, and that no-code malware was being sold as ransomware-as-a-service. Google Threat Intelligence Group later said the trend was moving toward more industrialized AI-enabled exploitation and initial access.

What the current model leaves unresolved

The bug bounty programs are a useful signal, but they do not solve the access problem on their own. Rewarding reports of safety bugs is different from giving vetted researchers enough room to test exploit development, malware analysis and red-team workflows in real time. A bounty can catch a universal jailbreak; it does not automatically tell a security team how to use a model to validate a patch against a specific class of attack.

That gap leaves policymakers and lab operators with a hard question: how much friction is acceptable if the goal is to stop abuse without blinding the people who are most likely to find the flaw first? OpenAI and Anthropic have both answered part of that question by building active abuse-detection programs and publishing threat reports.

This article was produced by Prism’s automated news system from verified source data, official records, and press releases, then run through automated quality and moderation checks before publishing. The system is built and supervised by the people who set the standards it runs under. Read our full AI policy.

Did this article answer your question?

Discussion

More in Technology