OpenAI slows Astra development over critical cybersecurity risk
OpenAI said Astra may hit a cyber “Critical” threshold, prompting slower development and tighter access. The warning raises fresh doubts about voluntary AI safety controls.

OpenAI has slowed development of its in-progress Astra model after warning that it may reach a “Critical” cybersecurity capability, a level the company says could materially shift the balance between defenders and attackers. The decision reflects a rare public pause from a frontier AI lab at the exact moment its systems are becoming more agentic, more autonomous, and more useful for tasks that can be turned toward intrusion as easily as defense.
The company’s concern is not abstract. OpenAI said Astra has shown major advances in agentic coding and cybersecurity work, the same traits that could help a security team spot weaknesses faster in bank networks, hospital systems, utilities, or government infrastructure. Those same capabilities could also be weaponized for phishing, malware creation, credential theft, and large-scale attack automation if they reach the wrong hands. OpenAI’s framing captures the central tension in frontier AI safety: a model that can chain together complex actions, inspect code, and find vulnerabilities can strengthen incident response while also lowering the cost of abuse.
OpenAI has been building policy around that risk for more than two years. Its preparedness framework beta was introduced on Dec. 18, 2023, and the company updated its Preparedness Framework on April 15, 2025, saying it was measuring and protecting against severe harm from frontier AI capabilities. The framework’s version 2, last updated the same day, said OpenAI was tracking and preparing for frontier capabilities that create new risks of severe harm. That backdrop helps explain why the company would delay a model that appears to be nearing one of its highest cyber-risk thresholds instead of pushing it out on schedule.
The company has also tightened access on the defensive side. On Feb. 5, 2026, OpenAI introduced Trusted Access for Cyber, describing it as a way to strengthen baseline safeguards while piloting trusted access for defensive acceleration. Two months later, on April 14, 2026, it said the program was being scaled to thousands of verified individual defenders and hundreds of teams responsible for defending critical software. That same day, OpenAI said it was preparing for more capable models over the next few months. In February, it also introduced GPT-5.3-Codex as its most capable agentic coding model to date, built for long-running tasks involving research, tool use and complex execution.
The move is likely to intensify pressure on regulators, researchers and competitors already debating how much frontier labs should disclose about internal tests and risk thresholds. The Frontier Model Forum has published work on risk taxonomy and thresholds, underscoring how quickly AI safety has moved from broad principles to formal gatekeeping. OpenAI’s decision to slow Astra before release puts a hard question at the center of the debate: whether a company can safely police models this powerful on its own.
This article was produced by Prism’s automated news system from verified source data, official records, and press releases, then run through automated quality and moderation checks before publishing. The system is built and supervised by the people who set the standards it runs under. Read our full AI policy.
Did this article answer your question?


