OpenAI has slowed development of its upcoming AI model Astra after internal safety tests raised concerns about its growing ability to perform sophisticated cybersecurity tasks.
The company’s evaluations found that Astra may be approaching a level of cyber capability that OpenAI classifies as “critical”. The finding has triggered additional safety measures and a slowdown in parts of the model’s development as researchers assess the potential risks.
The concern centres on Astra’s ability to carry out complex, multi-step cybersecurity tasks with limited human assistance. OpenAI’s preparedness framework defines a critical cyber capability as the ability to autonomously identify and exploit serious software vulnerabilities, including previously unknown flaws, or conduct sophisticated attacks against highly protected systems.
OpenAI has not said Astra has conducted a real-world cyberattack. The concern is based on its performance during controlled internal testing. The company is evaluating whether the model’s capabilities could be misused if it were given access to computer systems, software tools or other digital infrastructure.
The development comes as the AI industry moves towards models that can do more than generate text or answer questions. Newer AI systems are increasingly capable of writing and debugging code, analysing software, using digital tools and completing tasks across several stages without constant human direction.
These capabilities can be valuable for cybersecurity. AI can help security researchers examine large volumes of code, identify vulnerabilities, investigate suspicious activity and develop fixes more quickly. The same skills, however, could potentially be used to identify weaknesses in systems and automate parts of a cyberattack.
This dual-use problem has become one of the biggest challenges for developers of advanced AI models. A system capable of discovering a vulnerability can help defenders patch it, but the same capability could become dangerous if an attacker gains access to the model or manipulates it into assisting with malicious activity.
Astra’s development is therefore being assessed not only on what the model can accomplish, but also on how safely those capabilities can be deployed. OpenAI is examining additional controls that could limit the model’s access to sensitive systems and prevent it from carrying out high-risk actions without human approval.
The issue is particularly important as AI agents become more autonomous. Unlike traditional chatbots, agentic AI systems can break down a goal into multiple steps, use external tools and work through a task with less human intervention.
In cybersecurity, that could mean an AI system identifying a software weakness, analysing how it might be exploited and producing code to test the vulnerability. Such capabilities could dramatically speed up defensive security work but could also reduce the expertise and time required to launch sophisticated attacks.
OpenAI’s internal testing has highlighted the need for stronger safeguards as these capabilities improve. The company is expected to continue evaluating Astra while tightening controls around its cybersecurity functions.
The development also reflects a broader shift in how AI safety is being assessed. Earlier concerns around advanced models largely focused on misinformation, privacy, bias and misuse of generated content. As AI systems become more capable programmers and autonomous agents, cybersecurity has emerged as a more immediate technical risk.
For businesses, the implications extend beyond OpenAI. Companies are increasingly using AI tools in software development, cloud operations and security monitoring. Giving AI systems access to internal networks, databases or code repositories can increase productivity, but it also creates new security risks if permissions are not tightly controlled.
Security experts have increasingly advocated limiting AI systems to the minimum access required for a task. Isolated environments, strong authentication, continuous monitoring and human approval for sensitive operations can reduce the consequences of an AI system making an unsafe decision.
The Astra case also puts pressure on developers to improve their AI preparedness frameworks. Traditional software testing relies heavily on known attack scenarios, while frontier AI models can develop new combinations of skills as their reasoning and tool-use abilities improve.
That makes capability evaluations an important part of the development process. A model may appear safe during one stage of training but demonstrate substantially stronger abilities after further improvements.
OpenAI’s decision to slow Astra development demonstrates how such evaluations can influence the rollout of a frontier AI model. Instead of treating cybersecurity capability as simply another performance metric, the company is considering it as a core safety issue.
The wider technology sector is facing the same balancing act. More powerful AI could give cybersecurity teams tools that were previously available only to highly specialised experts. At the same time, those capabilities could make certain cyber threats faster, cheaper and easier to scale.
Astra’s development will therefore be closely watched as OpenAI works to determine the safeguards needed for its next-generation AI system. The key issue is no longer simply how capable the model can become, but whether those capabilities can be deployed without creating unacceptable cybersecurity risks.