• 4 mins read
  • Published

OpenAI Tightens Security on Astra After Internal Cyber Risk Tests

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

OpenAI Tightens Security on Astra After Internal Cyber Risk Tests Science.Report © science.report
OpenAI Tightens Security on Astra After Internal Cyber Risk Tests © science.report

OpenAI has placed its Astra AI model in the highest internal cybersecurity risk category following internal tests that suggest advanced offensive cyber capabilities, leading to stricter security controls and external review before any release

OpenAI has escalated internal security measures for its upcoming AI model Astra after internal testing indicated the system may possess advanced offensive cyber capabilities. According to the company, Astra is now classified under the "Critical" risk tier in OpenAI's Preparedness Framework, a designation reserved for models that could potentially develop or execute sophisticated cyberattacks without human assistance. This is the first time an OpenAI model has been assessed at this level, surpassing previous models such as GPT-5.6-Sol, which remained in the "High" category after similar evaluations.

The Preparedness Framework, introduced by OpenAI in late 2023, defines the "Critical" category as applying to models that can autonomously discover and exploit zero-day vulnerabilities in hardened real-world systems or conduct complex cyber operations from broad objectives. While OpenAI has not confirmed that Astra has fully crossed this threshold, the company's internal and expert reviews suggest the model's autonomous coding and cybersecurity performance have advanced enough to warrant the highest level of caution. OpenAI has stated that Astra is not connected to any recent cyber incidents, including the exploitation of Hugging Face.

Security Measures and Testing

In response to these findings, OpenAI has implemented a series of enhanced security protocols around Astra's development environment. These include isolated testing systems, stricter network segmentation, stronger encryption for model weights, and expanded monitoring tools. All internal work involving Astra must now comply with these upgraded requirements, and any non-compliant projects have been paused. Additionally, Astra's agentic applications are subject to universal monitoring, with systems in place to review the model's chain of thought during both training and evaluation. If potentially dangerous or misaligned behavior is detected, automated systems can trigger a security review and halt high-risk activities.

Before Astra is made available to external users, OpenAI plans to collaborate with government agencies and selected independent AI safety organizations. External evaluators will be provided with recommended security controls for higher-risk testing scenarios. The company has emphasized that ongoing testing is required to validate Astra's capabilities and to determine whether it meets the full criteria for the "Critical" risk category.

Numerical Context and Historical Comparison

OpenAI's Preparedness Framework was established in 2023 to track and manage emerging risks in advanced AI systems. Previous models, including GPT-5.6-Sol, were evaluated under this framework and remained in the "High" risk category after internal and expert review. Astra is the first model to prompt consideration of the "Critical" tier, which is defined by the ability to autonomously discover and exploit zero-day vulnerabilities or conduct sophisticated cyberattacks without human intervention. The company has not disclosed specific benchmark scores or the number of tests conducted, but the decision to escalate Astra's risk category was based on a combination of internal evaluations and expert assessments.

Safeguards and Intended Use

OpenAI has stated that the Preparedness Framework is designed to anticipate moments when AI systems approach sensitive capability thresholds, particularly those with dual-use or offensive potential. The company previously expanded testing and safeguards in 2025 when its models neared the "High" capability level for biological risks. Despite the increased restrictions around Astra, OpenAI maintains that its long-term goal is to develop advanced cybersecurity models that can help defenders identify and remediate vulnerabilities before they are exploited by malicious actors. Astra will not be released until it satisfies all necessary safety and security requirements, and OpenAI has committed to ongoing collaboration with external experts to ensure responsible deployment.

Understanding the Preparedness Framework is essential to interpreting OpenAI's actions. The framework is an internal risk management system that categorizes AI models based on their demonstrated or plausible capabilities, particularly in areas with significant safety or security implications. Models are evaluated through a combination of technical testing, expert review, and scenario analysis. The "Critical" category is reserved for systems that could independently conduct high-impact cyber operations, and entry into this tier triggers mandatory security upgrades, external review, and restricted access. The framework is intended to provide a structured approach to identifying and mitigating risks before advanced AI models are deployed beyond controlled environments.

Related articles