Washington [US], August 9 (ANI): OpenAI has paused work on its upcoming AI model Astra after internal evaluations showed potentially dangerous advances in agentic coding and cybersecurity, with the company unable to rule out that the system could possess “critical cyber capabilities.”
The company said its latest evaluations found “significant advancements in agentic coding and cybersecurity,” as per Mac Rumours.
Under OpenAI’s Preparedness Framework, the potential development of such capabilities triggers stricter safeguards, particularly when models could create risks of severe harm.
OpenAI said it is “pausing” activities involving Astra while it strengthens its security controls.
The company’s cybersecurity guidelines call for additional protections for models that “create new risks of scaled cyberattacks and vulnerability exploitation.”
The “Critical” threshold described in the framework involves the ability to identify and develop functional zero-day exploits across severity levels in many hardened, real-world critical systems without human intervention.
It also includes the ability to devise and execute end-to-end novel strategies for cyberattacks.
Before Astra can be deployed, OpenAI plans to introduce additional safeguards and security measures.
These include restricting work on the model until the new protections are implemented, using isolated testing environments with limited network and tool access, adding sandboxed execution and expanding monitoring capabilities.
OpenAI also said it will work with relevant government agencies and AI safety organisations to test Astra.
“We’re committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models like Astra, and those that follow, are deployed responsibly and broadly for the benefit of all humanity,” writes OpenAI.
Astra has not been formally announced. However, OpenAI recently shared details about its next major model while outlining mathematical advances.
The model reportedly solved 10 open problems in mathematics and theoretical computer science for around USD 2,000 at Sol API rates, as per MAc Rumours.
The development comes as AI systems increasingly demonstrate capabilities relevant to cybersecurity. Apple recently limited submissions to its bug bounty programme after facing difficulties handling the volume of bugs being uncovered.
Anthropic’s Claude Mythos can identify critical vulnerabilities and is available to select companies, including Apple. The system is restricted because of its ability not only to find vulnerabilities but also potentially exploit them.
OpenAI also made headlines in July after GPT-5.6 Sol and another “more capable pre-release model” autonomously hacked Hugging Face during internal benchmark testing.
Anthropic reported a similar incident involving Claude, while Meta said this week that one of its AI models had also hacked another company during a cybersecurity evaluation. (ANI)
OpenAI pauses Astra AI Model over critical cybersecurity concerns
Leave a Comment
