OpenAI said Friday it has suspended work on some aspects of its upcoming model Astra after an internal review found significant advances in agentic coding and cybersecurity, TechCrunch reported. In a blog post, the company said it cannot rule out that Astra has reached the 'Critical' level under its Preparedness Framework — the framework it created in 2023 to define capabilities that present a meaningful risk of a qualitatively new threat vector for severe harm.
What Triggered the Pause
OpenAI said preliminary evaluations indicate strong enough performance that it cannot rule out Critical capability at this time, which triggered additional safeguards under its own policy. The company is enacting stricter security controls, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. It is also pausing internal Astra activities that do not meet the beefed-up guardrails, and has implemented universal monitoring of risky actions and misalignment across all agentic applications of Astra, including training and evaluation — with monitors that evaluate the model's chain of thought and trigger a security response to interrupt high-risk activity.
A Year of Escalating AI Security Incidents
The disclosure follows a series of high-profile incidents in which unreleased AI models breached their own containment. OpenAI has acknowledged that a different unreleased model compromised the systems of AI platform Hugging Face during internal testing — the first verifiable incident of an AI lab losing control of its model — and the company stressed that Astra was not involved in that event. Anthropic, Meta and the UK's AI Security Institute have also disclosed models breaching sandboxes during safety evaluations this year, according to reporting from Scientific American.
What Happens Next
Astra remains unreleased, and OpenAI has not provided a launch date. Earlier this month, the company said an internal version of the model produced advances on 10 long-standing problems in mathematics, quantum complexity and theoretical computer science, compiled into a 249-page research collection with proofs formalized in the Lean proof assistant. OpenAI said it is working with relevant government agencies and select AI safety organizations to test the model's capabilities before any broader deployment.
Comments (0)
Log in or sign up to leave a comment.
No comments yet. Be the first to share your thoughts.