Brendan Smialowski | AFP | Getty Images
The company said Astra can find previously unknown security flaws and exploit them without step-by-step guidance from humans, which means the model falls under the most advanced category of its so-called Preparedness Framework. OpenAI said it still plans to make Astra available “soon,” but that access to its cybersecurity capabilities will be more limited.
OpenAI introduced its Preparedness Framework in 2023, and it serves as the company’s method for “tracking and preparing for advanced AI capabilities that could introduce new risks of severe harm.” In an update to the framework last year, the company outlined a “High” capability threshold, where models could amplify “existing pathways” to severe harm, and a “Critical” capability threshold, where models could introduce “unprecedented new pathways” to severe harm.
“We will share more details about our safety, security and alignment testing and evaluations in the model’s System Card at launch,” OpenAI said in a blog post on Tuesday.
OpenAI’s security and safety practices have been under intense scrutiny after the company disclosed that two of its models escaped their training environment, accessed the open web and breached Hugging Face’s systems last month. OpenAI characterized the attack as an “unprecedented cyber incident” and temporarily paused some of its internal training and research.
The company decided to delay parts of Astra’s development even though the model was not involved in the Hugging Face incident. After strengthening and testing protections, OpenAI said Tuesday that it believes the model’s safeguards “sufficiently minimize the risk of severe harm for release under our Preparedness Framework.”
Astra’s advanced cyber capabilities will be available to a select group of organizations that are part of its cybersecurity coalition called Daybreak, OpenAI said.