OpenAI has announced that it is pausing part of the development process for Astra, its forthcoming artificial intelligence model. The company says this new AI model is too capable in cyber, and is taking time to develop safeguards better suited to those abilities.
AI is advancing rapidly and some frontier models have now gained cyber capabilities that concern their own developers. This year, Anthropic made headlines when it unveiled Claude Mythos, a model so powerful that it could be used for hacking if it fell into the wrong hands. To prevent this, the company had to create Claude Fable 5, a safer consumer version whose cyber, biology and chemistry capabilities have been restricted. OpenAI is now also concerned about the excessive cyber performance of an AI model still under development.
OpenAI pauses Astra
Known internally as Astra, this AI model was involved in the widely reported Hugging Face security incident during testing. That is not the only concern: OpenAI also believes Astra could cross the critical threshold set by its AI risk-management system, known as the Preparedness Framework. According to an earlier OpenAI publication, that threshold is reached when an AI model “is capable of identifying and developing functional ‘zero-day’ exploits of every severity level across a broad range of real-world, highly secure critical systems, without human intervention, or if it can devise and implement end-to-end innovative cyberattack strategies against highly secure targets when given only a broad objective.”
In response to this risk, OpenAI says it has chosen to slow Astra’s development. “This notably involved a two-week pause in reinforcement learning (RL) training for our latest models intended for deployment, while we further strengthened the security of our research environments, carried out attack-simulation testing in them, and expanded the coverage of our monitoring systems,” the company said.
OpenAI keeps Astra but strengthens its safeguards
Despite the incidents that have already occurred and the risks identified, OpenAI is not abandoning Astra. However, slowing its development should give the company time to reinforce its safeguards, so that the new AI can be developed and used safely. Put simply, as AI capabilities progress, the safeguards must progress as well.
The pause in Astra’s development is intended to enable these extra measures to be introduced. The company’s work will focus on monitoring AI actions to spot suspicious behaviour, alignment to reduce the likelihood of harmful or prohibited actions, and security by strengthening the isolation of AI models.
What we think
These efforts to secure frontier AI models could nevertheless have a significant effect on OpenAI’s finances. The company is not yet profitable, even though ChatGPT now claims more than one billion people. “Developing methods that can keep pace with these capabilities will require sustained investment in model-assisted security, more effective monitoring, and continued progress in alignment research,” the company says.
In the same publication, OpenAI explains that “the monitoring burden accounts for around 20% of monitored inference compute”. However, this remains only an estimate, as the cost can vary from one situation to another.
Comments
No comments yet. Be the first to comment!
Leave a Comment