OpenAI says its upcoming Astra model can find previously unknown software flaws and turn them into working attacks without a human guiding each step, crossing a cybersecurity threshold that until recently belonged largely to expert hacking teams.
It is the first model OpenAI has classified as having “Critical” cyber capabilities under its Preparedness Framework, the firm wrote in a Tuesday post.
To qualify, a model must be able to find previously unknown software flaws, known as zero-days, and develop working exploits for them across hardened real-world systems without human intervention, or devise and execute an attack from little more than a high-level goal.
In testing, Astra scored 100% on a benchmark for developing exploits from known vulnerabilities and found two previously unknown flaws while building an exploit chain on a separate internal test.

It also broke out of a hardened browser sandbox and executed commands on the host computer, while separately finding and combining multiple flaws in an operating system to gain root access, OpenAI said.
The company has since delayed parts of Astra’s development while adding safeguards, and plans to initially restrict its most advanced cybersecurity abilities to selected testers.
That capability is particularly relevant to crypto, where a software flaw can be converted into money within minutes. CoinDesk reported in June that increasingly capable AI models could compress the work of searching code, finding misconfigurations and assembling attacks from days or weeks into machine-speed operations.
Read More: Crypto’s next billion-dollar hacker may move at superhuman speed
At the time, security researchers said the bigger change was not necessarily a new class of hack, but how quickly existing weaknesses could be found and exploited.
The advance follows other signs that frontier models are moving beyond answering questions and writing code. Anthropic’s Claude Fable 5 helped solve an 87-year-old mathematics problem in July, as CoinDesk reported.
coindesk.com