Why OpenAI slowed development of its Astra model over security concerns

OpenAI says its in-development Astra model has reached what the company calls a "critical cybersecurity threshold" — meaning, in its assessment, the model could independently identify and carry out cyberattacks against well-protected real-world systems.
The disclosure came as part of OpenAI's internal risk framework, which the company uses to track the capabilities of its own models. Under that policy, models that cross certain capability thresholds trigger a deliberate slowdown in development or additional safety measures.
The cybersecurity threshold measures how capable an AI model is at finding vulnerabilities, writing exploit code, and deploying that code against real systems without human involvement. That kind of capability can be valuable in defensive security research, but it also poses a serious risk if misused.
OpenAI says it hasn't halted Astra's development, but has deliberately slowed its pace to add extra security testing and safeguards before any release. The approach has become increasingly common among AI companies in recent months.
The move is part of a broader debate within the AI industry. As model capabilities advance rapidly, so does their potential for misuse — particularly in cybersecurity, where the line between a model used for defense and one used for attack is often blurry.
Security researchers have long noted that models with this kind of capability can be both useful and dangerous. On one hand, automated vulnerability scanning could help defensive teams strengthen their systems faster; on the other, the same capability could be devastating in the hands of a malicious actor.
OpenAI's disclosure signals that the company recognizes this risk and believes the model needs additional safeguards layered on before release. The company says it is working with outside security experts as part of preparing any such model for deployment.
The development could set a precedent for rival AI companies to adopt similar risk frameworks. Most major players in the industry have developed comparable threshold systems to track their own models' capabilities in cybersecurity, biosecurity, and other high-risk domains.
Critics note that disclosures like this can serve marketing purposes as much as transparency ones — announcing that a model is "dangerously powerful" can also function as an advertisement for its capabilities. Still, regulators and security experts hope such disclosures push the industry toward greater transparency overall.
When — and with what safeguards — Astra will eventually be released remains unclear. OpenAI says development will continue, but that any release decision will depend on the outcome of further safety evaluations.
Read next

Why Microsoft Edge is about to lock out older ad blockers
Microsoft Edge is ending support for the Manifest V2 extensions platform, the same move Google Chrome made earlier this year, which will disable the uBlock Origin ad blocker and others like it. The impact on most users will be limited, but the shift is part of a broader transformation of the browser extension ecosystem.

What is Copernicus, and how Europe's free satellite service tracks wildfires
Europe's free Copernicus satellite program has added wildfire visualization to its browser-based tool. The addition arrives during a record wildfire season and lets anyone track the spread of fires in near real time.

Half a million black holes: how astronomers built one of the largest sky maps ever made
An astronomy consortium has released a massive all-sky map cataloging more than half a million supermassive black holes. The project gives scientists studying the distribution of these mysterious objects at the centers of distant galaxies a vast new dataset to work with.

Voyager 2: how NASA keeps its 48-year-old probe running for one more year
Voyager 2, the spacecraft launched in 1977, will keep transmitting data for at least another year thanks to NASA power-saving measures. It's the latest step in extending the life of one of the only missions still sending signals back from the outer reaches of the solar system.

AI music watermarks: how Suno plans to label songs made by algorithms
According to Ars Technica, AI music platform Suno plans to add watermarks to the songs it generates and impose download limits. The company says the move is intended to curb 'large-scale abuse' of its platform.