OpenAI Prepares to Launch Astra With Enhanced Cybersecurity Safeguards

Open AI

OpenAI is preparing to release Astra, its next frontier artificial intelligence model, after delaying parts of its development to strengthen protections against cybersecurity misuse and unauthorised actions.

The company said Astra would become available soon, although access to its most advanced cybersecurity functions will initially be limited. The phased approach reflects concerns about the model’s ability to identify previously unknown vulnerabilities and develop methods for exploiting highly protected computer systems.

OpenAI has classified Astra at the “Critical” cybersecurity capability threshold under its Preparedness Framework—the first of its models to receive that designation.

According to the company, Astra can, when equipped with appropriate tools and access, discover unknown security weaknesses and create working exploits across protected systems without requiring human guidance at every stage.

Internal evaluations showed that the model could build complex exploit chains. In one assessment, Astra discovered and used two previously unknown vulnerabilities, which OpenAI said it was working to disclose responsibly to the relevant maintainers.

The company also reported that the model identified weaknesses in a hardened browser and operating system, combining them into working attack chains during controlled testing.

These capabilities prompted OpenAI to delay portions of Astra’s development and deployment while introducing stricter monitoring, network isolation and security controls.

Some training and evaluation activities were paused after a separate security incident involving OpenAI and Hugging Face. OpenAI clarified that Astra was not involved in that incident but said lessons from it informed the safeguards developed for the new model.

The company has trained Astra to refuse harmful cybersecurity requests more consistently and remain within the scope authorized by users and system policies.

OpenAI said Astra refused 91.5% of prohibited requests in its cyber-jailbreak evaluations, compared with 59% for GPT-5.6 Sol. It has also introduced additional monitoring capable of identifying and stopping potentially unauthorised model activity.

Despite the restrictions, OpenAI intends to make Astra broadly available for general use. Its most powerful cybersecurity tools will initially be offered to a small group of testers before access expands through the company’s Daybreak Blue program for defensive security work.

The additional safeguards may occasionally slow, pause or stop legitimate tasks if the system identifies them as potentially harmful. Users of ChatGPT or Codex may be asked to review certain actions before continuing, while flagged tasks conducted through the API may be stopped automatically.

OpenAI described Astra as a significant advance in alignment, saying its evaluations showed that the model was more likely than GPT-5.6 Sol to respect security restrictions and remain within its authorized scope.

Further information about Astra’s capabilities, limitations and safety evaluations will be published in the model’s system card when it launches. OpenAI has not yet announced a specific release date. Official OpenAI announcement

 

 

 

Source: Omanghana


About us

Omanghana is an online news portal that provides readers around the world with a greater focus on Ghana and other parts of Africa. Established in 2009, Omanghana regularly publishes articles related to News, Sports, and Entertainment.


CONTACT US