OpenAI Designates Astra as First Model with Critical Cyber Capabilities
OpenAI announced that Astra is the first model to meet the Critical cybersecurity capability threshold under its Preparedness Framework — meaning it can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without step-by-step human guidance. In internal evaluations, Astra scored a perfect 100% on ExploitBench and even discovered and used two zero-day V8 vulnerabilities as part of an exploit chain, which OpenAI is now disclosing to maintainers.
To prepare for release, OpenAI delayed parts of Astra’s development while strengthening safeguards: Astra refuses 91.5% of cyber jailbreak requests (vs. 59% for GPT-5.6 Sol), triggers fewer misaligned behaviors in honeypot tests, and is described as the company’s most aligned model to date. Advanced cybersecurity access will initially be limited to a small group of alpha testers, with Daybreak Blue expansion to follow for defensive use.
Path to Astra: critical capabilities and frontier safeguards
Related: OpenAI Slows Model Development Over Cyber Capabilities, OpenAI Opens Unfiltered Models to Cybersecurity Professionals via Expanded Daybreak Program, OpenAI’s Astra Solves 10 Longstanding Math Problems, OpenAI Confirms Its Models Behind Hugging Face Security Breach