OpenAI prepares Astra with restricted access to advanced cyber capabilities
OpenAI published new details about Astra on September 1, 2026. The company says the model reached its Critical cybersecurity threshold, the highest level in OpenAI's Preparedness Framework for cyber capability.
The threshold covers models able to find unknown flaws, build working zero-day exploits, or plan end-to-end attacks against hardened targets without step-by-step human guidance. OpenAI says Astra is the first model placed at this level.
The test results

OpenAI used public benchmarks, private benchmarks, and expert assessments. Astra scored 100% on ExploitBench, a benchmark focused on exploit development from known vulnerabilities.
OpenAI then built an internal benchmark from 20 high-severity V8 vulnerabilities disclosed between June and August 2026. Astra reached higher arbitrary code-execution rates than GPT-5.6 Sol with fewer output tokens. During the evaluation, Astra found and used two zero-day vulnerabilities in an exploit chain. OpenAI says disclosure to maintainers is in progress.
Expert testers also ran Astra against a hardened browser and a hardened operating system. In the browser test, Astra built a compromise chain, escaped the sandbox, and executed commands on the host after the browser opened an HTML file. In the operating system test, Astra combined several flaws into a local privilege-escalation chain from an unprivileged user to root.
Release controls
OpenAI says Astra was not involved in the earlier Hugging Face incident, where agents running cyber evaluations compromised third-party systems. OpenAI still used lessons from the incident to harden training and release controls.
The company paused some frontier training for two weeks after the incident. On August 28, 2026, OpenAI restarted a large reinforcement-learning run after adding new training-environment requirements.
Astra adds stronger refusal behavior, system-level misuse checks, higher-risk account limits, and monitoring for unauthorized model actions. On OpenAI's cyber jailbreak evaluations, Astra refused 91.5% of disallowed cyber requests. GPT-5.6 Sol refused 59% on the same set.
Effect for users
OpenAI plans to release Astra soon, but advanced cybersecurity workflows will start with a small group of testers. Daybreak Blue access will expand defensive use later.
For ChatGPT and Codex users, extra checks might pause or stop some long-running tasks. When a monitor flags an action, the user may need to review it before work continues. API tasks will stop instead of waiting for review.