OpenAI today unveiled its latest frontier model, GPT-6 Astra, calling it “the world’s most intelligent and aligned model.” The update brings major upgrades across computer use, software engineering, browsing, professional productivity, mathematics, science, and cybersecurity. The release is rolling out gradually, starting with select organizations and eventually opening to multiple subscription tiers and the OpenAI API on AWS.
What Makes Astra Stand Out
Astra represents a significant advance from previous models. It achieves top-tier scores on several elite benchmarks—FrontierMath Tier 4, ARC-AGI-3, and ExploitBench—with scores in the high 90s or even 100%. These metrics highlight progress in math, science, and especially cybersecurity. The model is now able to handle multistep tasks, produce polished work—documents, spreadsheets, presentations—and manage complex workflows with less oversight.
One of Astra’s most critical upgrades is its improved “computer use” capability. The model can operate across a desktop environment—including Macs in the background while users work—to execute multitool workflows and complete demanding tasks with speed and precision. For Codex users, Astra introduces a new approach to long-session context: instead of summarizing when context windows fill, it preserves detailed notes and makes earlier windows searchable, even if they haven’t been compressed yet.
Cybersecurity Elevated: Critical Threshold Reached
OpenAI states Astra is the first of its models to officially hit the “Critical” level under its Preparedness Framework for cybersecurity. This marks a shift from earlier models, which had been classified at the “High” level. The Critical designation means Astra can discover unknown vulnerabilities, develop functional zero-day exploits, and devise end-to-end attack strategies against hardened systems without step-by-step human guidance, in certain configurations.
In internal evaluations, Astra scored 100% on ExploitBench, and successfully worked with fewer output tokens than previous models on a newly assembled benchmark of 20 high-severity V8 vulnerabilities, including finding and chaining zero-day flaws. In expert-led tests, it uncovered novel weaknesses in hardened browsers and operating systems, developed exploit chains that broke out of browser sandboxes, and achieved privilege escalation from unprivileged accounts to root access. That said, these advanced capabilities were demonstrated only under Daybreak Blue access—not the default production setup.
Availability, Access & Safeguards
GPT-6 Astra is being made available first to limited organizations and will be offered in the coming days to users with Plus, Pro, Business, and Enterprise subscriptions, as well as via the OpenAI API and AWS. The subscription allowances cover Astra usage, though additional credits are available for heavier usage. Enterprise administrators must enable Astra—it’s off by default.
The most powerful cybersecurity tools within Astra are initially gated behind trusted-access programs and tighter controls. OpenAI also paused some internal development and reinforcement learning runs that did not yet meet its heightened security standards. Enhanced safety measures include stricter monitoring, access restrictions, and layered defenses during development—rather than waiting until after release.
GPT-5, the prior generation released in early August 2025, brought automatic routing, fewer factual errors, and improved reasoning. Subsequent interim versions—5.1 through 5.6—focused on tone, context length, and professional task performance. Astra builds on this trajectory with far greater ambition in autonomy, context preservation, and cybersecurity.
What this means: Astra sets a new bar for what foundation models can do—but its “Critical” status also raises important governance questions. As models capable of novel exploit development emerge, access control and safety frameworks become as vital as raw capability. Organizations adopting Astra should watch how controls evolve, whether auditability and transparency keep pace, and how the trade-offs between power and risk are managed in real deployments. This is a turning point in model safety, not just model strength.