OpenAI is introducing GPT-6 Astra, described as the world's most intelligent and aligned model. It represents the culmination of years of research and major investments across pre-training, reinforcement learning, and alignment.
Benchmark Performance
GPT-6 Astra achieves state-of-the-art results across computer use, browsing, software engineering, cybersecurity, science, and professional work. It saturates FrontierMath Tier 4 with a 98% score and has already helped solve long-standing open problems in mathematics. It also saturates ARC-AGI-3 with a 99.9% score and ExploitBench with a perfect 100% score. The model sets a new frontier on computer and browser use, handling demanding professional work with unmatched speed, accuracy, and judgment.
On Terminal-Bench Science 0.1, which tests scientific research workflows, GPT-6 Astra reaches 64.6% versus 52.6% for Claude Fable 5.1 at roughly 31% lower estimated API cost. On Terminal-Bench 4.0 for complex terminal-based tasks, it achieves 57.9% compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1.
Computer Use
GPT-6 Astra marks a new frontier in the speed, accuracy, and safety of computer use. It can handle tedious tasks such as filling out online forms, updating customer records in a CRM, and organizing calendars. It conducts online research, drafts summaries, analyzes scientific data, generates plots, creates websites, runs frontend QA checks, and helps autonomously install and test software.
On Agents' Last Exam, which tests complex professional tasks in real software, Astra scores 59.3% compared to 55.5% for Claude Opus 5 and 53.6% for GPT-5.6 Sol, while using approximately 65% fewer output tokens than Opus 5. In latency simulations on OSWorld 2.0, Astra achieves 72.6% in about 40 minutes per task, compared to GPT-5.6 Sol's 65.7% at roughly 75 minutes - a 47% reduction in time.
Alongside Astra, OpenAI is also updating the Codex harness to significantly improve computer use speed. Combined with Astra's efficiency, this results in 1.9x faster task completion compared to GPT-5.6 Sol on the Mind2Web benchmark.
Professional Work
GPT-6 Astra pairs computer use advances with targeted training for professional environments. It combines intelligence for complex problems with the ability to carry out multistep workflows and produce polished documents, spreadsheets, and presentations.
On BenchCAD, which tests 3D object reconstruction from multi-view renders via CAD code, Astra achieves a 95.9% geometric-overlap score versus 83.3% for GPT-5.6 Sol and 84.3% for Claude Fable 5.1, at approximately 43% and 86% lower estimated API cost respectively.
Astra is OpenAI's best model for adhering to existing templates and producing well-laid-out slides that convey key points with a structured narrative. It creates clear, well-structured documents, presentations, spreadsheets, and analyses that follow user templates and match writing and visual style. The model is also trained to pull only relevant context into outputs rather than repeating unnecessary information.
Astra also brings stronger visual judgment to websites, games, applications, and renderings it builds. With Sites in ChatGPT, it can create, host, and share websites, web apps, and games directly from a prompt.
When instructions leave room for interpretation, Astra is better than previous models at making the right call. It uses context to fill in routine gaps and asks focused questions when the answer could change the outcome. In Codex, it can ask asynchronously while continuing work that doesn't depend on the user's reply.
Coding
GPT-6 Astra is OpenAI's best model for software engineering to date. On Terminal-Bench 4.0, it reaches 57.9% compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1, at approximately 9% and 63% lower estimated API cost respectively.
OpenAI is also introducing a new way for Codex to preserve and retrieve context when the context window fills. Previously, models used compaction to summarize work during long sessions, which could lose important details. With Astra, Codex can keep notes across context windows, preserving accumulated details without repeatedly compressing them. Earlier context windows remain searchable, so Astra can find requirements or test results from previous messages and tool outputs.
Scientific Discovery
GPT-6 Astra is a major advance for scientific discovery, mathematics, and health. OpenAI is sharing two further results on gaps between prime numbers: Astra helped establish that infinitely many pairs of primes are at most 186 apart (improving the previous bound of 240), and it improved a term in a bound on unusually large gaps between primes that had remained unchanged for over 80 years.
On GPQA Diamond, which tests graduate-level scientific reasoning, Astra reaches 96.0%. It sets new records across math and science evaluations and can work directly in specialized software to inspect data and explore results.
Cybersecurity
Astra represents a significant jump in cyber capabilities and meets the Critical threshold in cybersecurity under OpenAI's Preparedness Framework. On ExploitBench, Astra achieved a perfect 100% compared to 78.5% for GPT-5.6 Sol. On ExploitGym, it reached 42.4% versus 30.3% for Sol.
OpenAI also built an internal "ExploitBench (June–August 2026)" evaluation using vulnerabilities from the previous three months. Astra achieved substantially higher arbitrary code-execution rates than Sol, and during evaluation it discovered two previously unknown zero-day vulnerabilities, which OpenAI is disclosing to maintainers.
On SRE-Bench for reverse engineering software binaries, Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared to 55.9% and 68.7% for Sol.
The launched version of Astra will refuse advanced cybersecurity tasks like creating proof-of-concept exploits. Through OpenAI Daybreak, expanded access with less restrictive safeguards is planned for the coming weeks.
Alignment and Safety
Astra is OpenAI's most aligned model. It excels at exercising care, respecting task boundaries, and communicating transparently. In an evaluation informed by the Hugging Face incident testing whether a model facing a difficult or impossible task will go beyond its intended scope, Astra did this in 0% of cases compared to 48% for GPT-5.6 Sol without production safeguards.
In computer use safety benchmarks, Astra causes fewer misaligned outcomes than any other frontier model tested. It never attempted to circumvent a Codex Auto-Review denial, even when Auto-review was deliberately configured to be evadable. Astra is also three times less likely than Sol to make inaccurate representations about its capabilities.
OpenAI's evaluations found Astra's written reasoning harder to monitor than Sol's, attributed to Astra's greater control over written reasoning and ability to solve problems with fewer written steps. Improving monitorability remains a research priority.
OpenAI is deploying misalignment monitoring in production for Astra-class models, involving classifiers that check reasoning and actions for unauthorized behavior and automatically stop potentially unauthorized activity.
Availability and Pricing
GPT-6 Astra is rolling out initially to a limited set of organizations and will become available over the coming days to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock. Usage is included within existing subscription allowances, with additional credits available for purchase. Pro, Business, and Enterprise users also get access to GPT-6 Astra Pro.
For developers, it is available as gpt-6-astra. Standard API pricing is $10 per million input tokens and $50 per million output tokens, with separate rates for cache reads and writes. A Fast mode delivers up to 2x speed at 2x the Standard price.
Astra supports Zero Data Retention for eligible API customers, and OpenAI is testing Private Safety Processing to strengthen safety monitoring while preserving customer privacy.