OpenAI is releasing the GPT-5.6 family of models for general availability following a limited preview. The lineup includes Sol, the new flagship model, Terra, a balanced model suited for everyday work, and Luna, the most cost-efficient option.
Efficient by Default, Maximum Performance on Demand
GPT-5.6 Sol establishes a new benchmark for both intelligence and efficiency, achieving state-of-the-art results across coding, knowledge work, cybersecurity, and science while outperforming previous and competing frontier models using fewer tokens at lower estimated cost. The result is stronger performance per dollar-more successful work for the same spend, or comparable results at reduced total cost.
OpenAI also introduces a new way to accelerate the most demanding work: ultra is the highest-capability setting, coordinating multiple agents across parallel workstreams to complete complex tasks faster. Enhanced computer use and design judgment make GPT-5.6 Sol OpenAI's most polished collaborator yet, able to inspect, refine, and deliver ready-to-use results.
GPT-5.6 was trained to extract more useful work from every token. On Agents' Last Exam, an evaluation of long-running professional workflows across 55 fields, GPT-5.6 Sol achieves a new high of 53.6, surpassing Claude Fable 5 (adaptive reasoning) by 13.1 points. Even at medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. GPT-5.6 Terra and GPT-5.6 Luna outperform Fable 5 at around one-sixteenth the cost. On the Artificial Analysis Intelligence Index, GPT-5.6 Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at roughly half the estimated cost.
GPT-5.6 launches with OpenAI's most robust safeguards to date, designed to resist determined and adaptive misuse without broadly limiting legitimate work. Before general availability, the models and safeguards underwent OpenAI's most extensive evaluation period yet, combining human red teaming with large-scale automated testing.
Coding Performance
GPT-5.6 Sol is OpenAI's best coding model yet. On the Artificial Analysis Coding Agent Index, GPT-5.6 Sol with max reasoning sets a new state of the art at 80, 2.8 points above Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one-third less. Terra performs just above Fable 5, while Luna outperforms Opus 4.8-each in roughly one-third the time, with about half as many output tokens, at approximately one-quarter the estimated cost. GPT-5.6 Sol also sets new state-of-the-art results on Terminal-Bench 2.1 and DeepSWE.
GPT-5.6 can write and run lightweight programs that coordinate tools, process intermediate results, monitor progress, and choose the next action as work unfolds. This lets tool-heavy tasks advance with fewer tokens, fewer model round trips, and less guidance. Programmatic Tool Calling in the Responses API can filter large amounts of intermediate data, retain only what matters, and adapt its workflow along the way.
For problems that reward greater investment of time and compute, GPT-5.6 can push beyond the efficient default. max gives the model even more time than xhigh to reason and explore alternatives, run checks, and revise its approach. ultra goes further by coordinating four agents in parallel by default, trading higher token use for stronger results and faster time-to-result on demanding tasks.
A Leap Forward in Design
GPT-5.6 delivers a step change in design judgment. With only high-level direction, it creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result-not just generate the underlying code or content-catching visual and functional issues and applying finishing touches before handing the work back.
GPT-5.6's frontend capabilities also turn natural-language requests into polished, interactive explanations and visualizations within ChatGPT Work.
End-to-End Knowledge Work
GPT-5.6 delivers improved results for professional tasks. It takes messy context from documents and everyday workflows like Slack, Notion, Microsoft 365, and Google Drive, and converts it into expert-level, shareable artifacts.
GPT-5.6 Sol sets new state-of-the-art results on BrowseComp at 92.2% and OSWorld 2.0 at 62.6%; on OSWorld, it surpasses Opus 4.8 while using 85% fewer output tokens. Luna nearly matches GPT-5.5's peak performance at less than half the estimated cost, while Terra surpasses it at lower cost.
GPT-5.6 Sol improves quality in presentations, documents, and spreadsheets, producing outputs that are more polished and accurate. It can create fully editable presentations from scratch, translating a prompt and source material into a coherent visual narrative with strong layouts, hierarchy, and design. The improvement is especially pronounced when following templates and reference decks-GPT-5.6 can infer a deck's design system (layouts, typography, spacing, colors, and recurring content patterns, including rules embedded in the Slide Master) and apply those conventions consistently to new material.
GPT-5.6 also creates more visually refined documents and spreadsheets, follows complex reference formats more faithfully, handles equations and financial models with greater precision, and makes better use of typography, spacing, hierarchy, and page or worksheet layout.
Pushing the Frontier on Cyber and Science
Cybersecurity
GPT-5.6 is OpenAI's strongest cybersecurity model yet, achieving frontier performance with significantly fewer tokens:
- On ExploitBench, it scores 73.5% versus GPT-5.5's 47.9% at a comparable output-token budget.
- On ExploitGym, it almost doubles GPT-5.5's peak pass rate, from 15.1% to 24.9% under the two-hour cap; with six hours, it reaches 33.7%.
- On SEC-Bench Pro, it scores 71.2% versus GPT-5.5's 45.8% at improved latency.
GPT-5.6 supports important defensive tasks such as secure code review, patching, threat modeling, and blue teaming. Qualified individuals and organizations in OpenAI Daybreak's Trusted Access for Cyber program can access more of its defensive capability through more precise safeguards for verified work in authorized environments.
Science
GPT-5.6 Sol shows broad gains across scientific research. On life sciences evaluations, GPT-5.6 demonstrates Pareto improvements over GPT-5.5 on real-world biology, life science research workflows, and chemistry.
GPT-5.6 Accelerates OpenAI
GPT-5.6 is OpenAI's strongest model yet for accelerating AI research. Inside OpenAI, researchers use it across the development loop: diagnosing failures, optimizing training systems, running experiments, and interpreting results. Average daily output tokens per active researcher were more than twice the highest level observed for GPT-5.5 during the internal testing period.
Over the past six months, the share of research compute devoted to internal coding inference grew 100-fold, while internal agentic token usage increased approximately 22-fold.
On a bundle of evaluations measuring progress towards recursive self-improvement, GPT-5.6 Sol represents a 16.2-point improvement over GPT-5.5, accelerating internal research across the board.
Scaling Safety and Security with Capability
As model capabilities increase, OpenAI strengthens its safety stack so advanced intelligence can remain broadly useful while applying greater scrutiny to the highest-risk uses. For GPT-5.6, OpenAI built its most robust safety system to date, calibrated to each model's capabilities and powered by more compute than ever before.
The GPT-5.6 models are more capable than earlier models in both biology and cybersecurity but do not cross the Critical threshold in either category. In cybersecurity, testing suggests GPT-5.6 is better at finding and fixing vulnerabilities than at reliably carrying out autonomous, end-to-end attacks against hardened targets. In biology, testing suggests GPT-5.6 can support legitimate research but does not provide the end-to-end capability needed to create, engineer, or synthesize a highly dangerous novel threat.
GPT-5.6's safeguards are layered for greater accuracy and redundancy, designed to adapt quickly as new attacks emerge. Protections trained into the model work alongside real-time checks, continuous monitoring, and account-level enforcement. OpenAI's approach adds a reasoning monitor that reviews the conversation to determine if there is potential for harm. Compared with previous models, GPT-5.6 Sol cyber safeguards block roughly ten times more potentially harmful activity.
Before general availability, OpenAI ran its most intensive safety evaluations to date, including extensive red teaming, robust capability and safeguard testing with external experts, and approximately 700,000 A100e GPU hours of black-box automated red teaming.
Availability and Pricing
GPT-5.6 spans three model tiers: Sol (flagship), Terra (lower-cost, competitive with GPT-5.5), and Luna (fastest and most affordable). The number identifies the generation, while Sol, Terra, and Luna are durable capability tiers that can advance on their own cadence.
GPT-5.6 is available starting July 9, 2026 across ChatGPT, Codex, and the OpenAI API, with global rollout continuing gradually over 24 hours.
- Chat: Plus, Pro, Business, and Enterprise users access GPT-5.6 Sol through medium and higher effort settings. Pro and Enterprise users can also select GPT-5.6 Sol Pro for the highest-quality results.
- ChatGPT Work and Codex: Free and Go users access GPT-5.6 Terra. Plus, Pro, Business, and Enterprise users can choose among Sol, Terra, and Luna with adjustable effort levels.
maxis available to all users with GPT-5.6 access.ultrais available to Pro and Enterprise users in ChatGPT Work, and to Plus and higher plans in Codex. - API: Developers can access Sol, Terra, and Luna through the OpenAI API. Programmatic Tool Calling in the Responses API is Zero Data Retention (ZDR) compatible. Multi-agent support is initially available in beta.
Pricing (per 1M tokens):
| Model | Input | Output |
|---|---|---|
| Sol | $5 | $30 |
| Terra | $2.50 | $15 |
| Luna | $1 | $6 |
GPT-5.6 also introduces more predictable prompt caching, including support for explicit cache breakpoints and a 30-minute minimum cache life. Cache writes are billed at 1.25x the model's uncached input rate, while cache reads continue to receive the 90% cached-input discount.