OpenAI announced that it shared how GPT-5.6 contributed to making itself more efficient to run. Now, the company is passing those efficiency gains on to customers through lower prices for GPT-5.6 Luna and Terra, along with faster performance for GPT-5.6 Sol in the API. Together, these updates help customers extract more value from every dollar spent on AI and move faster when time is critical.
As of July 30, GPT-5.6 Luna, OpenAI's fastest and most affordable model, sees an 80% price reduction, while GPT-5.6 Terra, the balanced model for everyday tasks, drops by 20%. These lower prices for Luna and Terra are also reflected in how usage counts against paid subscriptions when using Codex and ChatGPT Work. Luna offers businesses a far more cost-effective way to handle high-volume work at very high quality levels. It can use tools and complete multi-step workflows, making a broader range of AI applications practical to run at scale.
Bringing advanced intelligence to more people at lower cost is central to OpenAI's mission of ensuring AGI benefits all of humanity. These changes put that commitment into practice, reflecting years of improvements in how OpenAI's models are built, served, and deployed.
OpenAI is also introducing Fast mode in the API, replacing the previous Priority Processing offering. For GPT-5.6 Sol, Fast mode delivers up to 2.5× faster speeds than Standard processing at twice the price, with no change in intelligence. Fast mode is backward compatible: requests tagged as priority will automatically use Fast mode.
Matching Intelligence to the Outcome
Using AI efficiently starts with the desired outcome. The stakes, cost of error, urgency, and scale determine the right balance of intelligence, speed, reliability, and cost. That balance can shift from one step of a workflow to the next.
GPT-5.6 gives businesses much more room to optimize that equation. Luna delivers performance comparable to models that were frontier-class a year ago at roughly 6 cents on the dollar per task, and at nearly nine times the speed. On professional work, as measured by Agents' Last Exam, Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower.
In practice, businesses can define the outcome and quality standard they need, then use evaluations to determine where additional intelligence materially improves the result and where faster, lower-cost processing can deliver the same quality. A coding workflow, for example, might use Sol to resolve uncertainty and define the plan, then use Luna to implement well-specified changes, write and run tests, and evaluate the results.
The GPT-5.6 family expands the range of those choices. Businesses can apply the maximum useful intelligence at every stage while paying the right price for the value it creates.
How OpenAI Advances the Efficiency Frontier
OpenAI's efficiency edge comes from improving the models, the inference systems that run them, and the agentic harness that connects them to tools and context. GPT-5.6 models take a more direct path through work. Better routing keeps hardware productive, optimized production software generates tokens more efficiently, and smarter context management helps agents avoid repeating completed work. Together, these improvements let OpenAI complete more useful work with the same compute, reducing time, tokens, and cost per result.
GPT-5.6 Sol is increasingly helping OpenAI find and deliver the next round of gains. Within a human-led process, Sol autonomously rewrote and optimized production kernels, designed and ran hundreds of experiments to improve token generation, and monitored training while intervening when problems arose. The kernel work helped reduce the end-to-end cost of serving the model by 20%, while its experiments increased token-generation efficiency by more than 15%. This work continues, creating a tighter feedback loop: as OpenAI's models improve and are able to work more autonomously, the company's ability to improve efficiencies accelerates. Read more about the engineering behind GPT-5.6.
A Compute Strategy Built for Scale
Meeting demand for abundant intelligence requires both more compute and more productive compute. OpenAI is building a resilient infrastructure portfolio and matching each workload to the systems best suited to run it. That approach supports both ends of the price-performance curve. At the lower-cost end, the new Luna and Terra prices make high-volume work economical at much greater scale. At the frontier end, Fast mode gives API customers faster access to Sol when response time matters.
Enterprises can move more AI into everyday operations without sacrificing speed on their most consequential work. Large-scale document analysis, customer-interaction classification, and routine implementation can become economical to run broadly, while complex Sol workloads can move faster when the premium is justified.
The gains can compound. Within a human-led process, more capable models help OpenAI's technical team find the next generation of improvements, shortening the path to better performance and lower costs. OpenAI's strategy remains focused on advancing both capability and efficiency so each generation of intelligence can accomplish more work at a lower cost.
Availability and Pricing
GPT-5.6 Terra and Luna remain available in ChatGPT Work, Codex, and the OpenAI API. In ChatGPT Work and Codex, Free and Go users can access Terra, while Plus, Pro, Business, and Enterprise users can choose both Terra and Luna.
Starting July 30, API pricing is $2 per million input tokens and $12 per million output tokens for Terra, and $0.20 per million input tokens and $1.20 per million output tokens for Luna. Sol pricing remains unchanged. ChatGPT and Codex subscription prices and quota budgets remain unchanged, while Terra and Luna usage now consumes fewer credits. Pricing changes will begin rolling out in AWS later that day.
Fast mode for GPT-5.6 Sol replaces Priority Processing in the API and aligns with /fast in Codex. Existing API requests tagged priority will continue to work. View complete API pricing details.