GPT-5.6 Price and Performance Updates

OpenAI has reduced pricing for GPT-5.6 Luna and Terra and introduced a "Fast mode" for GPT-5.6 Sol, enabling developers to scale high-volume AI workloads more cost-effectively while accelerating high-priority tasks.

API Pricing and Performance Updates

OpenAI has implemented significant price reductions for the GPT-5.6 model family to lower the cost of high-volume AI applications.

  • GPT-5.6 Luna: The most affordable and fastest model in the family, now costs 80% less. API pricing is now $0.20 per million input tokens and $1.20 per million output tokens.
  • GPT-5.6 Terra: The balanced model for everyday work, now costs 20% less. API pricing is now $2 per million input tokens and $12 per million output tokens.
  • GPT-5.6 Sol: Pricing remains unchanged, but a new Fast mode replaces Priority Processing. Fast mode delivers up to 2.5× faster speeds than Standard processing at twice the price, with no change in intelligence.

These pricing changes are also reflected in how usage is counted against paid subscriptions for Codex and ChatGPT Work.

Model Selection and Workflow Optimization

Businesses can optimize AI costs by matching the specific intelligence requirements of a task to the model version.

  • Luna's Efficiency: Luna delivers performance comparable to frontier-class models from one year prior at approximately 6 cents on the dollar per task and nearly nine times the speed. On professional work (measured by Agents’ Last Exam), Luna outperforms Fable 5 at an estimated cost per task nearly 99% lower.
  • Hybrid Workflows: Organizations can use a tiered approach to intelligence. For example, a coding workflow may utilize GPT-5.6 Sol to define a plan and resolve uncertainty, then transition to GPT-5.6 Luna to implement changes, run tests, and evaluate results.

Technical Drivers of Efficiency

The price-performance gains in GPT-5.6 are driven by improvements across the model architecture, inference systems, and the agentic harness.

  • Infrastructure Improvements: Efficiency is increased through better routing to keep hardware productive, optimized production software for token generation, and smarter context management to prevent agents from repeating work.
  • AI-Driven Optimization: GPT-5.6 Sol was used within a human-led process to autonomously rewrite and optimize production kernels, which reduced the end-to-end cost of serving the model by 20%. Additionally, Sol-led experiments increased token-generation efficiency by more than 15%.

Availability

GPT-5.6 Terra and Luna are available in ChatGPT Work, Codex, and the OpenAI API. In ChatGPT Work and Codex, Free and Go users have access to Terra, while Plus, Pro, Business, and Enterprise users can access both Terra and Luna. Pricing updates are rolling out to AWS starting July 30, 2026.

Sources