GPT-5.6 Sol, Terra, Luna: Performance, Features, and Availability

GPT-5.6 Sol, Terra, Luna: Performance, Features, and Availability

TL;DR

OpenAI releases the GPT‑5.6 family—Sol as the flagship, Terra as a balanced model, and Luna as the most cost‑efficient—delivering higher intelligence per token, a new ultra setting that runs multiple agents in parallel, and the most robust safety system to date.

Model Family Overview

GPT‑5.6 consists of three tiers: Sol, Terra, and Luna. Sol is the flagship model for high‑intensity work, Terra offers balanced performance for everyday tasks, and Luna provides the fastest, lowest‑cost option. All three share the same generation number while the tier names indicate durable capability levels that can evolve independently.

Performance and Efficiency Gains

GPT‑5.6 Sol achieves state‑of‑the‑art results across coding, knowledge work, cybersecurity, and science while using fewer tokens and at lower estimated cost than previous and competing frontier models. On Agents’ Last Exam, Sol scores 53.6, exceeding Claude Fable 5 by 13.1 points; at medium reasoning it beats Fable 5 by 11.4 points at roughly one‑quarter the estimated cost. Terra and Luna each outperform Fable 5 at around one‑sixteenth the cost. On the Artificial Analysis Intelligence Index, Sol with max reasoning comes within one point of Fable 5 while completing tasks in 61% less time at about half the estimated cost.

Coding Excellence

GPT‑5.6 Sol sets a new state‑of‑the‑art score of 80 on the Artificial Analysis Coding Agent Index, 2.8 points above Claude Fable 5, while using less than half the output tokens, taking less than half the time, and costing about one‑third less. Terra performs just above Fable 5, and Luna outperforms Claude Opus 4.8; each does so in roughly one‑third of the time, with about half as many output tokens, and at approximately one‑quarter the estimated cost. Sol also sets new state‑of‑the‑art results on Terminal‑Bench 2.1 and DeepSWE, which evaluate complex command‑line workflows and long‑horizon engineering in real codebases.

Knowledge Work and Design Judgment

GPT‑5.6 Sol reaches 92.2% on BrowseComp and 62.6% on OSWorld 2.0, surpassing Claude Opus 4.8 while using 85% fewer output tokens. The model can turn natural‑language requests into polished interactive explanations and visualizations within ChatGPT Work. It extracts messy context from Slack, Notion, Microsoft 365, and Google Drive and converts it into expert‑level, shareable artifacts. In presentations, documents, and spreadsheets, Sol creates more polished and accurate outputs, can generate fully editable presentations from scratch, and infers a deck’s design system—layouts, typography, spacing, colors, recurring content patterns—to apply those conventions faithfully to new material.

Cybersecurity Capabilities and Safety

On ExploitBench, GPT‑5.6 Sol scores 73.5% versus GPT‑5.5’s 47.9% at a comparable output‑token budget. On ExploitGym, it almost doubles GPT‑5.5’s peak pass rate, rising from 15.1% to 24.9% under a two‑hour cap and reaching 33.7% with six hours. On SEC‑Bench Pro, it scores 71.2% versus GPT‑5.5’s 45.8% at improved latency. The model supports defensive tasks such as secure code review, patching, threat modeling, and blue teaming, with verified access via OpenAI Daybreak’s Trusted Access program. GPT‑5.6’s safeguards are the most robust to date, combining protections trained into the model with real‑time checks, continuous monitoring, and account‑level enforcement, augmented by a reasoning monitor that reviews conversation for potential harm. Compared with previous models, Sol’s cyber safeguards block roughly ten times more potentially harmful activity, while options to retry prompts on lower‑capability models mitigate friction for benign use.

Ultra and Multi‑Agent Workflows

The ultra setting coordinates four agents in parallel by default, trading higher token use for stronger results and faster time‑to‑result on demanding tasks. Charts show that adding parallel agents shifts the score‑latency frontier upward and to the left on BrowseComp, SEC‑Bench Pro, and Terminal‑Bench 2.1, with 16‑agent configurations also demonstrated. Developers can build ultra‑like experiences using the multi‑agent beta in the Responses API. The max setting gives more reasoning time than xhigh, allowing the model to explore alternatives, run checks, and revise its approach.

Availability, Access, and Pricing

GPT‑5.6 Sol, Terra, and Luna are available starting today across ChatGPT, Codex, and the OpenAI API, with a global rollout progressing toward full availability over the next 24 hours. In ChatGPT, Plus, Pro, Business, and Enterprise users access Sol through medium and higher effort settings; Pro and Enterprise users can also select Sol Pro for the highest‑quality results. Free and Go users access Terra in ChatGPT Work and Codex; Plus, Pro, Business, and Enterprise users can choose among Sol, Terra, Luna and set an effort level for each. The max toggle is available to all users with access in ChatGPT Work and Codex; ultra is available to Pro and Enterprise users in ChatGPT Work and to Plus and higher plans in Codex. API developers can access all three tiers via the OpenAI API; Programmatic Tool Calling in the Responses API lets the model write and run in‑memory programs that coordinate tools, and the multi‑agent beta enables concurrent subagents. Pricing per 1M tokens is: Sol $5 input / $30 output; Terra $2.50 input / $15 output; Luna $1 input / $6 output. Prompt caching is more predictable, with explicit cache breakpoints and a 30‑minute minimum cache life; cache writes are billed at 1.25x the uncached input rate, while cache reads receive a 90% cached‑input discount.

Conclusion

GPT‑5.6 delivers a family of models that raise the performance‑per‑dollar frontier, introduce ultra‑scale multi‑agent reasoning, improve coding, knowledge work, design judgment, and cybersecurity defenses, and pair these advances with the most extensive safety system OpenAI has deployed to date.

Sources