OpenAI introduces GPT-6 Sol and Luna

TL;DR

OpenAI launched GPT‑6 Sol and GPT‑6 Luna, cost‑efficient successors to GPT‑6 Astra that halve API prices compared with GPT‑5.6 and retain high performance in professional tasks, factual accuracy, coding, and computer‑use workloads.

Overview of the GPT‑6 family expansion

OpenAI introduced GPT‑6 Sol and GPT‑6 Luna to complement the flagship GPT‑6 Astra. All three models share the same training methodology, but Sol and Luna are optimized for lower latency and reduced inference cost. Improvements in caching and inference infrastructure enable a 50 % price reduction relative to the GPT‑5.6 promotional rates, making advanced AI more accessible for everyday applications.

API pricing reduction

Model Input price (per M tokens) Output price (per M tokens) Relative reduction
GPT‑6 Sol $4 → $2 $20 → $10 50 % cheaper
GPT‑6 Luna $0.20 → $0.10 $1.20 → $0.50 50 % cheaper

Prices are quoted per 1 million tokens. GPT‑6 Astra remains the top‑tier model for users who need the best possible results.

Professional‑work performance

Key result

GPT‑6 Sol outperforms Claude Opus 5 on the AutomationBench benchmark while costing only 9 % of Opus 5’s per‑task expense; GPT‑6 Luna improves over its predecessor by 5.4 percentage points at 58 % lower cost.

Detailed comparison

  • AutomationBench 1.0.6 evaluates AI agents on end‑to‑end workflows across 47 tools. GPT‑6 Sol (xhigh effort) scores 33.2 % with a cost of $0.27 per task, beating GPT‑6 Astra (low effort) and Claude Opus 5 (max effort) on both score and cost.
  • On Agents’ Last Exam, GPT‑6 Sol (max effort) achieves 56.4 %, surpassing Claude Opus 5 while delivering a 60 % cost reduction per task.

Factuality improvements

GPT‑6 Sol makes roughly half as many factual mistakes as GPT‑5.6 Sol on OpenAI’s internal factuality benchmark, approaching the reliability of Astra at a fraction of the cost. GPT‑6 Luna also shows substantial gains, matching GPT‑5.6 Sol’s factuality at about one‑hundredth of the cost when run at higher effort levels.

Coding capabilities and cost efficiency

Key result

On the FrontierCode and DeepSWE v1.1 benchmarks, GPT‑6 Sol delivers coding quality comparable to top competitor models while reducing per‑task cost by up to 80 %.

Benchmark details

  • FrontierCode: GPT‑6 Sol surpasses GPT‑5.6 Sol and matches Claude Fable 5.1 (xhigh effort) at a much lower expense.
  • DeepSWE v1.1: GPT‑6 Sol (max effort) scores 68.8 %, within 1.1 pp of Claude Fable 5’s best score, at ~80 % lower cost per task. GPT‑6 Luna (max effort) scores 66.6 %, comparable to Claude Opus 5 and Claude Fable 5 at medium effort, while costing 93 %–96 % less per task.

Computer‑use performance

GPT‑6 Sol (xhigh effort) reaches a 60.5 % score on OSWorld 2.0 offline, matching Claude Opus 5 (medium effort) with about 80 % lower cost. GPT‑6 Luna (max effort) exceeds GPT‑5.6 Sol (medium effort) while using only one‑tenth of the cost.

Enhanced collaboration style

Both Sol and Luna inherit Astra’s refined communication style: clearer explanations, reduced jargon, fewer irrelevant details, and slightly shorter answers without sacrificing substance. Example comparisons show Sol providing more measured, precise responses than its GPT‑5.6 predecessor.

Prompt‑caching advances for agents

OpenAI introduced higher default cache‑hit rates and a 90 % discount on cached input‑token reads. New developer tools include:

  • Prompt Caching Dashboard for visualizing cache usage.
  • Diagnostics tool to identify missed caching opportunities.
  • Controls for reasoning effort and tool availability that preserve cached context.
  • Explicit breakpoints to define cached prompt prefixes. GitHub reports that these changes cut the share of fresh prompt‑token processing by over 50 % across billions of requests, accelerating Copilot response times.

Alignment progress

Sol and Luna build on Astra’s alignment improvements. Internal evaluations show lower rates of misleading claims in coding tasks compared with GPT‑5.6 models. The reported metrics target deliberately challenging scenarios and do not reflect typical usage failure rates. Full results are available in the GPT‑6 Astra system card.

Availability

  • ChatGPT Work and Codex: GPT‑6 Sol and GPT‑6 Luna are live for Plus, Pro, Business, Enterprise, and Edu tiers.
  • Free and Go users: GPT‑6 Luna is accessible via the desktop app.
  • API: Models are reachable as gpt-6-sol and gpt-6-luna.
  • ChatGPT: Rollout will occur gradually throughout the day; users may need to retry if models are not yet visible.

All evaluations were performed in OpenAI’s research environment or via the API; production ChatGPT may differ due to system‑prompt and tool variations. Competitor scores are taken from publicly available reports.

Sources