OpenAI GPT‑6 Sol and Luna launch: 50% price cut and performance overview
TL;DR
OpenAI released GPT‑6 Sol and GPT‑6 Luna, each priced 50 % lower than their GPT‑5.6 predecessors, while offering measurable gains in professional‑task performance, factual accuracy, coding ability, and computer‑use efficiency.
Pricing breakthrough
| Model | Input price (per 1 M tokens) | Output price (per 1 M tokens) | Relative reduction |
|---|---|---|---|
| GPT‑6 Sol | $2 (down from $4) | $10 (down from $20) | 50 % |
| GPT‑6 Luna | $0.10 (down from $0.20) | $0.50 (down from $1.20) | 50 % |
OpenAI states the cuts are passed directly to API users and to ChatGPT subscription tiers. The cheaper rates are intended to make advanced AI practical for everyday automation and large‑scale applications.
Professional‑task performance
- AutomationBench (47‑tool workflow benchmark):
- GPT‑6 Sol (extra‑high effort) outperforms Claude Opus 5 (max effort) while costing only 9 % of Opus 5 per task.
- GPT‑6 Luna (high effort) improves its predecessor by 5.4 pp and reduces cost by 58 %.
- Agents’ Last Exam (complex workflow evaluation): GPT‑6 Sol (max effort) scores 56.4 %, delivering the same quality as Claude Opus 5 at 60 % lower cost.
These results show Sol and Luna dominate the cost–intelligence frontier for business‑process automation.
Factuality gains
OpenAI’s internal factuality test—based on real‑world conversations where users flagged errors—reports that:
- GPT‑6 Sol makes about half as many factual mistakes as GPT‑5.6 Sol, approaching Astra‑level reliability.
- GPT‑6 Luna matches GPT‑5.6 Sol’s factuality at ≈1 % of the cost when run at higher reasoning effort.
The evaluation notes that the test set is biased toward error‑prone queries, so absolute error rates in typical usage are lower.
Coding benchmarks
- FrontierCode (code‑change readiness): GPT‑6 Sol surpasses GPT‑5.6 Sol and matches Claude Fable 5.1 (extra‑high effort) at a fraction of the price.
- DeepSWE v1.1 (software‑engineering tasks): GPT‑6 Sol (max effort) scores 68.8 %, within 1.1 pp of Claude Fable 5’s best score, while costing ≈80 % less per task.
- GPT‑6 Luna (max effort) scores 66.6 %, comparable to Claude Opus 5, and is 93 % cheaper per task.
Community comments note mixed experiences: some users see Luna as a “free‑for‑all” pair‑programmer, while others report regressions in Luna’s coding scores compared with the 5.6 version.
Computer‑use efficiency
- OSWorld 2.0 offline (long‑horizon computer‑use workflows): GPT‑6 Sol (extra‑high) matches Claude Opus 5 (medium) with a ≈80 % cost reduction. GPT‑6 Luna (max) exceeds GPT‑5.6 Sol (medium) at one‑tenth the cost.
Alignment and communication style
OpenAI claims Sol and Luna inherit the alignment improvements of Astra, showing lower rates of misleading coding claims. Qualitative examples highlight a more concise, jargon‑free tone and fewer odd phrasing artifacts.
Prompt‑caching enhancements
- New Prompt Caching Dashboard and diagnostics let developers see cache hit rates and missed opportunities.
- Adjusting reasoning effort or tool availability now preserves cached prefixes, enabling up to 90 % discount on cached input‑token reads.
- GitHub reports a >50 % reduction in fresh token processing across billions of requests, accelerating Copilot responses.
Availability
- ChatGPT Work and Codex: Sol and Luna are live for Plus, Pro, Business, Enterprise, and Edu users.
- Free/Go users can access Luna via the desktop app.
- API identifiers:
gpt-6-solandgpt-6-luna. - Gradual rollout in ChatGPT; models may appear later in the day.
Community takeaways (selected HN comments)
- Pricing impact – Many commenters (e.g., @simonw, @Cu3PO42) emphasize the “insane” $0.10/M‑input price for Luna as a game‑changer.
- Model choice – Users debate when to prefer Sol vs. Luna vs. Astra; some note that higher reasoning levels can compensate for lower base model size.
- Performance variance – Several users report Luna’s coding scores slightly down from 5.6, while Sol shows strong gains at extra‑high effort but can be slower due to increased reasoning steps.
- Cache tooling – Developers like @apitman praise the new caching dashboard after fixing a broken proxy.
- Competitive landscape – Commenters compare OpenAI’s price war to Anthropic’s Opus 5.5, highlighting OpenAI’s transparent benchmark tables.
Bottom line
OpenAI’s GPT‑6 Sol and Luna expand the GPT‑6 family with half‑price models that retain most of Astra’s intelligence. Benchmarks demonstrate superior cost‑efficiency across automation, factuality, coding, and computer‑use tasks, while new caching tools further lower operational expenses. The launch positions OpenAI competitively against Anthropic and other frontier models, especially for high‑volume developer workflows.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch