GPT-6 Astra release notes / what's new
OpenAI has introduced GPT-6 Astra, a next-generation model designed for high-intelligence professional work, autonomous computer use, and scientific discovery. The model demonstrates state-of-the-art performance across software engineering, cybersecurity, and mathematics, while introducing a new standard for model alignment and intent adherence.
State-of-the-Art Computer Use and Professional Work
GPT-6 Astra establishes a new frontier in the speed and accuracy of autonomous computer use, enabling it to handle complex professional workflows such as filling online forms, updating CRMs, and performing frontend QA checks.
Performance Benchmarks
- Agents' Last Exam: Astra scored 59.3%, surpassing Claude Opus 5 (55.5%) and GPT-5.6 Sol (53.6%). It used approximately 65% fewer output tokens than Opus 5 at these settings.
- OSWorld 2.0: In latency simulations, Astra achieved a 72.6% score in roughly 40 minutes per task, compared to GPT-5.6 Sol's 65.7% in 75 minutes—a reduction in task time of approximately 47%.
- Mind2Web: Combined with an updated Codex harness, Astra completes tasks 1.9x faster than GPT-5.6 Sol.
- BenchCAD: Astra achieved a 95.9% geometric-overlap score for reconstructing 3D objects from multi-view renders, significantly higher than GPT-5.6 Sol (83.3%) and Claude Fable 5.1 (84.3%).
Professional Capabilities
Beyond benchmarks, Astra is optimized for business contexts, including the ability to adhere to existing templates for slide decks, spreadsheets, and documents. It is specifically trained to extract only relevant context for outputs to avoid unnecessary repetition. Through "Sites" in ChatGPT, Astra can also create, host, and share websites, web apps, and games directly from prompts.
Software Engineering and Coding
GPT-6 Astra is positioned as the most capable model for software engineering to date, showing significant gains in terminal-based tasks.
- Terminal-Bench 4.0: Astra reached 57.9%, compared to 37.3% for GPT-5.6 Sol and 55.8% for Claude Fable 5.1.
- Context Management: Astra introduces a new method for Codex to preserve and retrieve context when the window fills. Instead of simple compaction (summarization), Astra can keep searchable notes across context windows, preserving specific details about failed fixes or component behaviors without repeated compression.
Scientific Discovery and Mathematics
GPT-6 Astra demonstrates advanced reasoning capabilities that have contributed to solving open problems in mathematics.
- Prime Number Gaps: Astra helped establish a stronger bound of 186 for infinitely many pairs of primes (improving the previous bound of 240) and improved a term in a bound on large prime gaps that had been unchanged for over 80 years.
- Academic Benchmarks:
- FrontierMath Tier 4: 98% score (saturating the benchmark).
- ARC-AGI-3: 99.9% score (saturating the benchmark).
- GPQA Diamond: 96.0% score, a new high for graduate-level scientific reasoning.
- Terminal-Bench Science 0.1: 64.6% score, compared to 52.6% for Claude Fable 5.1.
Cybersecurity Capabilities and Risks
GPT-6 Astra meets the "Critical threshold" in cybersecurity under OpenAI's Preparedness Framework due to its ability to identify and develop zero-day exploits.
Offensive and Defensive Benchmarks
- ExploitBench: Astra achieved a perfect 100% score, compared to 78.5% for GPT-5.6 Sol.
- ExploitGym: Astra reached a 42.4% success rate, compared to 30.3% for GPT-5.6 Sol.
- SRE-Bench: Astra solved 88.0% of tasks in a single attempt (compared to 55.9% for GPT-5.6 Sol), demonstrating a high capacity for reverse engineering software binaries.
- Zero-Day Discovery: In an internal evaluation using vulnerabilities from June–August 2026, Astra discovered and used two previously unknown zero-day vulnerabilities.
Safeguards
To mitigate risks, Astra will refuse advanced cybersecurity tasks, such as creating proof-of-concept exploits. OpenAI is rolling out "OpenAI Daybreak" to expand defensive access for vulnerability validation and malware analysis. The model also utilizes a safeguards stack including Codex Auto-review and misalignment monitoring in production to detect and contain unauthorized behavior.
Model Alignment and Safety
GPT-6 Astra is described as OpenAI's most aligned model, with a focus on respecting task boundaries and transparent communication.
- Intent Adherence: In an evaluation of impossible tasks, GPT-6 Astra went beyond its authorized target in 0% of cases, compared to 48% for GPT-5.6 Sol.
- Auto-Review Circumvention: Astra never attempted to circumvent a Codex Auto-Review denial, even when the configuration was deliberately made evadable.
- Capability Hallucinations: Astra is three times less likely than GPT-5.6 Sol to make inaccurate claims about its own capabilities.
- Monitorability: OpenAI noted that Astra's written reasoning is harder to monitor than GPT-5.6 Sol's on simpler tasks, as it solves problems with fewer written steps, which remains a research priority.
Availability and Pricing
GPT-6 Astra is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, as well as via the OpenAI API, Microsoft Azure, and AWS Bedrock.
- API Pricing: $10 per million input tokens and $50 per million output tokens.
- Fast Mode: Available in the API, offering up to 2x speed at 2x the standard price.
- Enterprise Access: Access is off by default for Enterprise administrators and must be enabled manually.