Claude Opus 4.6 Release Notes
Anthropic has released Claude Opus 4.6, a significant upgrade to its smartest model designed to enhance agentic coding, long-horizon planning, and complex multidisciplinary reasoning. The update introduces a 1M token context window in beta and provides developers with granular control over the model's reasoning depth and cost through new effort settings.
State-of-the-Art Performance and Benchmarks
Claude Opus 4.6 leads the industry in several high-complexity evaluations, demonstrating superior capabilities in knowledge work and agentic tasks:
- Agentic Coding: It achieves the highest score on the Terminal-Bench 2.0 evaluation.
- Multidisciplinary Reasoning: It leads all other frontier models on Humanity’s Last Exam.
- Knowledge Work: On GDPval-AA, which evaluates performance in finance and legal domains, Opus 4.6 outperforms OpenAI’s GPT-5.2 by approximately 144 Elo points and its predecessor, Claude Opus 4.5, by 190 points.
- Information Retrieval: It outperforms all other models on BrowseComp, measuring the ability to locate hard-to-find online information.
Long-Context Capabilities and "Context Rot"
Opus 4.6 addresses the issue of "context rot" (performance degradation in long conversations). On the 8-needle 1M variant of the MRCR v2 needle-in-a-haystack benchmark, Opus 4.6 scored 76%, significantly higher than Sonnet 4.5's 18.5%. This allows the model to track information across hundreds of thousands of tokens with minimal drift and retrieve buried details more effectively.
Technical Enhancements for Developers
Anthropic has introduced several new features on the Claude Platform to optimize how developers implement long-running agents and manage costs:
Adaptive Thinking and Effort Controls
Developers can now manage the balance between intelligence, speed, and cost using four effort levels: low, medium, high (default), and max.
Additionally, adaptive thinking allows the model to autonomously decide when to use extended thinking based on contextual clues, rather than requiring a binary on/off toggle.
Context Management and Output
- Context Compaction (Beta): To prevent agents from hitting context window limits during long tasks, this feature automatically summarizes and replaces older context when a configurable threshold is reached.
- 1M Token Context (Beta): This is the first Opus-class model to support a 1M token window. Prompts exceeding 200k tokens are subject to premium pricing ($10/$37.50 per million input/output tokens).
- 128k Output Tokens: The model can now generate outputs up to 128k tokens, reducing the need to split large tasks into multiple requests.
Product Integration and Agentic Workflows
Opus 4.6 is integrated into several productivity tools to enable more autonomous work:
- Claude Code: Now features "agent teams" in research preview, allowing multiple agents to work in parallel and coordinate autonomously on tasks like codebase reviews.
- Office Integration: Claude in Excel has been upgraded to handle multi-step changes and unstructured data ingestion. Claude in PowerPoint is now in research preview for Max, Team, and Enterprise plans, allowing the model to generate decks based on layouts, fonts, and slide masters.
- Cowork: Opus 4.6 powers the Cowork research preview, enabling autonomous multitasking across research, financial analysis, and document creation.
Safety and Alignment
According to the system card, Opus 4.6 maintains a safety profile as good as or better than other frontier models. It shows low rates of misaligned behaviors, including deception and sycophancy, and has the lowest rate of over-refusals among recent Claude models.
To mitigate risks associated with the model's enhanced cybersecurity capabilities, Anthropic has implemented six new cybersecurity probes to detect harmful responses and is focusing the model's application on cyber-defensive uses, such as patching vulnerabilities in open-source software.
Sources
- OriginalClaude Opus 4.6
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch