Anthropic Claude Sonnet 4.6 release: upgraded capabilities and 1M‑token context
TL;DR
Claude Sonnet 4.6 is now the default model for Anthropic’s Free, Pro, and Claude Cowork plans, delivering markedly better coding, computer‑use, long‑context reasoning, and agent planning while offering a beta 1 million‑token context window—all at the same $3/$15 per‑million‑token pricing as Sonnet 4.5.
Core Technical Upgrades
Skill upgrades across the board – Sonnet 4.6 improves coding, computer use, long‑context reasoning, agent planning, knowledge work, and design. The model’s instruction‑following consistency and reduced hallucinations are reported to surpass its predecessor and even the older Opus 4.5 frontier model.
1 M‑token context window (beta) – The extended window can hold entire codebases, lengthy contracts, or dozens of research papers in a single request, and the model demonstrates effective reasoning over that length, as shown in the Vending‑Bench Arena evaluation.
Safety enhancements – Extensive safety evaluations indicate Sonnet 4.6 is as safe as or safer than recent Claude models, with a “warm, honest, prosocial” character and no major high‑stakes misalignment concerns. Prompt‑injection resistance is also markedly improved compared with Sonnet 4.5.
Computer‑Use Capabilities
From experimental to production‑grade – Since Anthropic’s first general‑purpose computer‑using model in October 2024, Sonnet models have steadily climbed the OSWorld benchmark. Sonnet 4.6 now achieves human‑level performance on tasks such as complex spreadsheet navigation and multi‑step web‑form completion.
Benchmark evidence – The OSWorld‑Verified scores (released July 2025) show Sonnet 4.6 outperforming earlier Sonnet versions across hundreds of real‑software tasks (Chrome, LibreOffice, VS Code, etc.).
Risk mitigation – Prompt‑injection attacks remain a concern; however, safety tests show Sonnet 4.6’s resistance is comparable to Opus 4.6. Anthropic provides mitigation guidance in its API documentation.
Benchmark and User‑Facing Performance
| Evaluation | Relative performance | Notable observation |
|---|---|---|
| Claude Code (coding) | Preferred over Sonnet 4.5 ~70% of the time; over Opus 4.5 59% of the time | Better context reading, less duplicated logic, fewer hallucinations |
| OfficeQA (enterprise document reasoning) | Matches Opus 4.6 | Significant upgrade for document‑comprehension workloads |
| Vending‑Bench Arena (long‑horizon business simulation) | Outperforms Sonnet 4.5 by investing heavily early and pivoting to profitability later | Demonstrates strategic planning over multi‑month horizons |
| OSWorld‑Verified (computer use) | Highest score among Sonnet models; best among evaluated models on a partner insurance benchmark (94% accuracy) |
Customer feedback – Across multiple partners (Databricks, Replit, Cursor, GitHub, Cognition, Windsurf, Hebbia, Box, Pace, Bolt, Rakuten, Zapier, Convey, Triple Whale, Harvey), Sonnet 4.6 is praised for:
- More polished visual outputs and design sensibility
- Faster resolution of complex code fixes across large codebases
- Improved bug‑detection and review throughput
- Higher answer‑match rates in financial‑services benchmarks
- Stronger multi‑step reasoning in contract routing and CRM coordination
Product and API Enhancements
- Adaptive and extended thinking are supported on the Claude Platform, with context compaction in beta to automatically summarize older conversation turns.
- Web‑search and fetch tools now generate and run filtering code, keeping only relevant content in context and improving token efficiency.
- Tool suite – code execution, memory, programmatic tool calling, tool search, and tool‑use examples are generally available.
- Claude in Excel – MCP connectors enable seamless data pulls from S&P Global, LSEG, PitchBook, Moody’s, FactSet, etc., without leaving the spreadsheet.
Availability and Migration
Claude Sonnet 4.6 is live on all Claude plans, Claude Cowork, Claude Code, the API, and major cloud platforms. The free tier now defaults to Sonnet 4.6 and includes file creation, connectors, skills, and context compaction. Developers can start using the model via the identifier claude-sonnet-4-6 in the Claude API.
Implications for Users and the Market
- Cost‑effective frontier performance – Sonnet 4.6 delivers near‑Opus reasoning quality at the lower Sonnet price point, expanding the feasible use‑cases for heavy‑weight AI workloads.
- Broader automation – The mature computer‑use ability reduces the need for custom connectors, enabling AI to interact with legacy software through the UI.
- Long‑context reasoning – The 1 M‑token window opens possibilities for full‑document analysis, extensive code‑base queries, and multi‑turn planning that were previously impractical.
- Safety posture – Continued emphasis on safety testing and prompt‑injection mitigation reinforces Anthropic’s positioning as a responsible AI provider.
Related Announcements
- Improving Fable 5's biology safeguards – see Anthropic’s blog for details.
- Mariano‑Florentino (Tino) Cuéllar joins Anthropic as Chief Global Affairs Officer – corporate leadership update.
Sources
- OriginalIntroducing Sonnet 4.6
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch