Claude Opus 4.8 Release Notes
Anthropic has released Claude Opus 4.8, an upgrade to Opus 4.7 that improves benchmarks across coding, reasoning, and practical knowledge work. The update introduces a more reliable collaborator for agentic tasks and is available immediately at the same price point as its predecessor.
Enhanced Model Capabilities and Reliability
Claude Opus 4.8 demonstrates improved judgment and reliability, particularly in agentic and autonomous workflows. According to early testers and internal evaluations, the model is more effective at catching its own mistakes and proactively flagging uncertainties.
Key Performance Gains
- Honesty and Accuracy: Opus 4.8 is approximately four times less likely than Opus 4.7 to allow flaws in its own written code to pass unremarked.
- Agentic Performance: The model is reported as the only one to complete every case end-to-end on the Super-Agent benchmark, surpassing prior Opus models and GPT-5.5 at cost parity.
- Browser and Computer Use: Opus 4.8 scored 84% on Online-Mind2Web, outperforming both Opus 4.7 and GPT-5.5.
- Legal and Professional Work: The model achieved the highest recorded score on the Legal Agent Benchmark, becoming the first model to exceed 10% on the all-pass standard.
Industry Feedback
Industry partners have highlighted specific operational improvements:
- Databricks: Reported a "step change in agentic reasoning" for their Genie AI agent, with multimodal reasoning over unstructured content being 61% cheaper in token cost than Opus 4.7.
- Devin: Noted that Opus 4.8 fixes comment-verbosity and tool-calling issues found in Opus 4.7, enabling more consistent unattended autonomous engineering.
- Hebbia: Observed better citation precision and improved token efficiency on retrieval for financial-document workflows.
- Cursor: Reported that tool calling is more efficient, requiring fewer steps for the same level of intelligence.
New Features and Tooling
Alongside the model release, Anthropic has introduced several features to give users more control over model behavior and scale.
Dynamic Workflows in Claude Code
Available in research preview for Enterprise, Team, and Max plans, dynamic workflows allow Claude to manage large-scale problems by planning work and running hundreds of parallel subagents in a single session. This enables codebase-scale migrations across hundreds of thousands of lines of code from kickoff to merge.
Effort Control
Users on claude.ai and Cowork can now select the amount of effort Claude puts into a response:
- High Effort (Default): Balanced quality and user experience; performs better than Opus 4.7's default on coding tasks while using similar tokens.
- Extra/Max Effort: The model spends more tokens to achieve better results, recommended for difficult tasks and long-running asynchronous workflows.
- Low Effort: Faster responses that consume rate limits more slowly.
Messages API Update
The Messages API now supports system entries inside the messages array. This allows developers to update instructions, permissions, token budgets, or environment context mid-task without breaking the prompt cache or requiring a user turn.
Alignment and Safety
Opus 4.8 has undergone a detailed alignment assessment. The Anthropic Alignment team found that the model reaches "new highs" in prosocial traits, specifically in supporting user autonomy and acting in the user's best interest. Rates of misaligned behavior, including deception or cooperation with misuse, are substantially lower than in Opus 4.7 and are comparable to the Claude Mythos Preview model.
Pricing and Availability
Claude Opus 4.8 is available via the Claude API (claude-opus-4-8) and claude.ai.
| Mode | Input Price (per M tokens) | Output Price (per M tokens) |
|---|---|---|
| Regular | $5 | $25 |
| Fast Mode (2.5x speed) | $10 | $50 |
Fast mode for Opus 4.8 is three times cheaper than fast mode was for previous models.
Future Roadmap
Anthropic is working on two primary directions: developing Opus-level capabilities at a lower cost and creating a new class of higher-intelligence models. As part of Project Glasswing, the company is currently testing Claude Mythos Preview for cybersecurity work and expects to bring Mythos-class models to all customers in the coming weeks following the implementation of necessary cyber safeguards.
Sources
- OriginalIntroducing Claude Opus 4.8
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch