Claude 3.5 Sonnet release notes / what's new
Anthropic has launched Claude 3.5 Sonnet, the first release in the Claude 3.5 model family. This model delivers intelligence that outperforms competitor models and the previous top-tier Claude 3 Opus across a wide range of evaluations, while maintaining the speed and cost profile of a mid-tier model.
Performance and Intelligence Benchmarks
Claude 3.5 Sonnet sets new industry benchmarks in graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval). The model demonstrates significant improvements in grasping nuance, humor, and following complex instructions, and produces high-quality content with a natural tone.
Agentic Coding Capabilities
In internal agentic coding evaluations, Claude 3.5 Sonnet solved 64% of problems, compared to 38% solved by Claude 3 Opus. These tests measure the ability to fix bugs or add functionality to open-source codebases based on natural language descriptions. When equipped with the relevant tools, the model can independently write, edit, and execute code, facilitating tasks such as code translations and the migration of legacy applications.
Speed, Cost, and Availability
Claude 3.5 Sonnet operates at twice the speed of Claude 3 Opus. This combination of increased performance and cost-efficiency makes it suitable for multi-step workflows and context-sensitive customer support.
Pricing and Technical Specifications:
- Input Cost: $3 per million tokens
- Output Cost: $15 per million tokens
- Context Window: 200K tokens
Availability:
- Free Access: Available on Claude.ai and the Claude iOS app.
- Paid Access: Claude Pro and Team plan subscribers have higher rate limits.
- API/Cloud Access: Available via the Anthropic API, Amazon Bedrock, and Google Cloud’s Vertex AI.
Visual Reasoning and Vision Capabilities
Claude 3.5 Sonnet is Anthropic's strongest vision model to date, surpassing Claude 3 Opus on standard vision benchmarks. The model shows marked improvements in visual reasoning, specifically in interpreting charts and graphs and accurately transcribing text from imperfect images. These capabilities are particularly applicable to the retail, logistics, and financial services sectors.
Artifacts: Collaborative Workspace
Anthropic has introduced "Artifacts" on Claude.ai, a feature that transforms the interface from a conversational AI into a collaborative work environment. When Claude generates code snippets, website designs, or text documents, they appear in a dedicated window alongside the conversation. This allows users to see, edit, and build upon AI-generated content in real-time.
Safety, Privacy, and External Evaluation
Claude 3.5 Sonnet remains at AI Safety Level 2 (ASL-2) according to red teaming assessments. To ensure safety, Anthropic engaged with the UK’s Artificial Intelligence Safety Institute (UK AISI), which conducted pre-deployment safety evaluations and shared results with the US AI Safety Institute (US AISI).
Safety and Privacy Measures:
- External Expertise: Anthropic integrated feedback from subject matter experts, including child safety experts at Thorn, to refine classifiers and fine-tune models.
- Data Privacy: Generative models are not trained on user-submitted data without explicit user permission.
Future Roadmap
Anthropic plans to release Claude 3.5 Haiku and Claude 3.5 Opus later this year to complete the model family. Future developments include new modalities, enterprise application integrations, and a "Memory" feature to allow Claude to remember user preferences and interaction history.
Sources
- OriginalIntroducing Claude 3.5 Sonnet
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch