Anthropic Claude 3.5 Sonnet and Claude 3.5 Haiku Release
Anthropic has introduced an upgraded version of Claude 3.5 Sonnet and a new model, Claude 3.5 Haiku, while launching a public beta for a new capability called "computer use." This update enables AI models to interact with computer interfaces—moving cursors, clicking buttons, and typing text—much like a human user would.
Computer Use: General Computer Interaction in Public Beta
Claude 3.5 Sonnet is the first frontier AI model to offer a public beta for "computer use," allowing it to translate high-level instructions into a series of computer commands to navigate software and the web. Rather than relying on specific tools for individual tasks, this capability focuses on general computer skills.
Technical Implementation and Performance
Developers can integrate an API that allows Claude to perceive and interact with computer interfaces. In evaluations on OSWorld, which measures the ability to use computers like humans, Claude 3.5 Sonnet achieved a score of 14.9% in the screenshot-only category, outperforming the next-best AI system's score of 7.8%. When allowed more steps to complete a task, its score increased to 22.0%.
Current Limitations and Safety
Anthropic describes the capability as experimental, noting that it can be cumbersome and error-prone. Common human actions such as scrolling, dragging, and zooming remain challenging for the model. To mitigate risks associated with spam, misinformation, and fraud, Anthropic has developed new classifiers to identify when computer use is active and whether harm is occurring.
Upgraded Claude 3.5 Sonnet: Advanced Software Engineering
The updated Claude 3.5 Sonnet provides across-the-board improvements over its predecessor, maintaining the same price and speed while significantly enhancing its agentic coding and tool use capabilities.
Coding and Tool Use Benchmarks
- SWE-bench Verified: Performance increased from 33.4% to 49.0%, surpassing all publicly available models, including reasoning models like OpenAI o1-preview.
- TAU-bench: In the retail domain, performance improved from 62.6% to 69.2%; in the airline domain, it improved from 36.0% to 46.0%.
Industry Adoption
Companies such as GitLab, Cognition, and The Browser Company have reported significant gains. GitLab noted stronger reasoning (up to 10% across use cases) with no added latency, while Cognition observed improvements in coding, planning, and problem-solving.
Claude 3.5 Haiku: Speed and Affordability
Claude 3.5 Haiku is the next generation of Anthropic's fastest model, designed for low latency and high efficiency. It matches or surpasses the performance of Claude 3 Opus (the previous generation's largest model) on many intelligence benchmarks while maintaining speeds similar to the original Claude 3 Haiku.
Performance and Use Cases
Claude 3.5 Haiku scores 40.6% on SWE-bench Verified, outperforming the original Claude 3.5 Sonnet and GPT-4o. Due to its low latency and improved instruction following, it is optimized for user-facing products, specialized sub-agent tasks, and processing large volumes of data such as inventory or pricing records.
Pricing and Availability
As of December 3, 2024, Claude 3.5 Haiku is priced at $0.80 per million input tokens and $4.00 per million output tokens. It is available via the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI, initially as a text-only model with image input support to follow.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch