Claude 3 Model Family Release
Anthropic has announced the Claude 3 model family, consisting of three models—Opus, Sonnet, and Haiku—designed to provide a scalable balance of intelligence, speed, and cost. This release marks a significant advancement in general intelligence, introducing multimodal vision capabilities and reducing unnecessary refusals across the entire family.
Model Tiering and Capabilities
Claude 3 is divided into three models tailored for different operational needs:
- Claude 3 Opus: The most intelligent model in the family, designed for highly complex tasks. It outperforms peers on benchmarks including MMLU (undergraduate level expert knowledge), GPQA (graduate level expert reasoning), and GSM8K (basic mathematics). It is intended for task automation, R&D, and advanced strategic analysis.
- Claude 3 Sonnet: Positioned as the balance between intelligence and speed. It is 2x faster than Claude 2 and 2.1 and is optimized for enterprise workloads such as RAG, sales automation, and code generation.
- Claude 3 Haiku: The fastest and most cost-effective model. It is designed for near-instant responsiveness in live customer chats, content moderation, and extracting knowledge from unstructured data.
Technical Advancements
Vision and Multimodality
All Claude 3 models now possess sophisticated vision capabilities, allowing them to process visual formats including photos, charts, graphs, and technical diagrams. This modality is specifically targeted at enterprise users with knowledge bases stored in PDFs, flowcharts, or presentation slides.
Accuracy and Hallucination Reduction
Claude 3 Opus demonstrates a twofold improvement in accuracy on challenging open-ended factual questions compared to Claude 2.1. The model also exhibits reduced levels of incorrect answers (hallucinations) and a higher frequency of admissions of uncertainty when it does not know an answer.
Context Window and Recall
The family launches with a 200K context window, though all three models can accept inputs exceeding 1 million tokens for select customers. In 'Needle In A Haystack' (NIAH) evaluations, Claude 3 Opus achieved over 99% accuracy in recalling information from vast corpora, occasionally identifying when the test "needle" was artificially inserted into the text.
Instruction Following and Structured Output
Claude 3 models show improved adherence to complex, multi-step instructions and brand voice guidelines. They are also more proficient at producing structured outputs, such as JSON, which facilitates natural language classification and sentiment analysis.
Safety and Responsible Design
Anthropic has implemented several safety measures to ensure the models remain trustworthy:
- Refusal Rates: The models are significantly less likely to refuse harmless prompts that border on system guardrails compared to previous generations.
- Bias Reduction: According to the Bias Benchmark for Question Answering (BBQ), Claude 3 exhibits fewer biases than previous models.
- Safety Level: The family remains at AI Safety Level 2 (ASL-2) per Anthropic's Responsible Scaling Policy. Red teaming evaluations concluded that the models present negligible potential for catastrophic risk.
- Constitutional AI: The models continue to utilize Constitutional AI to improve transparency and safety.
Pricing and Availability
| Model | Input (per M tokens) | Output (per M tokens) | Context Window |
|---|---|---|---|
| Opus | $15 | $75 | 200K* |
| Sonnet | $3 | $15 | 200K |
| Haiku | $0.25 | $1.25 | 200K |
*1M tokens available for specific use cases.
Availability: Opus and Sonnet are available via the Claude API (generally available in 159 countries) and claude.ai (Sonnet for free users, Opus for Pro subscribers). Sonnet is available via Amazon Bedrock and in private preview on Google Cloud’s Vertex AI Model Garden, with Opus and Haiku following soon.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch