Claude Sonnet 5.5 Release Notes

Claude Sonnet 5.5 delivers significant gains in speed, cost, and agentic coding performance

Claude Sonnet 5.5 is a high-performance model designed for well-scoped everyday tasks, bug fixing, and document creation. It is 30% faster and up to 30% cheaper per task than its predecessor, Sonnet 5, while offering a substantial leap in agentic capabilities, particularly in coding and professional knowledge work.

Performance Benchmarks and Capabilities

Sonnet 5.5 demonstrates dramatic improvements over Sonnet 5 across several key evaluations, with some performance levels approaching those of the more capable Claude Opus 5.5.

Agentic Coding and Computer Use

Sonnet 5.5 shows a massive jump in agentic coding, scoring 70.6% on Terminal-Bench 4.0 compared to Sonnet 5's 10.3%. It also performs strongly on OSWorld 2.1 (80.1% partial) and CursorBench 4.0 (55.5%), placing it within two points of Opus 5.5 on the latter.

Knowledge Work and Reasoning

On the GDPval-AA benchmark, which tests real-world tasks across 44 occupations, Sonnet 5.5 scores 1844, nearly matching Opus 5.5 (1846) and significantly exceeding Sonnet 5 (1449). It also outperforms GPT-6 Sol on long-horizon knowledge work and visual chart recognition (61.6% on Chartography).

Comparison Table

Benchmark Sonnet 5.5 Sonnet 5 Opus 5.5 GPT-6 Sol
Terminal-Bench 4.0 70.6% 10.3% 66.4%¹ —
FrontierCode 1.1 (Main) 46.2% (Max)² 42.4% 54.4% 49.3%
CursorBench 4.0 55.5% 34.1% 57.8% —
GDPval-AA v2.1 1844 1449 1846 1487⁴
Humanity's Last Exam 64.5% (tools) 54.9% (tools) 67.7% (tools) —
OSWorld 2.1 80.1% (partial) 57.0% (partial) 81.8% (partial) —
Chartography 61.6% (no tools) 15.6% (no tools) 64.4% (no tools) 53.6%⁴

¹ Note: The gap in Terminal-Bench is partially attributed to a higher rate of safeguard fallbacks for Opus 5.5 (10%) compared to Sonnet 5.5 (1.5%).

Cost and Speed Efficiency

Sonnet 5.5 maintains the same per-token pricing as Sonnet 5 but reduces the total cost per task by requiring fewer tokens to achieve the same results.

Pricing Structure

Price per 1M tokens Claude Sonnet 5.5 Claude Opus 5.5
Cache reads $0.20 $0.20
Cache writes $2.50 $5
Input tokens $2 $4
Output tokens $10 $20

Efficiency Gains

  • Speed: Generates outputs 30%+ faster than Sonnet 5.
  • Task Cost: Costs up to 30% less per task than Sonnet 5 due to increased token efficiency.
  • Effort Levels: Users can adjust effort levels (Low, Medium, High, Xhigh, Max) to balance speed and cost against quality. Medium is the default for Claude apps.

Safety, Alignment, and Safeguards

Sonnet 5.5 introduces advanced safeguards to match its increased capabilities, particularly in cybersecurity.

  • Cybersecurity: Due to capabilities comparable to Opus 5.5, Sonnet 5.5 is the first Sonnet model to launch with cyber safeguards. High-risk requests will fall back to Sonnet 5.
  • Distillation Prevention: To prevent the extraction of model capabilities via industrial-scale attacks, Sonnet 5.5 includes safety classifiers to prevent reasoning extraction and expands "preserved thinking."
  • Alignment: Automated behavioral audits show that Sonnet 5.5 matches or improves upon Sonnet 5 in honesty, resistance to misuse, and alignment.

Community Insights and Counterpoints

While official benchmarks are strong, community feedback from Hacker News highlights several practical concerns and observations:

  • Value Proposition vs. Opus: Several users questioned the utility of Sonnet 5.5 at high effort levels, noting that Opus 5.5 often provides better accuracy for a similar or lower cost per task.
  • Thinking Token Limits: Some users reported that at "Max" effort, the model can burn through its 128,000 thinking token limit without producing a final response.
  • Safeguard Friction: Users in the Cyber Verification Program reported that safeguards can be overly aggressive, flagging authorized bounty work as Cyber violations.
  • Performance Regressions: Some users reported regressions in specific niche areas, such as esoteric programming languages, where Sonnet 5.5 was noted to be more reluctant to persevere than Sonnet 5.
  • Competitive Landscape: Commenters noted that while Sonnet 5.5 is a leap over Sonnet 5, it remains significantly more expensive than competing models from DeepSeek or OpenAI's Luna.

"In Epic’s early testing, Claude Sonnet 5.5 cleared the same quality bar you’d expect from a higher-tier model... The new model managed tens of thousands of lines of code for gameplay system architecture, kept responses snappy, handled multi-hour tasks, and delivered with less prescriptive prompting." — Daniel Vogel, COO, Epic Games

Sources

Related