Claude Sonnet 5.5 is the next-generation product in Anthropic’s Sonnet model series; Claude Sonnet 5.5 offers the best balance between speed and intelligence. The following five major changes will affect code currently running on Claude Sonnet 5:
1. Using "between_tools" can disable up-front thinking.
2. Forcing the use of tools will result in an error.
3. Thinking content blocks are bound to the model and the conversation.
4. On the Claude API and Google Cloud, the older "computer_20251124" computer tool usage is no longer accepted.
5. The advisor tool will refuse to accept Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisor models.
Use between_tools to disable up-front thinking.
Forcing tool use will return an error.
Thinking blocks are bound to the model and the conversation.
On the Claude API and Google Cloud, early computer_20251124 computer tool usage is no longer accepted.
The advisor tool no longer accepts Claude Opus 4.8, Claude Opus 4.7, and Claude Sonnet 5 as advisors.
Pricing
Input Modalities
- Text
- Vision
Output Modalities
- Text
Context length
- 1M tokens
Providers
AWS claude-sonnet-5-5
Pricing$2$10
Cache Read$0.2/M tokens
Web Search$0.01/request
Cache Write$2.5/M tokens
Cache Write 5 Minutes$2.5/M tokens
Cache Write 1 Hour$4/M tokens
Context1M
Max output0
Latency4.1S
Throughput102.5TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
100.00% uptime today
VertexAI claude-sonnet-5-5
Pricing$2$10
Cache Read$0.2/M tokens
Web Search$0.01/request
Cache Write$2.5/M tokens
Cache Write 5 Minutes$2.5/M tokens
Cache Write 1 Hour$4/M tokens
Context1M
Max output0
Latency4.1S
Throughput113.5TPS
Uptime
0.00% uptime 2 days ago
0.00% uptime yesterday
100.00% uptime today
Performance for claude-sonnet-5-5
Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).
Uptime
Loading...
Latency
Loading...
Throughput
Loading...
Try this model
Python