Benchmark data from the official Anthropic Sonnet 5 launch page. Use this for internal justification when switching your team to Sonnet 5.
Key benchmark results
Sonnet 5 vs Sonnet 4.6: substantially better on every benchmark tested — full improvement across reasoning, tool use, coding, and knowledge work.
Sonnet 5 vs Opus 4.8: doesn't beat Opus on any specific benchmark, but comes very close. Slightly better on knowledge work. On computer use and agentic search at high effort levels, it can match Opus 4.8.
BrowseComp (agentic search): Sonnet 5 at high effort approaches Opus 4.8 performance. Sonnet 4.6 fell well short of Opus 4.8 at any effort level.
OSWorld-Verified (computer use): same pattern. Sonnet 5 covers a much wider range of cost-performance options than Sonnet 4.6, and at high effort can match Opus 4.8 on some tasks.
Humanity's Last Exam: Sonnet 4.6 scores 34.6% without tools, 46.8% with tools (grader updated).
Safety: lower misaligned behaviour rate than Sonnet 4.6. Higher than Opus 4.8 and Mythos Preview.
Cybersecurity: Sonnet 5 was never able to develop a full working exploit in Firefox 147 vulnerability testing. It shows slightly higher partial success than Sonnet 4.6, due to general intelligence improvements rather than specific training.
Cost-performance comparison for GTM agents
For a similar agentic task at medium effort:
| Model |
Cost |
Notes |
| Opus 4.8 |
$8 |
Example cost from the YouTube breakdown |
| Sonnet 5 |
~$4 |
At introductory pricing, roughly half |
Performance difference: roughly 5% lower pass rate on the same task.
For GTM teams spending $1,000–2,000/month on Claude tokens through Hermes or OpenClaw:
- Switching from Opus 4.8 to Sonnet 5 for execution work: roughly 50% cost reduction
- Performance impact: minimal for most GTM use cases — outreach, research, inbox, proposals
- Remaining Opus usage: planning, complex reasoning, tasks where quality is critical
Recommended split: 80% Sonnet 5 for execution and everyday agent work, 20% Opus 4.8 for planning and hard reasoning.
The introductory pricing window
Introductory pricing is $2/MTok input and $10/MTok output, available through August 31, 2026. Standard pricing from September 1, 2026 is $3/MTok input and $15/MTok output.
Sonnet 5 uses an updated tokenizer that can map the same input to 1.0–1.35x more tokens depending on content type. Introductory pricing is set to be roughly cost-neutral versus Sonnet 4.6 at these rates.
If you're currently paying for Sonnet 4.6 API usage, your bill should be roughly the same on Sonnet 5 through August 31, with meaningfully better output quality. After September 1, evaluate whether the quality improvement justifies the price difference for your specific workflows.