Anthropic released Claude Sonnet 5 on Tuesday at $2 per million input tokens and $10 per million output tokens, undercutting its own Opus 4.8 while landing within 6 points of it on agentic coding. Sonnet 5 scores 63.2% on the benchmark, against Opus 4.8 at 69.2% and the February-vintage Sonnet 4.6 at 58.1%. Starting Tuesday, the model is the default for free and Pro users across every Claude subscription tier.
The promotional pricing runs through August 31, after which input tokens move to $3 per million while output stays at $10. That keeps Sonnet 5 cheaper than Opus 4.8, OpenAI's GPT-5.5, and Google's Gemini 3.1 Pro, though still more expensive than Gemini 3.5 Flash, which Google launched in May.
“It can make plans, use tools like browsers and terminals, and run autonomously at a level that, just a few months ago, required larger and more expensive models.”— Anthropic, company blog post
Anthropic's framing is that mid-tier models can now do what only flagships could a few months ago — plan multi-step work, drive browsers and terminals, and run without a human checking each step. That is the same pitch OpenAI used for GPT-5.6 Sol, which entered preview last week with subagent splitting, and the same pitch Google used for Gemini 3.5 Flash.
Key facts
- 01Claude Sonnet 5 launches at $2 per million input tokens and $10 per million output tokens through August 31, then rises to $3 per million input.
- 02Sonnet 5 scores 63.2% on an agentic coding benchmark, versus Opus 4.8 at 69.2% and Sonnet 4.6 at 58.1%.
- 03Sonnet 5 becomes the default model for free and Pro plans starting Tuesday.
- 04On a knowledge work benchmark, Sonnet 5 slightly outperforms Opus 4.8, Anthropic's flagship for hardest-problem reasoning.
- 05Anthropic says Sonnet 5 cuts hallucinations, sycophancy, and prompt-injection susceptibility compared with Sonnet 4.6.
The implication is straightforward. Agentic capability is no longer the differentiator at the top of the price band — it is the floor across every tier. What labs now compete on is cost per successful task and reliability without supervision, which is exactly where Sonnet 5 is aimed.
On the knowledge work benchmark Anthropic cites, Sonnet 5 slightly beats Opus 4.8, the model Anthropic still positions as its first choice for subtle judgment calls and deep research. Anthropic told developers to mix the two: "Opus 4.8 is still the model of choice for higher accuracy on these tasks, but Sonnet 5 provides developers with lower-priced options that are of much higher quality than what was previously available." The company describes the choice as an effort dial between cost and accuracy.
Early enterprise testers point to fewer mid-task stalls as the real unlock. Zapier ran a two-part workflow — update Salesforce account tiers, then send an enterprise launch announcement — that previous models would abandon halfway through.
On safety, Anthropic reports lower rates of cooperation with misuse, deception, hallucination, and sycophancy than Sonnet 4.6, plus better refusal of malicious prompts and prompt-injection hijacks. The company also says Sonnet 5 has a much lower ability to perform dangerous cybersecurity tasks than current Opus models — a deliberate ceiling, given how widely Sonnet 5 will now be deployed as the default chat experience.
Lovable, which exposes coding tools to a mass developer audience, flagged refusal behavior as the differentiator that matters to it, calling Sonnet 5 a model that knows when to say no. For consumer-facing builders deploying agentic features at scale, refusal reliability is closer to a load-bearing feature than a safety footnote.
The gaps are worth naming. Sonnet 5 still trails Opus 4.8 and Anthropic's Claude Mythos Preview on misaligned-behavior evaluations, and the 6-point agentic coding gap to Opus 4.8 is real for teams whose workflows fail intolerably at the margin. Anthropic's own guidance is to route hard problems to Opus and volume work to Sonnet, which only works if developers actually build the routing logic.
The economics of agents are starting to bend in a way that favors whoever can hold quality at the lower price point. A model that finishes a Salesforce update and a customer email end-to-end at $2 per million input tokens changes what kinds of workflows survive a build-versus-buy review inside a mid-market company. Anthropic is betting that the cost curve, not the capability ceiling, is what determines whose agents get deployed at scale through the back half of 2026 — and on the pricing math alone, that bet looks well placed.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



