Anthropic released Claude Opus 4.8 today, holding Opus 4.7's price of $5 per million input tokens and $25 per million output tokens while cutting fast-mode pricing to $10 and $50 — three times cheaper than previous fast modes, and running at 2.5x the speed. The model posts 84% on Online-Mind2Web, a browser-agent benchmark where Anthropic says Opus 4.8 beats both Opus 4.7 and OpenAI's GPT-5.5. It also becomes the first model to clear 10% on the all-pass standard of the Legal Agent Benchmark.
Pricing parity is the strategic choice. Most frontier launches raise prices alongside capability, but Anthropic is keeping Opus's headline rate flat for a second generation and absorbing the gains in fast mode, where output is now a fifth of the equivalent quality token-for-token compared with prior releases. Developers can access claude-opus-4-8 through the Claude API starting today, and the model is live across claude.ai, Claude Code, and Cowork.
The reliability gains are the more consequential story. Anthropic's own evaluations show Opus 4.8 is roughly four times less likely than Opus 4.7 to allow flaws in code it has written to pass unremarked — a direct response to the failure mode that has dogged autonomous coding agents since GPT-4-class models first started shipping pull requests. The model defaults to high effort, spending a similar token budget to Opus 4.7 but converting it into better outcomes.
Key facts
- 01Claude Opus 4.8 scores 84% on Online-Mind2Web, ahead of both Opus 4.7 and GPT-5.5 on browser-agent tasks.
- 02API pricing holds at $5 per million input tokens and $25 per million output; fast mode runs 2.5x faster at $10/$50, 3x cheaper than prior fast modes.
- 03Opus 4.8 is roughly 4x less likely than Opus 4.7 to let flaws in its own code pass unremarked, per Anthropic's evaluations.
- 04Claude Code's new dynamic workflows let Opus 4.8 run hundreds of parallel subagents and handle codebase migrations spanning hundreds of thousands of lines.
- 05Databricks reports Genie running Opus 4.8 at 61% cheaper token cost than Opus 4.7 on multimodal document reasoning.
Anthropic also published a roster of testimonials from partners running early access. Cursor's Michael Truell reports tool calling that uses fewer steps for the same intelligence on CursorBench. Cognition's Scott Wu, whose Devin agent runs unattended engineering work, said Opus 4.8 fixes the comment-verbosity and tool-calling regressions seen in Opus 4.7. Hebbia's Aabhas Sharma flagged better citation precision on dense financial filings.
“Claude Opus 4.8 has noticeably better judgment. In Claude Code, it asks the right questions, catches its own mistakes, pushes back when a plan isn't sound, and builds up confidence around complex, multi-service explorations before making big changes. It's a great model to build with.”— Tom Pritchard, Staff Engineer
Claude Code is getting the second headline feature: dynamic workflows, available in research preview, lets Claude plan a task and run hundreds of parallel subagents in a single session before verifying outputs and reporting back. Anthropic says the combination of Opus 4.8 and dynamic workflows can carry codebase-scale migrations across hundreds of thousands of lines of code from kickoff to merge, gated by the existing test suite. The feature is rolling out to Enterprise, Team, and Max plans.
Two smaller changes round out the release. Effort control sits next to the model selector on claude.ai and Cowork, letting users dial how much thinking time Opus puts into a response — lower settings burn rate limits more slowly, while extra and max settings spend more tokens on harder problems. The Messages API now accepts system entries inside the messages array, so developers can update Claude's instructions mid-task without busting the prompt cache or routing through a user turn.
Enterprise traction is the other tell. Databricks integrated Opus 4.8 into Genie, its agent for data and knowledge work. Thomson Reuters subsidiary CoCounsel Legal's Joel Hron said the model delivers meaningful consistency and reasoning gains for fiduciary-grade legal workflows. Reka co-founder Kay Zhu called Opus 4.8 the only model to complete every case end-to-end on her firm's Super-Agent benchmark, beating prior Opus models and GPT-5.5 at parity on cost.
“Claude Opus 4.8 sets a new bar for enterprise AI. In Genie, Databricks' AI agent for data and knowledge work, the new Opus model unlocks a step change in agentic reasoning, tackling deeper, multistep questions faster than any prior Opus.”— Hanlin Tang, CTO, Neural Networks at Databricks
Anthropic's Alignment team concluded Opus 4.8 reaches new highs on prosocial traits like supporting user autonomy and acting in the user's best interest, with rates of misaligned behavior substantially lower than Opus 4.7 and comparable to Claude Mythos Preview — Anthropic's best-aligned model to date. A small set of organizations is already using Mythos Preview for cybersecurity work under Project Glasswing, but Anthropic says the model class requires stronger cyber safeguards before broader release, expected in the coming weeks.
The caveat Anthropic itself flags: this is described as a modest but tangible improvement, not a generational leap. The bigger story arrives when Mythos-class models clear safeguards, and when Anthropic ships the cheaper Opus-capability tier it says is in development. Opus 4.8 also remains expensive relative to mid-tier competitors on raw token cost, which matters for high-volume agent deployments even with the fast-mode discount.
For the AI market, the pricing stance is the most strategically loaded move in this release. By holding Opus's flagship rate flat across two generations and pushing the discount into fast mode, Anthropic is signaling that the margin pressure in agentic workloads is coming from inference economics, not from sticker price — and that whoever wins long-horizon coding and browser-agent reliability gets to dictate the next round of enterprise contracts. The 4x reduction in unflagged code flaws and the 84% Online-Mind2Web score are the numbers Cursor, Cognition, and Databricks will quote back when their own customers ask why the bill stayed the same.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.



