AI

Anthropic launches Claude Sonnet 5.5, a mid tier model that beats Opus 5.5 on coding at half the price

Anthropic's Claude Sonnet 5.5 undercuts the flagship Opus 5.5 on price while claiming a higher agentic coding score, but independent testing finds a steep token cost at maximum effort.

T
By TechQuire Daily Staff TechQuire Daily Staff
September 29, 2026 / 7 min read

Anthropic released Claude Sonnet 5.5 on September 28, 2026, a mid-tier model that the company says runs more than 30% faster than the Sonnet version it replaces and costs up to 30% less per task, while matching or beating its own flagship, Claude Opus 5.5, on several agentic coding evaluations. Sonnet 5.5 is the second model in the Claude 5.5 family, arriving six days after Opus 5.5, which Anthropic launched on September 22 as its most capable system and priced at $4 per million input tokens and $20 per million output tokens.

For most of the past two years, Anthropic's tiers have been easy to tell apart. Opus carried the top scores and the top price, Sonnet handled the high-volume work at a fraction of the cost, and Haiku sat below both for cheap, fast tasks. Sonnet 5.5 blurs that line. Anthropic says the model scores 70.6% on Terminal-Bench 4.0, a test of whether an AI agent can complete complex professional tasks by typing commands, which is above Opus 5.5's 66.4% and far above Sonnet 5's 10.3%. On GDPval-AA, a benchmark that grades real-world work across 44 occupations using an Elo rating, the two 5.5 models finish effectively level at 1844 and 1846.

The release is also a business story. Reuters reported on September 28 that Anthropic rolled out Claude Sonnet 5.5 as the AI lab expands its product lineup ahead of a planned IPO, noting that enterprise customers account for about 80% of its business. That mix explains the pitch: Anthropic describes the new model as suited to everyday corporate work such as coding, document creation and spreadsheet tasks, the sort of unglamorous, high-volume jobs that generate steady revenue rather than headline demos.

It arrives into a price war. Rival OpenAI cut GPT-6 Sol to $2 per million input tokens and $10 per million output tokens in the week before the launch, exactly the pricing Anthropic is now asking for Sonnet 5.5. Safety remains part of the backdrop too. Anthropic chief executive Dario Amodei earlier in September called on the global AI community to slow the pace of releasing new capabilities to address safety concerns, a message that sits alongside a product cadence of roughly one major model launch per week.

Key Facts

Anthropic said the pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads. That is half the price of Opus 5.5 at $4 and $20. The company also claims the model typically needs fewer tokens to complete the same work, which is where the claim of up to 30% savings per task comes from. Anthropic announced on September 28 that Sonnet 5.5 runs more than 30% faster than its predecessor.

On the benchmarks Anthropic published, Terminal-Bench 4.0 comes in at 70.6% for Sonnet 5.5, against 10.3% for Sonnet 5 and 66.4% for Opus 5.5. On GDPval-AA, Sonnet 5.5 scores 1844, Opus 5.5 scores 1846, and OpenAI's GPT-6 Sol scores 1487. At the High effort setting on FrontierCode, Anthropic says Sonnet 5.5 matches GPT-6 Sol's best score for roughly a fifth of the cost per task. The model is also the first Sonnet to beat the video game Pokemon Red using only screenshots as input.

SiliconANGLE reported on September 28 that the model is positioned as the most capable mid-tier offering in Anthropic's lineup, and that it includes invisible text watermarking to comply with global regulations including the EU AI Act. It is also the first Sonnet model to launch with cyber safeguards and fallbacks covering cybersecurity and biology, the same foundation used for Anthropic's most capable systems, because its cybersecurity capabilities are comparable to Opus 5's. Anthropic says most software development and life sciences work is unaffected by those guardrails.

Decrypt reported on September 28 that independent tester Artificial Analysis measured 63.6% for Sonnet 5.5 on Terminal-Bench 4.0, against 59.6% for Opus 5.5 and 59.1% for OpenAI's GPT-6 Astra. The same independent testing found the catch. At maximum effort, Sonnet 5.5 wrote about 193,000 tokens per test task, the highest figure Artificial Analysis has measured and roughly 60% more than Opus 5.5, working out to about $7.60 per task, some 50% above Sonnet 5. That runs against the headline claim of up to 30% savings.

Anthropic's answer is that the savings depend on effort settings. At Medium effort, the company says Sonnet 5.5 beats Sonnet 5's best coding score for less than a tenth of the cost. Zendesk director of AI Abhinay Kathuria, an early customer, said tickets were processed 20% faster after testing hundreds of real support use cases.

Analysis

The bigger picture here is that the mid-tier model, not the flagship, has become the most contested product in enterprise AI. Sonnet 5.5 is priced squarely at the volume business that Anthropic says generates about 80% of its revenue, and it is priced to match OpenAI's freshly cut GPT-6 Sol rates rather than to undercut them. When two leading labs land on the same $2 and $10 numbers within a week of each other, that is a commodity market forming in public, and the differentiators shift from raw capability to cost per finished task, latency, reliability and compliance features.

On the numbers, the honest judgement is that Sonnet 5.5 is a genuine step change at the low end of Anthropic's family and a more ambiguous proposition at the high end. Going from 10.3% to 70.6% on an agentic coding benchmark in one generation is an enormous jump, and it is the kind of improvement that changes what teams are willing to automate. But the independent measurements are less flattering than Anthropic's own, and the token consumption at maximum effort is a real warning sign. A model that is cheaper per token but spends far more tokens per job is not cheaper per job, and that distinction is exactly where procurement teams get burned.

What this really means is that buyers should treat the possibility of up to 30% savings per task as a best case tied to a specific effort setting rather than a default. Anthropic's Medium and High effort claims, where the model beats Sonnet 5's best score for a tenth or a fifth of the cost, are the ones that matter for production deployments, because that is where most enterprise traffic will actually run. The max effort mode, at roughly $7.60 per task, is closer to flagship economics and should be reserved for the hardest problems rather than left as a default.

There is also a strategic read on the timing. Launching Sonnet 5.5 six days after Opus 5.5, with the mid-tier model beating the flagship on a marquee coding benchmark, is a confident move for a company heading toward a public listing. It signals a repeatable release process and pricing pressure that comes from its own product line rather than from competitors alone. It also raises a question investors will ask: if the cheaper model is good enough for most work, what holds the premium tier up?

Why It Matters

The stakes are concrete for the enterprises Anthropic names as customers, including Salesforce, Databricks, Goldman Sachs and Novo Nordisk. Coding assistants, document drafting and spreadsheet work are the tasks that consume the most tokens across large organizations, and small changes in price per task scale into large budget lines. A mid-tier model that lands within two points of the flagship on a real-world work benchmark at half the listed token price reshapes how those teams allocate AI spend, and it makes the effort setting a budget control rather than a technical footnote.

The launch also matters for how labs handle risk while shipping fast. Sonnet 5.5 is the first Sonnet to carry cyber safeguards and fallbacks comparable to those built for Anthropic's most capable models, and the first to carry invisible text watermarking aimed at the EU AI Act and similar rules. That combination, strong capability at a low price with safety controls attached by default, is the pattern regulators have been asking for, and it arrives from a company whose chief executive spent September asking the industry to slow down.

Consumers of these benchmarks should also note the gap between vendor-reported and independent scores. Anthropic's 70.6% and 66.4% for Sonnet 5.5 and Opus 5.5 compare with Artificial Analysis's 63.6% and 59.6%. The gap between the two models is narrower in independent hands. Neither result is dishonest, but for anyone choosing a model on a benchmark number, the source of that number matters as much as the number.

Next Up

Claude Haiku 5.5, described by Anthropic as built for high-volume and cost-sensitive applications, is due in the coming weeks, according to both the company and Reuters. That would complete the Claude 5.5 family and give Anthropic a full ladder from cheap and fast to expensive and capable, all released inside a single month.

The next real test is not another benchmark. It is whether enterprises can get the promised savings without the token blowout that independent testing found at maximum effort, and whether Anthropic's IPO story holds up when its own mid-tier model is the strongest argument for spending less with it rather than more.

Tagged

Comments (0)

No comments yet. Be the first to share your thoughts.