Anthropic has released Claude Sonnet 5.5, a mid-tier model that outperforms the company's own flagship Opus 5.5 on a key coding benchmark while costing roughly half as much per token — though early independent testing suggests the savings may be less dramatic than they appear.
A Mid-Tier Model That Punches Above Its Weight
Sonnet 5.5 sits below Anthropic's premium Opus line in the company's product hierarchy, yet it managed to top the flagship on Terminal-Bench 4.0, a benchmark that measures how well AI systems handle command-line and software engineering tasks. That result is notable because it upends the usual assumption that bigger, pricier models automatically deliver better performance on demanding technical work.
The pricing gap is the headline selling point. Sonnet 5.5 costs about half as much per token as Opus 5.5, positioning it as an attractive option for developers and businesses who want strong coding capabilities without paying flagship rates.
A mid-tier model beating the flagship at coding for half the price rewrites the usual rules of AI pricing.
For teams running large volumes of automated coding tasks, the combination of higher benchmark scores and lower per-token costs could translate into meaningful budget relief — at least on paper.
The Catch: Token Consumption
An independent tester who evaluates AI models found a complication that undercuts the simple price story. According to the tester's measurements, Sonnet 5.5 burns through more tokens than any other model they have benchmarked, meaning it consumes more of the units that determine cost to complete comparable work.
Because AI usage is billed by the token, a model that is cheaper per token but uses far more of them can end up costing about the same — or even more — depending on the task. The real-world price advantage therefore depends heavily on how efficiently the model handles a given workload.
- Sonnet 5.5 beat Opus 5.5 on Terminal-Bench 4.0
- It costs roughly half as much per token as the flagship
- Independent testing flagged unusually high token consumption
The takeaway for potential users is that headline pricing and benchmark scores tell only part of the story. Anyone weighing Sonnet 5.5 against rival models — or against Anthropic's own lineup — will need to factor in token efficiency to understand the true cost of running it at scale.
