FALCONINTERNET

Claude Sonnet 5.5 Out-Codes Opus 5.5 at Half the Token Price

Artificial Intelligence
Claude Sonnet 5.5 Out-Codes Opus 5.5 at Half the Token Price

Anthropic shipped Claude Sonnet 5.5 on September 28, and the headline number is the one that surprised even close watchers of the API space: the model scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, against Sonnet 5's 10.3%. That is a different class of capability on the benchmark that matters most to developers who let AI models run shell sessions and drive multi-step tool chains. And it does this while keeping the same per-token price as Sonnet 5.

Terminal-Bench 4.0 and CursorBench results

Terminal-Bench 4.0 isn't a multiple-choice reasoning test. It measures a model's ability to complete agentic software engineering tasks in a real terminal: writing code, running it, reading failure output, fixing it, and iterating without a human in the loop. That is the work a Claude Code or Codex session does in production.

Sonnet 5 scored 10.3% on Terminal-Bench 4.0, roughly consistent with a tool that needs heavy prompting and supervision to complete these tasks reliably. Sonnet 5.5 scores 70.6%, which also clears Opus 5.5's 66.4% on the same benchmark. For the first time, the mid-tier model in the family outperforms the flagship on an agentic coding task at default effort settings.

On CursorBench 4.0, which recreates real-world Cursor IDE coding sessions from actual session data, Sonnet 5.5 scores 55.5%, up from Sonnet 5's 34.1%, and just two points below Opus 5.5's 57.8%. On GDPval-AA, a broader evaluation across realistic occupational tasks, Sonnet 5.5 lands two points below Opus 5.5. These numbers paint a consistent picture: for most structured work, the gap between the mid-tier and flagship model is now small enough to matter only at the margins.

Where the 30% per-task savings come from

The pricing announcement sounds straightforward: $2 per million input tokens and $10 per million output tokens, unchanged from Sonnet 5 and half the cost of Opus 5.5 ($4/$20). But Anthropic's claim that tasks cost up to 30% less than Sonnet 5 isn't about lower token prices. Those didn't move.

The savings come from two places: the model runs more than 30% faster, which reduces wall-clock time and increases throughput at the same cost, and it burns fewer tokens to complete the same task by reaching conclusions in fewer steps, with fewer retries, and (in agentic contexts) with fewer unnecessary tool calls. When your billing is effectively "tasks completed," efficiency like this translates directly into lower costs even with an unchanged rate card.

Anthropic is concrete about the magnitude: at medium effort (the default setting in the Claude apps and Claude Code), Sonnet 5.5 exceeds Sonnet 5's best score for less than a tenth of the cost per task. That changes what the mid-tier model can do at standard settings without extra investment.

At xhigh effort, Opus 5.5 is cheaper per task

There's a nuance that early reviews are glossing over. At maximum (xhigh) effort, the economics flip. Sonnet 5.5 at xhigh matches Opus 5.5's score, but the per-task cost works out to roughly $10.67 against $4.86 for Opus 5.5, because at the top of the effort scale, Sonnet burns far more tokens to match Opus's deliberate, careful reasoning.

The practical rule: route to Sonnet 5.5 at low-to-high effort for well-scoped tasks: bug fixes, document generation, agentic coding with clear success criteria, high-volume classification or extraction. Use Opus 5.5 when a task is ambiguous, long-running, or where a wrong answer costs more than a few extra cents per call. The token price tells you the rate; the effort setting determines the actual bill.

What changes for teams using Claude APIs

The model is available now under the identifier claude-sonnet-5-5 on the Anthropic API, Amazon Bedrock, Google Cloud Vertex AI, and Microsoft Azure AI. It ships with zero data retention (API calls are not logged for training), consistent with Sonnet 5 and Opus 5.5. The knowledge cutoff for Sonnet 5.5 is June 2026.

Teams currently pinned to claude-sonnet-5 will need to update their model identifiers. Anthropic does not automatically redirect version-pinned calls. If you're using the claude-sonnet-latest alias, you're already on it. Cache pricing for prompt caching remains at $0.20 per million for cache reads and $2.50 per million for cache writes.

For most API integrations this is a drop-in upgrade worth making soon. The main reason to hold is if your production system has been tuned against Sonnet 5's behavior and you haven't budgeted for regression testing. At 30%+ faster output, timing-sensitive pipelines may behave differently at the margins.

Agentic coding benchmarks versus general benchmarks

The industry has largely settled on a handful of benchmarks that measure reasoning, mathematics, and trivia-adjacent recall. Those are useful for comparing base capabilities, but for developers building software the more meaningful evaluations are the ones where the model runs code, observes failure, and fixes itself without handholding. Terminal-Bench 4.0 is that test, and a jump from 10.3% to 70.6% between consecutive Sonnet generations suggests Anthropic made a deliberate architectural bet on agentic reliability rather than general scaling.

Whether that holds across the full range of real-world codebases and tool environments is an open question. Benchmarks measure what benchmarks measure. But the CursorBench numbers, built from real session data rather than synthetic tasks, are harder to dismiss, and they trend the same direction.

At Falcon Internet, the models we use in our own tooling get evaluated primarily on one criterion: does the agent finish the job without supervision? By that measure, Sonnet 5.5 moved the bar considerably.

Need this handled instead of explained?

We do this for a living — talk to an engineer about your setup.