← Back to Blog
29 September 2026 /// AI / CLAUDE / LLM

Claude Sonnet 5.5 Arrives: 70.6% on Agentic Coding, 30% Faster, Up to 30% Cheaper Per Task

Anthropic's new Claude Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 where Sonnet 5 scored 10.3%, runs 30%+ faster and costs up to 30% less per task — at exactly the same $2/$10 token price.

By Guruji Corporation
Bar chart of Terminal-Bench 4.0 agentic coding scores: Sonnet 5 at 10.3%, Opus 5.5 at 66.4%, Sonnet 5.5 at 70.6%
Original graphic by Guruji Corporation.

The workhorse of the Claude family just went from 10.3% to 70.6% on the same coding test — and its price tag did not move.

On 28 September 2026 Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family and the cheaper, faster sibling to the flagship Opus 5.5 that arrived on 22 September [2]. It is not a new frontier model — Anthropic is explicit that it does not push the capability ceiling — but on the benchmarks that measure everyday work it is a large jump, and on cost per task it is the rare release that gets cheaper instead of dearer.

What actually happened

The headline number. On Terminal-Bench 4.0, an evaluation where a model has to complete multi-step professional tasks inside a command-line environment, Sonnet 5.5 scores 70.6%. Sonnet 5 scored 10.3% on the same test. Anthropic's own published table also puts Opus 5.5 — a model that costs twice as much per token — at 66.4%, so the mid-tier model now tops the family on this particular test [1].

It is a broad upgrade, not one lucky score. Anthropic's numbers show Sonnet 5.5 at 55.5% on CursorBench 4.0 (built from real coding sessions) against Sonnet 5's 34.1%, 46.2% on FrontierCode 1.1 at maximum effort against 42.4%, and 80.1% partial on the OSWorld 2.1 computer-use test against 57.0%. On GDPval-AA, which scores real-world tasks across 44 occupations, it reaches 1,844 — roughly 400 points above Sonnet 5 and two points below Opus 5.5 [1].

Same price, lower bill. Token pricing is unchanged from Sonnet 5: $2 per million input tokens, $10 per million output, $0.20 per million cache reads. Because the model typically needs far fewer tokens to finish a job, Anthropic says real tasks cost up to 30% less than on its predecessor [1]. Output generation is more than 30% faster, which the company calls its quickest Sonnet yet [1].

Who it is for. Sonnet 5.5 is positioned as the model for well-scoped, everyday work: fixing bugs, producing documents, slides and spreadsheets, and iterating quickly on less complex problems. Opus 5.5 ($4 input / $20 output per million) stays the recommendation for complex work that needs sustained judgement — Anthropic is careful to say that gap has not closed [1][2]. A Claude Haiku 5.5 for high-volume, cost-sensitive workloads is promised for the coming weeks [1].

Availability and migration. The model is live on the Claude Platform under the ID claude-sonnet-5-5, plus Amazon Web Services, Google Cloud and Microsoft Azure, with zero data retention [1]. One migration gotcha matters for anyone automating with the API: if you currently run Sonnet with thinking disabled, you must move to the new between_tools setting before switching models, otherwise the behaviour you relied on changes [3]. Inside Claude apps and Claude Code the default effort level is Medium, on the Claude Platform it is High, and dropping the effort level is what buys the biggest cost saving on routine work [1].

Safety controls moved down a tier. Sonnet 5.5 is the first Sonnet to ship with cybersecurity safeguards and fallbacks of the kind Anthropic reserves for its most capable models: its cyber capabilities are described as comparable to Opus 5, and higher-risk security requests visibly fall back to Sonnet 5 while ordinary bug-fixing work is unaffected [1]. Anthropic has also opened an expanded Cyber Verification Program so qualified defenders can apply for tiered access to the stronger capabilities [4]. It is likewise the first Sonnet released with classifiers designed to block reasoning extraction through distillation attacks, alongside expanded preserved-thinking controls [1]. Its biology safeguards are unchanged from Sonnet 5, aimed at a narrow band of genuinely high-risk requests [1].

Why this matters for Indian SMBs and mid-market teams

For most Indian businesses the relevant number is not 70.6% — it is that the bill per finished job goes down while the price list stays where it was.

  • Cost per task, not cost per token, is what hits the invoice. An SMB running document drafting, invoice parsing, ticket triage or report generation buys the same $2/$10 tokens and gets roughly 30% more done with them. That is a budget line that improves without renegotiating anything [1].
  • Latency is a product feature. If your AI touches a customer — a WhatsApp intake assistant, a support responder, an internal search bot — a 30%+ faster output speed is felt directly as response time, at no extra charge [1].
  • The capability gap between mid-tier and flagship pricing narrowed sharply. A 70.6% agentic coding score at half the token price of Opus 5.5 means a lean engineering team can run agent-style code tasks, codebase navigation and test fixes on a mid-tier budget. That is the difference between "we experiment with AI" and "we run it in the pipeline".
  • Do not swap blindly, though. Anthropic still rates Opus 5.5 as clearly stronger on complex, open-ended work needing long judgement calls, and the new cyber guardrails mean high-risk security tasks quietly downgrade to the older model [1]. Anyone migrating needs to test on their own workload — an effort-level or model change can alter output quality in ways a benchmark average hides.

The practical move before switching: run your real task set through both models for a week, compare token spend and quality on your documents, then decide. Chasing a benchmark is not a migration strategy.

Where Guruji Corporation fits in

This is exactly the kind of change that is easy to read about and hard to act on. We build AI-assisted workflows — document pipelines, support automation, internal assistants — and a model upgrade touches every one of them: prompt behaviour, token spend, latency, and the guardrails around what the system is allowed to do. We benchmark on your own tasks before recommending a switch, wire the migration (including details like the between_tools change) so nothing silently degrades, and put usage dashboards in place so the claimed 30% saving shows up as a number you can check at the end of the month.

If you are already paying for an AI model in production, the cheapest hour you can spend this week is seeing whether this release lets you cut the cost of the work you already do. Talk to us and we will tell you straight whether upgrading is worth it for your workload — or whether it is not.


Sources

Want to see how we can build your project?

Send us your project outline or chat directly on WhatsApp.