HomeArticle

Sonnet 5.5 is released, with its intelligence level surpassing GPT-6 Astra. Where is the promised deceleration?

字母AI2026-09-29 10:25
Sonnet 5.5 is released, Amodi claims to slow down in words but actually steps hard on the accelerator.

Amid, who has been repeatedly calling for deceleration, has slammed the throttle all the way down.

On September 26, 2026, Anthropic released Sonnet 5.5. It not only delivers performance close to Opus 5.5, but also costs only half of what Opus 5.5 does.

In the intelligence level ranking released by Artificial Analysis, Sonnet 5.5 even outperforms GPT-6 Astra as well as Anthropic's own flagship model Claude Fable 5.1.

The core advantage lies in cost-effectiveness. At present, Sonnet 5.5 handles almost all medium-difficulty tasks, especially in development, scientific research and mathematics. Excluding those extremely simple tasks, no other model can match Sonnet 5.5 in terms of cost-effectiveness.

So where is the promised deceleration?

Why is it so powerful yet so cheap?

Sonnet 5.5 boasts extremely strong performance. It scores 70.6% on Terminal-Bench 4.0, while the previous generation Sonnet 5 only scored 10.3%.

On GDPval-AA v2.1, it gets 1844 points, just 2 points short of Opus 5.5's 1846 points. On CursorBench 4.0, it reaches 55.5% against Opus 5.5's 57.8%, which is also very close to Opus 5.5.

Anthropic states that although Sonnet 5.5 has performance close to Opus 5.5, Opus 5.5 delivers better results when dealing with complex work that requires continuous judgment and has unclear objectives.

Sonnet 5.5 is more suitable for daily tasks, such as debugging, iterating functions, as well as creating documents, slides and spreadsheets.

Sonnet is the lightweight model in the Claude family, and its performance was supposed to lag far behind the flagship model. But why has Sonnet 5.5 suddenly become so capable?

According to public materials, the most noticeable change is not that it has memorized more knowledge, but that its ability to complete tasks has been greatly enhanced.

Starting from Sonnet 5, Anthropic has strengthened reasoning, tool usage and programming, making the model more willing to call tools and check results when working. By version 5.5, the improvements are mainly reflected in links requiring continuous actions, such as agent programming, computer operation and visual understanding.

Its score on OSWorld 2.1 has risen from 57.0% of the previous generation to 80.1%, and its score on the tool-free chart recognition test Chartography has increased from 15.6% to 61.6%.

In other words, Sonnet 5.5 has mastered the ability to collect information from the screen and then advance the task accordingly.

Take debugging as an example. Reading and understanding the code is only the first step. For the model to complete the entire debugging task, it also needs to find relevant files, judge where the changes will affect, call tools to make modifications, and then confirm whether the problem is solved.

If any step goes wrong, the whole task may fail.

This is why Sonnet 5.5's score on Terminal-Bench has increased significantly. It is not a simple surge in "coding IQ", but the information collection ability enables the model to complete the entire task chain.

One speculation is that Sonnet 5.5 may also be distilled from Opus 5.5. During testing, users found that Sonnet 5.5 would say some catchphrases that only Opus 5.5 would use, such as "You are right!"

When Anthropic released Opus 5.5, it explicitly stated that it had improved its writing and communication styles, with "You are right!" being the most representative. Fable 5.1 never uses this catchphrase, but Sonnet 5.5 has started to use it.

Then why is it cheap? Let's look at the pricing first: Sonnet 5.5 charges $2 per million input tokens and $10 per million output tokens, exactly half of Opus 5.5. The 5-minute cache write price is $2.50 against $5, and the cache read price is $0.20 for both.

The so-called 5-minute cache write refers to the behavior that when you send a piece of reusable prompt content through the API for the first time, you let the system cache it, and this write operation is billed at the "5-minute cache available" tier.

Meanwhile, Anthropic also states that "each task can be up to 30% cheaper", because completing the same task usually consumes fewer tokens. With the same price and fewer detours, the bill will naturally be lower.

Still take programming as an example.

Early testers observed that Sonnet 5.5 is more inclined to batch tool calls than Sonnet 5. This reduces the total number of steps, thus saving more tokens.

Anthropic claims that on the High tier of FrontierCode, it is about 10% higher than the previous generation at the same tier, while the cost per task is only 1/15 of the latter.

In some evaluations, using the Low or Medium tier of Sonnet 5.5 can exceed the best performance of Sonnet 5 at about 1/10 of the task cost.

The key to low cost does not only lie in the price per token, but also in how many tokens the model actually spends to complete the task and how many tool calls it bypasses.

In addition, the output speed of Sonnet 5.5 is more than 30% higher than that of Sonnet 5.

The following is the speed and performance comparison between Sonnet 5.5 and Sonnet 5.

What else is there beyond the model?

Compared with performance, Anthropic now pays more attention to engineering improvements.

For example, Fable 5.1 has introduced Preserved Thinking, a mechanism to prevent reasoning content extraction. Now, this mechanism has been added to Sonnet 5.5 and further optimized.

Anthropic has added a security classifier to Sonnet 5.5 to prevent reasoning extraction, and binds the thinking blocks it generates to the account that creates it or associated accounts.

If you replay the content with an unrelated account, the API will discard the thinking block, and the model cannot see the previous reasoning. However, ordinary developers do not need to worry about this protection system at all, because it does not affect your access to the entire thinking process. Anthropic's main goal is to prevent enterprises that want to distill the model on a large scale and call it repeatedly to extract the model's capabilities.

The signature of the thinking block is also bound to the previous conversation.

If developers modify the existing system prompts, tool definitions or historical messages, and then insert the old thinking block back, the new account may receive a 400 error by default.

Normal additional conversations are usually not affected, but those agent frameworks that are used to silently rewriting history in the background will need to reprocess the sessions.

The safety guardrails have also evolved from "Should I reject the user's excessive request?" to "Who should I assign this request to for response?"

Sonnet 5.5's cybersecurity capabilities have been improved, so high-risk cybersecurity requests may trigger the classifier and be explicitly rolled back to Sonnet 5.

Normal bug finding and debugging can still use Sonnet 5.5.

The server-side automatic rollback on Claude API is currently in beta and requires developers to enable it.

When Sonnet 5.5 encounters some requests blocked by the security system, it will not necessarily end the process directly; if the developer enables "automatic rollback", the system will try to switch to Sonnet 5 to respond.

This article is from the WeChat Official Account "Letter AI", author: Miao Zheng, published with authorization from 36Kr.