Claude Sonnet 5.5 made a surprise late-night debut, outperforms Opus at half the price, and developers have seen a sharp surge in their actual measured usage bills.
According to reports from Zhidx on September 29, in the early hours of today, Anthropic released the second model in the Claude 5.5 series, Claude Sonnet 5.5, which features high speed and excellent cost performance.
Anthropic disclosed in its blog that compared with Claude Sonnet 5, the new model has a speed improvement of over 30% and a cost reduction of nearly 30%, and it is better at processing daily tasks such as fixing errors, creating documents, making slides and spreadsheets.
First, let's look at several key indicators. Sonnet 5.5 is comparable to Opus 5.5 in multiple metrics, with performance several times higher than Sonnet 5. Anthropic claims that this is the first Sonnet model that can win the game *Pokémon Red* only through screenshots. On Artificial Analysis, Sonnet 5.5 scored 56 on the AI Analysis Index, second only to Opus 5.5's 58 points, and the Claude series models take the top three spots on this list.
In terms of pricing, Claude Sonnet 5.5 is priced the same as the previous generation, which is $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache read operations. However, the blog mentions that when the model actually completes tasks, the cost per task is 30% lower than that of Claude Sonnet 5.
The official released case shows that when Sonnet 5.5 restores a Rubik's Cube, it only takes 10 seconds and costs $0.11 to complete the task.
The Artificial Analysis list shows that Sonnet 5.5 outputs more tokens per task, with a per-task cost of $7.60, about 50% higher than the per-task cost of Sonnet 5, which is close to Opus 5.5's $7.63.
Many developers found in actual tests that the cost of Sonnet 5.5 is very high, and "it is recommended to stick to Opus 5.5".
Claude Sonnet 5.5 is now available on all platforms, including AWS, Google Cloud and Microsoft Azure. Developers can access and use it on the Claude platform through the model identifier claude-sonnet-5-5. If you previously disabled the thinking function on the Sonnet model, before migrating to Sonnet 5.5, you need to switch to the new between_tools parameter configuration, which can keep the pre-thinking function disabled.
Anthropic also revealed that Claude Haiku 5.5, which is specially designed for high-concurrency and cost-sensitive applications, will be released in the coming weeks.
01. Multiple developers' actual test comparison: The effect is amazing but unaffordable, a shooting game costs more than a thousand yuan
Many developers have tested and compared Sonnet 5 and Sonnet 5.5, and the performance improvement between the two generations of models is visible to the naked eye.
The well-known developer @_re_pete compared the effects of Sonnet 5 and Sonnet 5.5 generating a falling leaves simulator. It can be seen that the falling leaf effect generated by Sonnet 5.5 is more natural, and the overall picture is more harmonious.
In less than a minute, developer @kevin_t_ngo asked Sonnet 5.5 to show the history of dinosaurs through JavaScript animation, and he also compared the generation results of Sonnet 5 and Sonnet 5.5. The animation generated by Sonnet 5.5 has a clear theme image, can add supporting text, and generate complete illustrations with a sense of narrative.
Another developer compared the effects of Sonnet 5.5 with GPT-6 Sol and GPT-6 Astra, and the gap between Sonnet 5.5 and GPT-6 Astra is not large. Sonnet 5.5 and GPT-6 Astra do not have major flaws when generating the action of riding a bicycle.
However, behind the amazing effect, users found in actual tests that the cost of Sonnet 5.5 is not low.
BridgeMind used Sonnet 5.5 to make a shooting game. It said the new model performed excellently, and the generated game is a masterpiece, but the construction cost reached $177, which took 49 minutes to complete.
In the following city simulator case made by the developer, the effect of Sonnet 5.5 is significantly better than that of Sonnet 5. Sonnet 5.5 outputs 380,000 tokens at a cost of $9.23, while Sonnet 5 outputs 120,000 tokens at a cost of $4.85.
Another developer used Sonnet 5.5 to make a jam watermelon browser animation, and the final API usage cost was about $8.20, of which $2.80 for code and reasoning, $3.6 for cached context data, $1.8 for cache write operations, and $0.02 for regular input tokens, 2/3 of the resources are used to process context-related tasks.
02. Multiple tests show it is comparable to Opus 5.5, Anthropic suggests developers adjust gears to reduce costs
Judging from the benchmark tests released by Anthropic, Sonnet 5.5 has surpassed the Opus 5.5 model with higher performance in multiple indicators.
Specifically, this benchmark test compares the performance of the four models Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol in agent programming, knowledge work, multidisciplinary reasoning, computer operation, chart recognition and other aspects.
Compared with the previous generation Sonnet 5, Sonnet 5.5 has scores increased several times in programming and chart recognition; compared with GPT-6 Sol, Sonnet 5.5 is better in agent programming, knowledge work and chart recognition; Sonnet 5.5's score in the active programming field even exceeds that of Claude Opus 5.5, and the other items are comparable to it.
The figure below shows the comparison curve of the score of each model under different input levels and the cost per task. The closer the data point is to the upper left corner of the chart, the stronger the model capability that can be obtained for every dollar spent.
In multiple benchmark tests, Sonnet 5.5 under low/medium input gears can exceed the highest score of Sonnet 5, while the per-task cost is only about 1/10 of the latter. When Sonnet 5.5 runs at a lower input gear, the per-task cost is lower, which forms a complement with Opus 5.5; after raising the input gear, it can reach performance close to Opus, while the cost is at a similar level.
Programming is a prominent scenario where Sonnet 5.5's performance is improved. Under the high input gear of the agent programming benchmark FrontierCode, its score is about 10 points higher than that of Sonnet 5 at the same gear, and the per-task cost is only about 1/15 of the latter.
CursorBench conducts tests based on Cursor's real programming session tasks. On this benchmark, the highest score of Sonnet 5.5 is only about 2 points away from Opus 5.5.
Many early testers said that in direct comparison tests, Sonnet 5.5 can merge tool calls together, thereby reducing operation steps and lowering costs.
Sonnet 5.5 has achieved performance improvements on multiple knowledge work tasks. The GDPval-AA benchmark covers real business tasks in 44 occupations and 9 major industries. In this test, Sonnet 5.5's score is almost the same as Opus 5.5, about 400 points higher than Sonnet 5. In terms of computer operation and chart recognition capabilities, its performance is also very close to Opus 5.5; while in long-cycle knowledge work scenarios, its performance is significantly better than Sonnet 5 and GPT-6 Sol.
In terms of pricing, developers can adjust the input gears to trade off between cost, response speed and overall task quality. In Claude Code and Anthropic's official applications, the default input gear is Medium; the default gear on the developer platform Claude Platform is High.
Anthropic officially states that when the gear is lower, Claude responds faster and consumes fewer tokens, which is suitable for routine transaction processing; after the gear is raised, Claude will perform longer reasoning and conduct more comprehensive self-inspection and verification of the output results.
03. Most alignment indicators are comparable to Sonnet 5
In its automated behavior audit, Sonnet 5.5 is better than or at the same level as Sonnet 5 in most alignment indicators. Because its cybersecurity capability is comparable to Opus 5, it is the first Sonnet model with cybersecurity protection measures and fallback mechanisms. Its biosafety protection mechanism is the same as that of Sonnet 5. Both of these safety protection mechanisms are aimed at a small number of high-risk operations; conventional software development and most operations in the life sciences field are not affected.
Anthropic researchers revealed in the blog that Sonnet 5.5 does not break through the cutting-edge upper limit of its model capabilities, so they focus the model's alignment assessment on a set of specific risks applicable to models of all capability levels, including harming user interests, misleading users, and causing synergistic harm when used in high-risk malicious scenarios.
Researchers conducted automated behavior audits, testing Claude in about 1850 scenarios.
The results show that in most indicators of alignment, anti-abuse capability and honesty, Sonnet 5.5 has improved or remained flat compared with Sonnet 5. In its newly launched isolation protection assessment, the probability of Sonnet 5.5 trying to break through the sandbox is close to Opus 5.5, the best performer in this test; and among all Anthropic models, it has the lowest probability of testing the boundary of container protection. They did not find that Sonnet 5.5 will execute goals that conflict with user intentions.
In terms of cybersecurity, Sonnet 5.5's cybersecurity-related capabilities have been improved compared with Sonnet 5, so researchers have deployed a security protection mechanism similar to Opus 5.5 for it. Users can still use the model to find and fix code vulnerabilities in the conventional software development process;