Just now, the king of cost-performance, Sonnet 5.5, was released.
Sure enough, Sonnet 5.5 is really here!
The day before the opening of OpenAI's annual developer conference, Anthropic dropped a blockbuster overnight — Claude Sonnet 5.5.
As the perfect partner of Opus 5.5, the mission of Sonnet 5.5 is to become the "all-around work assistant" that is the fastest, most cost-effective, and most familiar with your daily work.
Compared with Sonnet 5, it runs 30% faster, and the cost of completing the same task is reduced by up to 30%.
In addition, its price is only half of the large-sized Opus 5.5, while its performance is almost on par.
Not only can it write code at an extremely fast speed, but it has even become the first Sonnet model in history that can beat the game Pokémon Red only by looking at the screen screenshots.
More impressively, in the extremely hardcore agent programming test Terminal-Bench 4.0, this mid-sized model even outperformed the large-sized one!
Sonnet 5.5 scored as high as 70.6%, which is nearly 7 times the original score, and even higher than the 66.4% of Opus 5.5.
Within one generation iteration, it directly rushed from the bottom to the top of its own product line.
Shortly after its launch, the authoritative third-party evaluation agency Artificial Analysis's intelligence index list was completely dominated: the top three positions were all taken by Claude products, and Sonnet 5.5 ranked second.
And the third place, the extra-large-sized Fable 5.1 launched by Anthropic four weeks ago, is priced at a full 5 times that of Sonnet 5.5!
In just one month, the large-sized level Claude products have directly dropped to extremely affordable prices, sold at the price of mid-tier models.
Moreover, users can freely adjust the "Effort Level".
Select Low/Medium: Claude responds extremely fast, consumes very few Tokens, which is suitable for casual use and routine tasks.
Select High/Maximum: Claude will enter the thinking mode, double-check its work repeatedly, and focus on solving complex difficult problems.
Sonnet 5.5 breaks the "Impossible Triangle"
This time, the release of Sonnet 5.5 directly broke the impossible triangle of "fast speed, low cost and high quality".
Its output speed has increased by more than 30% compared with the previous generation, making it undisputed the fastest Sonnet model to date.
The pricing of Sonnet 5.5 is nominally the same as that of Sonnet 5: $2 per million input Tokens, $10 per million output Tokens, and $0.20 per million Tokens for cache reading.
But the trick lies in efficiency!
Anthropic found that as the model becomes smarter, the number of Tokens required by Sonnet 5.5 is greatly reduced, and the actual cost of each task is reduced by up to 30%.
It can accomplish more things with fewer Tokens.
The following are the results of Sonnet 5 and Sonnet 5.5 under the same prompt, with a very clear contrast.
Half the price to match Opus, even the top-tier self-owned model is surpassed
On the Artificial Analysis list, Sonnet 5.5 scored 64%, higher than the 60% of Opus 5.5 and GPT-6 Astra.
On Terminal-Bench-Science, it scored 53%, ranking third, only losing to GPT-6 Astra and Opus 5.5.
It also performed extremely well in the other two code writing tests.
In the FrontierCode test, Sonnet 5.5 scored up to 52.1%, easily exceeding OpenAI's GPT-6 Sol (49.3%).
On CursorBench, which simulates real Cursor programming sessions, Sonnet 5.5 scored 55.5%, only 2 percentage points lower than the large-sized model, while the previous generation only scored 34.1%.
Developers in the early beta test said: "Its speed of understanding huge code bases is incredibly fast!"
Its tool invocation efficiency is also amazing: in back-to-back tests, Sonnet 5.5 learned to "package" tool invocations for batch processing, with fewer steps, fewer errors, and lower costs.
Not only code writing, but its office capabilities are also on par with top models.
GDPval-AA measures the model's ability to directly deliver finished products across 44 occupations. Sonnet 5.5 scored 1844 points, and the large-sized Opus 5.5 scored 1846 points.
In the AA-Briefcase test that examines long-form knowledge work, Sonnet 5.5 scored 1811 points, almost tied with Opus 5.5's 1822 points, and more than 300 points higher than GPT-6 Sol.
The most surprising part is its demonstrated "design talent".
This time, Sonnet 5.5 understands aesthetics! It can actively beautify UI, and even perfectly follow slide templates to generate PPT that hardly needs human modification.
There is a hardcore case inside Anthropic: testers threw Sonnet 5.5 the quarterly financial report materials of a listed company, conference call records, and a PPT template, asking it to generate a 10-page operation review slide deck.
As a result, the first draft it generated was identified by two human experts as: "Perfect, can be sent directly."
In the "final human exam" with tools, it scored 64.5%, nearly 10 percentage points higher than the previous generation.
On another third-party list Vals Index, Sonnet 5.5 ranked second on its first appearance, only 0.47 points lower than Opus 5.5.
Four weeks ago, when Fable 5.1 was just released, it was still the world's strongest model for programming and knowledge work.
Now, Sonnet 5.5 is nearly 15 percentage points higher than it on Terminal-Bench 4.0, 109 points higher in GDPval-AA, and 3 points higher in intelligence index.
Tech blogger @synthwavedd posted on the release day that Anthropic is launching new products at an extremely fast pace.
In the complex office test Box, the overall accuracy of Sonnet 5.5 reached 65%, while that of Sonnet 5 was 61%.
Its speed of delivering completed results is also about 2.4 times faster, and the total token consumption is reduced by 12%. In agent work, speed and accuracy usually restrict each other, so a model that can take both into account at the same time is a real tangible progress.
The San Francisco it built can automatically dispatch fire trucks when a fire breaks out
Actual tests show that Sonnet 5.5 is incredibly powerful.
AI blogger Matthew Berman got the early access qualification for Sonnet 5.5, and released a series of practical tests on the release day.
He asked Sonnet 5.5 to build a city of San Francisco from scratch in Unreal Engine. The streets are full of flowing traffic, and there are even cyclists passing through the intersections.
Input the sentence "The Transamerica Pyramid is on fire", the system immediately rated the fire as level 4.6 (full score 10), and then independently dispatched 5 fire trucks and 3 police cars.
Thick black smoke rose from the building in the distance, and the police cars with flashing lights arrived at the intersection and set up a cordon on the spot.
Even the Starship V3 was made into an interactive exploded view by it. The entire spacecraft is disassembled layer by layer, all the way down to the Raptor 3 engine at the bottom.
He also used Claude Sonnet to create an online Lego game.
Matthew Berman lamented that 3D simulation has been conquered by Claude.
But some people expressed different opinions.
Robbert van Empel maintains a LEGO building block planet challenge, which requires creating an interactive game with a South Park plot:
Create a 3D world as a globe with a South Park theme. I want to be able to walk through the world using WSAD, interact via 'e', and jump using the spacebar. I also want to be able to rotate the mouse and move objects. I want the world to be made of Lego bricks. Also, make it a game where the characters make typical South Park jokes, and the game must really have a South Park storyline.
Claude Sonnet 5.5 was tested this time, but it failed: it could not build the entire world at one time. While GPT Luna succeeded in doing so.