HomeArticle

"Prices are slashed to rock-bottom levels", Claude Haiku 5.5 pushes the competition of small AI models into the 1-cent era.

36氪的朋友们2026-10-08 08:25
Small Model Price Race

It has been almost a year since the release of the previous generation Haiku, and Anthropic has finally updated its smallest product line.

On October 7 local time in the United States, Claude Haiku 5.5 was officially launched. The most notable highlight is its pricing: only $0.1 per million input tokens and $0.5 per million output tokens, which is directly one-tenth of the price of the previous generation Haiku 4.5, and on par with GPT-6 Luna released by OpenAI last month.

But this time, Anthropic is clearly not just joining a price war.

As a model positioned for lighter weight and faster speed, Haiku 5.5 has this time made up for the shortcomings of its predecessor in programming, computer operation and knowledge work scenarios.

Meanwhile, it also adds an adjustable mode, supports a 1 million-token context window, brings lower cache prices for Sonnet and API credits for subscribers. This release is more like an across-the-board price adjustment from Anthropic.

The competition for small language models has moved from the stage of competing for lower prices to the stage of competing for both low cost and high usability.

01

Small Models Are No Longer "Barely Usable"

Haiku 4.5 was released nearly a year ago. At that time, it was already Anthropic's small model optimized for speed and cost, but in today's new generation of model competition, its shortcomings have become increasingly obvious. It only scored 735 points in the GDPval-AA knowledge work test, 15.7% in the OSWorld computer operation test, and even 0 points in the Terminal-Bench programming test.

Haiku 5.5 has almost completely addressed all these shortcomings.

In tests covering knowledge work, computer operation and programming, Haiku 5.5 has seen significant improvements compared to the previous generation. Among them, it scored 1620 points in GDPval-AA v2.1, more than double that of Haiku 4.5 and exceeding the 1437 points of GPT-6 Luna.

In the AA-Briefcase knowledge work test, Haiku 5.5 reached 1578 points, also higher than Luna's 1336 points. It scored 72.4% on the OSWorld 2.1 offline subset, compared with 48.9% for Luna and only 15.7% for Haiku 4.5.

The improvement in programming capability is particularly notable. In Terminal-Bench 4.0, Haiku 4.5 previously scored 0%, while Haiku 5.5 has been boosted to 39.2%.

Artificial Analysis, a third-party evaluation institution, gave it an intelligence index of 43 points, 26 points higher than Haiku 4.5, slightly higher than 42 points of GLM-5.3 Flash and 41 points of Gemini 3.8 Flash, on par with Kimi K3's 44 points, and lower than its own Sonnet 5.5's 56 points.

In intelligent knowledge work scenarios, Haiku 5.5 (Max mode) reached 1578 Elo on AA-Briefcase, leading models including Kimi K3 and GLM-5.3, and comparable to Muse Spark 1.3 (Max mode).

However, a small model is still a small model after all, and there are still gaps in factual knowledge. Its AA-Omniscience accuracy is only 36%, compared with 55% for Gemini 3.8 Flash and 44% for GPT-6 Luna.

There is another test where Haiku 5.5 did not perform well: it only got 35% on AutomationBench-AA, while Luna, Gemini 3.8 Flash and GLM-5.3 Flash all scored between 53% and 60%. Artificial Analysis pointed out that this may be caused by the model's over-rejection due to safety policies, and Anthropic is working on fixes that are expected to raise the score after implementation.

In addition, Haiku 5.5 is the first Haiku model that supports adjustable mode settings. It uses Medium mode by default, and offers five tiers in total: Low, Medium, High, xhigh and Max. The higher the mode tier, the more computing resources the model will invest, which usually brings better performance, but the token consumption will also increase accordingly.

02

Cheap Enough to Use Freely

Parameters and benchmark scores are one thing, actual usage experience is another. On the day Haiku 5.5 was released, many developers tried it out and shared various interesting test results.

Independent developer Simon Willison tested it by generating an SVG image of a pelican riding a bicycle. In Low mode, Haiku 5.5 spent 0.0936 cents in 7 seconds, and the bicycle frame of the pelican was relatively complete. In Max mode, it took 5 minutes and 9 seconds, but only cost 3.3826 cents.

In the generated description, the white pelican wears a red hat, its long orange legs step on the orange pedals, one gray wing stretches out to hold the handlebar, its large beak points forward, there are white speed lines behind it, and the sun and clouds above its head. For comparison, the pelican drawn by Haiku 4.5 a year ago was "round and tan in body, pink in beak", which was basically incorrect.

AI analyst @NFT_Chen conducted a comparative test of SVG generation for zebras on the grassland.

The zebra generated by Haiku has stripes that follow the direction of its body, its front and rear legs are staggered when running, the nearby grass has layers, the distant trees are pressed on the ridge, and there is dust and shadow under the zebra's feet as it steps into the grass. The zebra generated by Luna, by contrast, "has messy lines piercing through its belly, its belly is bulged into an ellipse, and its legs poke out like random sticks", the background trees look like pasted silhouettes, and the zebra floats in the middle of the frame, misaligned with the grassland.

@NFT_Chen concluded: "For the same grassland scene, one zebra is running, the other is floating. Competition makes the world a better place, provided that the drawing is correct first!"

But Haiku did not win all the tests.

The AI benchmark evaluation account @bridgemindai did a lava lamp test, and Haiku 5.5 "completely failed to pass": the generated result "has no glass, no lava, just a yellow cylinder". It took 4 minutes and 34 seconds, costing $0.03. In comparison, GPT-6 Luna produced a real lava lamp in 42.8 seconds, costing less than 1 cent, and was more than 6 times faster.

@bridgemindai's rocket launch test also got similar results. Haiku spent $0.05 and 6 minutes and 58 seconds building a clean scene, while Luna completed it for less than 1 cent in 1 minute and 5 seconds, 6 times faster.

Feedback from enterprise customers is more focused on real-world scenarios. Aaron Vinh, Senior Software Engineer at Asana, said they deployed Haiku 5.5 on their AI Teammates product to handle tasks such as bug classification, project setup, and searching across large project portfolios, which reduced task completion latency by more than 30% and increased the inference speed per agent round by up to 2.5 times.

Yashodha Bhavnani, Vice President of AI Products at Box, stated that Haiku 5.5 scores 11 points higher than Haiku 4.5 with only about half the latency, making it suitable for large-scale analytical work such as cost reporting, financial summarization, and weekly regular reviews.

03

The Logic Behind the 90% Price Cut

The most eye-catching part of Haiku 5.5 is still its pricing. For requests within 100,000 tokens, the price is $0.1 per million input tokens and $0.5 per million output tokens, only $0.01 per million tokens for cache reads, and $0.125 per million tokens for cache writes. Compared with Haiku 4.5's $1 for input and $5 for output, the price cut reaches 90%.

However, there is a threshold here: after exceeding 100,000 tokens, the price of Haiku 5.5 rises 5 times: $0.5 for input, $2.5 for output, $0.05 for cache reads, and $0.625 for cache writes. Even so, it is still 50% cheaper than the previous generation model.

Anthropic explained that about 90% of Haiku 4.5's requests are under 100,000 tokens, so for most users, it is equivalent to a direct 10% of the original price. Considering that the new tokenizer uses about 25% more tokens, the overall average cost saving is about 75%.

In addition to the price cut of Haiku itself, Anthropic also adjusted the price of Sonnet 5.5. The cache read price for Sonnet was cut from $0.2 per million tokens to $0.1, exactly half of the original. Since cache reads account for a large proportion of agent work, Anthropic estimates that the cost of most agent tasks can be reduced by about 20%.

There is also a new benefit of API credits for subscribers: $100 per month for Max 5x users, $200 per month for Max 20x users, and up to $500 per month for Team subscribers (shared across the team). The credits can be used for any model on the platform, and their total value is exactly equal to the subscription fee, which is equivalent to a full refund of the subscription fee in the form of API quota. Willison commented that this is "really generous", which greatly lowers the threshold for subscribers to use the API. The credits are valid for the current month and do not roll over to the next month.

04

The Small Model Battle Has Officially Started

The release of Haiku 5.5 has pushed the competition in the small model track to white-hot intensity.

Prior to this, OpenAI's GPT-6 Luna, with its low price of $0.1 and good performance, had almost no rivals in the small model field. Now Anthropic has directly matched its price, and has a slight edge in performance.

But if we look at the global market, the situation is more complicated. Several small models from China have even lower prices. Qwen3.7 Flash on Alibaba International Station charges only $0.03 for input and $0.13 for output within 32,000 tokens; between 32,000 and 256,000 tokens, the price is $0.1 for input and $0.4 for output, which is cheaper than Haiku. Zhipu's GLM-5.3-Flash scored 1647 points on GDPval-AA v2.1, slightly higher than Haiku 5.5's 1620 points.

Willison mentioned that when Haiku 4.5 was first released, its price was not expensive, but a year later, competing products have cut their prices to one-tenth of its original level. This time Haiku 5.5 matched Luna's price, which is more like a passive response rather than an active lead.

From a positioning perspective, Anthropic has assigned a clear role to Haiku 5.5: to handle all kinds of miscellaneous tasks. Document summarization, information classification, database query, acting as a sub-agent for large models, real-time customer service, browser operations, these tasks with large volume, high speed sensitivity and no extremely high requirement for ultimate accuracy are Haiku's main battlefield. Complex coding work is still left to Sonnet and Opus.

Frédéric Lardinois, a reporter from TechCrunch, pointed out that decision-making models (such as Jev) are also disrupting this price range with even lower prices. The use cases of small models are expanding from traditional summarization and classification to various lightweight tasks in agent workflows.

For developers, the real test is whether Haiku 5.5 can stably and reliably complete tasks with clear goals and well-defined boundaries, and withstand repeated calls. The price has been cut so much that the threshold for usage has indeed been greatly lowered, but there is still