DeepSeek Officially Announces Steep Price Cuts, With Discounts of Up to 60%
On Tuesday evening, the DeepSeek Open Platform released an announcement, planning to adjust the pricing of its large model APIs once again, and this time it is a widely welcomed price cut.
The unified off-peak pricing is as follows: Input (cache hit) 0.02 yuan, Input (cache miss) 1 yuan, Output 4 yuan (all units are yuan per million tokens). The peak-hour pricing remains twice that of off-peak hours, i.e. 0.04 yuan for cache hit, 2 yuan for cache miss, and 8 yuan for output.
This adjustment only applies to the V4-Flash series, while the pricing of the higher-positioned V4-Pro remains unchanged. The new pricing will take effect at 12:00 on September 10, Beijing Time.
The official also clarified that the peak hours are 9:00–12:00 and 14:00–18:00 from Monday to Friday Beijing Time, and the rest are off-peak hours.
Previously, DeepSeek's price increase sparked widespread public discussion. The official explanation was to allocate resources more reasonably and encourage users to adjust their task execution time. The general public believes that this is because DeepSeek V4 Flash 0731 has extremely high cost-effectiveness, attracting a large number of calls from users around the world, and DeepSeek does not have sufficient computing power.
For longitudinal comparison, before the price increase at zero hour on August 17, the unified pricing of Flash was 0.02 yuan (cache hit) / 1 yuan (cache miss) / 2 yuan (output). Putting the new pricing into comparison:
This means that the two prices on the input side have returned to the level before the price increase, while the output price is still twice as high as before the price increase. The gap is even larger during peak hours: the new peak pricing is 0.04 / 2 / 8 yuan, which is 2 times, 2 times and 4 times the previous pricing of 0.02 / 1 / 2 yuan respectively.
Therefore, whether users can truly return to the "unrestricted usage" state in the past depends on three variables of their own applications: cache hit rate, input-output ratio, and the degree to which tasks can be scheduled for off-peak execution. For scenarios with high cache hit rates such as long contexts, fixed system prompts, and Agent loop calls, the price reduction will be very significant; for scenarios dominated by generation output, the impact will be limited. As is known to all, the cache hit rate of DeepSeek API is relatively high, so the current public feedback is relatively optimistic.
This price cut is not an isolated event. Yesterday afternoon, DeepSeek announced in its official communication group that the intermediate version of V4.1 Flash has started internal testing: the model name is deepseek-v4.1-flash-expires-on-0910, the base_url remains unchanged, the rate limit is 20 concurrent requests per account, and the billing is temporarily the same as that of V4 Flash.
In addition to native multimodality and extremely fast inference speed, the official also stated that the new model has lower costs through architecture upgrades. It can be seen that through engineering and technical optimization of the large model itself, DeepSeek has been continuously trying to solve the problem of inference cost.
The model name expires-on-0910 indicates that the new model will automatically expire and go offline on September 10 — the same day the price cut takes effect.
This month, leading AI companies have successively updated their large models. If the official release of DeepSeek's V4.1 can bring a significant capability leap, together with this price cut, it may set off another big wave.
This article is from the WeChat Official Account "Machine Heart" (ID: almosthuman2014), author: Machine Heart focusing on large models, edited by the Editorial Department of Machine Heart, published by 36Kr with authorization.