HomeArticle

With a peak valuation of 500 billion yuan, DeepSeek and Kimi are being wildly sought after by investors.

凤凰网科技2026-08-05 15:46
China's large models are embarking on their next journey.

Abstract: The two fastest runners on their respective tracks have almost simultaneously opened the financing floodgates once again. Liang Wenfeng and Yang Zhilin have also become the two most popular fellow Guangdong natives this summer.

The primary market in August has rarely seen such an explosive surge of enthusiasm.

According to reports from Fortune China magazine, DeepSeek's second round of financing has been restarted, with a planned fundraising of 500 billion yuan and a pre-money valuation of about 5 trillion yuan. Another person involved in the transaction said that the actual scale of funds that expressed investment intentions around DeepSeek in the first round reached 1 trillion yuan. If the second round of financing goes smoothly, the total financing of DeepSeek's two rounds will reach exactly 100 billion yuan. At the same time, the Series G financing of Moonshot AI (Kimi) is also in progress, with a valuation of 500 billion US dollars (about 3.373 trillion yuan). After the launch of K3, the share of this round of financing has suddenly become extremely sought-after.

The two fastest players on their respective tracks have almost opened the financing floodgates at the same time.

There are many familiar faces in the shareholder lists of these two companies. Those who placed bets on both of them a year ago or even earlier have now become the most stable winners in this AI feast.

Who Is Betting On Both At The Same Time

Cao Xi is probably one of the most enviable VC partners this year.

This former HSG partner and founder of Monolith Capital holds two tickets to the AI era: one for DeepSeek and the other for Moonshot AI. What's more, he bought both of them very early and placed heavy bets on each.

The story dates back to a dinner at Tsinghua University in early 2022. At that time, large models had not yet become popular, and Yang Zhilin had not founded Moonshot AI. Cao Xi met this young man a few years younger than himself at the dinner, and the strongest feeling he had was two words: "unconvinced". Among many founders, Yang Zhilin is also highly recognizable. He had a dream of becoming a rock star in his early years, there is a piano in his office, and the company was named Moonshot because of his love for Pink Floyd's album.

After that meeting, Monolith Capital became one of the earliest investors in Moonshot AI. From the angel round to Round A and Round C, Cao Xi continuously increased his positions in multiple rounds. It can be said that Monolith Capital gave Yang Zhilin the most attention in the very early stage.

If betting on Moonshot AI depends on vision and early market positioning, then investing in DeepSeek is more like a proof of strength.

In June this year, DeepSeek completed its first round of external financing since its establishment, with a total size of 51 billion yuan, which can be called the largest single-round financing in the history of China's AI industry. The seats at the main table were almost all occupied by industry giants: Tencent invested 10 billion yuan, CATL invested 5 billion yuan, NetEase and JD each invested 3 billion yuan. There are very few spots left for VCs.

But Cao Xi got a seat. Monolith Capital and IDG Capital each took a share of about 3 billion yuan. Now Moonshot AI's valuation has rushed to 500 billion US dollars, and DeepSeek's pre-money valuation is 5 trillion yuan. The paper returns of Monolith Capital on these two investments have reached a figure that is widely admired by peers in the industry.

IDG Capital is another typical case. It placed heavy bets during the growth stage, led the Series C financing of Kimi, increased its stake in Series E, and later became the largest institutional shareholder. In DeepSeek's first round of financing, IDG invested another 3 billion yuan.

Compared with VCs, the bets placed by large tech giants have more strategic value.

Tencent's strategy is more like a strategist who casts a wide net. Tencent's 10 billion yuan contribution in DeepSeek's first round of financing makes it the largest external investor; and Tencent also took a seat in Moonshot AI's shareholder list very early, entering the game in the Series B round in August 2024, and continued to increase its investment thereafter.

Tencent's logic is very clear: do not put all eggs in one basket, which is not unprecedented in the mobile Internet era. In addition to social networking and games, Tencent has invested in half of the mobile Internet market. In the AI era, it has its own Hunyuan large model, and has simultaneously invested in many independent large model companies such as Moonshot AI, DeepSeek, MiniMax, and Zhipu AI.

The dual layout of the National Artificial Intelligence Industry Investment Fund is of greater vane significance. In DeepSeek's first round of financing, it is the only external institution that directly invested and obtained voting rights, contributing about 980 million yuan; in Moonshot AI's Series F round, it is also one of the co-lead investors. The dual bet of the national team, to some extent, is a recognition of the entire track.

If the double bets of institutions are out of the rational consideration of investment portfolios, then the cross-border bets of some traditional enterprises are more like an unexpected pleasant surprise.

Andon Health is the most typical example. This A-share company, which is famous for its home blood pressure monitors and blood glucose meters, has seen its stock price rise in the past six months in a way that many people cannot understand. Its semi-annual report stated that "the layout gains in the science and technology innovation investment sector have emerged, and the valuation of underlying targets has risen". In other words, the valuations of the AI companies it invested in have skyrocketed, bringing huge profits.

Andon Health's two core bets fall exactly on DeepSeek and Moonshot AI respectively. For DeepSeek, Andon Health Hong Kong contributed 750 million yuan to indirectly participate in the shareholding through Tianjin Shixiang Industrial Fund, with a penetrating shareholding of about 0.21%; for Moonshot AI, Andon Health made an even earlier move - it invested 10 million US dollars in August 2023, added 20 million US dollars in March 2024, and participated in the leading investment again in February 2026, with a total investment of 30 million US dollars.

At that time, Moonshot AI's valuation was only more than a dozen billion US dollars, and DeepSeek had not even officially opened financing to the public. Roughly calculated, the floating paper profit of Andon Health from these two investments alone is nearly 3 billion yuan, which is several times the annual profit of its main business.

Some people have thus given Andon Health a nickname - the most hidden mini Berkshire in A-share market. A medical device company has earned several years of profits from its main business in half a year relying on AI investment.

By-health took an even bolder approach. As the leading health supplement company, it poured a large amount of money earned from selling protein powder into the AI track. For DeepSeek, it indirectly holds 0.04% of the shares by contributing 130 million yuan through the Lisi Xingling Fund; for Moonshot AI, it made three moves within this year - subscribing to Series D preferred shares with 10 million US dollars in May, adding 5 million US dollars through Hong Kong Bairui in July, and announcing a planned additional investment of 70 million yuan at the end of July, with a total investment of about 175 million yuan.

Chinarion Co., Ltd. did not want to fall behind either. This Anhui listed company that makes luggage has indirectly participated in the investment of three large model companies including Zhipu AI, Moonshot AI, and DeepSeek through multiple funds such as Hangzhou Chengli and Lisi Xingling. Although the amount of each investment is small and the shareholding ratio is extremely low, the valuations of all three companies have risen - on the day the news was announced, its stock price hit the daily limit directly.

Cost-Effectiveness Slasher and Intelligence Challenger

If DeepSeek pried open the door to the global market with its price advantage, then Kimi K3 sounded the alarm for closed-source giants with its strength.

On the day the official version of DeepSeek V4 Flash went online, its popularity overseas continued to rise. Data from Openrouter shows that DeepSeek V4 Flash had 7.11 trillion Token calls in the past week, which directly pushed it to the top of the global ranking. All the top five are Chinese large models - domestic AI has completed a group leading run in the global developer market.

Why is it so popular? It is cheap, extremely cheap.

Data from US research institution Artificial Analysis shows that DeepSeek V4 Flash charges $0.14 per million input Tokens and $0.28 per million output Tokens, with an average cost of about $0.03 to complete each test. In contrast, the operating cost of OpenAI GPT-5.6 Sol is $1.86, and Anthropic Claude Fable 5 is even more expensive. After calculation, the price of DeepSeek is only a few tenths of that of the top closed-source models. However, this comparison is not entirely fair, after all, the intelligence levels of the two sides are not in the same gear.

But overseas developers still gave it a nickname - price slasher. Data from the OpenCode platform shows that from June 10 to August 4, in less than two months, the total usage of DeepSeek V4 Flash reached 90 trillion Tokens, while the total cost was only 920,000 US dollars - the comprehensive cost of 100 million Tokens is only 1.02 US dollars.

The evaluation of the overseas developer community is that we should thank Chinese open-source large models for forcing giants like OpenAI to cut prices and lower their stance.

But low price is only one side of DeepSeek. What really makes Silicon Valley nervous is that it is constantly optimizing its performance while being low-cost.

The official version of V4 Flash scored 50 points in the Artificial Analysis evaluation, second only to Zhipu GLM-5.2 in China. V4 Flash is only a lightweight version, and many benchmark tests even exceed the preview version of its own flagship V4 Pro - which means DeepSeek's optimization of model efficiency is still ongoing.

DeepSeek has even redefined the kill line of cost-effectiveness and the Pareto curve. It does not pursue to surpass closed-source giants in every indicator, but to push the price to the extreme under the premise of "being sufficient for use", and build barriers with scale effect. The end of this road is to become the "utility infrastructure" of the AI era - everyone uses it, and everyone can afford it.

And Kimi K3 chose another more radical path - directly challenging Anthropic at the intelligence level.

At the end of July, Moonshot AI fully open-sourced Kimi K3, with a total parameter count of 2.8 trillion and a context window of 1 million Tokens. As soon as this model was released, the entire Silicon Valley was shocked. On Artificial Analysis's Intelligence Index, Kimi K3 ranked third with 57 points, second only to Claude Fable 5 and GPT-5.6 Sol - this is the closest an open-source model has ever been to the top closed-source level in history.

"Kimi K3 has made Anthropic's trillion-dollar market cap meaningless." Comments like this are not uncommon in overseas developer communities.

What is more dramatic is the contrast between the attitudes of Nvidia and Anthropic. Jensen Huang personally opened an X account to endorse open-weight models, and hundreds of technology companies jointly urged the US government not to cut off access to Chinese open-source models. But Anthropic remained vigilant throughout the process, did not sign the joint letter, and even accused Kimi K3. Independent researchers believe that this accusation is untenable from the timeline point of view.

Zhang Tong, senior research director at Gartner, made an accurate distinction of the reversal moment brought by these two Chinese large models in a recent interview including interviews with Phoenix Tech.

"The DeepSeek moment essentially only changed the rhythm of OpenAI releasing models. At that time, R1 and O1 were at the same level, but OpenAI already had O3, which had not been released yet. After R1 appeared, it accelerated the release process of O3. That is to say, we did not really catch up with the most mainstream models," Zhang Tong said, "But it is different now. Less than half a month after Fable was released, K3 reached a considerable level comparable to it, proving that the generational gap is indeed narrowing greatly."

In Zhang Tong's view, the deeper change lies in customer mentality. "When DeepSeek appeared, most overseas enterprises were still afraid and dared not use it, thinking there were security risks. But after K3 came out, many enterprises have to agree: Chinese open-source models do have the potential to replace some expensive closed-source models. Especially customers in Europe and Southeast Asia, at the end of last year, they felt that they could not rely on the United States and needed to make new choices. K3 gave them this choice."

Zhang Tong told us that Kimi's technical report describes in detail many issues that need attention during the deployment process, which is of very high level. "DeepSeek did not cover so many details when it released its model. If it is an engineer-oriented team, after carefully reading these deployment details, it is easy to quickly master the technology through the technical report, or carry out local deployment on its own."

Two paths: one fights from the bottom up, spreading volume with extreme cost-effectiveness; the other charges from the top down, leveraging the enterprise market with top open-source capabilities. What they have in common is that they have both broken through in the overseas market.

The Next Exam For Chinese Large Models

Although K3 has set off a new wave of celebration, everyone knows that this is only a small victory. Zhang Tong told us that he believes Kimi K3's leading advantage window is about 3 months.

"No one wants to lose this competition. Domestic players including Zhipu and Alibaba are all training better models. There is a lot of personnel exchange, and technical features will be shared quickly," Zhang Tong judged, "Including Anthropic and OpenAI, they still have the best capabilities and the most GPUs, and they will not stop the competition."

3 months is roughly a model iteration cycle. In this cycle, the question Kimi needs to answer is not only "can it maintain its technological lead", but also "can it convert its technological advantage into commercial barriers".

Zhang Tong believes that Kimi's moat mainly lies in two places: one is data, Moonshot AI has done a lot of work on data engineering; the other is engineering capability, the technical recognition including the attention mechanism is the core advantage of Yang Zhilin's team.

"Kimi first proposed the concept of long text, which has now become a point that all models must pay attention to, because only long text can support more long-term tasks," Zhang Tong said.

DeepSeek is facing a different test. Its cost-effectiveness advantage has been established, but the official version of V4 Pro has not been released for a long time, and the original release time in mid-July has been delayed again and again. For a company famous for its fast technical iteration speed, this is somewhat unusual.

The deeper challenge is computing power. The common pain point of both companies is the shortage of GPUs.

"I think Kimi's biggest short board is still the shortage of GPUs and computing power," Zhang Tong said bluntly, "If Kimi had more GPUs, its K4 would definitely be launched faster." According to public information, Kimi has already started training K4, which is said to be a large model with 50,000 to 100,000 scale parameters. Just like the slogan Moonshot AI put up at its celebration party, "K3 expands and upgrades, K4 expands wildly, rushing to the moon".

This is not a problem unique to Kimi. Computing power is a major bottleneck for China's AI industry. Although domestic industries are catching up from storage chips to network equipment, when training the best models, Nvidia's GPUs are still unavoidable.

But Zhang Tong also saw a positive side. "I believe that in a few years, we may really see large models trained with China's own chips."

The exploration of business models is also equally critical.

Zhang Tong specially mentioned Kimi's licensing model - it learned from Meta