HomeArticle

Yang Zhilin is about to achieve a major achievement.

字母榜2026-09-03 08:42
The Second Coronation of the Tech Genius

On September 2, as reported by LatePost, Moonshot AI has submitted its A1 form to the Hong Kong Stock Exchange in confidential form this week. This is the official application form for enterprises applying for Hong Kong listing, marking that Moonshot AI has officially launched the Hong Kong stock IPO process.

AlphaRank verified the situation with Moonshot AI, and the official responded that "we have no relevant information or response to release". There were previous rumors that Moonshot AI closed its Series G financing on August 27.

In fact, before that, the entire primary market was scrambling to get a ticket to the final game of Kimi.

"Secondary share transfer directly from Yang Zhilin's own holdings", "The quota price rises at any time, you will miss it if you make a slow decision"... All kinds of unverifiable equity share news circulated widely in the investment circle. It was not until Moonshot AI issued an official statement to clarify that all secondary share transfers must be approved, otherwise they will be deemed invalid.

After the launch of K3, the financing pace of Moonshot AI has continued to accelerate. Its Series F financing was oversubscribed by three times and the quota was closed in advance, with a post-money valuation reaching 35 billion US dollars. For this reason, the Pre-IPO round originally scheduled to start in August had to be advanced, and the pre-money valuation has reached 50 billion US dollars.

At present, Yang Zhilin is only one step away from the major milestone that belongs to Moonshot AI, but the process has not been all smooth sailing.

Back to the beginning of 2025, DeepSeek R1 was launched, and Liang Wenfeng replaced Yang Zhilin as the new protagonist of the "technical genius" narrative overnight. What lay in front of Yang Zhilin were multiple dilemmas: the collapse of the commercialization path, the loss of public opinion momentum, and the stagnation of financing.

At the most difficult time, Yang Zhilin made two key strategic choices in a row. One was to completely abandon the C-end traffic acquisition strategy that had made him famous, and go all in on foundational model R&D. The other was to seize the high ground of Agent after DeepSeek took the leading position of deep thinking, making Moonshot AI one of the domestic players closest to Anthropic's technical paradigm.

When most giants and peers scaled back their general model business, this long-term strategy that adhered against the trend finally gave birth to K3, achieved technological breakthroughs, and accelerated commercial implementation.

In an interview, Yang Zhilin compared the technical journey of building large models to "climbing a snow mountain". "It's a bit like driving on the road, there are continuous snow mountains ahead, but you don't know what's inside, you are moving forward step by step."

The snow mountain of AGI has no end, Yang Zhilin is still climbing up, and IPO is just an unprecedented altitude on this snow mountain.

A

K3 has become a phenomenal hit, not only ranking third in the world in the Artificial Analysis comprehensive evaluation list, but also being recognized as having the strength to catch up with Claude Fable 5, which has put Moonshot AI in the global first tier of large model developers.

But just a year and a half ago, it was still labeled as "traffic spending maniac", with its growth path questioned and its model strength doubted.

In October 2023, Kimi Chat was first launched, focusing on lossless context of 200,000 Chinese characters, making it one of the C-end AI products with the largest context window in the world at that time.

What really made it a rising star in the industry was the product update in 2024. Kimi expanded its context to 2 million characters, which made Kimi a huge hit. It not only ranked top 3 in the AI App list, but also became the only enterprise among the "Six AI Dragons" that had successfully run a profitable general C-end AI product.

Only two months after its establishment, Moonshot AI secured a nearly 2 billion yuan angel round. After the Kimi product became popular, capital continued to pour in with large-scale financing. Top institutions and giants such as HSG, Alibaba, and Tencent flocked to it, and the valuation reached 3.3 billion US dollars after the Series B round, making it the most sought-after AI startup for a time.

At that time, Moonshot AI was learning from OpenAI, following the ChatGPT-style growth path, and building a C-end super application in a closed-source way. With sufficient capital reserves and in the fleeting competitive window of the AI C-end market, Moonshot AI began to spend heavily on user acquisition to pursue growth.

According to data from AppGrowing, in the fourth quarter of 2024, Kimi's advertising investment reached 530 million yuan; the monthly advertising expenditure in October was about 222 million yuan, which indeed exceeded the total advertising expenditure of the entire third quarter.

Everyone saw the surge in traffic, but they were also aware that domestic users' willingness to pay for AI tools is relatively weak, the inference cost of large models remains high, and there is no truly profitable C-end business model that has been proven feasible.

No one could have predicted the rise of DeepSeek. On January 20, 2025, DeepSeek R1 was launched. Less than two weeks after its release, the daily active users of the DeepSeek App exceeded 30 million, making it the fastest national application in history to reach this milestone.

It overturned the inherent perception of the industry, breaking the consensus that top-tier models require massive computing power, and proving that with algorithm and engineering optimization, top reasoning models can also be built at lower costs. At the same time, it challenged the growth logic of acquiring users through advertising for large models: DeepSeek barely had any advertising investment, and achieved viral word-of-mouth communication and explosive user growth only relying on its model strength.

The most impacted party was Kimi, which focused on C-end business. In the month R1 was released, DeepSeek's monthly active users reached 33.7 million, 1.7 times that of Kimi (19.43 million). In the following two months, the top three domestic native AI App rankings changed from "Doubao, Kimi, Wenxiaoyan" to "DeepSeek, Doubao, Yuanbao".

At the same time, under the dual "encirclement and suppression" of large tech companies and DeepSeek, the "Six AI Dragons" changed their attitudes towards foundational model training. Limited "talent, capital, and computing power" cannot support the two-wheel drive of model and application development, leaving model manufacturers with a choice to make.

01.AI and Baichuan fell behind one after another, giving up general foundational large model training and shifting resources to vertical application tracks. In March 2025, there was even news that Baichuan Intelligence laid off its B-end business team to streamline operations. The remaining players, Zhipu AI, MiniMax, Moonshot AI and Stepfun, became the new "Four AI Titans".

After that, Yang Zhilin not only stopped C-end advertising investment, but also suspended other C-end applications such as Ohai and Noisee that had been tested before. Moonshot AI's advertising spending dropped by more than 70% compared with the fourth quarter of 2024, and its cooperation with multiple Android channel platforms and third-party advertising platforms were all terminated.

After the rise of DeepSeek, the primary market investment logic changed completely. Capital no longer paid for the pure foundational model story, but valued implementable commercial performance more. Moonshot AI voluntarily gave up the C-end traffic growth narrative, and had to bear the huge sunk cost of computing power brought by the trillion-parameter MoE model. After the completion of Series B financing, the company fell into a 16-month financing gap.

At this point, Moonshot AI entered its "Odyssey period".

During that period of time, Yang Zhilin continued to iterate the foundational model on the one hand, and tried to improve capabilities in high-threshold scenarios such as healthcare and law on the other hand, and launched membership subscription, starting to collect fees from C-end users and developers to promote the commercialization process. But the priority of commercialization was obviously lower than foundational model R&D.

First, it reached a cooperation with Caixin Media, and later media reported that Moonshot AI began to form an AI medical product team and increased recruitment of talents with medical-related backgrounds.

The outside world speculated that Moonshot AI wanted to enter vertical fields, but the official soon responded that the move was to enhance Kimi's competitiveness and optimize the quality of search sources in professional fields such as finance, law, and medicine.

In the following nearly half a year, Moonshot AI held no high-profile press conferences and carried out no intensive marketing. This once overnight-famous AI star company suddenly became quiet.

B

It was not until half a year later that Kimi K2 was released, bringing Yang Zhilin back to the spotlight.

As the world's first open-source MoE model with trillion-level parameters, K2 performed amazingly in coding, Agent, and mathematical reasoning tasks, and achieved SOTA results in various benchmark performance tests, which allowed Kimi to enter the global developer community for the first time.

What surprised the outside world the most was that K2 not only focused on high performance, low cost and open source just like DeepSeek, but also maintained performance close to the mainstream Claude models.

K2 was originally scheduled to be released in the first half of 2025, following the reinforcement learning route consistent with DeepSeek. But after being unexpectedly "preempted" by R1, Yang Zhilin no longer competed head-on with DeepSeek in deep thinking, and shifted the optimization focus of K2 to Agentic and coding capabilities.

At that time, Anthropic had already made a name for itself overseas. This company relied on its extremely strong Agent and programming capabilities to make Claude hold a stable position among developers and B-end users, and obtained steady revenue.

K2's performance in the Agent scenario is close to Claude, so the outside world began to compare the two companies, and some views believed that Kimi is shifting from "China's OpenAI" to "China's Anthropic". More than a month later, Zhang Xiaojun raised this question to Yang Zhilin in an interview.

Yang Zhilin did not agree with the simple labeling logic. He said frankly that "the idea of 'being the Chinese version of a certain company' does not make much sense", believing that the industry environment in China and the United States is different, and the company should pursue long-term technological climbing from a global perspective.

But only a few months later, his attitude changed subtly. In the internal letter at the end of 2025, Yang Zhilin directly set surpassing Anthropic as the company's most important goal, aiming to become a world-leading AGI company, and no longer avoided this benchmarking relationship.

Around the Spring Festival this year, the industry established the AI Coding technical route with Anthropic as the benchmark. Zhipu and Kimi quickly rose because they bet on this track.

After the model route adjustment, Moonshot AI iterated three versions K2, K2Thinking and K2.5 continuously within half a year, continuously strengthening its Agent, long-range reasoning and multi-modal capabilities.

When K2.5 was launched, it just caught up with the "lobster craze" trend.

As an "all-round" model integrating visual understanding, code, and multi-modal input, Kimi K2.5 introduced the "Agent Cluster" capability for the first time. Its coding, tool invocation and multi-step task planning capabilities highly match the underlying requirements of the OpenClaw framework.

Moonshot AI naturally became one of the fastest-responding companies.

On February 18, Kimi Claw was launched. Users do not need complex configuration, and can associate OpenClaw with one click on the web page to automatically configure the Kimi K2.5 Thinking model, greatly lowering the usage threshold.

The call volume of Kimi K2.5 on the OpenRouter platform once ranked top 3 in the world. Less than 20 days after its release, its revenue exceeded the total of 2025, and one month later, its ARR exceeded 100 million US dollars.

This also led to a breakthrough in the long-standing commercialization problem that plagued Moonshot AI. Its API revenue achieved leapfrog growth, accounting for more than 70% of total revenue, and overseas developers became an important source of customers.

The outbreak in the market side quickly spread to the capital level, accelerating the company's financing. Within three months, Moonshot AI completed three rounds of financing, with its valuation soaring from 4.3 billion US dollars to 18 billion US dollars, setting a historical record for the domestic large model company to break through the 10 billion US dollar valuation at the fastest speed.

The highlight moment belonging to Yang Zhilin came again.

At the 2026 NVIDIA GTC Conference, Yang Zhilin became the only founder of an independent Chinese large model to deliver a speech on the main stage, sharing the stage with technical leaders from giants such as Tesla and DeepMind, and Jensen Huang also mentioned the Kimi model many times during the conference.

In the speech titled "How We Scaled Kimi K2.5", Yang Zhilin demonstrated a very clear technical route. Instead of catching up with closed-source cutting-edge technologies by simply stacking parameters and data, the company replaced all the "foundations" that had been used in the Transformer era for nearly a decade. For the optimizer Adam, attention mechanism and residual connection, Moonshot AI provided three alternative solutions, all of which are open sourced.

In the speech, Yang Zhilin mentioned that K2.5 adopted native image-text joint training. The team only let the model practice tasks such as counting, image recognition, and visual question answering, and as a result, its pure text reasoning ability also became stronger.

While the model practiced "seeing", its "thinking" ability improved accordingly. Then Yang Zhilin "predicted" that the next generation of models will evolve from a single Agent to a dynamically generatable Agent cluster.

K2.5 is the first implementation of Yang Zhilin's prediction. By introducing the Orchestrator, it can create multiple sub-Agents according to task requirements, and split complex tasks into parallel sub-tasks for execution.

At that time, many big names in the AI circle including Andrej Karpathy and Elon Musk spoke highly of Moonshot AI's technology.

C

Four months later, Moonshot AI released Kimi K3, fulfilling all the roadmaps mentioned in the speech. With a parameter scale of 2.8 trillion, it was the open-source model with the largest number of parameters in the world at that time.

The hybrid linear attention and attention residuals mentioned in the speech were also implemented as scheduled. KDA hybrid linear attention increased the decoding speed of million-token long text by 6.3 times, and Attention Residuals improved training efficiency by about 25% with an additional cost of less than 2%.

After its release, Kimi K3 caused a sensation in Silicon Valley, and topped the trend list of the international open-source AI model community Hugging Face only 30 minutes after its launch.

As mentioned earlier, in the Artificial Analysis comprehensive evaluation, K3 ranked third in the world with a score of 57, second only to Claude Fable 5 and GPT-5.6 Sol. In the Code Arena front-end programming blind test, K3 scored 1679 points and ranked first in the world.

Back to November 2024, Moonshot AI launched the mathematical reasoning model K0-Math, which can do verification and correct mistakes like humans, benchmarking OpenAI's strongest reasoning model o1 at that time.

Yang Zhilin also put forward a judgment at the same time that the most important capability of AI in the next stage is deep reasoning, and the next Scaling Law is called Reinforcement Learning Scaling (RL Scaling).