HomeArticle

Kimi's Valuation Surges to 330 Billion, Uncovering Yang Zhilin and His Trillion-Parameter Model

中国企业家杂志2026-07-24 16:34
China's large models, reclaim the pricing power.

At 9 p.m. on July 16, less than 24 hours before the opening of WAIC (World Artificial Intelligence Conference), Moonshot AI (Kimi) suddenly made a surprise launch of Kimi K3, seizing the spotlight.

The most disruptive label of Kimi K3 is its parameter scale of 2.8 trillion. As the world's first trillion-level open-source large model, K3 supports a 1 million-token context window, and has outperformed GPT-5.6 and Fable 5 in scores across multiple mainstream benchmarks. Many frontline AI practitioners commented: In multimodal, long-context, and complex task scenarios, Kimi K3 holds a dominant, unmatched lead in China.

The matching pricing strategy has also sparked widespread discussion. Kimi K3 is priced at 20 yuan per million tokens for uncached inputs and 100 yuan per million tokens for outputs, ranking in the highest price tier among domestic models — 20 times more expensive than DeepSeek V4 Pro; 30% of the price of the overseas model Fable 5, and 60% of GPT-5.6 Sol.

Goldman Sachs gave a qualitative conclusion in its research report: "This marks a brand-new milestone where Chinese models are moving toward pricing power."

"Impressive." Elon Musk praised K3 on X. Later, while teasing Grok 4.5, he stated: "It is expected to surpass Kimi." In response, Huang Zhenxin, co-head of enterprise business at Moonshot AI, openly responded to media including *China Entrepreneur* on July 21: "We'll wait and see, and hope they can compete with us on equal footing."

The impact brought by K3 has dramatically stirred the capital market. On July 17, the stock prices of four tech giants — NVIDIA, Google, Meta, and Intel — trended downward. U.S. media reports stated that while it was not caused by a single factor, K3 was undoubtedly a key catalyst.

In the Hong Kong stock market, the day after K3's release, Zhipu's stock price fell by 28.49% in a single day, and MiniMax fell by 15.62%; a Huachuang Securities research report pointed out that although the Hong Kong stock index was sluggish that day, the release of K3 was still a major reason for the stock price adjustments of Zhipu and MiniMax.

The shock caused by K3 overseas also pushed Yang Zhilin to the center of public attention. His defense video for the Tsinghua University Special Scholarship and his papers published during his PhD studies at Carnegie Mellon University (CMU) have gone viral one after another. A Silicon Valley venture capital partner posted a question on X that garnered tens of millions of views: Why didn't such a talent stay in the United States after graduation? This prompted Russ Salakhutdinov, Yang Zhilin's PhD supervisor at CMU and Apple's first AI director, to step forward to defend his beloved student.

Recently, *China Entrepreneur* exclusively contacted Professor Russ, who stated: Yang Zhilin performed exceptionally well, completing his PhD in less than four years, which is extremely rare at CMU. "Zhilin is one of the top students at CMU. He can not only think about research problems, but also has strong technical capabilities, and can do low-level programming — usually these capabilities are separated, and it is rare to see a student with such a combination of strengths."

Russ said that Yang Zhilin's research work during his PhD was very outstanding, as he completed a large number of foundational machine learning tasks. His papers on XLNet (a generalized autoregressive pretraining model) and Transformer-XL (an ultra-long context language model) have become part of the theoretical foundation of AI models.

When Yang Zhilin graduated, many companies extended job offers to him. Russ recalled that he could easily have stayed in the U.S. to teach at MIT or Princeton, but he insisted on returning to China to start his own business immediately. "For someone of Zhilin's caliber, the typical path is an academic career, and starting a business is a high-risk endeavor, but he excels in both business and research capabilities."

Although the strength of Kimi K3 has been recognized by the industry, its expensive invocation price has also sparked considerable controversy.

"The reason for the high price lies in the high token consumption and the structural high cost of long-running agent tasks," said Li Jialong, an AI system R&D engineer. Jia Zijian, founder of TalkMe, an AI English training product company, gave an example to *China Entrepreneur*: "Rough estimates show that deploying DeepSeek V4 Pro requires about 8 H100 or A100 GPUs, while running K3 with the same logic requires 20 H100s."

In response, Huang Zhenxin said: "Open-source models and Chinese large models should not be labeled as low-priced. We have developed a SOTA-level model, and it deserves a reasonable commercial pricing."

The massive parameter size has also slowed down K3's operation and delivery speed. Li Jialong noted that when K3 processes complex tasks, a complete programming task can take more than 30 minutes. Jia Zijian shared a similar feeling: "As a C-end product, users cannot wait that long, so we need to balance capability and speed."

Under the pressure of computing power, K3 had to suppress demand. Less than 48 hours after its release, Kimi announced that K3 subscriptions were "temporarily closed to new users." But in Russ's view, this is just a short transition period. He believes Yang Zhilin will find better architectures and computing technologies to solve this problem in the future.

He said that Yang Zhilin once mentioned that he wanted to research "infinite context" models. "Right now, Kimi's context length is around the 1 million level, and maybe in the future he will give the model truly infinite long context."

The release of Kimi K3 has also directly boosted Moonshot AI's valuation in the capital market. On July 22, media reports stated: Moonshot AI plans to launch its final pre-IPO financing round in August, with a target pre-money valuation of up to 500 billion U.S. dollars (about 3.38 trillion yuan), and will list on the Hong Kong stock market as soon as within 6 months.

Built for long-range complex tasks

What main capabilities have been enhanced with the release of K3? Peng Mingliang, a large model application development engineer, summarized it as: The biggest experience is pushing conversational AI further toward project-based AI. Russ praised it highly, saying that it is basically on par with models from Anthropic and OpenAI. Especially the open-source nature of the model will profoundly change the entire AI landscape, making "K3 extremely unique and extremely valuable."

Engineer Li Jialong found through testing that K3's front-end performance is 10% to 15% stronger than GPT-5.6 Sol, with "lower rework rates and higher aesthetics." In actual web development scenarios, Kimi K3 has a higher first-pass rate for generated code than OpenAI's corresponding model, which means Kimi K3's coding capabilities have been greatly improved. "The path of scaling up parameter counts for K3 has truly succeeded."

In terms of multimodality, Kimi K3 has made tremendous progress, achieving a SOTA score of 91.2% on the BrowseComp benchmark. It can not only understand complex cross-modal instructions, but also integrate text, images, and code into deliverable front-end products. Jia Zijian described it: "K3 has lifted the front-end capabilities of domestic large models to a very high level. The generated web pages can almost go online directly, with visual completeness far exceeding that of other models."

In terms of user experience, Peng Mingliang summarized three specific enhanced capabilities of K3: stronger task persistence, with a 1 million-token context window large enough to hold an entire codebase, allowing the model to complete file-crossing engineering-level modifications in a single session; outstanding front-end and visual delivery capabilities; and significantly improved parallel agent sub-cluster capabilities.

These three dimensions also correspond to the three focus areas of Kimi that Yang Zhilin once publicly mentioned: token efficiency, long context, and agent clusters.

Technically, the two core supports of K3 this time come from the self-created KDA mixed linear attention mechanism (Kimi Delta Attention) and the Attention Residuals architecture. Simply put, Kimi can use full attention calculation at key nodes of a task to ensure accuracy, and switch to linear attention on non-critical paths to reduce inference costs. Kimi's official data shows that its million-context decoding speed is up to 6.3 times faster, training efficiency is improved by about 25%, with additional costs less than 2%.

Based on this, K3 essentially serves three types of scenarios — long-range programming, end-to-end knowledge work, and deep reasoning.

From the release of K2 in July 2025 to the launch of K3 a year later, Kimi has iterated 6 times. Among the four key technical iterations, K2 solved the problem of "scale," K2.5 solved the "multimodality" problem, K2.6 and K2.7 Code solved the "programming specialization" problem — and K3 integrates all these capabilities into a sufficiently long, sufficiently complex task continuum.

With every version upgrade, Kimi has always been answering one question: Can the model maintain stable reasoning capabilities in longer task chains?

Huang Zhenxin clearly defined K3's technical route: "The performance leap of K3 comes from underlying original architectural innovations, not distilling small models." At the same time, he emphasized that Kimi's adherence to the open-source path will not change: "The core underlying technologies and papers are all publicly disclosed."

After the release of Kimi K3, Dean Bowers, OpenAI's Head of Future Strategy, posted on X that open-source is "decelerationism" that will weaken the return on investment for commercial AI companies, thereby "hindering the massive capital investment required by the entire industry." However, Russ strongly supported his beloved student's open-source strategy, arguing that Kimi's open-source initiative is remarkable and crucial.

"Open weights and open-source are always the harder path — because your competitors can take the model, study it, analyze it, use it to generate data, and then train their own models. This is giving advantages to competitors," Russ told *China Entrepreneur*.

But open-source opens the door to convenience for more users. "The current high price of models is mainly because they are large. But since K3 is open-source, many people will perform distillation optimizations on it. After the model is open-sourced, as user data gradually accumulates, many people will distill it into smaller models," Russ said.

He noted that many labs now train a large model first, and then train small models based on the outputs of the large model — making the small models as close as possible to the performance of the large model while becoming more efficient. Based on this, he believes that K3's token consumption will also be greatly reduced in the future.

The "Genius" Yang Zhilin: Dual Foundations of Technology and Business

The success of Kimi K3 has also made the public more curious about the source of Yang Zhilin's talent and his growth trajectory.

"Brilliant, hardworking, humble." When *China Entrepreneur* interviewed Russ, he thought for a long time when asked to "describe Yang Zhilin in three words," saying it was hard to summarize him in a few words, and finally gave the above description — talented, extremely diligent, and humble.

Yang Zhilin completed his bachelor's and master's degrees at Tsinghua University, and later pursued his PhD at the LTI (Language Technologies Institute) in the School of Computer Science at CMU. As a senior professor in the Machine Learning Department at Carnegie Mellon University, Russ has supervised dozens of PhD students, many of whom later took on core roles at Google Brain, OpenAI, and DeepMind. But Yang Zhilin still left him an extremely unusual impression.

"He has strong technical aspirations, his own ideas, and great autonomy and execution. Zhilin's research basically starts with his own ideas, which are quickly implemented after discussing with me," Russ mentioned. During his PhD studies, Yang Zhilin was always passionate about natural language processing (NLP), and was keen on thinking about how to build systems and models with generalization capabilities that excel at understanding and generating language.

"He has great ambitions, and always takes on the most difficult tasks."

Russ revealed that Yang Zhilin interned at Google Brain during his PhD, "The work he did was basically optimizing TPUs (Google's custom AI chips for machine learning)," which is low-level system work that very few PhD students can participate in.

More than one person has made similar evaluations. Tang Jie, founder of Zhipu AI, was Yang Zhilin's teacher during his undergraduate years at Tsinghua. They are not only teacher and student, but also rivals in the business world now.

"There is no doubt that Yang Zhilin is the most outstanding and talented student I have seen in recent years," Tang Jie, as his recommending supervisor, commented at Yang Zhilin's defense meeting for the Tsinghua Special Scholarship.

During his studies at Tsinghua, 90% of Yang Zhilin's professional course scores were above 95, he got full marks in all programming courses, and ranked first in his grade. In the first semester of his sophomore year, he joined Tang Jie's lab and quickly became a core member, proposing a series of optimization algorithms that improved computing performance by dozens of times, which were adopted by companies including Tencent, Sina, and Huawei.

In his junior year, Yang Zhilin was selected for Stanford University's Undergraduate Visiting Scholar Program for his cancer prediction research based on big data. In a global cancer prediction competition, the algorithm he proposed based on protein pathways defeated teams from top research institutions such as Stanford and Columbia University. As of July 2025, Yang Zhilin's paper citations on Google Scholar have exceeded 20,000 times.

Beneath his reputation as a recognized genius, business capability is a lesser-known label of Yang Zhilin.

"Yang Zhilin was very comfortable discussing with professors from Stanford University, the University of Michigan, and MIT." At that time, Yang Zhilin was only 20 years old, but Tang Jie believed that he had already demonstrated strong communication skills. Yang Zhilin was not a "bookworm" either — he formed a band at Tsinghua, serving as the drummer and main songwriter, and had a large number of fans on campus.

"His uniqueness lies in the fact that he is very technically savvy, an outstanding researcher, and also an excellent businessman. It's very hard to find someone who has all these traits at the same time," Russ said. When he graduated from CMU, Yang Zhilin expressed his wish to return to China to Russ. "I want to start a business. If I don't even try it, I will regret it."

In Silicon Valley circles, Russ has seen too many technical geniuses fail in business, and too many CEOs who know nothing about technology, "but Zhilin can perfectly combine technology and business together."

His technical leadership and business rhythm control have also allowed Moonshot AI to maintain a leading position in the fierce talent war. "When the CEO is someone who deeply understands technology, employees will trust and respect him," Russ said. A former Moonshot AI employee once told *China Entrepreneur*, "Yang Zhilin has a unique charisma, and he recruits talents by 'hiring entire teams' from university labs."

Domestic AI regains pricing power

While reaching the peak of performance, Kimi K3's pricing strategy has also sparked huge discussions — its price is about 20 times that of DeepSeek V4 Pro, 5 times that of Zhipu GLM-5.2, and 4 times that of Qwen 3.8.

"After using K3 for less than two days, the quota of my 199-yuan package was exhausted, and my friend's 699-yuan membership package quota was also quickly used up," said an AI development engineer.

48 hours after launch, Kimi admitted that "user demand for K3 is approaching the limit of current GPU capacity." To prioritize the usage experience of paying users, Kimi urgently restructured its membership system, splitting into two independent paid product lines: one type of membership for Kimi Web, App, and Work; another type of Kimi Code membership for coding workflows.

Huang Zhenxin frankly stated in the interview: Kimi is facing a dilemma. Opening registration to all new users will seriously damage the usage experience of existing paying users; prioritizing paying customers will require temporarily restricting new user registrations. "The company ultimately chose to temporarily give up new revenue in order to maintain the bottom line of the usage experience for paying customers."

Source: Official website screenshot

As an enterprise-side user, Jia Zijian believes, "The