HomeArticle

Moonshot AI Kimi: Aggressive Ambition, Restrained Expansion

中国企业家杂志2026-06-29 15:20
The uncharted technological territory and uncompromising dedication to aesthetic perfection — this uniqueness is the very reason for Moonshot AI Kimi's existence.

In the past year, Dark Side of the Moon Kimi (hereinafter referred to as "Kimi") has made a leap from being one of the "Six AI Tigers" to an object pursued by global capital.

In December 2025, Kimi completed a Series C financing of $500 million, with a post - investment valuation of $4.3 billion. In May this year, Kimi completed a Series D financing of $2 billion, and its valuation reached $20 billion. In June, it was rumored that Kimi was in talks for a new round of financing of up to $2 billion, with a pre - investment valuation of $30 billion.

Technically, the release of Kimi K2.5 in January 2026 became a key turning point. This flagship model supporting full - modality processing was launched less than a month ago, and Kimi's cumulative revenue within 20 days exceeded the total for the whole year of 2025. Its ARR (Annual Recurring Revenue) exceeded $200 million.

Since then, Kimi has accelerated its product output rhythm: In April 2026, Kimi K2.6 was released, featuring programming and Agent cluster capabilities. It can schedule up to 300 specialized sub - agents for parallel collaboration in a single task. In mid - June, Kimi Work, a general local Agent product for the computer, and the K2.7 Code programming - specific model were released.

It is reported that Kimi K2.7 Code has made great improvements in benchmark tests compared with the previous generation. In long - range tasks, the average Token consumption of K2.7 is directly reduced by 30%, and the multi - language code generation ability is greatly improved.

"Programming ability is the starting point for the improvement of AI productivity. From the perspective of Token consumption, programming accounts for 90%. But this is just the beginning of the adoption of AI intelligence. The productivity revolution brought by general Agents will expand from 30 million programmers to 1 billion knowledge workers." said Yang Zhilin, the founder of Kimi.

In his view, the vast expanse of the large - model field is more worthy of expectation. "The paradigm of large - model R & D is changing. From the second half of 2026 to 2027, AI will play a more leading role in the research aspect."

01

Aggressive Technological Ambition

Looking back at the beginning of 2025, DeepSeek emerged suddenly and disrupted Kimi's R & D rhythm. In July 2025, Kimi K2, the world's first open - source MoE model with a trillion parameters, was released, which once again showed Yang Zhilin's technological cards.

After the release of Kimi K2, an internal researcher at Kimi wrote in a blog: "At the reflection meeting at the beginning of (2025), I put forward some rather radical suggestions. Unexpectedly, Zhilin's subsequent actions were even more radical than I thought. For example, he stopped updating the K1 series of models and concentrated resources on basic algorithms and K2."

People close to Kimi told China Entrepreneur that K2 was born at a critical moment for the company, and Yang Zhilin's decision to abandon K1 and start working on K2 was crucial for the company.

The release of the multi - modality model K2.5 marked the evolution of Kimi's technical route. Its native multi - modality architecture first integrated text and visual input at the bottom layer. The Agent cluster supports 100 sub - agents for parallel collaboration, and a single task can call 1500 steps. After the launch of this one - trillion - parameter model, it was in short supply, directly pushing Kimi's ARR to exceed $100 million.

In April, K2.6 expanded the Agent cluster to 300 sub - agents and supported 4000 - step coordinated execution. Its programming ability achieved a leap - forward improvement - it defeated GPT - 5.4 with 58.6% on SWE - Bench Pro, and could code continuously for 13 hours and modify more than 4000 lines of code. In June, K2.7 Code further focused on vertical scenarios, with a 30% reduction in inference Tokens and a 21.8% improvement in Kimi Code Bench V2. The multi - language code generation was greatly optimized.

With three iterations in half a year, Kimi's product path has gradually become clear: open up the ability boundary through architectural innovation, and then approach the precision limit of professional scenarios through vertical B - to - B iterations.

According to the sharing of Kimi's algorithm researchers, behind the accelerated product iteration speed, there is a key technological breakthrough: visual reinforcement learning training has fed back the pure text ability. Yang Zhilin calls it "a discovery that breaks the industry's perception." "Previously, it was generally believed that introducing visual ability would reduce text ability, but we found that the two can improve each other."

In Yang Zhilin's view, in the past 10 years, the Transformer architecture, Adam optimizer, residual connection, etc. have formed the technical foundation of deep learning and were once regarded as the industry's consensus infrastructure. However, with the continuous expansion of model scale and the increasing complexity of tasks, these once "standard configurations" may become obstacles to model evolution.

Therefore, Kimi's technical route also shows distinct characteristics - it focuses on the underlying layer. It not only makes engineering optimizations on the existing architecture but also goes back to the most basic components of the AI system to solve problems one by one. It optimizes the optimizer, attention mechanism, residual connection, etc. one by one to improve algorithm efficiency and gain a higher intelligence ceiling.

For example, the MuonClip optimizer used in K2 has doubled the Token processing efficiency compared to AdamW. The Kimi Linear hybrid linear attention architecture has achieved a 5 - to 6 - fold increase in decoding speed in the ultra - long context from 128K to 1M. The Attention Residuals technology, which improves the neural network architecture layer in K2, redesigned the core residual connection mechanism in the neural network. On the premise of similar effects, the training computation is reduced by about 20%, which is equivalent to a 1.25 - fold efficiency advantage.

"MuonClip, Kimi Linear, and Attention Residuals are essentially all for efficiency. Through algorithm innovation, we can make full use of existing resources to achieve higher Token efficiency and model intelligence levels." a Kimi researcher said.

Regarding the next - generation model K3, Yang Zhilin said that the next - generation model will adopt a new model architecture. One of the goals is to make the model more suitable for the long - range task ability of Agents, because this is the most critical ability.

"In the future, Kimi will continue to research and reconstruct the underlying technology, and a large amount of underlying technology will be rewritten in the next 2 to 3 years. I hope that K3 can become a more distinctive model, allowing users to experience new abilities that other models have not defined." Yang Zhilin said.

02

Restrained Organizational Expansion

In sharp contrast to its aggressive technological ambition, Kimi is quite restrained in organizational expansion.

Internally, Kimi maintains a flexible "small - team" combat state. As a unicorn with a valuation of over $30 billion, Kimi has about 300 employees in the whole company, making it the one with the fewest employees among the leading large - model startups.

The "elite troops" model is also deliberately adopted by Yang Zhilin. He publicly said: "Among these large - model startups, we always keep the smallest number of employees. It is very important to maintain the highest ratio of cards to people. We don't want the team to expand too much, as expansion is fatally harmful to innovation."

China Entrepreneur learned that there is no OKR system in Kimi, no departmental barriers, and even no departments in the traditional sense. The company has cancelled various position labels such as directors and vice - presidents. Kimi's organizational structure is extremely flat, and several co - founders directly communicate with dozens of team members. Yang Zhilin's WeChat signature only has four words: Direct Communication.

Yang Zhilin, the founder of Kimi

People close to Kimi told China Entrepreneur that when talking about models, Yang Zhilin often repeatedly mentions a word - "taste". In the increasingly homogeneous competition of computing power and data, "taste" has become the core driving force for Kimi to build a differentiated barrier.

Zhang Yutong, the president, explained Kimi's talent concept like this: Kimi prefers people with "abstract ability", "a bit of paranoia", and "a willingness to do things crazily". "If you have a good idea, will you try it 1000 times? Most people may think it can't be done after trying 10 times. But there are also very few people who believe more in their own ideas and form new cognitions in the process of trying."

"In the early days of Kimi's establishment, it gathered the inventors of many core AI technologies, and these people later found more like - minded people." In Yang Zhilin's view, technology itself is still the biggest variable in large AI models, and Kimi's attractiveness to technical talents is the key to its competitiveness.

03

An Undefined LLM

At the end of December 2025, when MiniMax and Zhipu successively finalized their IPO progress, the market turned its attention to Kimi. Yang Zhilin was once indifferent to this. He said in an all - staff email that the company had sufficient cash flow and was not in a hurry to go public.

However, the large - model industry is changing rapidly, and the number of players is still shrinking sharply.

On May 7th and 8th, 2026, the Chinese large - model track announced more than $10 billion in financing within 48 hours. The media commented: "The money is not flowing into the industry but to the last few players." Kimi has proven its technical strength and still needs to prove its commercialization ability to the market.

In March after the release of K2.5, Kimi's ARR exceeded $100 million; in April, this figure reached $200 million. "For a long time, K2.5 was in short supply." People close to Kimi told China Entrepreneur.

On June 12th, Kimi released the desktop AI Agent product Kimi Work, which supports 300 concurrent Agents and has a built - in Cron scheduler. At the same time, Kimi Work also achieved direct connection to financial data. Agents can directly read and write files on the computer, and all operations are completed locally without data leaving the device.

In addition to accelerating commercialization efficiency, Kimi also needs to address issues such as computing power, talent, resources, and the competitive - cooperative relationship with Internet giants.

Yang Zhilin's stance has always been clear. "Our company was not established for the sake of competition." he said in an early interview. At the end of 2025, he further explained his judgment: "The industry has entered a new stage. At the beginning, many companies were involved, and now there are fewer. In the future, what everyone does will gradually become different."

The flow of technology in the open - source community also provides an explanation for this "harmonious yet different" pattern. When DeepSeek released V4, its technical report clearly thanked Kimi for the innovative and open - sourced Muon optimizer. Yang Zhilin's response was calm and honest: "This is the meaning of open - source. We have benefited from open - source technology, and we also hope to bring our contributions to the community."

Who will Kimi become in the future? In an internal letter at the end of 2025, Yang Zhilin expressed his clear and firm stance, along with his unique "taste" and confidence.

"In 2026, Kimi will become a 'distinctive' and 'undefined' LLM (Large Language Model). Whether it is the technological uncharted area that others dare not bet on or the aesthetic persistence that requires a bit of paranoia, I believe that more Kimi - defined innovations can make unique contributions to the accelerated development of human civilization. This uniqueness is the greatest meaning of our existence."

This article is from the WeChat official account "China Entrepreneur Magazine" (ID: iceo - com - cn), author: Sun Xin, published by 36Kr with authorization.