Is Kimi also about to "enter the customs"?
In the past month, Moonshot AI has simultaneously taken a seat at two major "tables" in the United States.
One is the policy table of the US government, where discussions are held on imposing restrictions on it.
On July 22, US Treasury Secretary Scott Bessent publicly stated that the administration is considering adding Moonshot AI to the trade blacklist. On the same day, Michael Kratsios, a technology official under the Trump administration, accused Moonshot AI of developing Kimi K3 by distilling Anthropic's Fable and obtaining NVIDIA GB300 servers.
Moonshot AI denied the relevant accusations, stating that K3's performance improvement comes from original architectural optimizations.
The other table is the commercial negotiation table of leading US tech giants.
On local time August 26, Reuters exclusively revealed that Moonshot AI is negotiating with Microsoft, Amazon and Google at the same time, hoping to make Kimi K3 available on Azure, AWS and Google Cloud. According to people familiar with the matter, Moonshot AI expects to take up to 30% of the revenue generated by K3-related services on the three major cloud platforms.
If the negotiations are successfully closed, this is likely to become the first significant revenue-sharing agreement between a Chinese AI company and large US cloud vendors.
Prior to this, domestic Chinese AI models have long been gradually "entering the US market".
In February last year, IBM added two distillation models based on DeepSeek-R1 to the on-demand deployment catalog of its enterprise AI platform watsonx.ai. By August this year, IBM signed a multi-year $240 million agreement with Together AI to build an AI inference cluster on IBM Cloud, and the models provided by Together AI include domestic Chinese models such as DeepSeek, Kimi, and MiniMax.
The first chapter opened by DeepSeek is that Chinese models, with the help of enterprise IT giants like IBM, cross the trust and channel thresholds of large US enterprises.
The second chapter that Kimi is opening now is that Chinese models not only need to gain access, but also need to clarify the revenue distribution rules with the largest US cloud vendors after entering the ecosystem.
But this also raises a new question:
For a model with open weights, what makes it eligible to take up to 30% of the revenue from Microsoft, Amazon and Google?
01
This share of revenue is still hanging on the negotiation table for now.
According to Reuters' report on August 26, Moonshot AI is separately discussing revenue sharing for Kimi K3 with Azure, AWS and Google Cloud, hoping to obtain up to 30% of the revenue from K3-related services on the three major cloud platforms.
The negotiations are still in the early stage. How to calculate the revenue, to what extent data can be opened, and how many Tokens K3 actually consumes are all unresolved issues at present. The last item in particular is directly related to how revenue will be distributed in the future: cloud platforms process massive amounts of model requests every day, and a set of mutually recognized auditing methods is required to confirm exactly how much invocation volume K3 contributes.
If we go back more than ten days, this 30% figure did not appear for the first time.
On August 7, Reuters already revealed that Moonshot AI is designing a revenue-sharing mechanism for large commercial users of K3, with the maximum ratio also set at 30%.
Looking further back, the license of K3 has also reserved an entry for such an arrangement.
After the release of Kimi K3, Moonshot AI opened the model weights, allowing developers to download and deploy them on their own. But the license explicitly sets a commercial boundary: if a company and its affiliates operate MaaS (Model as a Service) business and generate a total revenue of more than 20 million US dollars for 12 consecutive months, they need to reach a separate agreement with Moonshot AI before using K3 for commercial purposes.
This set of rules has already been applied in real commercial collaborations.
Citing people familiar with the matter, Reuters reported that Moonshot AI has previously signed similar revenue-sharing agreements with some smaller cloud platforms, but the specific partners and sharing terms have not yet been disclosed.
One publicly confirmed case occurred on July 20. Chinasoft International announced via the Hong Kong Stock Exchange that it has signed a Token revenue sharing and joint innovation cooperation agreement with Moonshot AI. The two parties will jointly develop enterprise-level Agents for industries including energy, power and finance; the Token consumption revenue generated by Kimi and other cooperative revenues during the collaboration will be shared at an agreed ratio.
In terms of division of responsibilities, Chinasoft International is in charge of industry scenarios and enterprise delivery, while Moonshot AI provides Kimi model capabilities. The two parties will jointly package and sell AI services to enterprises, and then share the resulting revenue.
By August 26, the names on the other side of the negotiation table had changed to Microsoft, Amazon and Google, and the stakes of the matter have also risen accordingly.
Azure, AWS and Google Cloud hold a large group of the most important enterprise AI customers around the world, and the scale of K3 makes large cloud platforms particularly significant for Moonshot AI.
It has 2.8 trillion parameters. Although its weights are open, running it stably requires extremely large computing power infrastructure. The official Kimi recommendation is that deploying K3 is best done on a supernode composed of 64 or more accelerators.
64 accelerator cards far exceed the inference environment that an ordinary enterprise can set up casually. Citing analysts, Reuters said that due to the huge computing cost, few customers will choose to run K3 entirely on their own infrastructure.
This leads to a very interesting situation: The weights can be downloaded freely, but large-scale Token consumption is continuously flowing to the cloud.
Moonshot AI now hopes to take up to 30% of this part of the revenue.
If the final negotiations are successful, Reuters believes that this will likely become the first important revenue-sharing agreement between a Chinese AI company and large US cloud vendors.
02
Chinese models' access to the US market first relies on a realistic economic calculation.
In February last year, the distillation model of DeepSeek-R1 entered IBM's official deployment catalog; in August this year, IBM signed a multi-year $240 million agreement with Together AI to build an AI inference cluster on IBM Cloud, and in Together AI's model menu, Chinese models including DeepSeek, Kimi and MiniMax have already taken their place.
Nowadays, domestic Chinese models are increasingly appearing on US developer platforms, cloud services and enterprise-level model catalogs.
The reason why these models can be adopted in the US market is not complicated: their performance is already competitive enough to enter the arena, and their cost is sufficiently competitive.
For enterprises that consume billions or even more Tokens every day, a few dollars cheaper per million Tokens for the same task can eventually add up to a very large amount of money.
At the same time, the capability gap between domestic open-weight models and leading US models is also narrowing. In its August 26 report on Kimi, Reuters cited evaluations from Artificial Analysis, stating that K3's performance on complex, multi-step tasks is already comparable to OpenAI GPT-5.5 and Anthropic Claude Opus 4.8.
If the model capabilities are far apart, low price is just a compromise. But when the capabilities enter the same competitive range, price will become a real procurement reason.
In our previous article "ByteDance Retains Microsoft China", we mentioned a very interesting mirror situation:
On the one hand, Chinese technology companies need US advanced models, global cloud infrastructure and compliance capabilities. The overseas businesses of companies like ByteDance and SHEIN have already become an important commercial reason for Microsoft to stay in China. Reuters disclosed that by the mid-2020s, helping Chinese companies go global has become Microsoft's largest China-related business.
On the other hand, US cloud vendors also need a sufficient supply of models that are abundant and cheap enough.
On July 2, Reuters observed that Western startups and developers' interest in Chinese open-weight models is rising — one of the most direct reasons is that their performance is getting closer and closer to leading US models, while the price is significantly lower.
For Azure, AWS and Google Cloud, the model itself is also part of cloud services. The richer the model choices and the more competitive the prices, the more opportunities there are to attract developers and enterprises to keep their Token consumption on their own platforms.
Chinese companies need US cloud platforms, and US cloud vendors are starting to need Chinese models too.
When it comes to Kimi, the issue moves further forward.
Now that Chinese models can already appear in the model menus of US enterprises, Moonshot AI wants to figure out: can this demand be further turned into bargaining power in revenue distribution?
The three major clouds are not short of models, nor are they short of computing power. For Moonshot AI to obtain a maximum 30% share, it ultimately depends on whether K3 can bring enough invocation volume and customers to these platforms.
If customers are indeed willing to choose Kimi, cloud vendors will have reasons to include it in their model services, and Moonshot AI will have room to continue negotiating for revenue sharing.
If this revenue-sharing mechanism can finally be negotiated successfully, it means that Moonshot AI, as well as more Chinese model vendors, can turn model demand into real commercial bargaining power.
However, it is still too early to draw this conclusion. The three major clouds have not yet reached an agreement with Moonshot AI, and 30% is only the maximum revenue-sharing requirement disclosed by people familiar with the matter. There is currently no answer as to what the final agreed ratio will be, or even whether the deal can be closed.
It is worth noting that this negotiation takes place exactly at the critical node when Moonshot AI is preparing for its public listing.
For a company whose latest round of financing valuation has reached 35 billion US dollars and is advancing its IPO at a valuation of up to 50 billion US dollars, what Moonshot AI needs to prove now is not only model capabilities and user demand, but also a business model that can generate sustainable profits.
If Azure, AWS and Google Cloud can accept revenue sharing, Kimi will have an additional revenue source that grows along with overseas invocation volume. This is particularly important for Moonshot AI before its listing.
03
How to continue making profits after opening model weights is not a question that only Moonshot AI needs to answer.
On August 7, Reuters simultaneously disclosed that Alibaba also plans to adopt a similar revenue-sharing mechanism for the next version of Qwen. According to people familiar with the matter, the next version of Qwen will still open the model weights, but if large commercial users generate revenue relying on the model, Alibaba plans to take a share of it, and the specific ratio is still under discussion.
Different from Moonshot AI, Alibaba itself is a cloud vendor. When Qwen is called on Alibaba Cloud, Alibaba can directly charge for Token usage; but once the open weights are downloaded to the customer's own server, the subsequent commercial use of the model may bypass the billing system of Alibaba Cloud.
Now Alibaba is preparing to push the billing boundary further outward: the model remains open, and even if large commercial users deploy Qwen on their own infrastructure, as long as they build a sufficiently large business around it, they may need to split the revenue with Alibaba again.
To understand why both Kimi and Qwen are considering similar issues at this time, we need to go back one more year.
Over the past year, open weights have almost become the fastest path for domestic Chinese models to go global.
From DeepSeek to Qwen, and then to Kimi, model weights are directly placed on platforms such as Hugging Face, so that cloud vendors, inference platforms and developers can deploy them on their own. Coupled with the significantly lower Token price, domestic models do not need to knock on the door of each company one by one, and can quickly appear in various APIs, cloud services and enterprise model menus.
Kimi K3 is a very typical example. According to Reuters' calculation of the public input and output Token prices on August 7, the usage cost of K3 at that time was only about one third of Anthropic Fable. DigitalOcean has already provided K3, and US inference platforms such as Together AI are also providing open model services including Chinese models.
Low price and openness together have pushed the distribution speed to a level that was hard to imagine in the past; but the more widely the model spreads, the more paths the revenue starts to disperse through.
A cloud vendor that obtains the weights can build its own inference service and charge customers by Token; the inference platform can make profits by improving GPU utilization and optimizing Token processing efficiency; application companies then package the model into coding, search or Agent products, and continue to charge users at the next layer.
Although the usage volume of the model has increased, the money may not flow back to the company that trained the model along the original path.
Facing this situation, domestic model companies have started to move in different directions.
Simply put, DeepSeek bets on maximum dissemination; Zhipu AI focuses more on monetization through its own services; Moonshot AI and Alibaba are experimenting with revenue sharing based on open weights.
DeepSeek has the most lenient rules for commercial usage. The DeepSeek V3 series explicitly allows commercial use. According to its model license, third parties can deploy the model on their own, or even make it into SaaS or API to charge customers, and DeepSeek will not automatically participate in the revenue sharing for that.
At present, the most clear official charging entry for DeepSeek is its own API: users who are willing to directly use the inference service provided by DeepSeek will pay by Token.
Zhipu AI's approach is different. Although it has also opened the GLM model weights, in terms of charging methods, Zhipu AI is closer to Anthropic: it tries its best to package model capabilities into its own services and sell them layer by layer.
Developers can call the BigModel API by Token, or purchase