Amazon has started to become Kimi's money printer.
A month ago, Moonshot AI approached Microsoft, Amazon and Google at the same time, hoping to integrate K3 into Azure, AWS and Google Cloud, and take a share of the revenue generated by related services.
A maximum 30% share was put on the negotiation table, but no one knew at that time whether the deal would finally be closed.
Now the first answer is out: Kimi has actually built its own "toll booth" into AWS.
On September 18 US time, AWS officially announced the launch of Kimi K3 on Amazon Bedrock. Global enterprise developers can directly call this model within AWS and pay on a per-token basis.
Moreover, Moonshot AI is not the only Chinese model vendor that has started collecting revenue from overseas cloud platforms. On September 16, Zhipu AI also disclosed during a conference call that it has signed revenue-sharing agreements with a number of leading cloud service providers at home and abroad.
Chinese model developers have collectively started to cash in on this overseas market.
A
In fact, K3 was already accessible via Microsoft Foundry earlier, but Microsoft adopted a third-party path: underlying inference is provided by Fireworks AI, while Foundry is mainly responsible for model integration, deployment and enterprise governance.
Users call K3 on Microsoft's platform, but the inference service is still backed by Fireworks.
This time it's different for AWS, which itself acts as the provider of K3's hosting and inference services.
K3 is now an officially managed model on Amazon Bedrock, with AWS responsible for handling all calls and billing.
AWS's listed pricing for K3 is $3 per million input tokens and $15 per million output tokens for the global standard tier, which is fully consistent with the official API pricing of Moonshot AI; for the US region, the prices are $3.3 and $16.5 respectively.
A month ago, this matter was still only on the negotiation table.
On August 26, Reuters cited three people familiar with the matter as reporting that Moonshot AI was negotiating with Microsoft, Amazon and Google at the same time, hoping to bring K3 to Azure, AWS and Google Cloud and take up to 30% of the revenue generated by K3-related services on the three major clouds.
But the "30%" figure was only Moonshot AI's negotiation term at that time, and the talks were still in the early stage. Moonshot AI and the three major cloud vendors had not yet worked out how to split the revenue specifically, to what extent data could be opened up, and how to audit the actual number of tokens consumed by K3.
Now, among the three major cloud providers, AWS has taken the lead in putting K3 on its shelves.
Although the agreement between the two parties has not been made public, the specific revenue sharing ratio and even the authorization method cannot be confirmed at present, one thing is certain: AWS and Moonshot AI have at least negotiated a separate set of commercial terms.
K3 clearly states in its license that Moonshot AI allows ordinary developers to download, deploy and use K3 for commercial purposes, but if an enterprise and its affiliates operate MaaS business and their total revenue exceeds 20 million US dollars for 12 consecutive months, they must reach a separate agreement with Moonshot AI before using K3 for commercial purposes.
AWS has obviously far exceeded the threshold of "total revenue of more than 20 million US dollars for 12 consecutive months", and Bedrock is a typical MaaS service - AWS is responsible for hosting the models, and enterprises call them via APIs and pay by tokens.
By the way, K3 has also been integrated into Alibaba Cloud Bailian at present.
Bailian currently provides two types of K3 services at the same time: one is directly supplied by Moonshot AI, and the other is deployed and inferred by Alibaba Cloud itself. The latter follows a similar logic to AWS: Alibaba Cloud, as a MaaS platform that commercializes K3, also needs to comply with the separate authorization requirements in the K3 license.
That means AWS is not the only cloud provider that has reached this commercial deal with Moonshot AI.
Next, it remains to be seen whether Microsoft and Google will follow suit.
B
Bedrock is by no means short of open-weight models.
In February this year, AWS integrated six models into Bedrock at once: DeepSeek V3.2, MiniMax M2.1, GLM 4.7, GLM 4.7 Flash, Kimi K2.5 and Qwen3 Coder Next. AWS itself also stated that since 2025, Bedrock has added dozens of open-weight models from vendors such as DeepSeek, MiniMax, Qwen, Mistral, and NVIDIA.
The primary reason why AWS is willing to negotiate K3 separately with Moonshot AI is that K3 itself has sufficiently strong practical value.
With 1 million-token context window and native vision capabilities, AWS explicitly positions K3 for long-duration coding, knowledge work, and agent tasks.
It is also the first open-weight model on Bedrock that supports explicit Prompt Caching, which allows developers to actively control which prompt content is cached. This helps reduce input costs and latency when the same codebase, tool instructions or reference documents are repeatedly used in long tasks.
Moreover, as we mentioned in the previous article, K3 is perfectly suited to be sold via cloud platforms.
Enterprises that want to use K3 generally have three options: self-deployment, direct purchase of Kimi's official API, or invocation via cloud platforms.
For self-deployment, K3 has a total of 2.8 trillion parameters, and the official recommendation is to use a supernode composed of 64 or more accelerators. For direct purchase of the official API, domestic enterprises need to go through certification, recharge and quota systems, while overseas users also face independent account systems, network stability issues and sales communication barriers.
AWS is targeting the third path: integrating all these access, procurement and operation and maintenance costs into the cloud platforms that enterprises are already using.
The license only preserves the negotiation right, not the bargaining chip.
Whether cloud vendors are willing to accept additional authorization depends on a very simple logic: whether the model has sufficient demand. If not offering K3 will drive users and token traffic to other platforms, AWS has a good reason to close this business deal with Moonshot AI.
K3 has now at least proved that open weights and commercial charging are not conflicting. As long as the model itself is attractive enough, the original developer can still retain the right to charge.
C
Over the past year, Chinese models have been increasingly appearing on overseas developer platforms, enterprise clouds and API marketplaces.
Data from Reuters on September 16 shows that as of September 14, the combined market share of DeepSeek, Zhipu AI, Tencent and Alibaba on OpenRouter has reached 45%; approximately 60% of MiniMax's sales come from overseas. With low pricing, open weights and capabilities that are increasingly close to top US models, Chinese models have captured a real share of overseas demand.
This pressure has even been felt by the top US model companies.
On September 19, Reuters pointed out in its report on Anthropic's consideration of launching new models ahead of schedule that apart from OpenAI's newly released GPT-6 Astra, low-cost, flexibly deployable open-weight models are also becoming increasingly important competitors for both Anthropic and OpenAI.
In other words, at this stage of development, Chinese models are no longer a matter of "whether US developers are willing to give them a try". The models have a large user base and generate overseas revenue, and even Anthropic has begun to treat Chinese open models as competitors that must be taken seriously.
The next step is naturally to ensure that more of the revenue generated by this demand flows back to the teams that train the models.
In the past, when overseas cloud platforms obtained Chinese open models, they could host, sell APIs and charge token fees on their own. The more popular the model is, the better the platform's business performs, but the original developer may not be able to continuously share revenue from this part of the business.
Kimi is changing this situation. It keeps the model's weights open, but leaves a charging clause for large MaaS platforms in its license. Now that AWS has accepted this threshold, Moonshot AI has taken the very first meaningful step.
Chinese model developers are not only going overseas to seize users, but also starting to generate revenue directly from overseas cloud platforms.
Interestingly, the price of K3 in AWS's US region is 10% higher than that of the global standard tier. US enterprises that want to call Kimi on AWS's home turf even have to pay a slight premium.
And this trend has begun to spread among Chinese model vendors.
On September 16, Zhipu AI disclosed during a conference call for analysts and investors that the company has signed revenue-sharing agreements with a number of leading domestic and foreign cloud service providers. The GLM series of open-weight models will be launched on overseas cloud platforms in the form of hosted APIs, with revenue split between the two parties at an agreed ratio, and relevant revenue will be recognized starting from October.
This incremental revenue has even been included in the growth forecast: Zhipu AI has raised its year-end ARR guidance from 2.4 billion US dollars to 3 billion US dollars, and the current full-business-caliber ARR is about 1.8 billion US dollars.
Alibaba has also been revealed by Reuters that it plans to adopt a similar mechanism for its next-generation Qwen3.8-Max: the model will remain open-weight, but large commercial users that profit from it will also need to share part of their revenue with Alibaba.
Kimi has already received revenue, Zhipu AI has signed the agreements, and Qwen is also preparing for the move.
I would not be surprised at all if the next versions of DeepSeek, MiniMax and more Chinese open-weight models also start redesigning their commercial licenses to negotiate revenue sharing terms with large cloud service providers.
After all, since overseas cloud vendors can already profit from selling tokens using Chinese models, there is no reason for Chinese model companies to keep all that money on other people's books.
This article is from the WeChat Official Account "Alphabet Ranking" (ID: wujicaijing), written by Yuan Xinyue, edited by Wang Jing, and republished with authorization from 36Kr.