For every $100 a model company earns, cloud providers take nearly $40.
Barclays' latest research shows that for every $100 in revenue generated by artificial intelligence (AI) model companies, approximately $35 to $40 flows to the three major cloud giants, namely Amazon AWS, Microsoft Azure and Google Cloud Platform (GCP), in the form of inference computing power fees.
Of this portion of revenue, cloud service providers can obtain an operating profit of approximately $10 to $20, corresponding to an operating margin of around 35% to 45%.
The above conclusion comes from an AI industry unit economic model research report released by Barclays on August 28.
The report points out that the profit margin of paid inference business for AI labs has risen sharply, from the low double-digit level in 2025 to 50% to 65% or even higher in 2026, with the adjusted gross margin rising by 30 to 50 percentage points year on year.
The main reason for the substantial improvement in profit margin is that enterprise customers and agent workflows have become "must-buy" products in the market.
Barclays analysts believe that the actual profit margin may even be higher than the estimate in the report, but as competition for cutting-edge models intensifies and the supply of computing power continues to increase, the profit margin is expected to gradually fall back.
Two Distinctly Different Financial Structures
To break down the profit differences between different AI labs, Barclays built two hypothetical cutting-edge AI lab models.
Among them, "Lab A" generates approximately 70% of its revenue from the API business and 30% from the subscription business; while "Lab B" is the opposite, 80% of its revenue relies on the subscription business, and the API revenue accounts for only 20%.
The API business naturally has a higher inference profit margin than the subscription model. Coupled with the differences in training cost allocation methods and partner revenue sharing mechanisms, the adjusted gross profit margin of the two labs has a gap of as high as 17 percentage points: around 55% for Lab A and approximately 38% for Lab B.
The revenue recognition method further amplifies this difference.
Lab A recognizes indirect API revenue using the gross method, while Lab B adopts the net method, or simply does not recognize indirect API revenue operated by strategic partners.
Barclays compares this difference to the difference in financial statements between the two ride-hailing giants Uber and Lyft: even if their core businesses are highly similar, the final disclosed revenue and profit data may vary greatly simply due to different accounting treatment methods.
Profit Margin Analysis by Product Line
Subscription products, such as Anthropic's Claude Code and OpenAI's Codex, are expected to have an inference profit margin of around 70%, the lowest among the three main product lines.
The reason is that AI labs are willing to subsidize Token costs to improve user retention. Subscription products usually adopt a model of fixed monthly fee with a usage cap, and AI labs have significantly increased the frequency of resetting usage quotas recently, which may reflect both the pressure on user retention and the impact brought by improved model efficiency.
Direct API is the first business of AI labs to form a large-scale business model, and it is also the product line with the highest profit margin.
Developers using tools such as Cursor and Figma pay according to the actual consumption of Tokens. Barclays estimates that the inference profit margin of the API business has exceeded 80% at present.
The improvement of model Token efficiency, the rise of nominal API prices, and the continuous optimization of inference services at the infrastructure level, including quantization technology, speculative decoding, and new-generation computing platforms, are all continuing to unlock room for profit margin growth.
Barclays points out that the API inference profit margin in the second quarter of 2026 has actually been much higher than the level shown in its public charts, but the bank expects that this profit margin will still fall back to a certain extent at some stage in the future.
The usage experience provided by indirect APIs to end users is basically the same as that of direct APIs, but the billing relationship occurs between users and cloud service providers.
As the proportion of indirect API revenue continues to increase, the differences in revenue recognition methods among different AI labs will further expand, and the comparability between financial statements will become lower and lower.
Profit Structure of Cloud Vendors
Barclays' model shows that for every $100 of revenue generated by an AI lab, the corresponding cloud service provider revenue for Lab A is approximately $35. After deducting infrastructure costs, cloud service providers can obtain a profit of around $11.8, corresponding to an operating margin of 34%.
Lab B can generate more revenue for cloud service providers, due to the strategic partner revenue sharing mechanism: which accounts for about 20% of revenue and is subject to a cumulative cap. In this case, the corresponding revenue for cloud service providers is approximately $41, with a profit of around $19.1, and the operating margin reaches 47%.
However, Barclays emphasizes that the revenue sharing mechanism will artificially inflate the apparent profit margin of cloud service providers. If the revenue sharing portion is excluded, the actual profit that cloud service providers obtain per unit of Token does not differ. This revenue sharing arrangement is expected to be gradually phased out after 2028.
Agent subscription products can also create additional value for cloud service providers.
Agent-based subscription products usually need to continuously save task status and call other cloud services such as databases and storage, so they can bring additional revenue beyond inference computing power to cloud vendors.
Ebb and Flow: The Three Major Cloud Vendors' AI Computing Power Share May Decline After 2028
Barclays predicts that the total global revenue of AI labs will grow from $7 billion in 2024 to $137 billion in 2026, and further reach $690 billion in 2028.
If calculated based on the annualized recurring revenue (ARR) at the end of the year, the growth curve is even more aggressive: by the end of 2026, the ARR of AI labs is expected to reach approximately $2000 billion; by the end of 2028, it will further reach around $7820 billion.
At present, training expenditure still accounts for about 48% of AI labs' revenue, which means that for nearly every $1 of revenue obtained by AI labs, roughly $1 flows to cloud service providers.
However, the proportion of training costs in revenue is declining rapidly. This proportion was as high as 96% in 2024, and it is expected to drop sharply to 35% in 2027, and further to 30% in 2028.
In other words, as the profit generated by the inference business gradually exceeds the training expenditure, the overall profitability of AI labs will continue to improve.
At the same time, the ratio of cloud service providers' AI revenue to AI labs' revenue is also declining. This ratio is expected to drop from 153% in 2024 to 90% in 2026, and further to 73% in 2028.
Barclays expects that in the next two years, the market share of AWS, Azure and GCP in AI labs' computing power expenditure will remain roughly stable.
However, starting from 2028, the exclusive computing power infrastructure that AI labs have previously signed contracts for or locked in will be launched one after another, and gradually become a more important source of computing power for them. This may mean that the three major cloud giants will gradually lose part of their market share in both the AI training and inference markets in the future.
This article is from the WeChat official account "CLS", author: Xia Junxiong, published with authorization from 36Kr.