Closed source and open source are colliding.
The most counterintuitive trend in the AI circle recently is that closed-source models are getting cheaper, while open-source ones are becoming increasingly expensive. This is not a typo, no reversal at all.
Does that sound the other way around?
Our intuition tells us that open-source weights are free to download and run whenever you want, so how expensive can they be? Closed-source model APIs are billed by tokens, and enterprises pay a monopoly premium. You say closed-source is getting cheaper? No one would believe that.
But data does not lie. BenchLM has a dedicated index that tracks the market price of AI inference.
From March 2023 to the present, the actual transaction price of all cutting-edge models has dropped by 88%. Note that this is the market price, the actual amount people pay in real money.
The median price of cutting-edge models is now $4.50 per million tokens. For mid-tier models? It is $1.93. The low-budget tier has been directly pushed down to $0.50.
Just in three years, the industry has shifted from "feeling distressed every time you use it" to "use it as you like, it's not expensive anyway".
Sounds like the whole trend is unidirectional: everything is getting cheaper. Open-source models must be cheaper, that's a sure thing.
Hold on. There is a key concept called expenditure-weighted price, which measures how much users are actually willing to pay.
Bloomberg and Silicon Data jointly launched the LLM Token Expenditure Index for exactly this purpose, and the results are very interesting.
The overall index is around $1.29 per million tokens, which sounds acceptable and not too outrageous. If you break it down, there are two lines moving in completely opposite directions.
Closed-source prices have been falling all the way from $3.07, while open-source weighted prices have been rising steadily from $0.66.
Yes, you read that right. Open-source is getting more expensive. The weird part of this is that the weight files of open-source models are indeed free to download, so why do we say their cost is rising?
The reason is simple. What is free is the weight file, not the cost of running it.
When you run inference with an open-source model, you need GPUs, engineers to tune parameters, optimize quantization, fix video memory overflow, modify prompt templates, and even get up at 2 a.m. to locate bugs when the enterprise work group alerts. Every second of these tasks adds to your cost ledger.
More importantly, the scale is expanding, and more and more people are using open-source models.
According to data from OpenRouter, DeepSeek's token share has doubled from 10% at the beginning of this year to 18%; the overall share of Chinese open-source models has surged from less than 2% at the beginning of 2025 to 46% in June 2026, nearly half of the total.
When everyone is running open-source models, and the token volume jumps from 34% to 65%, the total bill can only go up. This is a physical law, which has little to do with business models.
What about the closed-source side? It is moving in the opposite direction.
Anthropic has raised its gross margin for inference from -94% to over 60%. What does -94% mean? It means for every dollar of revenue, it lost 94 cents; now for every dollar of revenue, it earns 60 cents.
How did they achieve that? Through economies of scale, inference optimization, and falling hardware costs. But the most critical driving force is competition.
OpenAI, Google Gemini, and Claude are fighting fiercely in the same price range; if your API is twice as expensive as others, the enterprise CIO will send an email the next day asking why.
Therefore, the price of closed-source models has been forced down by fierce competition.
The two curves are moving towards each other's territory: closed-source price is sliding down from $3.07, while open-source price is climbing up from $0.66.
Here is a side proof: DeepSeek V4 Flash is priced at $0.09 for input, and even cheaper at $0.18 per million tokens for output. If you look at the price of closed-source cutting-edge models at $4.5, the gap is about 25 times.
This is the current situation, and this gap is narrowing at a speed visible to the naked eye.
Every time I think about this trend, I think of the mobile phone industry from 2012 to 2015. The pattern back then is exactly the same as the AI industry today.
Apple took the high-end market with closed-source iOS, selling an iPhone for six to seven thousand yuan, while Android manufacturers focused on cost performance with open-source systems, selling phones for one or two thousand yuan. Then the script began to reverse.
Apple was under pressure from Android's market share, so it launched the iPhone SE and mid-tier models to pull down the average price of its products.
Android manufacturers were trapped by the user label of "your product is only worth 2,000 yuan", and tried their best to move to the high-end market: Samsung launched the Galaxy S series, Huawei pushed the P series, and their product prices kept rising.
The two lines collided in the middle.
In today's mobile phone market, it is hard to say "Android is cheaper" or "Apple is expensive". The price of iPhone can go as high as ten thousand yuan, and also as low as three thousand yuan. Android flagship phones can also be priced at seven or eight thousand yuan.
Brand premium has disappeared, the substitution logic has become mainstream, and you have to give a reason if your product is even one dollar more expensive.
Today's AI industry is running on the exact same script.
Closed-source models are moving down the price ladder, while open-source models are squeezing upward. They will meet at a certain price range, most likely the $1.9 to $4.5 range defined by BenchLM.
On that day, the two stories of "open-source is cheaper" and "closed-source is better" will no longer hold water.
Because the cheap one is not necessarily really cheap, and the expensive one is not necessarily really better than others. The game will no longer be about "who has stronger model capabilities", but about "who can deliver the product to customers stably".
.......
Okay, the two lines are moving towards each other. Then what?
Then the most scary thing comes. When open-source and closed-source are squeezed into the same price range, the model itself will no longer be valuable. "Model capability", as a commodity that can be sold separately for profit, is being emptied of its value.
There are already some signals of this.
Mozilla released a 2026 July report on the state of open-source AI, and there is a number that stunned me after reading it.
Open-source models support about 33% of active AI applications worldwide, but they only take 4% of the total revenue. 33% of the work is done by open-source, but only 4% of the money goes to it.
Where did the remaining 96% go?
It went to cloud vendors, system integrators, and those who help enterprises "run" the models. The model is free, but deployment and implementation cost money.
This data is even more striking in Wall Street statistics: a16z released a CIO survey in January this year, covering 100 companies from the Global 2000. The share of expenditure on open-source models dropped from 19% a year earlier to 11%. For the same group of enterprises, the share of closed-source expenditure accounted for 89%.
Note that this means for the same group of enterprises, the money they spend on open-source is relatively less, and the money is flowing to closed-source delivery solutions.
Enterprise CIOs are not stupid. Downloading open-source weights is only the first step. The subsequent work including deployment, operation and maintenance, compliance, and SLA can be fully taken care of by closed-source vendors. Even if closed-source vendors charge several times more, enterprises still feel that they have saved trouble.
Ultimately, the extra money earned in this process is not for model capabilities, but for "peace of mind".
Here is another detail:
I checked that Alibaba's AI-related revenue this year is close to 9 billion RMB, with an annualized revenue of 35.8 billion RMB. Its AI MaaS (Model as a Service) annual recurring revenue exceeded 10 billion RMB in June, and the management said it is targeting 30 billion RMB by the end of the year.
Look at how Alibaba sells AI: Tongyi Qianwen is fully open-source, weights are free to get; but you have to run the inference on Alibaba Cloud. Open-source is the bait, and cloud service is the fishhook.
This is not an isolated case. Chinese vendors including DeepSeek, Alibaba, ByteDance, Tencent, Zhipu AI, and Moonshot AI are all adopting a strategy internally called "strategic open-source".
Basic models are open-source, high-end capabilities are closed-source, basic functions are free, value-added services are paid, you get a free taste, and you will stay after you are full.
What pushes this strategy to the extreme is Moonshot AI's Kimi K3.
With 2.8 trillion parameters, fully open-source under the MIT license, its programming capability directly outperforms GPT-5.6. The pricing is 100 RMB per million tokens. Within 48 hours after release, its computing power was completely overwhelmed, and new user registration was directly suspended.
Goldman Sachs used three words in its report to evaluate this event: pricing power.
Chinese model vendors are proving one thing: when open-source weights can achieve 99% of what closed-source models can do, with cost more than 60% lower, the pricing power will slip from the hands of closed-source vendors to open-source vendors.
After the release of K3, the combined expected valuation of OpenAI and Anthropic was cut by about 314 billion USD, while Moonshot AI's own valuation rose to 330 billion RMB.
The money did not disappear, it just moved to a new place.
However, this is not the most critical point. The most fatal factor is Jevons Paradox.
William Stanley Jevons, a 19th century British economist, found that the higher the efficiency of the steam engine, the less coal it burns; but as a result, Britain's coal consumption did not drop, but rose instead. Because burning coal became cheaper, everyone started to use steam engines, and the total consumption exploded.
AI is replicating this script. According to data from the Mozilla report, the inference cost of GPT-level models has dropped 50 times in 36 months, from 20 USD to 0.40 USD per million tokens.
It is obvious that with such a low cost, no one will only run it once.
In the past, you would feel distressed after running it once. Now if one prompt does not satisfy you, you can run it ten or twenty times, then let the Agent iterate on its own.
A MIT study calculated a more striking number: if enterprises switch to open-source models, their inference cost can be cut by 70%, saving about 25 billion USD per year in the United States.
The 25 billion USD saved will not lie idle in accounts, it will be reinvested in more AI applications, generating more tokens and consuming more computing power.
Therefore, after checking so much data, I draw this conclusion: the more the end-user cost drops, the larger the total market size grows. This is the logic of profit pool shifting: the model layer can no longer get premium, and money flows to both sides along the pipeline.
One part flows to the inference bills of cloud vendors, the other part flows to the deployment and maintenance work inside enterprises.
Enterprises are willing to pay this sum of money, because for the 28 percentage point gap between 51% and 79%, every percentage point lost means more overtime work for engineers.
Next time someone tells you "open-source models are free", you can reply to him: Yes, the weight is free. To run it, the bill is waiting for you on the cloud.
.......
So closed-source is doomed, right?
If the story ends here, Anthropic should not even exist. But it not only survives, but also performs extremely well.
Anthropic's annualized revenue has surged from 9 billion USD to over 60 billion USD.
In the same a16z survey, its enterprise penetration rate increased by 25 percentage points within one year, and 44% of Global 2000 enterprises are using its products. This is not a story about "better model performance".
As we mentioned earlier, the performance gap is only 3%, and coding tasks are already on par. It wins in other areas.
Where does it win?
IP indemnification. SLA guarantee. No data cross-border. Compliance audit. Enterprise-level permission management. You will find that none of these things have anything to do with model capabilities themselves. They are all about "delivery".
What Anthropic charges for is the service of "you can rest assured, I will take responsibility if anything goes wrong".
At the same table, players are playing completely different games. Open-source vendors are using model capabilities to grab token share. Closed-source vendors are using delivery capabilities to grab budget share.
Then look at the data again:
In the a16z survey, 78% of enterprises are running OpenAI models in production environments, and 81% of enterprises are using more than three model families at the same time. The average enterprise large model budget has risen from 4.5 million USD to 7 million USD within two years.
The budget is growing, but the growth is for the cost of "making AI work properly".
Morgan Stanley's report describes three end scenarios: closed-source wins, open-source wins, and mixed coexistence.
It does not state it explicitly, but the data points to the third scenario: mixed coexistence. Enterprises will assign their most core, most risk-sensitive workloads to closed-source models, and assign high-frequency, low-sensitivity long-tail tasks to open-source models. Both sides will survive, they just shift their positions in the industrial chain.
Then who is losing money?
The ones stuck in the middle. They do have strong models, but no delivery capabilities; no enterprise-level support, no stable supply of computing power.
Your model is very strong, but enterprises dare not use it. You are the world champion, but no one invites you to play in the league.
Once this logic holds, the entire AI industry chain will split into two layers.
The upper layer is the model layer, competing for parameters, benchmark scores, and academic papers. The lower layer is the delivery layer, competing for stability, security, and taking responsibility for customers. The profit of the upper layer is evaporating, while the profit of the lower layer is expanding.
A CMU study calculated that for a small-scale deployment with parameters below 30B, the payback period can be as short as three months.
But once you go for large-scale deployment with parameters from 200B to 1T, the payback period is six years. Every second in these six years is burning delivery costs.
Therefore, as long as closed-source vendors hold their ground in delivery, they will always have a place in the market. The ones that really need to worry are those who only have models but no delivery capabilities.
.......
After talking so much, we have not touched on a core premise. What exactly do we mean by "open-source" today?
MIT and HuggingFace jointly conducted a study at the end of last year, tracking all large models that claimed to be "open-source" between 2022 and 2025, and screened