With valuation skyrocketing 5 times within 3 months, this article compares and deconstructs the wealth creation vs capital burning of "Token Factories" in China and the United States.
In the current narrative of large language models, the spotlight has long been fixed on the parameter competition of models and massive computing power infrastructure; but stepping out of this framework, an "AI intermediary" that neither builds models nor hoards computing power cards has reached a valuation of over 7 billion US dollars (equivalent to 47.2 billion RMB), corresponding to a P/S ratio of approximately 50 times (referring to its 140 million US dollars ARR).
According to reports from Bloomberg citing people familiar with the matter, payment giant Stripe has finalized an agreement to acquire OpenRouter at a transaction price of over 7 billion US dollars. Back in May, OpenRouter just completed its Series B financing, and the post-money valuation disclosed by the media at that time was only between 1.3 billion and 1.5 billion US dollars.
In other words, in less than three months, this company has achieved an astonishing 5x premium increase.
What exactly makes OpenRouter so valuable? To put it simply, it connects major models of both open-source and closed-source vendors such as OpenAI, Anthropic, DeepSeek, GLM to the same API, allowing developers to switch freely and route automatically — it is essentially a "traffic intermediary" that provides model scheduling services.
This type of business is not unfamiliar in China.
SiliconFlow, which is currently sprinting for an IPO, also provides unified API, multi-model invocation and inference services. The number of registered users on its platform has exceeded 10 million, and the Token invocation volume is also growing rapidly. However, looking through its prospectus, we see another set of financially contrasting statements: the overall gross profit margin in 2025 has dropped to -24%, and the gross profit margin of the "public cloud Token service", which is closest to OpenRouter's business form, is as low as -119%.
On one hand, the more you sell, the more overwhelmed you are by computing power costs, and on the other hand, there is an annualized gross profit margin of about 70%. Both seem to be operating as "model transfer stations", why is the benefit gap between the two business models in China and the US so huge?
Clarifying this issue is also the key to understanding why OpenRouter was acquired at a valuation of 7 billion US dollars.
Three Months, Valuation Surge from 1.3 Billion to 7 Billion
OpenRouter was founded in 2023. Its founder Alex Atallah was best known previously as the co-founder of NFT trading platform OpenSea.
Figure Note: The core team of OpenRouter and its investors, from left to right: Chris Clark, Chief Operating Officer of OpenRouter, Alex Atallah, Founder and CEO of OpenRouter, Deedy Das, Managing Director / Senior Partner of Menlo Ventures, and Matt Murphy, Senior Partner of Menlo Ventures
This Web3 experience seems to have nothing to do with AI at first glance, but the underlying logic is quite continuous.
At the peak of the NFT bull market, Atallah experienced the typical "traffic beating" of the Internet: sudden explosion of demand, server overload, and search index crash. He recalled that his most important task at the time was to prevent OpenSea from becoming the "Fail Whale" that frequently went down.
This obsession with sudden traffic, capacity and stability was later almost entirely brought to OpenRouter, and also shaped OpenRouter's product logic: do not heavily bet on a single supplier, but prepare redundancy in advance for unpredictable demand and traffic peaks.
What is more interesting is that both of his two entrepreneurial ventures are essentially operating a bilateral market Marketplace.
OpenSea does not produce NFTs, it connects buyers and sellers; OpenRouter does not produce models, it connects computing power supply and model demand.
Without touching heavy assets, its scale flywheel spins extremely fast. When Menlo Ventures invested in OpenRouter in early 2025, it only had about 2.5 million developers; by May this year, this number had exceeded 8 million.
In the same period, its weekly processing volume rose from 5 trillion Tokens to 25 trillion Tokens within half a year, a 5-fold increase, and it has supported access to more than 500 models.
At present, OpenRouter has started to make real profits. A report by The Information in July stated that OpenRouter's annualized ARR reached approximately 140 million US dollars, with annualized costs of about 40 million US dollars, which means the annualized gross profit is as high as 100 million US dollars, corresponding to a gross profit margin of about 71%.
According to a rough estimate based on the 7 billion US dollar acquisition price, it is equivalent to approximately 50 times the annualized revenue, which is obviously not cheap. But what Stripe values is not its current revenue, but its extremely special "positioning" in the industrial chain.
One Set of API, Two Lines of Business
Many people simply regard OpenRouter as a "large model API supermarket", which is too superficial. Because in real scenarios, developers are not faced with one multiple-choice question, but two: first, which model to choose? Second, on which service provider (Provider) to run this model?
The same open-source model may be hosted by dozens of inference service providers at the same time. Different providers have uneven capabilities in quantization technology and cluster scheduling, which ultimately leads to huge differences in price, latency and reliability.
What OpenRouter does is to abstract these two layers of supply together. When a request comes in, it automatically selects the optimal Provider according to cost and speed; in case of downtime, it automatically switches (Failover).
Here comes an extremely cruel industry reality: SiliconFlow itself is actually one of the underlying Providers of OpenRouter.
If we split the industrial chain positioning of the two companies, the differences between these two Sino-US Token factories will be very clear.
Providers like SiliconFlow, although they do not own large-scale GPU assets, must rent computing power with real money, deploy and optimize models, and convert computing power into Tokens to sell to customers. They bear the most direct hardware cost in the production link.
As disclosed in its prospectus: in 2025, the gross loss of public cloud services exceeded 34 million. Generating each Token consumes computing power. If the GPU is rented at a high price or the model price is reduced, it will greatly affect the gross profit fluctuation.
But OpenRouter is facing a completely different set of mathematical problems.
It does not need to be the "cheapest person to produce Tokens", it only needs to constantly find out "who is the cheapest and most stable producer at the moment", then direct the traffic there, and steadily draw a platform fee of about 5.5% for itself.
In this case, one side seems to be an athlete stuck in the mud racing against GPU costs, while the other side is a referee sitting in the stands picking winners.
The Cheaper the Token, the More Valuable It Is
And in the "price war" of large models, this model difference is further amplified.
In an interview with 20VC, Atallah repeatedly emphasized that he does not believe that AI will eventually move towards "one takes all". This trend is becoming more and more obvious now.
Enterprises will not throw every task to the most expensive cutting-edge model. Complex reasoning can be handed over to powerful models such as Claude and GPT, but deterministic tasks such as classification, extraction, and format conversion can completely be handed over to small models that are ten times cheaper. In the Agent era, "temporarily assembling a multi-model team for one task" will become a rigid demand.
The more models there are, the cheaper the Token is, and the greater the difficulty of selection and scheduling.
In the interview, Atallah mentioned a typical case of the "Jevons Paradox": the price of a certain OpenAI model on OpenRouter dropped by 90% cumulatively within two weeks, and the Token usage volume subsequently increased by about 13 times. That is, the drop in unit cost has instead stimulated greater total demand.
But when we put OpenRouter and SiliconFlow together, this case has another meaning. Facing the same 13-fold surge in traffic, the two companies have completely different fates:
For underlying Providers like SiliconFlow, price reduction is first and foremost an extreme cost war. If the market price is cut in half and their GPU utilization and inference efficiency cannot keep up, the more they sell, the more they lose;
But for OpenRouter, as long as the total demand stimulated by price reduction is growing, the transaction flow through the platform will continue to expand, and the tolls it charges will only increase.
So the underlying logic is that Providers bet on "whether the Tokens I produce can be cheaper than others", while Router bets on "whether more and more Tokens will need to be allocated in the future".
Manufacturing Tokens is increasingly like an involution efficiency business; while the upper-layer Token distribution has become a high-gross-margin transaction business.
What exactly is the 7 billion US dollar acquisition paying for?
Stripe spent sky-high money, of course, not to buy a "better Token factory", but to completely monopolize the capital flow of "intelligent consumption" in the upcoming Agent era.
Looking closely at Stripe's moves this year, the clues have long been laid. In January this year, Stripe just acquired the usage-based billing platform Metronome, and in the same month, it publicly disclosed its in-depth cooperation with OpenRouter.
Putting these three business puzzles together, a larger commercial closed loop emerges:
OpenRouter: Decide what model a task calls, which node it goes through, and how many Tokens it consumes;
Metronome: Responsible for extremely complex usage-based billing;
Stripe: In the most underlying layer, responsible for pricing, invoicing, taxation and payment verification.
If the puzzle is completed smoothly, Stripe can basically be regarded as the dual role of "Alipay for global developers" + "computing power UnionPay in the AI era".
In the past, it was "people" who bought software and paid a fixed monthly subscription fee; in the future, it is "machines (Agents)" that execute tasks, which may trigger dozens of model calls in one second behind the scenes.
Users don't even know how much "intelligence" their Agent has just purchased. So whoever controls the Routing will see the consumption flow of these machines first; whoever masters this flow will become the Visa and UnionPay in the AI era.
Conclusion
The sharp contrast between SiliconFlow and OpenRouter precisely reflects the two stratified forces in the current AI inference market:
One side is working hard on chips and scheduling at the bottom layer, trying to squeeze more intelligence out of one graphics card; the other side is floating above all models, trying to become the "distribution hub" of the whole industry.
The former determines how cheap the underlying computing power of AI can eventually be; the latter points out at what price and where these cheap computing power will eventually flow.
As long as Tokens keep flowing, no matter who wins, someone has to be responsible for billing and collecting money. If model manufacturers are fighting for the winner at the poker table, then what Stripe wants to buy is the poker table itself.
This article is from WeChat official account "AI Front" (ID: ai-front), written by Siyue, authorized by 36Kr to release.