The "unlimited monthly subscription" never comes, and AI can only become the next "dial-up internet".
"High cost" is only the superficial phenomenon, and it is time to reflect on the current AI business model.
Recently, I have been following various cutting-edge discussions in the AI circle. People are keen to explore the boundary of model capabilities, the future form of Agent, and the rise of local AI. Among them, one judgment from Chen Danqing struck me deeply: One of the core pain points that local models try to solve, in addition to security and privacy issues, is another important pain point: "high cost".
Chen Danian is the founder of Shanda Networks and LianShang Networks. He started his business at the age of 20, and is a well-known first-generation Chinese programmer of the same period as Zhang Xiaolong and Lei Jun. His views are really worth reading.
The phrase "high cost is a pain point" sounds extremely counter-intuitive at first. After all, in the past two years, the most fierce battle in the large model industry has been the price war. The cost of one API call is now much cheaper than it was two years ago. However, the flip side of the coin is: More and more enterprises and individuals who really integrate AI into their business find that the comprehensive cost of AI is not as low as imagined.
For many practitioners, the model subscription fee of only a few hundred dollars per month already makes people feel pressured; for enterprises, once AI changes from an occasional chat tool to a long-running production system, the costs generated by model calls, Agent execution, tool usage and repeated reasoning will be further amplified.
However, is the "high cost" of AI really an inevitable outcome of technological development?
The answer is no. In the era of traditional software, we purchase and download an Office suite, install it on the computer and get long-term right to use it. Microsoft will not charge you again every time you open Word or tap the keyboard. The same is true for the App ecosystem on smartphones. After the hardware is purchased at one time, the software can run continuously in the background, and users will not have to recalculate the cost for every tap on the screen.
Why is it that in the AI industry, a suffocating model has gradually evolved: Every conversation generates a cost; every step of reasoning corresponds to a new computing cost?
Therefore, what we really need to think about today may no longer be how much higher the IQ of AI can be, but the commercial operation logic of AI. Only by understanding why AI is expensive and how this high cost has, to a certain extent, inhibited the initiative of AI, can we see clearly where the next generation of artificial intelligence should go.
Is the "high cost" of AI really an inevitable outcome of technological development?
I. High cost is superficial, the problem lies in the model
The reason why AI is considered "too expensive" is ostensibly an arithmetic problem about the unit price of computing power, but in deeper analysis, it is actually a problem of business model.
The vast majority of mainstream AI services today are essentially still tightly bound to the cloud computing track. Users put forward demands on the terminal, and the huge remote data center immediately starts to schedule GPU clusters. The model with hundreds of billions of parameters completes a heavy matrix calculation, and finally returns the results through the optical cable. The larger the model, the longer the context, and the more complex the reasoning process, the higher the costs of chips, power, servers and networks behind it.
If AI is only occasionally used to summarize an article, revise an email, or provide some inspiration, then this cloud call mode does not have major problems. For C-end individual users, paying a subscription fee of tens of dollars per month is not unacceptable.
But when this logic is moved to the real enterprise production scenario, the nature has undergone fundamental changes.
The AI that is truly revolutionary in the future is by no means just helping employees make beautiful PPTs, but needs to be deeply embedded in the core business flow of the enterprise, and participate in work for a long time, stably and continuously. It may need to read tens of thousands of internal documents overnight, call different external tools at high frequency, and repeatedly verify, test errors and re-execute itself when encountering logical dead ends. This means that a macro task instruction issued by humans often no longer corresponds to "one call" in the black box of AI, but will be disassembled into hundreds of consecutive reasoning calculation processes.
At present, the mainstream commercial large models generally adopt the billing mode based on Token consumption. This means that once AI changes from "occasional use" to "long-term operation", the cost will accumulate rapidly.
As a result, the entire technology industry has witnessed a seemingly absurd but extremely real paradox: On the one hand, major model manufacturers are frantically announcing that API prices have plummeted to extremely low levels; on the other hand, when enterprises really use AI on a large scale, the total bill is getting higher and higher. According to previous reports from *Caijing*, in June 2026 alone, thousands of employees of Kingsoft Office consumed about 4 trillion Tokens, resulting in a computing power cost of millions of yuan.
What is really noteworthy about this number is not "whether millions of yuan is expensive or not", but that after AI enters the production process, it is increasingly difficult for enterprises to accurately predict usage costs in advance as they do when purchasing traditional software.
In the traditional IT era, enterprise CIOs can accurately predict the annual budget by planning the number of purchased accounts. But after entering the Agent era, how many rounds of chained reasoning does a complex business require AI to perform? How many times do you need to call external plugins halfway? These variables can hardly be predicted mathematically in advance. Moreover, the more complex and advanced work AI tries to undertake, the more difficult it is to control its hidden costs.
High cost itself is not a dead end that blocks transactions. If an enterprise invests 50,000 yuan in AI every month and can get a net profit growth of 500,000 yuan in return, few entrepreneurs will complain that it is too expensive. The real anxiety of enterprises at present is: Facing the ever-increasing AI bills every month, they cannot find equivalent value feedback on the balance sheet.
A latest survey of global CEOs by PwC shows that in the past year, as many as 56% of surveyed CEOs admitted that the introduction of AI has not brought significant revenue growth or cost reduction to the enterprise; only a handful of 12% of enterprises have truly seen a win-win situation of revenue increase and cost reduction.
This is also why the industry has begun to re-discuss how AI should be priced.
Because what users really want to buy is never the Token itself, nor a single calculation, but the completion of a task. However, today's AI, in many cases, still follows the rule of "one calculation, one additional cost generated". Therefore, "high cost" is ostensibly a price issue, and at a deeper level, it is actually an operation mode issue.
When AI is only called occasionally, this mode has no big problem; but if the future AI needs to be online for a long time, work continuously, and test errors repeatedly, this mode will become more and more cumbersome. It is precisely at this point that "high cost" is connected with another issue: the initiative of AI.
II. "Initiative" trapped in the billing sheet
If "pay-per-use" only makes enterprise CFOs feel distressed, the far-reaching impact of this model is that it inhibits the "initiative" of AI at both the physical and psychological levels.
This touches on the second core logic we must clarify today. Today's AI, no matter how realistic its output is and how rigorous its logic is, in most application scenarios, it is still only a passive responsive "talking tool". It is like a smart but lazy answering machine, which moves only when humans give instructions. It can perfectly answer your questions, but it will never take the initiative to wake up at 3 a.m. to patrol the whole network data for you, identify potential risks in the supply chain, and actively generate early warning emails to relevant departments.
Initiative inherently means extremely high fault tolerance, massive redundant calculations and continuous background operation. A truly AI system that can "take the initiative to get things done for humans" is destined not to follow a straight working path. Faced with an open commercial problem, proactive AI needs to test constantly: it will first spend thousands of Tokens to retrieve information on the Internet; if it finds that the information does not match, it needs to overthrow itself, change the angle and retrieve again, which may consume thousands more Tokens; then, it needs to integrate massive fragmented information into an intermediate draft, which may consume tens of thousands more Tokens; then, it needs to start the self-review mechanism to judge whether this draft is relevant to the topic, and Tokens continue to be consumed... After several cycles, it will output the final result.
Trust in the business world is always built on controllability. When every operation and every external call of AI affects the gears of the billing meter, users will inevitably erect a high wall in their psychology. They simply dare not let AI run freely. This is like you sitting in a taxi that charges by the second, watching the meter jump wildly. Even if the driver swears that he is taking a detour to find a shortcut that will never be congested, your first reaction will still be extreme anxiety, and you can't help shouting "stop immediately".
The same is true for AI. Initiative requires space, and this space inevitably includes a lot of "invalid calculations": trial and error, rollback, inspection, rework, and even an exploration path that does not produce results in the end.
If every attempt requires additional payment, enterprises will naturally set call upper limits, cycle times and budget boundaries for Agents. Individual users will also subconsciously reduce the frequency of use, and dare not let it run in the background for a long time.
Just imagine, if AI can completely get rid of the tight hoop of pay-per-use, become cheap enough, or even its operating cost approaches a fixed fee, what kind of reversal will the situation have? Only in that environment without any financial pressure can users truly let go of their guard and entrust those complex, long-cycle, and error-prone trivial tasks to the machine completely.
So, is there a way to let AI run at low cost for a long time?
Some enterprises are already exploring a new AI large model method: local model.
What is a local model? Unlike opening a web page every time you ask a question, the data of the local model will not go out of your computer or your mobile phone. It belongs to you, and only belongs to you. It can be used without network, and presents different results for different users.
For the cost problem, the local model provides another idea: instead of continuing to lower the price of each call, change where the calculation occurs and how the cost is generated. Chen Danian has a vivid summary of this: "Tokens cost nothing, you only need to pay for electricity." Of course, this does not mean that local AI has no cost. Computers, chips, power and model maintenance all require investment, but when the model runs on your own device, part of the expenditure that originally increases with the number of calls may be converted into relatively fixed hardware and operating costs.
Therefore, "high cost" is ostensibly a price issue, and at a deeper level, it is a cost structure and operation mode issue. At a further level, what it really affects is whether humans dare to hand over things to AI. Only when AI is cheap enough can it get enough space for trial and error; only by getting enough space for trial and error can it gradually change from a "talking tool" to a system that can truly take the initiative to complete tasks.
III. Crossing the gap of business model: from pay-per-use to industrial prosperity
In the past technology industry, similar changes are not uncommon. Very often, the real popularization of a technology is not only because the technology itself has become stronger, but also because the cost structure of using it has changed.
In the dial-up Internet era, users will control their online duration.
The PC Internet is the most typical example. In the early days, Internet access was billed by duration and traffic. With dial-up Internet access, users would naturally control their online time; after broadband, Wi-Fi and mobile traffic gradually became popular, people began to get used to staying online for a long time, and no longer calculated the cost every time they opened a web page or played a video. When "Internet access" becomes cheaper and even unnoticeable, applications that require continuous online such as short videos, live broadcasts, and mobile payments have ushered in a real outbreak point.
Cheap and even unnoticeable Internet access costs have promoted the outbreak of short video and other applications
The software industry has also experienced similar changes. Traditional ERP often requires enterprises to purchase expensive software licenses at one time, and then bear the costs of deployment, customization and maintenance, which is unaffordable for small and medium-sized enterprises. After the emergence of SaaS, enterprises can subscribe monthly and use it on demand, splitting one-time large investment into continuous small expenditures, so that software can enter more enterprises.
Behind these two examples is actually the same thing: When the usage cost of a technology drops, or the billing method is more in line with users' usage habits, users will dare to use it more frequently and continuously. Today's AI is also facing a similar problem.
The current Token billing system for cloud large models is in the "dial-up Internet era" of the entire AI industry. Although it is powerful, it makes people count every penny and walk on thin ice.
For AI to usher in a real big outbreak, it must experience its own model leap from "dial-up billing by the minute" to "unlimited broadband monthly package". Therefore, the entire industry needs to re-examine and even reconstruct the AI operation route. Whether it is through the fierce competition of end-side chip computing power to make local models capable of handling daily complex reasoning for free, or through disruptive system architecture innovation to reduce cloud reasoning costs exponentially, the core goal is only one: completely change the cost structure of AI.
Only when the operation of AI on the device is as natural as the resident Office software in our computer; or as unnoticeable as the fixed broadband fee we pay every month, can the real AI revolution kick off.
Therefore, when we discuss the "high cost" of AI today, what is really worth paying attention to is not how much the API of a certain model company has been reduced, or how much discount it has offered.
At one level, this is a price issue; at a deeper level, it is an operation and business model issue; and at a further level, it ultimately relates to whether we dare to let AI truly obtain initiative.
In the past few years, the entire AI industry has been asking: How much smarter can the model become? This question is of course still important. But more importantly: Can we make this intelligence cheap enough, sustainable enough, and free enough from usage burdens, so that people are finally willing to hand over things to it?
When users no longer need to plan carefully for every thought of AI, Agents can truly have space for repeated attempts; when AI can work continuously at low cost for a long time, it has the opportunity to gradually change from a "talking tool" to a system that can take the initiative to complete tasks.
In a sense, this may be the threshold that AI needs to cross before it is truly popularized.
This article is from the WeChat official account "Digital Intelligence and Humanism View", author: He Fei, published with authorization from 36Kr.