Why are corporate bills getting increasingly expensive even as the marginal cost of AI continues to decline?
Full reading takes approximately 22 minutes
The artificial intelligence industry is evolving a distinct pricing logic: capabilities of older-generation models are rapidly turning into commoditized public goods, while cutting-edge models consistently occupy the high-value ground. Michael Watkins will thoroughly break down how this mechanism reshapes the underlying logic of corporate competitive advantages.
Key Takeaways
- The cost of infrastructure-level AI capabilities is declining exponentially, yet models with industry-leading exploration capabilities maintain high prices due to massive computing power consumption and extremely fast performance iteration rhythms.
- Commoditized general AI cannot build lasting competitive moats, and the core proposition for decision-makers is to precisely identify: which capabilities are about to become universal industry standards, and which ones can truly create tangible business gaps.
- Long-term corporate competitive advantages rest on three pillars: targeted bets on cutting-edge capabilities, consolidation of complementary assets including proprietary data, trust systems and distribution channels, and agile dynamic restructuring of organizations and resources amid the cycle of large model accelerated depreciation.
At BlackRock's 2026 US Infrastructure Summit in March 2026, OpenAI CEO Sam Altman painted a picture: "In the future, intelligence will become a universal utility like water and electricity, with customers paying by usage and accessing it on demand."
Facing the continuously rising computing power budget curve, pinning the AI pricing logic on "utility-ization" undoubtedly gives corporate executives a peace of mind. Utilities inherently mean low marginal costs, and affordable intelligence seems to indicate that AI will eventually move out of management's cost agenda — just like the internet broadband and long-distance calls back in the day: with early layout and waiting for technology popularization, the cost problem will naturally resolve itself.
But the reality is the opposite. The development of broadband has long proven this cognitive misunderstanding: even though the cost of transmitting one megabyte of data has dropped sharply, enterprises' overall network expenditure has never decreased. When broadband is no longer scarce, corporate demands quickly upgrade from simple text emails to ultra-high-definition video streams.
The expansion speed of computing power consumption always far outpaces the decline of unit price. In the AI computing power field, this rule not only holds true, but the iteration pace is even more rapid.
When OpenAI first opened the GPT-3 API at the end of 2021, the charge for 1 million Tokens (approximately 750,000 English words) was 60 US dollars; only three years later, the cost of a model with equivalent performance processing the same amount of text has dropped to 6 cents, a drop of as much as 1000 times.
But at the same time, the expenditure of leading enterprises on calling the latest cutting-edge models has never decreased. The popularization of intelligence does not land synchronously, but presents a distinct "staggered peak price reduction" feature: the capabilities of the previous generation of models are rapidly approaching free, while current cutting-edge capabilities remain highly priced. This law of business evolution is unlikely to change in the foreseeable future.
Uber's CTO previously revealed that the company's 2026 annual AI budget was exhausted in only four months — the root cause being that 5000 engineers fully adopted Agentic programming tools, with the per capita monthly computing cost reaching $500-$2000.
Breaking Down Two Critical Curves
The first curve is the unit price trend of fixed-performance AI capabilities. Industry data provides clear evidence: computing power prices are plummeting. After evaluating six core benchmarks, AI cost research institute Epoch AI found that the computing power cost to reach specific performance standards drops by 9 to 900 times per year, depending on the task type.
Taking solving doctoral-level scientific problems as an example, the unit cost of reaching GPT-4 level performance is reduced by about 40 times per year; estimates from well-known venture capital firm Andreessen Horowitz also show that the average market price of models with equivalent performance drops by 90% every year.
The core driving force behind the rapid price decline is fierce market competition, among which the impact from China's open source forces is particularly prominent. Open source weight models launched by DeepSeek, Alibaba, Zhipu AI, Moonshot and Meta have anchored a solid price bottom line for the entire industry: once a high-level capability is successfully replicated by an open source model, enterprises can deploy it locally on their own hardware, and the high-premium pricing power of closed-source vendors will instantly collapse.
What deserves more attention is that this price bottom line is shifting upward at an accelerated pace — by the middle of 2026, the technology generation gap between top open source models and proprietary cutting-edge models has narrowed from the initial "several years" to "several months". It is this continuously rising price bottom line that provides practical support for Altman's judgment that "intelligence will become abundant".
At the same time, the second curve is rising sharply in the opposite direction. When old technical capabilities become universal, the "cutting-edge capabilities" representing the industry's top level still firmly hold pricing power. When OpenAI launched its first model with cutting-edge reasoning capabilities, the initial pricing remained at a high of $60 per million Tokens. More severely, the computing power overhead to complete a single complex task grows exponentially.
Traditional question-and-answer interactions usually only trigger one model call; while an Agentic System needs to go through multiple rounds of iterations including goal decomposition, tool invocation, result verification and self-correction to complete a task, and a large amount of context information needs to be resubmitted at each step.
According to Gartner's estimates, the Token consumption of agentic workflows is 5 to 30 times that of traditional chat interactions; a software engineering benchmark study on eight cutting-edge models in April this year shows that the Token overhead of agentic programming tasks is 1000 times that of equivalent dialogue requests, with costs mainly coming from frequent repeated input contexts rather than generated results. Even for different running instances of the same task, the fluctuation of computing power consumption can reach 30 times; the models themselves also have a tendency to systematically underestimate overhead.
Combining the two curves, the cost paradox plaguing management gets its answer. An enterprise's AI bill is the product of "unit price" and "total consumption". Assuming that the single cost of a certain AI task was $1 last year and drops to 10 cents this year (a 90% drop), if the agentic workflow makes the invocation frequency of a single task surge 50 times, the actual bill will rise from $1 to $5. The cost dividend brought by price drops is completely offset by the explosive growth in usage.
This also explains why at the moment when the AI industry is witnessing the biggest price cut in history, enterprises' model expenditure keeps hitting new highs. Data from Menlo Ventures shows that in 2025, enterprises' investment in large language models doubled within half a year; Gartner predicts that global enterprises' expenditure on AI models and platforms will exceed $64.3 billion in 2026.
The micro-level impact is equally significant: Uber's CTO admitted in April this year that the company's full-year 2026 AI budget only supported four months, the core reason being that after agent programming tools covered 5000 engineers, the per capita monthly cost climbed to $500 to $2000. "Unit price plummeting, total expenditure soaring" is no longer an isolated phenomenon, but a common problem faced by executives across the industry.
The differentiated barrier that is still very valuable this quarter may become a free general capability next quarter. The technological watershed is always evolving unidirectionally and never stops moving.
The Continuously Shifting Watershed
Only by examining the two curves at the same time can we see the true face of the AI business landscape: model capabilities will never stay at the cutting-edge high ground, but fall towards the price bottom line at an accelerated rate with a very short half-life.
Differentiated barriers that are of high value this quarter may become free general capabilities next quarter, and the new generation of cutting-edge capabilities will immediately take their place at the top. This dynamic shift of the watershed never stops, and always evolves unidirectionally.
The "generational iteration mechanism" in the automotive industry is the best reference to understand this phenomenon: old models depreciate and reduce in price year by year, while new models always maintain high prices, and the cycle repeats. But it must be made clear that the "model year" in the AI field is not measured by the natural year — the actual technology iteration cycle has been compressed to the quarterly level.
If corporate executives still follow the traditional annual planning rhythm, their decision-making rhythm will actually lag behind the market by three full technology cycles.
When Will AI No Longer Constitute a Competitive Advantage?
Classical strategic management theory points out that for a resource to build long-term competitive advantage, it must meet the criteria of value, rarity and non-substitutability (namely the VRIO framework). Although previous-generation AI models have extremely high value, they have lost rarity and imitation barriers. When a certain capability falls to the free price bottom line, it naturally becomes the general infrastructure for all competitors and even new entrants.
AI applications are rapidly degrading from "differentiated tools" to "industry entry thresholds"; those enterprises that were complacent about successfully deploying AI in 2024 to 2025 will eventually find that these measures do not build lasting competitive moats.
Strategic academia has long been familiar with this evolutionary path. When an innovation is of great value but easy to replicate, the subject that ultimately captures commercial value is often not the direct adopter of the innovation, but the entity that masters the "complementary assets" required for value creation.
In the AI era, these irreplaceable complementary assets include: proprietary data that cannot be crawled, strict regulatory access qualifications, mature distribution channels, physical infrastructure, institutional-level trust, and the capital strength to maintain continuous investment during the critical period of cutting-edge technology evolution.
However, the "model generational iteration mechanism" adds a more brutal dynamic attribute to the classic strategic framework. When the external environment re-prices AI capabilities on a quarterly basis, even if an enterprise has abundant static assets, if it cannot complete agile organizational restructuring around the new generation of AI capabilities, its original advantages will continue to be eroded.
The truly lasting competitive advantage comes from the "dynamic capability" defined in strategic management: the insight to continuously identify the "generation" to which AI capabilities belong, the technical execution to roll out commoditized capabilities faster than competitors, and the organizational flexibility to precisely allocate talent and capital as the price bottom line shifts downward. Static moats determine your current competitive position, and restructuring iteration speed determines whether you can maintain long-term leadership.
The cognitive framework of the "model generational iteration mechanism" will fundamentally reshape the decision-making logic of executives in all industries.
Strategic Implications for Corporate Executives
The cognitive framework of the "model generational iteration mechanism" will fundamentally reshape the decision-making logic of executives in all industries. The healthcare industry is a typical sample that confirms this evolutionary logic, as it has all structural features including proprietary data, strong regulation, licensed professionals and low profit margins.
Imagine a regional non-profit healthcare group with annual revenue of $10 billion, operating more than ten hospitals and a large outpatient network, while 40 AI projects are competing for limited capital budgets.
Traditional decision-making processes usually focus on calculating return on investment (ROI), but today's executives should rather ask: "Which projects will become free universal capabilities in 18 months?"
Foundation Model Capabilities: For example, AI medical note assistants that listen to consultations and automatically generate medical records, tools that automatically match billing codes, systems that perform intelligent triage on patient messages, and call centers that automatically handle routine requests. Such scenarios exist in all industries: primary contract review in the legal industry, front-line after-sales support in software development, and automatic reconciliation in the financial field.
Although such capabilities have business value (a multi-center study covering six major US healthcare systems shows that the related administrative burden is significantly reduced), the technical threshold has been completely broken through, and competitors can catch up quickly at very low cost. The efficiency dividend brought by commoditized capabilities will be reflected in operating profit margins and employee retention rates, but cannot change the competitive landscape.
Therefore, the procurement of such capabilities must follow the "short-term" principle: insist on signing short-term agreements, never lock in the current high price for capabilities whose prices keep falling, and always anticipate that the price bottom line will drop further.
Cutting-edge Model Capabilities: In the medical field, this includes handling complex insurance claim denial appeals, optimizing cross-hospital patient flow scheduling, assisting in clinical diagnosis of difficult diseases, and accurately matching clinical trials. In other industries, this corresponds to multi-step composite transaction analysis, dynamic supply chain optimization, and cutting-edge engineering design.
Such tasks are highly dependent on the reasoning capabilities of the latest generation of models, each call generates high actual costs, and the output stability still has room for improvement.
But this is exactly the core strategic battlefield where enterprises can create differentiated competitive gaps. In this regard, executives should adopt a "lean and focused" time-limited pilot strategy, preset clear hard clinical or financial indicators; quickly incorporate projects that meet the standards into the long-term business system, and decisively terminate projects that fail to meet the standards, never falling into the "infinite pilot loop" — because after the next generation of models is launched, the evaluation conclusions of the previous generation of pilots will immediately become invalid.
Lasting Complementary Assets: In the medical scenario, a typical representative is a 20-year longitudinal patient medical record database, as well as its in-depth associated data with specific populations' clinical prognosis. Such assets cannot be obtained through public data crawling, and competitors are difficult to replicate in the short term. In other fields, this corresponds to rich customer historical behavior data, franchise licenses, exclusive distribution networks and physical service outlets.
In addition, "legal and ethical accountability" is an extremely solid trust asset. No patient or regulator will allow a model to completely decide a patient's discharge, and the model itself cannot bear the legal liability for medical malpractice as an entity. The same applies to audit opinions, engineering safety certifications, and fiduciary legal consultations.
In knowledge-intensive service industries, accountability itself is the core value of the product. It is even the most lasting barrier against technology depreciation: even if AI capabilities are rapidly iterating and depreciating with "model years", the demand of regulators, judicial systems and patients for "human expert endorsement" is extremely rigid. This is also the core confidence for established enterprises to still hold a solid market position when facing challengers with abundant capital and access to the same AI models.
The core input factors driving the evolution of AI capabilities in the next few years have long been determined at the physical level: computing chip procurement agreements have been signed, long-term power supply contracts have been completed, and capital for data center construction has been in place.
Three Potential Scenarios of the Evolution Cycle
The real uncertainty lies in the evolution of the "generational gap": how much leading advantage cutting-edge capabilities can maintain, and whether this gap will continue to expand at an accelerated pace.
The hardware investment for AI capability evolution in the next few years is basically determined: chip procurement is finalized, power contracts are in place, and data center capital expenditures are arranged.
The certainty of hardware investment is very high, and the real unknown is how much effective intelligent output these inputs can be converted into. The following three scenarios outline the possible directions for the future.
Scenario 1: Commoditization Convergence
Open source models completely close the capability gap with cutting-edge models within 12 months, and the marginal improvement of cutting-edge capabilities in actual business scenarios is very limited. At that time, intelligence will completely become a general tool, and competitive advantages will return to traditional operational efficiency, proprietary data and strategic execution.
- Strategic Direction: Minimize the budget for cutting-edge capabilities, and fully focus on deeply implementing commoditized AI capabilities within the enterprise.
- Monitoring Signals: Open source models quickly top the cutting-edge leaderboards; cutting-edge vendors sharply reduce API prices under competitive pressure; performance gaps cannot be converted into commercial returns in actual business processes.
Scenario 2: Cutting-edge Compound Interest
The technology generation gap remains at