The AI price war has staged a comeback. What exactly is going on behind the phenomenon of far more powerful yet much cheaper offerings?
Are large language models starting a new round of price war?
Lei Technology (ID: leitech) has noticed that over the past few days, OpenAI, Anthropic, Xiaomi and SpaceX AI have launched new models almost one after another. GPT-6 Sol/Luna, Claude Opus 5.5, MiMo-V2.6 and Grok 4.7, different in names and positioning, all make intelligence far more "affordable".
OpenAI took the most direct move: the price of GPT-6 Sol has been almost halved compared to the previous generation, and Luna is further pushing prices down. The unit API price of Claude Opus 5.5 has also been cut by 20%, and Anthropic claims that the actual cost for typical tasks has been reduced by about 40%.
Of course, it is not all about "price cuts". Xiaomi's MiMo-V2.6 and SpaceX AI's Grok 4.7 both maintain the price of the previous generation while further improving performance. But for developers, the difference is not that significant:
The same amount of money is buying far more machine intelligence at an increasingly fast pace.
Image source: Lei Technology @GPT
This phenomenon is nothing new. Over the past few years, the price of large language model APIs has been on a continuous decline, especially with the fierce competition among domestic open-source large model vendors, the price per million tokens has been lowered significantly. What makes this round different is that these changes are taking place at a very special point in time.
On one hand, OpenAI and Anthropic have started to publicly discuss slowing down the development pace of cutting-edge models, worrying about the safety alignment risks brought by the excessively fast growth of AI capabilities. On the other hand, several model vendors have unanimously accelerated the process of making their already available capabilities cheaper and more efficient.
At the same time, Agents are moving from chat windows to programming, research and office scenarios. A single task may no longer consume thousands of tokens, but hundreds of thousands, millions of tokens, or even continuous model calls for several hours.
More powerful and cheaper: where exactly do the "cost savings" come from?
Why can large language models become more powerful while getting more affordable? According to the explanations given by several vendors, the direct reason lies in the improving efficiency of both model training and inference.
Compared with the past when vendors simply relied on expanding the pre-training scale, there are significantly more areas where they can "optimize efficiency" today.
OpenAI directly attributes the price cuts of GPT-6 Sol/Luna to the improvements in caching and inference efficiency. The two models inherit part of the training methods and capabilities of GPT-6 Astra, but further reduce costs on the inference side. OpenAI states that the saved costs can thus be directly reflected in the API prices.
Anthropic has gone a step further. The unit API price of Claude Opus 5.5 has only decreased by 20% compared to Opus 5, but Anthropic says the actual operating cost for typical tasks has dropped by 40%. In addition to the lower cost per single token, the new model requires fewer tokens to complete the same task, and the cache read price has also decreased by 60%.
Image source: Anthropic
The model becoming smarter itself has also started to act as a form of "price reduction".
Xiaomi has put more efforts into the training end. MiMo-V2.6 does not lower the API price, but further expands the scale of Agentic reinforcement learning on the basis of V2.5.
During less than 6 days of live-streaming reinforcement learning training, Flash and Pro generated a total of about 750,000 trajectories. Xiaomi states that through a larger Batch, more complex task environments and more refined reward mechanisms, the model is enabled to complete tasks with shorter paths and fewer tokens.
Eventually, V2.6 Flash has comprehensively surpassed the previous generation V2.5 Pro, while the price remains unchanged.
SpaceX AI has taken a slightly different path. Grok 4.7 does not shrink the model, but instead uses a larger base model, conducts longer-term reinforcement learning for more difficult tasks, and strengthens the model's capabilities of self-verification, long context and Agent.
Finally, while maintaining the same API price as Grok 4.6, it further improves the performance in coding and Agent tasks.
The four technical paths are not completely identical, but they lead to the same outcome. Today, from post-training reinforcement learning, inference optimization, prompt caching, to reducing invalid tokens and tool calls by the model, model vendors are exploring efficiency across the entire chain.
This is also the most direct technical background for this round of "price war".
Large language models are certainly still getting stronger, but more and more capability improvements no longer need to be exchanged for simply higher inference costs. For developers, the final manifestation is a very simple fact: the same task can be completed by the model with less cost.
While stepping on the brake, why accelerate "price reduction"?
However, the improvement of technical efficiency alone is not enough to fully explain this round of "price war".
Just a month ago, OpenAI announced that it would temporarily slow down model scaling, including pausing the reinforcement learning training of the latest deployed models for two weeks, and the originally planned largest-scale cutting-edge reinforcement learning training will also be put on hold. There are two direct reasons: one is the internal Agent safety incident, and the other is that GPT-6 Astra has reached the threshold of "key cybersecurity capabilities".
The core reason is that the capabilities of cutting-edge models are growing too fast, and monitoring, alignment and safety measures need time to catch up.
Dario Amodei, CEO of Anthropic, went a step further and publicly proposed "Pace the Frontier", calling on the entire industry to slow down the growth rate of capabilities of cutting-edge AI models. Anthropic even specifically mentioned this again in the newly released Claude Opus 5.5.
Image source: wikimedia
At first glance, this is somewhat contradictory to the current "price war".
But one point that is easily overlooked here is that OpenAI and Anthropic are not calling for slowing down the technological progress of the entire AI industry, but slowing down the speed at which the cutting-edge capability boundary continues to expand outward.
According to Amodei, "Pacing" does not mean stopping model training or technological progress. On the contrary, he believes that today's models themselves are huge "research gold mines", which are worth investing more time to understand, align and utilize.
This just provides another perspective to understand the current round of price war.
OpenAI has provided a very intuitive set of data. After imposing stricter security restrictions on Astra in August, the GPU allocation for Astra-class models once dropped by 59.2%, but the GPU allocation for other model categories immediately increased by 17.2%, offsetting about 85% of the previous decline. The computing power did not sit idle, but flowed to other models and experiments.
Of course, this cannot prove that the price reduction of GPT-6 Sol/Luna is the direct result of Astra's slowed-down development. But it at least shows that the competition among top model vendors is not limited to the single path of "training more powerful models".
The same amount of computing power and R&D investment can be used to break through the next capability upper limit, or to optimize existing models, making the already obtained capabilities cheaper and more reliable through reinforcement learning, inference optimization, caching, distillation and engineering improvements.
The rise of Agents has further amplified the value of the latter path.
In the era of chatbots, a single response may only consume thousands of tokens. But in the Agent era, models start to work continuously for several hours, repeatedly reading contexts, calling tools, performing self-checks and even scheduling other Agents. Every drop in the price of the model may directly determine whether an Agent product can operate on a large scale.
But Agents have another side. When Amodei explained why he began to support slowing down cutting-edge models, he mentioned a core change, which is RSI recursive self-improvement, in layman's terms:
AI is increasingly participating in the R&D of the next generation of AI.
Anthropic disclosed in August this year that Claude has been in a "dominant" position in 26% of internal AI R&D tasks, and more than 90% of related work has at least entered the state of human-AI collaboration. OpenAI has also observed a similar trend: Coding Agents are increasingly participating in writing code, building experiments, analyzing results and advancing model research.
Cheaper and more user-friendly models mean that more Agents can be run. More Agents can in turn participate in model R&D, improve the efficiency of researchers, and accelerate the training and iteration of the next generation of models.
The cheaper intelligence is, the easier it is to produce more intelligence. The accelerating diffusion of intelligence may be the most noteworthy part of this round of "price war".
Final Notes
Over the past few years, the price reduction of large language models has almost become a taken-for-granted thing. With more and more models, and the continuous improvement of computing power and algorithm efficiency, the price is naturally going down all the way.
But when OpenAI and Anthropic start to seriously discuss how to slow down the cutting-edge capability growth, another competition has become clearer: top-tier machine intelligence is being optimized, transferred and diffused at a faster speed.
As William Gibson wrote in Neuromancer: "The future is already here — it's just not very evenly distributed." But now, this "uneven distribution" of machine intelligence is also disappearing rapidly.
This article is from WeChat official account "Lei Technology AGI", author: 3721, published with authorization from 36Kr.