GPT-5.6 Sol sees a steep price cut, OpenAI brings the price war to its flagship product lineup.
BREAKING: GPT-5.6 Sol Launches Significant Price Cuts
Over the next three months, the API price and credit price (the pay-as-you-go usage beyond the plan-included quota) will be reduced by more than 20% in total.
At present, the price adjustment on the API side has taken effect immediately, and eligible ChatGPT Work and Codex credits are being pushed out in succession.
However, the three tiers of subscriptions, namely Pro, Plus and Business, see no price reduction at all. The included usage volume in the packages has not been increased either.
All the price cuts apply to the part billed by tokens. In other words, the more API calls a user makes, the more they can save.
How Much Can You Actually Save?
The previous API pricing of Sol was $5 per million input tokens and $30 per million output tokens, making it the most expensive model in the GPT-5.6 family.
Now check the pricing page, the new price is $4 for input and $20 for output.
The input price has dropped by 20%, which is a reasonable cut. But the output price has been slashed from $30 to $20, a full 1/3 reduction.
Therefore, the official statement that the price cut is "over 20%" is relatively conservative. For the most expensive output tokens, the price has been reduced by 33% at one go.
Assuming a large Agent consumes 100 million input tokens and 20 million output tokens per month, the original cost would be $1100, while now it only costs $800, with a comprehensive reduction of about 27.3%.
The exact amount you can save depends entirely on your input-output ratio. For batch processing scenarios that are heavy on input like document processing, you can save about 20% more; for scenarios where output takes the majority share like intensive coding in Codex, the savings can reach over 30%.
In short, the more heavily you use Sol as a productivity tool, the more you can save this time.
The caching tier pricing has also been adjusted accordingly: cached input price is reduced from $0.5 to $0.4, and cache write price is reduced from $6.25 to $5. The long context window pricing is also cut, with input price reduced from $10 to $8 and output price reduced from $45 to $30.
OpenAI Developers later commented that the price cut is made possible by the significantly improved operational efficiency of Sol.
Final Showdown of the Two ASI Giants: Pricing Becomes the Second Battlefield
During OpenAI's last round of price cuts, Sol was the only model whose price remained unchanged.
The price of Luna was cut by 80% at one go, and the price of the main model Terra was reduced by 20%. Only the flagship Sol kept its original price, with an additional Fast mode that offers faster speed at no extra cost.
At that time, OpenAI had a very clear plan: the cheaper models are responsible for driving high volume, while the most expensive flagship model supports the brand premium.
This time, it's Sol's turn to cut prices. And this timing is clearly no coincidence.
Chubby, a well-known tech influencer, put forward a straightforward judgment: This move is clearly targeted at Anthropic.
Anthropic is not in a good situation at the moment. It has not yet digested the negative feedback on Opus 5, and users are complaining about the rate limit issues in the Fable 5 subscription plan.
By launching this large price cut now, OpenAI is aiming to poach Anthropic's users.
Tech journalist Tae Kim made a more pointed comment: OpenAI is using commercial competition tactics during Anthropic's IPO roadshow.
This statement is not groundless.
Anthropic submitted a confidential S-1 filing to the SEC on June 1 this year, officially launching its IPO process. According to multiple media reports, underwriting banks have arranged meetings between the management and potential investors since mid-July, and the roadshow window is exactly from August to September.
The target listing timeline: as early as October, on the Nasdaq, with a valuation targeting $2 trillion.
What a company fears most during the roadshow is disruption from competitors.
After the benchmarking competition comes to an end, pricing has become the second battlefield in the final showdown of the two ASI giants.
OpenAI's aggressive price cuts are not an impulsive decision.
Its early aggressive self-built computing power investment and bets on more efficient model architectures are now starting to show their advantages.
A few days ago, OpenAI just announced that it has secured about 8 IT-GW of computing power at the PORTS-Pike campus in Pike County, Ohio, with a 20-year lease signed; NVIDIA will exclusively provide the AI computing infrastructure.
Previously, OpenAI also reached a 5-year cooperation agreement with Oracle worth more than $300 billion, corresponding to a maximum of 4.5GW of new capacity, and signed a 6GW GPU deployment agreement with AMD.
OpenAI confirmed in April this year that its locked-in AI infrastructure capacity exceeds 10GW, with a target of expanding to 30GW by 2030.
The more complete the infrastructure is, the lower the marginal cost will be, and the greater the room for price cuts. This is a positive flywheel: improved model efficiency → reduced inference cost → released pricing space → growing user base → increased revenue → further investment in infrastructure.
It can be said that OpenAI's flywheel is taking off. The company itself also stated that GPT-5.6 Sol has participated in optimizing its own production inference kernel, cutting the end-to-end service cost by 20% and increasing the token generation efficiency by more than 15%.
However, in addition to actively creating competitive pressure, this price cut is also partially a forced move.
DeepSeek V4 Flash was officially released on July 31, with an extremely low price that is almost negligible.
What's more stimulating is a mysterious model codenamed "Ox Alpha" that emerged two days ago. It has a 1M context window, supports text, image and video input, and is available for free for one week.
Now China's open-source models are not only cheap, but their capabilities are also constantly approaching the frontier models.
Therefore, this round of price cuts for Sol is not only an active offensive against Anthropic, but also a defensive counterattack against low-cost models from China.
Codex Monthly Active Users Exceed 20 Million
There is another set of data announced on the same day.
Tibo, head of Codex, announced that the weekly active users of Codex have exceeded 20 million this week.
To celebrate this milestone, the team has issued a one-time BANKED quota reset to all Codex and ChatGPT Work users.
However, regarding the community feedback that the quota is consumed too fast, Tibo said no abnormalities have been found so far, but a formal investigation has been launched, and results will be shared as soon as they are available.
He also made it clear that the official does not support the act of converting subscription accounts into API traffic through sub2api for resale or multi-person sharing. Such usage will be directly marked by the risk control system.
The practice of reselling one monthly subscription account to a hundred different users is now completely blocked.
Normal usage through the official client, or open-source clients such as Pi and OpenCode, is completely allowed.
References:
https://x.com/OpenAI/status/2090885187634905500
https://x.com/OpenAI/status/2090885188897460249
https://x.com/kimmonismus/status/2090890956287492564
https://x.com/rohanpaul_ai/status/2090885980849050036
https://x.com/rohanpaul_ai/status/2090891370663989566
https://x.com/thsottiaux/status/2090766694897619318
This article is from WeChat Official Account "AI Era", Author: ASI Insights, Editor: Solomon, Published with authorization from 36Kr.