GPT-6 Sol and Luna are launched, offering even steeper discounts than Liang Wenfeng.
Tonight is destined to be a sleepless night.
Opus 5.5 was released just 90 minutes ago, and OpenAI has launched GPT-6 Sol and Luna together.
Right after Anthropic cut the price of Opus, OpenAI slashed the price even deeper.
GPT-6 Sol costs only $2 per million input tokens and $10 per million output tokens, which is exactly half the price of Opus 5.5 and one fifth of Astra's price.
Luna is even more cost-effective, with prices of only $0.1 for input and $0.5 for output, directly dropping to 1% of Astra's price.
Throughout this night, before model capabilities could be fully verified, a price war has already broken out.
01
Low Price Without Performance Degradation
Astra remains the top-tier model in the GPT-6 series. OpenAI explicitly states that if your tasks are critical, highly complex, and you refuse to make any compromises on capability, you should still choose Astra.
The mission of the newly launched Sol and Luna is to bring Astra-level capabilities to much more affordable models.
OpenAI says the two models adopt a training approach similar to Astra, and "downstream" this generation's advances in professional work, factual accuracy, Coding, Computer Use and Alignment.
Among them, GPT-6 Sol is positioned as the main model for complex Coding and Agent workflows, while Luna is responsible for high-frequency, large-scale, relatively focused tasks, with core advantages of speed and low cost.
Both models feature a 1.05 million token context window and a maximum 128K token output, which is on par with Astra. They both support text and image input, and can normally call tools such as web search, file search, code interpreter, Computer Use, MCP, and Skills.
They even have one more reasoning mode option than Astra. Astra only has low, medium, high, xhigh and max, while Sol and Luna add an extra "none" mode that can turn off reasoning. The default mode for both is medium.
Sol's knowledge cutoff date is April 20, 2026, while Luna has a more recent cutoff date of May 18, 2026. Interestingly, Astra's knowledge cutoff date is April 30, 2026, falling between the two new models. It seems that the knowledge cutoff date is not directly equivalent to the model's "tier level".
Then look at the prices. GPT-6 Sol charges $2 per million input tokens, $10 per million output tokens, and only $0.2 for cached input. GPT-6 Luna is even more competitive, charging $0.1 per million input tokens, $0.5 per million output tokens, and cached reading is as low as $0.01 per million tokens.
And this is not the lower limit of the price. The Batch and Flex modes continue to offer 50% off, and input tokens that hit the cache are also charged at only 10% of the regular input price.
However, the million-level context cannot be calculated at this price from start to finish.
OpenAI stipulates that once the input of a single request exceeds 272,000 tokens, the entire request will be charged at the long-context price: the input and cache prices will be doubled, and the output price will be increased by 50%.
Therefore, after Sol uses the ultra-long context, the price becomes $4 for input and $15 for output; for Luna, it becomes $0.2 for input and $0.75 for output.
The 1.05 million context window is real, and the $0.1 price is also real, but the two numbers cannot be multiplied directly.
The price has dropped dramatically, but the performance has not been compromised.
OpenAI uses Sol to prove that tasks that previously required flagship models to complete at high cost can now be done at a much lower cost.
Take the cross-application Agent test AutomationBench for example — since OpenAI's documentation does not include the newly released Opus 5.5, we refer to Zapier's public leaderboard.
In this test, GPT-6 Sol at xhigh mode scored 33.2%, with a single-task cost of only $0.27.
If we only look at absolute capability, it still cannot beat the newly released Opus 5.5, which achieved 40.0% at max mode, second only to Astra max's 41.4%. But each task of Opus 5.5 costs $1.28, almost 4.7 times that of Sol; Astra costs $1.73, 6.4 times that of Sol.
Although Sol did not take the first place, it has squeezed into the same cutting-edge Agent leaderboard with a price of less than $0.3.
The same applies to Coding.
In the real codebase software engineering test DeepSWE v1.1, GPT-6 Sol max scored 68.8%, only 1.1 percentage points lower than the previous highest score of 69.9% from Fable 5, but OpenAI estimates that the single-task cost is about 80% lower.
Luna is even more impressive. This model that sells for only $0.1 per million input tokens also reached a 66.6% score at max mode.
OpenAI states that this performance is already close to the medium mode of Opus 5 and Fable 5, but the single-task cost is 93% and 96% lower respectively.
In the OSWorld 2.0 test for Computer Use, Sol xhigh scored 60.5%, which is basically on par with Opus 5 medium's 60.3%, but the cost is reduced by about 80%. Luna at max mode can also outperform the previous generation GPT-5.6 Sol medium, while the cost is only about one tenth of the latter.
However, it should be noted that although Opus 5.5 announced by Anthropic that day achieved a score as high as 81.8%, this set of results comes from Anthropic's own OSWorld 2.0 evaluation setup. The 60.5% announced by OpenAI is counted based on partial completion scores on the offline task set of the August 8, 2026 version. The two sides do not use the same public cross-test, so direct comparison is not applicable.
To determine which of Sol and Opus 5.5 has stronger Computer Use capability, we have to wait for a unified third-party test.
GPT-6 Sol and Luna have now been integrated into ChatGPT Work and Codex, and are available for Plus, Pro, Business, Enterprise and Edu users. Free and Go users can also use GPT-6 Luna in the desktop App.
Developers can call them directly through the API, with model IDs gpt-6-sol and gpt-6-luna respectively.
Currently, the two models are not yet available in regular Chat. The rollout for Work and Codex will be completed gradually on the same day, if you don't see them now, it may just not be your turn yet.
The familiar Tibo reset is also present as expected.
New models, permanent price cuts, plus a quota reset, this launch night has fully implemented the concept of "low cost".
02
No Explosive Benchmark Surge, But User Experience Is Changing
From the perspective of third-party organization Artificial Analysis, the intelligence of this generation of models is actually not much different from the previous generation.
In the latest Intelligence Index, GPT-6 Sol max scored 48 points, compared to 47 points for the previous generation GPT-5.6 Sol. GPT-6 Luna max scored 37 points, while the previous generation GPT-5.6 Luna even scored one point higher.
The overall capabilities of the two models have basically stayed the same, but the cost has been completely different.
Artificial Analysis measured that the average single-task cost of Sol max dropped from $1.99 of the previous generation to $1.06, and Luna's average single-task cost dropped from $0.18 to $0.07.
This is not even because the two new models have learned to "save tokens". The average number of tokens generated by Sol per task increased from 29,000 to 31,000, and Luna's increased from 41,000 to 51,000. They are simply sold at a much lower price.
Speaking of low prices, how can we not mention Chinese models?
The latest DeepSeek V4.1 Flash scored 39 points in the Artificial Analysis Intelligence Index, with an average single-task cost of about $0.27. Although the score is two points higher than Luna, Luna's single-task cost is only $0.07, almost a quarter of V4 Flash's.
But if it is called a "price killer", it still falls short — just yesterday, Xiaomi released the MiMo-V2.6 series. Among them, MiMo-V2.6-Pro scored 46 points in the same version of the Artificial Analysis Intelligence Index, only 2 points away from Sol max's 48 points. But its average single-task cost is only $0.13, while Sol's is $1.06.
If we only look at API prices, MiMo-V2.6-Pro charges $0.435 per million input tokens and $0.87 per million output tokens. The lower-tier MiMo-V2.6-Flash only charges $0.14 for input and $0.28 for output.
Unfortunately, Artificial Analysis has not yet given MiMo-V2.6-Flash a complete comprehensive index, so it is temporarily impossible to make a full-caliber cost-performance comparison with Luna.
This price war is far more than just a matter between OpenAI and Anthropic. OpenAI has pushed GPT-6's price into this range, but Chinese models have been competing fiercely here for a long time.
After talking about prices and benchmarks, it is worth noting that OpenAI has also specially adjusted the model's "response style" this time.
The official has specially listed a "collaboration mode" for Sol and Luna, which inherits Astra's improvements in communication methods: speaking more clearly, using less jargon, reducing strange wording and low-value details, the overall response will be slightly shorter, but without sacrificing truly useful information.
OpenAI believes that this change is particularly obvious in Coding and technical conversations.
The Opus 5.5 released 90 minutes ago has also made upgrades in almost the same direction. Anthropic says the new Opus will present important information earlier, reduce jargon and some Claude-specific strange expressions, and better follow the writing style specified by users.
On the same night, both OpenAI and Anthropic started to "remove the AI flavor" from their models.
Another very direct upgrade is the reduction of factual errors in the model — to be honest, this gives me a feeling of going back to last year.
Cutting-edge large models rarely take "not making up nonsense" as a selling point anymore. Although it is still important, it seems less "cutting-edge" compared to new capabilities such as Coding, Agent and Computer Use. Sol and Luna have brought this capability back this time, probably because for these models that focus on low prices and large-scale deployment, doing the most basic capabilities well is the key.
OpenAI conducted tests using a batch of real ChatGPT conversations that were previously marked with factual errors by users. The results showed that the number of factual errors made by GPT-6 Sol was reduced by about half compared to GPT-5.6 Sol, and it is beginning to approach Astra's level.
Luna has also made significant progress. At higher effort levels, its factual accuracy can reach