HomeArticle

GPT-6.1 Sol Extreme Edition is launched, and its $500 entry threshold has angered users: you pay for faster speed, yet your usage quota is burned through far more drastically?

AI前线2026-10-09 12:05
In the era of intelligent agents, speed has begun to be priced separately.

OpenAI has added a new speed tier for agents.

According to the official update log, on October 8, OpenAI added the Ultrafast mode for GPT-6.1 Sol in the Responses API, which is open to all API users and still subject to rate limits. Developers do not need to change the model name, and can enable this mode by setting the service tier to ultrafast when calling gpt-6.1-sol.

This update is mainly focused on generation speed.

The GPT-6.1 Sol model itself was released on September 29. Therefore, the more accurate meaning of the "launch of the ultrafast version" is that a faster service tier has been added to the existing model. OpenAI positions Ultrafast as its fastest API service tier, designed for latency-sensitive workloads where users are willing to pay a premium for speed.

The price has also increased accordingly: the unit price of API Tokens for the GPT-6.1 Sol Ultrafast mode is 6 times that of the standard mode. This shifts the focus of this update from "how smart the model is" further to "how much value is there in cutting down waiting time".

Officials claim performance is close to Astra

OpenAI positions GPT-6.1 Sol as a model that delivers performance close to GPT-6 Astra at a lower cost in complex programming, computer operation and professional work scenarios.

In fact, the GPT-6.1 Sol model does have many outstanding performances in some benchmark tests.

Compared with GPT-6 Sol, GPT-6.1 Sol has significantly improved performance in complex professional tasks, covering code writing and debugging, document understanding, and execution of multi-step business workflows. In many of these evaluations, it achieves performance close to GPT-6 Astra at a significantly lower cost.

In terms of programming, in the DeepSWE v1.1 test, GPT-6.1 Sol reaches a level comparable to GPT-6 Astra at about one-fifth of the cost, and outperforms the highest score of GPT-6 Sol by 6.4 percentage points with lower inference intensity and lower cost.

In terms of professional work, under all tested inference settings of GDP.pdf, the score of GPT-6.1 Sol is higher than that of Opus 5.5 with fallback mechanism, while the cost per task is less than half of it. It also approaches the performance of GPT-6 Astra at about one-fifth of the per-task cost.

AutomationBench measures whether agents can correctly complete multi-step business workflows. In this evaluation, GPT-6.1 Sol scores 2.2 percentage points higher than Opus 5.5 at medium inference intensity, with the cost being about one-third of the latter. This score is also 4.8 percentage points higher than GPT-6 Sol under the same settings.

On X, Tibo, the head of Codex, posted: "We have improved several features to make the model respond faster, so that it can respond to user adjustments more quickly, allowing you to correct the direction in real time and prevent the model from wasting resources. At the same time, we have also released the ultrafast version of GPT-6.1 Sol. The two work extremely well together."

At the same time, the official still positions Astra as the highest-capability model for handling the most difficult tasks, and recommends users to compare the two on their own tasks to judge the trade-off between quality and cost.

Judging from the public specifications, Sol and Astra have the same context capacity and maximum output length, while there is a clear gap between the unit prices of input and output in the standard mode.

Source: OpenAI official model comparison page.

Calculated by the same number of Tokens, the standard input and output fees of Sol are 80% lower than those of Astra. However, this proportion cannot be directly converted into "80% cheaper to complete the same task": the actual number of inference and output Tokens generated by the model, whether retries are needed, and how many tools are called will all affect the final bill.

This Ultrafast update also needs to distinguish two types of "performance": one is the quality of task completion, and the other is the speed of task completion.

How is the price calculated?

The standard mode, Fast mode and Ultrafast mode of GPT-6.1 Sol are charged at different unit prices respectively. For requests with input not exceeding 272,000 Tokens, the API prices are as follows:

Source: OpenAI official API pricing page.

This means that the unit price of the ultrafast mode is 6 times that of the standard mode and 3 times that of the fast mode. It needs to be emphasized that 6 times the price does not mean 6 times the speed, and the two cannot be directly mapped.

Take a simplified bill as an example: assuming a request consumes 100,000 uncached input Tokens and 10,000 billable output Tokens, excluding other costs such as tool usage, the cost of the standard mode is about $0.30, the fast mode is about $0.60, and the ultrafast mode is about $1.80. This is a calculation example based on the official unit price, not an actual task test.

Long context will further increase the price. When the input exceeds 272,000 Tokens, the input price of GPT-6.1 Sol Ultrafast mode rises to $24 per million Tokens, and the output price rises to $90. The official stipulates that the long context multiplier applies to the entire request, not just the part exceeding the threshold.

Another easily overlooked comparison is: the input and output prices per million Tokens of Sol Ultrafast mode are $12 and $60 respectively, which are already higher than the $10 and $50 of Astra standard mode. Calculated by the same amount of Tokens, the former is 20% higher. Therefore, the "lower cost" advantage of Sol needs to be judged in combination with the service tier; after choosing the ultrafast mode, the value users purchase comes more from reduced waiting time.

On the API side, the Ultrafast mode of GPT-6.1 Sol is open to all API users, with independent rate limits separate from the standard and fast modes. It supports global processing, as well as US and EU data residency.

The scope of opening on the ChatGPT product side is narrower. According to the official documentation, GPT-6.1 Sol is available in Work and Codex, and this model is not available in the normal Chat mode. The initial launch scope of the Sol model covers Plus, Pro, Business, Enterprise and Edu, but the ultrafast mode is only open to the $500 Pro tier and eligible Enterprise and Edu plans. The actual availability also depends on the client, workspace settings and rollout progress.

The consumption of subscription quotas cannot directly copy the price multiplier of the API. The official pricing description shows that the GPT-6.1 Sol Ultrafast mode consumes the included usage quota in the package at 8 times the speed of the standard mode; for purchased extra credits and enterprise pay-as-you-go usage, a 6x multiplier is applied.

Users: We need more quota, you give us speed

Around the GPT-6.1 Sol Ultrafast mode, users have launched heated discussions on platforms such as X and Reddit. The discussions mainly focus on two issues: whether speed improvement can make agents really practical, and whether users should bear the significantly higher cost for this.

Supporters believe that the significance of speed for agents goes far beyond "answering questions faster".

On X, a user with the ID HellFlow Studios said that when AI starts to actually perform tasks, faster inference combined with strong capabilities can improve the practicality of automation, especially in programming, agent and real-time workflows.

Brian Gastón and Shrilaxmi focus on the work cycle that agents execute repeatedly: the latency in multi-round inference and tool calls will accumulate continuously, and speed determines whether such products can become daily tools.

However, Brian also pointed out that it remains to be seen whether the ultrafast mode can maintain its performance in tasks that call tools continuously for a long time.

Some other users said that what everyone needs is longer usage time, not faster models. If you never run out of your quota, what's the point of ultra-fast speed???

Clarity Today reminds that the billing standards of API, Codex and Work are different, and the budget cannot be directly copied; whether the speed improvement is cost-effective depends on whether the user's waiting time is mainly consumed in model generation or external tool execution.

The access threshold and quota consumption of the ultrafast mode have also caused dissatisfaction among some paid users.

On X, user DoubleStraddle targeted the package tiering, sarcastically saying that this release makes users who pay $200 per month feel left out, as if only users on the $500 tier can get new features.

Wei Jia said that the Steering function, which allows users to adjust the direction of the agent during task execution, is the improvement he can feel every day; as for the ultrafast mode, he will consider it when it is no longer limited to the $500 package.

The quota issue is another point of controversy. Dotimus believes that under the existing usage limits, the ultrafast mode consumes the package quota at 8 times the speed, which does not solve the actual difficulties of users in completing their work. He called on OpenAI to prioritize improving quota limits instead of continuing to sell higher speed tiers.

Luca Hagenmayer, who claims to subscribe to the $500 monthly package, also said that his quota can't even last a day, so it's hard for him to understand the value of enabling the ultrafast mode.

Lorenzo believes that if the base speed already makes users feel slow, a faster experience should be a basic requirement, not a privilege of high-priced packages.