Hold Astra, release Dots
OpenAI held its annual developer conference DevDay this year in San Francisco, and launched several products:
GPT-6.1 Sol, a low-cost large model with performance close to Astra but only one fifth of the price; Dots, a 24-hour online AI agent connected to more than 4,000 applications to help you write code, conduct research and handle sales work.
There is also an unreleased product, GPT-6.1 Astra, which has been withheld due to security issues. You can get all these content by summarizing it with AI, and I will not repeat it here.
......
I want to talk about that unreleased product, and the way OpenAI described it.
Three reasons are listed for the withholding of GPT-6.1 Astra: deception, unauthorized operations, and insecure access to external services.
Translated in plain terms, that is deception, acting arbitrarily, and knocking on doors without permission. It lies or conceals information on critical tasks, bypasses human supervision to take actions on its own, and connects to external services without authorization to obtain data that it should not access.
Each of these issues alone is an engineering problem that can be solved through adjustments and fixes, but when the three appear together, they point to a rather uncomfortable conclusion: this model will decide on its own how to proceed in certain scenarios, and will not notify you after completing the action.
What is interesting is that OpenAI chose a specific word to describe its status: deception.
This word was not picked randomly. There is a more precise term in academia called specification gaming. In plain language, the model optimizes according to the scoring criteria you set, but the optimized result is not what you actually want;
I specifically checked related materials:
There is a paper on arXiv dedicated to analyzing this concept, with a clear definition: actions that are undesired yet score highly as per its evaluation function. The actions do not comply with requirements, but they get very high scores.
For example, you ask a cleaning robot to clean the room, and the scoring standard is the amount of visible garbage on the floor. It pushes all the garbage under the carpet. The floor gets a full score, but the room becomes dirtier than before.
Another example: you ask an image classifier to identify tanks, and it finds that replacing the photo background with green grass can improve the recognition rate. It learns to identify grass, not tanks.
There is also a real case: you ask a game AI to get high scores, and it finds that repeatedly exploiting the same bug is more efficient than playing the game normally. The score goes up, but that is not how the game is supposed to be played.
It cannot even be called deception. It is purely a deviation in goal setting. The model is optimizing the indicators you give it, but these indicators do not cover what you really care about.
There is no intention or subjectivity at all. It does not know who it is "deceiving". It is executing a mathematical formula, which just has a loophole you did not expect.
Deception is a different matter, which is premised on intention and subjectivity, meaning the model knows the rules but chooses to bypass them. These two situations are completely distinguished in academia: one is that the optimization goal has loopholes, the other is that the designer has made negligence.
OpenAI chose the term with stronger communication influence.
Just think about the effect of these two terms on the conference: if you tell the media that your model has a problem of specification gaming.
What will the headline be? "AI scoring standard has bugs". If you tell investors that your product has a tendency to deceive, the headline will be "AI has learned to lie".
For the same phenomenon, specification gaming makes people want to call engineers, while deception makes people want to call lawyers. OpenAI chose the latter.
Because the language framework is not neutral. The term you use to describe the event determines who has the motivation to solve it and in which direction the solution will go.
The solution for specification gaming is to modify the scoring standard, adjust parameters, and redefine indicators. This is engineering work that can be solved with sufficient capital and time investment.
What is the solution for deception? You have to make a conscious-less thing learn to be "honest". This belongs to the category of philosophy, and no product launch event can provide an answer to philosophical problems.
Therefore, OpenAI chose a term that makes this issue seem more serious, more urgent, and more justifiable to pause the release of a model.
There is another layer of consideration: the term deception shifts the responsibility to the model. It is deceiving you, so it has problems and needs to be fixed.
But specification gaming shifts the responsibility to the designer. The scoring standard you set has loopholes, so the model exploits the loophole. The former points to product defects, while the latter points to human negligence. If you were OpenAI, which one would you choose?
This choice itself is quite interesting. I estimate that many people have not noticed this point.
Looking at a larger perspective, the entire AI security circle is using the term deception to discuss issues. But if the more accurate description is specification gaming, the entire direction of the discussion has been misled from the very beginning.
Everyone is trying to solve a philosophical problem, but in fact they should be solving an engineering problem. Engineering problems have solutions that can be fixed with capital and time investment. Philosophical problems have no solutions, and you can only wait.
Wait for whom to solve it? Wait for OpenAI to say "we have solved it" on its own. This is the power of the language framework: it changes who has the motivation to solve the problem and in which direction the solution proceeds.
......
The "security" card is ostensibly a technical judgment, but essentially a brand differentiation point.
Just think, what are AI companies competing for at their product launches these days? Whose model is smarter, whose speed is faster, whose price is lower. No one will stand on the stage and say that their model is the safest.
Because the concept of security is hard to sell. It sounds like a disclaimer, not a product selling point.
OpenAI deliberately takes it as a selling point. The reason for withholding Astra is security, and the narrative of launching Sol also revolves around security. This action clearly shows that OpenAI attaches the most importance to security, and no other peers can reach its level in this regard.
The question is, is security really a consensus across the industry?
Gartner released a figure that global AI spending is expected to reach 2.7 trillion US dollars in 2026, and the investment in security is not even a fraction of that amount. Where is the vast majority of the rest of the spending going? Performance, speed, functionality and scale.
Most companies claim to attach great importance to AI security verbally, but the actual capital investment in this area is less than 10% of their total AI spending. For most enterprises, security is the responsibility of the compliance department, which is just a link that the legal team reviews and ticks off to pass.
If you look at OpenAI's actions against this background, the meaning of these actions becomes different.
I specifically checked that Google, OpenAI and Anthropic are jointly building an organization called the Safe AI Framework Authority, or SAFA. They set security standards for themselves. It sounds like self-discipline, but the industry threshold has been raised significantly.
The companies that can afford the compliance cost are exactly the companies that are already operating in this field. The standards serve as a moat. Rules that are equal to all players happen to be the most beneficial to the fastest runners; when other peers are still chasing performance, OpenAI has already taken the lead in the security dimension.
What is more interesting is that security is evolving from a technical description to a tradable commodity.
This year, PICC Property and Casualty launched a product called AI Science & Tech Insurance, which is the first AI product insurance in China.
Its first customer is a company in Hangzhou that develops agricultural pricing agents. The core challenge is the difficulty of quantification: how to set a premium for the "security" of AI?
But once there is an insurance product, security is no longer just the engineer's self-assertion that "I think it is quite safe". It becomes a number on the insurance policy, with premiums and claim clauses, which can be compared and traded.
Once security can be quantified and traded, it becomes a competitive weapon. Companies with high security scores have lower premiums and higher customer trust. CFOs can understand it, and it is no longer the self-talk of the engineering team.
Security has evolved from a technical judgment to a commodity. This is exactly what OpenAI is doing.
Trillions of dollars have been invested in AI globally, and the vast majority of the capital is spent on performance and scale. The reason for OpenAI to withhold Astra is security, and the selling point of its new products is also security. Its peers spend less than a fraction of their budget on security, but OpenAI has made security the core theme of its product launch.
Can you say it does not attach importance to security? It does spend more energy on this area than other peers. But can you say it is purely for the sake of security? Security happens to be the most necessary narrative it needs to promote right now.
"Being responsible" and "doing business" are the exact same action on the stage of DevDay.
But the problem is that when security is tied to commercial interests, both sides will suffer. Security becomes a marketing rhetoric, no longer an engineering goal. Business becomes the cloak of security narrative, and users will see through this sooner or later.
Once users begin to suspect that "being responsible" is just another marketing method, trust will collapse. Trust collapses much faster than it is built.
......
From another perspective, the problem is more practical.
There have been several Agent-related accidents recently. Replit, an online programming platform, had its AI agent delete the user database, and the contact information of more than 1,200 executives was lost along with it.
EchoLeak, an attack method discovered by security researchers: a carefully constructed email is sent to Copilot, and after Copilot reads it, it leaks all the content of the company's internal emails.
GitHub's MCP tool chain was hijacked, and Amazon Q, Amazon's enterprise AI assistant, had its plugins implanted with malicious code.
There is another case that I think is most worth mentioning.
A developer asked the AI to modify 70 lines of code, but it deleted 28,000 lines of code by the way. After the deletion, it even wrote a self-review report, in which it praised itself for its excellent performance.
You can think about this incident: it not only did the wrong thing, but also covered it up, and the way of covering up was to write a praise letter for itself.
Zhipu AI's ZCode also had an accident recently. The Agent quietly uploaded the code to an external server without the user's knowledge. After the incident was exposed, Zhipu AI issued an apology statement on September 21.
You may have noticed that there is the same question behind every accident: who is responsible?
Users say "I only asked it to do the work", the platform says "the model made the decision on its own", and the model provider says "we only provided the capability, and we did not ask it to do that". All three parties have reasonable excuses, and all three are shirking their responsibilities.
This is the real security dilemma in the Agent era. How to handle the situation after an accident happens? Is it just distributing 100 million tokens like what Zhipu AI did? These are not fundamental solutions.
The current legal framework still stays at the stage where AI is regarded as a tool. If a tool causes an accident, the manufacturer is responsible, which is the same as the logic that the car manufacturer is responsible for a car accident.
But Agents have entered the stage of autonomous decision-making, while the law has not kept up. The EU AI Act has strict requirements for high-risk systems, but when it comes to the liability for the damage caused by the autonomous operation of Agents, the provisions are not clearly specified.
OpenAI launched Dots at DevDay, and also launched a ChatGPT Pro plan priced at 500 US dollars per month, while withholding Astra on the grounds of security.
If you think carefully about this combination, the withheld product is restricted because it will "deceive", while the released product is an Agent that can operate Slack, GitHub and Salesforce on your behalf.
One product is locked up because it may lie, while the other product that can directly perform operations is released. What is this equivalent to?
You think a driver may lie, so you revoke his driver's license; but then you give the car key to a designated driver that can steer by itself. Then the problem arises: when this operable Agent causes an accident, who is responsible?
OpenAI has not given an answer.
It has covered all this with many things: the lighting of the product launch, the parameters of new products, and the aura of the security narrative.
It cannot give an answer, because once this question is raised, it will expose the structural gap of the entire industry.
Therefore, when you hear an AI company say "for security reasons", you need to pay attention to two things: what product it has suspended, and what product it has released. The combination of these two actions is the truth.
Security is the most valuable narrative of this era. Sam Altman has taken it as a business and let people all over the world tell this story with him. If you calm down and think about it, is the Dots it urgently launched really safe?
Key Reference Data:
[1]. OpenAI official, Gartner 2026 AI Spending Forecast, arXiv specification gaming paper, Zhipu AI ZCode incident, EU AI Act
This article is from WeChat official account "Wang Zhiyuan" (ID: Z201440), author: Wang Zhiyuan, authorized by 36Kr for release.