HomeArticle

Jensen Huang announced AGI on behalf of OpenAI, and the question setter refuted him publicly on the spot: after switching to a different test scenario, the performance plummeted from 99.9% to 62.7%

新智元2026-09-11 17:49
OpenAI has released GPT-6 Astra, and there is constant controversy among all parties over whether AGI has been achieved.

Last Thursday, OpenAI released GPT-6 Astra. 

The official website defines it as "the most intelligent and most aligned model in the world". 

Shortly after, Jensen Huang posted a post on X that pushed the incident to another level: 

AGI has arrived. Congratulations to the OpenAI team.

In contrast, you cannot find any phrase like "AGI has been achieved" on the official release page of GPT-6 Astra. 

A GPU seller has stated the claim to the fullest extent for the model developer. 

What is even more interesting is the party that designed the test. 

OpenAI's release page states that Astra scored 99.9% on ARC-AGI-3, "saturating" this AGI-named benchmark. 

After checking the test score, ARC Prize, the evaluation organization responsible for designing the test, directly refused to endorse it: "It does not count." 

The ARC Prize Foundation released an article on the same day to break down the situation: 

This score comes from the test environment customized by OpenAI. When switched to the standard environment that treats all manufacturers equally, the score is only 62.7%. 

Their conclusion is only one sentence: Astra has made substantial progress in generalization, "but we do not claim that it is AGI". 

The chip supplier directly gave its full endorsement, the model manufacturer left leeway, and the benchmark organization responsible for designing the test directly refused to sign off on the result. 

This is probably the most surreal scene in the AI circle in the past few days: AGI has not been truly "achieved" this time, but has been "announced" by spokespersons of all parties on social platforms. 

What did OpenAI itself say?

At the press conference call on the day of the release, President Greg Brockman said a highly provocative statement: 

Welcome to the AGI era.

He even predicted that if you look back a few years later to see when AGI was born, "it is probably this period, and possibly this model." 

Sounds like an official announcement? 

But he immediately added two qualifiers: first, this is his "personal judgment"; second, AGI is a gray, vague concept, "it is up to users to decide whether it counts or not." 

At its most important product launch event, the president even used "personal opinion" to set the tone for a new era, and even passed the final ruling power to netizens. 

This can only prove one thing: OpenAI itself cannot come up with a clear, publicly available AGI recognition document. 

What is even more interesting is the attitude of CEO Sam Altman. 

Before the release of Astra, he publicly complained that "AGI" is a poorly defined term that is close to marketing jargon. 

The CEO thinks this term is too vague, the president treats it as a "personal opinion", and the official website simply does not mention it at all. 

Jensen Huang's "endorsement" and congratulations on AGI just fell right into this subtle blank space. 

99.9% and 62.7%, two sets of scores for the same test paper

Is Astra really that powerful? 

It is written on OpenAI's release page: 

Astra got an incredible high score of 99.9% on the AGI-named benchmark test ARC-AGI-3, "saturating" this test.

But the slap in the face came very quickly. 

On the same day, the ARC Prize Foundation, which is responsible for designing the test, released an article publicly breaking down this "99.9%" score. 

It turns out that Astra took the test twice. 

On the ARC-AGI-3 leaderboard, Astra's scores in the two environments are marked separately; Claude Opus 5 on the same list gets 30.2%, and GPT-5.6 Sol gets 7.8%. 

In the deeply customized "built-in tools" environment of OpenAI (allowing the model to retain reasoning status and use exclusive technologies to manage long conversations), it indeed scored 99.9%, with a cost of nearly 19,000 USD. 

However, if it is switched to the "bare test" standard environment that treats all manufacturers equally, its highest score is only 62.7%, and the cost even surges to 26,000 USD. 

The attitude of ARC Prize is very firm: the future AGI must be able to solve problems under standard conditions. The 99.9% score obtained with tools and the 62.7% score obtained in the bare test are completely different things. 

Of course, the test designer also admitted that Astra brought a "step-change" surprise. In 96% of the levels, Astra uses fewer actions than humans, with an average of 51.7% fewer actions. 

This is the first time that AI has comprehensively outperformed humans in action efficiency. 

But ARC Prize has already stated it clearly in advance: on the day ARC-AGI-3 was launched, they wrote clearly that getting full marks on this benchmark does not mean "proving that AGI has been achieved". 

The reason is very simple: no matter how difficult the test environment is, it is closed, with determined mechanisms and closed targets, which cannot represent the complexity and openness of the real world. 

Just like a top student who gets full marks in the test paper, he may not still perform outstandingly in the real workplace. 

100,000 GPUs for model training, 400,000 GPUs for selling future expectations

Since the test designer does not recognize the result, why was Jensen Huang the first one to rush to announce that "AGI has arrived"? 

If you carefully break down Jensen Huang's post, you will find that its information structure is far more substantial than the superficial excitement. 

The full text of the post is: 

GPT-6 Astra was trained on about 100,000 NVIDIA Grace Blackwell NVLink72. It only took 4 years from ChatGPT to o1 and then to Astra. AGI has arrived. Congratulations to the OpenAI team. Another 400,000 GPUs will be launched soon.

The first number: 100,000. 

Astra is OpenAI's largest training run to date, consuming more than 100,000 GPUs. 

The second number: 4 years. 

From ChatGPT to o1 and then to Astra, Jensen Huang used a timeline to connect conversational models, reasoning models and the so-called AGI into a straight line. 

The third number: 400,000. 

"Another 400,000 GPUs will be launched soon." 

Who do these 400,000 GPUs belong to? What model are they? What are they used for? Jensen Huang left a suspense here. 

But when these three numbers are put into the same post together with "AGI has arrived", a commercial narrative will appear in the reader's mind: 

100,000 GPUs have created Astra (which is AGI), so what will the next 400,000 GPUs create?

In this context, AGI is no longer an abstract technical end point, but has become the most powerful proof of computing power demand. 

Once AGI is universally recognized as "already arrived", the sky-high cost of computing power will no longer be an enterprise's expense item, but a necessity that no one dares to quit the game for. 

AGI announced on X cannot be written into commercial contracts

In October last year, OpenAI and Microsoft's new agreement included a clause: after OpenAI announces AGI, it must be verified by an independent expert panel. 

In February this year, the two parties jointly reaffirmed that the definition of AGI and the recognition process in the contract "remain unchanged". 

In April, the agreement was revised again, and Microsoft's revenue sharing was changed to "not related to OpenAI's technical progress". The outside world generally interprets that the triggering effect of the AGI clause on commercial interests has been greatly weakened. 

In other words, whatever Brockman said at the press conference, whatever Jensen Huang wrote on X, and whatever score ARC-AGI-3 runs, do not constitute AGI recognition in the contractual sense. 

Now it is the key node for OpenAI to prepare for IPO, and its arch-rival Anthropic is also rushing to go public. 

The term "AGI" has long been written into the investment agreements between OpenAI and Microsoft, Amazon, and what it affects has gone far beyond technical beliefs, including equity pricing and commercial control rights. 

This explains why Brockman only dares to call it a "personal judgment", and why OpenAI's official website strictly guards against mentioning AGI. 

You can welcome the AGI era as much as you want in public opinion, but when it comes to the contract, not a single extra word can be written. 

Who is qualified to announce AGI?

A long article on Bloomberg sorted out this question, and the answer is no one has the final say. 

First look at the definition. 

OpenAI's version is "a system that surpasses humans in most economically valuable tasks". 

Google DeepMind's 2023 paper does not focus on economic value, but on versatility, and whether it can learn new skills with very little data. 

ARC Prize is more straightforward, stating that economic value "is the wrong scale for measuring intelligence", and their definition is "a system that can efficiently learn new skills outside the training data". 

Even the terms used are different. 

Dario Amodei of Anthropic does not like to say AGI, and calls it "powerful AI" instead. Mustafa Suleyman of Microsoft coined the term "artificial capable intelligence". 

Meta and OpenAI recently simply skipped AGI and directly talked about "superintelligence". 

The timelines are even more chaotic. 

In September 2024, Sam Altman said "systems pointing to AGI are emerging", and even said that superintelligence may appear within "a few thousand days". 

A year later, when GPT-5 was released, he admitted that "we are still missing something quite important", but did not say what it was. 

Another year later, Brockman said at the Astra launch event "Welcome to the AGI era". 

On the DeepMind side, Shane Legg has bet since 1999 that the probability of achieving AGI before 2028 is 50%; Demis Hassabis says it will take another five to ten years, and the timeline seems to be shortening only because "the definition of AGI is being diluted". 

So the current situation has emerged. 

The same term is defined in four different ways by four groups of people. 

Model companies hold the product narrative, only showing test scores without drawing conclusions;

Chip giants hold the computing power narrative, rushing to announce the arrival of the new era;

The benchmark organization holds the test ruler, firmly sticking to the bottom line and refusing to recognize the result;

Big capital providers like Microsoft hold the contract terms, sitting back and calculating accounts according to the rules.

Over the past four years, the entire industry has been competing to be the first to create AGI. 

Recently, the tense of the narrative has changed, directly jumping from "one to three years left" "at least half a year at the fastest" to "it has already arrived". 

However, no one has come up with a universally recognized AGI acceptance standard so far. 

References:

https://x.com/JensenHuang/status/2096700264569090384

https://www.bloomberg.com/news/features/2026-09-04