HomeArticle

What counts as the true arrival of AGI?

深流研究所2026-09-09 10:28
Technological progress is continuous, yet historical narratives always tend to single out a specific moment.

On September 3, Greg Brockman, President of OpenAI, released a statement after launching the new model GPT-6 Astra: "Welcome to the age of AGI."

Three days later, Jensen Huang put it more directly on social platforms: "It only took four years from ChatGPT to o1, and then to GPT-6 Astra. AGI has arrived."

Has the long-awaited AGI suddenly been announced to be realized?

Obviously, things are not that simple.

Shortly before the release of GPT-6 Astra, Sam Altman, CEO of OpenAI, was still saying that AGI is a term with "extremely vague definition", and almost called it an "irrelevant marketing term".

The ARC Prize team, which is responsible for the key evaluation of GPT-6 Astra, also specifically reminded: GPT-6 Astra does represent an important advancement in general capabilities, but "we do not claim that it is AGI".

On the one hand, it is "welcome to the age of AGI", and on the other hand, it is "we do not say this is AGI".

Today, everyone is talking about AGI, but no one can make it clear: What exactly does AI need to achieve to truly cross the line of AGI?

Can the high score on the ranking be equated with AGI?

First of all, the most staggering achievement of GPT-6 Astra is that it scored 99.9% on ARC-AGI-3.

ARC-AGI is different from ordinary evaluations. It does not directly tell the model the rules, but puts the model in an unfamiliar environment, letting the model judge the goal by itself, understand the feedback, and then find the solution.

What this evaluation tries to measure is exactly the most questioned capability of current AI: whether it can learn on the spot when facing problems it has never seen before.

A score of 99.9% is certainly amazing. But this number has a premise that is often omitted.

In the standard test framework uniformly provided by ARC Prize, the score of GPT-6 Astra is 62.7%; after accessing the dedicated adaptation framework provided by OpenAI, the score rises to 99.9%.

The latter allows the model to retain hidden reasoning states across requests, and uses context compression to continue previous exploration experience.

We cannot simply classify this operation of OpenAI as "cheating". After all, AI in the real world never runs on bare metal. Memory, tools and operation frameworks may all become part of the intelligent system.

The problem is that 62.7% measures something closer to the capability of a single model; 99.9% measures a complete system that has been specially optimized.

Both scores are real, but they cannot be confused. It is even more impossible to convert the 99.9% test score into "99.9% AGI" casually.

In addition, the performance of GPT-6 Astra in other evaluations does not present an impeccable generalist image.

It reached 97.6% on the high-difficulty math test FrontierMath Tier 4, and 72.6% on the computer operation test OSWorld 2.0; but when it came to AutomationBench, which is closer to the real office workflow, the score dropped to 41.4%, and it only got 59.3% on the complex professional task test Agents' Last Exam.

What these numbers really illustrate is that when the task rules are clear and the answers can be verified, Astra has been able to approach or even surpass humans; once it enters the real environment with lengthy processes and vague goals, it will still fail frequently.

It is not new for AI to perform well in closed exams.

When Deep Blue beat the world chess champion, Watson won *Jeopardy!*, and AlphaGo made its "divine move", people once briefly had a similar illusion: since machines have conquered such complex tasks, general intelligence is probably not far away.

The world with clear rules, immediate feedback and automatic referees is not the same as real life.

Is the AGI that everyone is talking about the same thing?

AGI still lacks a unified consensus so far.

OpenAI adopts the most economically oriented definition. Its articles of association released in 2018 describe AGI as "a highly autonomous system that surpasses humans in most economically valuable work".

This definition includes three key conditions: covering a wide range of work, performing better than humans, and being able to operate with a high degree of autonomy.

The output in real work constitutes the core scale. Consciousness, emotion and human-style thinking are not listed as necessary conditions.

This standard is highly consistent with OpenAI's product roadmap in recent years.

Reasoning models improve the processing capability of complex tasks, agents enhance tool invocation and computer operation, and products such as Codex further enter the software development process.

All capabilities eventually converge to the same goal: to let AI undertake more complete work, and convert model capabilities into productivity that can be deployed at scale.

Anthropic's approach to AGI is more cautious. Dario Amodei, CEO of Anthropic, believes that this term carries too many vague meanings, so he more often uses the term "powerful AI".

In *Machines of Loving Grace* published in October 2024, Amodei described a specific type of system:

It reaches or exceeds the level of top experts in most important fields, can independently complete tasks that last for hours, days or even weeks, and can be copied into millions of instances to work at the same time at a speed much faster than humans.

This picture includes three variables: capability, autonomy and replication scale. The three jointly determine the real impact of AI, and will also amplify the risks of abuse, loss of control and systematic risks synchronously.

Based on this judgment, Anthropic binds capability expansion to risk governance. The company's strategic focus can be summarized as: maintain the competitiveness of cutting-edge models, and at the same time let security guarantees upgrade synchronously with the capability level.

What Google DeepMind focuses on is the measurement standard of AGI.

In its article *Levels of AGI* released in 2023, taking "performance" and "generality" as two main axes, general intelligence is divided into five levels: "Emerging", "Competent", "Expert", "Virtuoso" and "Superhuman", and autonomy is treated as an independent dimension. This framework requires researchers to examine both capability intensity and coverage at the same time.

In March 2026, DeepMind further released a cognitive classification framework, which splits general intelligence into ten dimensions: perception, generation, attention, learning, memory, reasoning, metacognition, executive function, problem solving and social cognition.

The evaluation process uses undisclosed test sets, representative human samples, and the position of the model relative to the human capability distribution.

DeepMind also set up a $200,000 prize with Kaggle to solicit test schemes for five of the dimensions with weak evaluation.

This line of thinking transforms AGI from a single label into a comparable map of cognitive capabilities, and also reflects DeepMind's research orientation: to establish a unified scale, identify capability gaps, and then organize evaluation and R&D accordingly.

The definition of AGI is not just a conceptual issue.

What standards you adopt will determine what capabilities you regard as key progress; and what capabilities you regard as key progress will determine a company's product roadmap, resource allocation and security priorities.

The definition seems to be describing the future, but in fact it is also shaping the future.

Why has "whether AGI has arrived" become a narrative dispute?

People are always used to imagining the technological revolution as a clear moment.

The Wright brothers' plane left the ground, the first atomic bomb exploded, and Apollo 11 landed on the moon.

History needs dates, and it also prefers iconic images.

Returning to the current dispute over the arrival of AGI, "GPT-6 Astra is AGI" and "we have entered the age of AGI" are actually two expressions of different natures.

The former is a technical conclusion that can be questioned and tested: since it is AGI, it should explain what definition is adopted, what conditions are met, and in which tests it has been independently verified.

The latter is more like a judgment of the times. What Brockman said is exactly "Welcome to the age of AGI". This expression not only creates a historical moment, but also retains enough space for interpretation.

For model companies and the AI industry chain, "AGI" is first and foremost a powerful product narrative.

The real progress of model capabilities is usually scattered and uneven. To accurately describe these changes, a long list of evaluations, conditions and qualifiers are required.

"The age of AGI is coming" can compress everything into one sentence.

It turns a model release from a product upgrade into a historical node, and also organizes computing power investment, technical routes and business prospects into the same story.

For companies competing for users, enterprise customers, talents and capital, the value of this narrative may be no less than a major breakthrough in model capabilities.

Jensen Huang's statement cannot be understood without this narrative framework either.

In September 2025, NVIDIA and OpenAI announced that they had reached a strategic cooperation intention. According to the plan announced at that time, OpenAI will deploy at least 10 GW of NVIDIA systems, equivalent to millions of GPUs; NVIDIA is prepared to invest up to $100 billion in OpenAI as the infrastructure is gradually put into use.

This is a letter of intent, not a final agreement that has been fully implemented, and there were even rumors later that the plan might be adjusted.

But in the original cooperation vision, NVIDIA is positioned as the preferred computing and network partner when OpenAI expands its AI infrastructure. The two sides directly stated in the announcement that these facilities will serve OpenAI's path of "moving towards the deployment of superintelligence".

Recently, while announcing that "AGI has arrived", Jensen Huang specifically mentioned the NVIDIA computing system used to train Astra, and the next batch of GPUs that are about to go online.

The whole statement implies a complete causal chain: Large-scale computing power investment brings intelligence leap, so larger-scale computing power investment is still necessary.

As the main beneficiary of the global AI computing power expansion, NVIDIA certainly has reasons to emphasize that this is a turning point of the times.

In fact, this is not the first time Jensen Huang has announced the arrival of AGI.

In 2024, he once said that if AGI is defined as being able to pass almost all exams designed by humans, this goal may be achieved within five years.

By March 2026, he said on the Lex Fridman podcast: "I think we have achieved AGI."

In this light, AGI is not only an end point waiting for technological breakthroughs, but also a narrative resource that can affect capital, market, policy and industrial expectations.

As for how AGI can be considered to have really arrived? Is it a model release, an evaluation record, or an announcement from a company?

Maybe this is a process that everyone is in and imperceptibly experiencing.

As Andrej Karpathy, a well-known AI researcher and founding member of OpenAI, put forward a judgment, AGI will not bring a sharply rising curve, it will blend smoothly into the existing growth, and you can't even find where the inflection point is afterwards.

This article is from the WeChat official account "Deep Flow Institute", author: Wu Jiangfeng, published with authorization from 36Kr.