GPT-6 is already so powerful, why are we still reluctant to call it AGI?
On the afternoon of September 3, near the end of the closed-door media briefing, OpenAI President Greg Brockman made a remark:
Welcome to the AGI era.
A reporter followed up and asked: Are you officially announcing that you have achieved AGI?
He responded that the term AGI is no longer tied to the agreement between the company and Microsoft. It now functions more as a concept at the mission and spiritual level.
This seemingly cliché remark is actually the key to the entire matter.
-- Foreword
01
Who coined this term
The birth of the term AGI dates back much earlier than most people assume.
In 1956, John McCarthy created the term "artificial intelligence" in the proposal for the Dartmouth Conference. Back then, the group aimed to endow machines with the full range of human intelligence.
Yet over the next 50 years, the term became overused. Chess-playing programs were called artificial intelligence, spam filters were called artificial intelligence, and voice shopping guides in malls were also called artificial intelligence. By around 2000, "artificial intelligence" had been reduced to a catch-all phrase that could mean anything.
In 2001, a group of researchers grew frustrated with this state of affairs. They decided to return to their original ambition and co-wrote a book. The critical question was what to name the book. They tried titles like *True Artificial Intelligence*, which was rejected, and *Synthetic Intelligence*, which no one found satisfying.
Later, Shane Legg, who had just obtained his master's degree, proposed in the mailing list: Since we are talking about machines with general capabilities, let's call it Artificial General Intelligence, abbreviated as AGI, which is easy to pronounce.
Other participants in that round of discussion included Wang Pei, Peter Voss, and Eliezer Yudkowsky, who later became known for his concerns about AI risks. The name was finalized in 2002, published as a book title by Springer a few years later, and developed into a series of seminars of the same name starting in 2006, before spreading through the research community.
The interesting part comes later. Around 2005, in an online discussion in the AGI community, a stranger suddenly appeared claiming he had used the term back in 1997. After checking, the group found it was true: physicist Mark Gubrud used the term in an article discussing automated military equipment and technological risks, which had gone unnoticed for eight years. Legg later recalled this plainly: Suddenly someone popped up saying "I invented this term", and we asked "Who are you?"
So the origin of AGI goes like this: A physicist jotted it down casually, most people had no idea about it, then a group of researchers re-coined the term to justify their own ambitions.
As for Shane Legg, who named the term, he co-founded DeepMind with Demis Hassabis five years later. You will see him again later.
And this term, which has never had a definitive meaning, was written into the articles of association of a company in 2015.
02
The term worth 100 billion US dollars
When OpenAI was founded in 2015, it was a non-profit organization. There is a sentence in its articles of association that can be regarded as both its mission and definition: to ensure that Artificial General Intelligence, a highly autonomous system that outperforms humans at most economically valuable work, benefits all of humanity.
Pay attention to the phrase "benefit all of humanity". At least literally, OpenAI assumed from the very beginning that once AGI is built, it will be too important to be privately owned by any single company.
But building AGI burns huge amounts of capital, and a non-profit structure cannot raise funds of that magnitude. In 2019, OpenAI switched to a "capped-profit" structure: investors can earn returns, but their returns have a cap, and the excess part goes to the non-profit parent entity, with the board of directors remaining in control by the non-profit side.
In July of the same year, Microsoft invested its first 1 billion US dollars, and added another 10 billion US dollars in January 2023, with a total committed investment of 13 billion US dollars. Microsoft obtained extremely considerable returns in exchange: Azure became OpenAI's exclusive cloud service provider; Microsoft obtained exclusive commercial licenses for these models, and integrated them into Bing, Office, and GitHub Copilot; plus a revenue share: between 2023 and 2025, Microsoft generated about 30 billion US dollars in revenue from this partnership, 23 billion of which came from computing power sales.
But the key point is not how much money Microsoft earned, but a clause that OpenAI insisted on adding to the negotiation table: If the OpenAI board of directors announces that the company has achieved AGI, Microsoft's rights to the technology thereafter will terminate immediately.
Why was such a clause added? Because it is the safety rope for the promise of "benefiting all of humanity". To put it plainly: You can make as much profit as you want before reaching the AGI threshold, but once the threshold is crossed, the technology will be far too important to be the private property of a single company.
The problem is that the contract never clearly defines what AGI is from beginning to end.
Probably to make up for this flaw, the 2023 round of investment added a second trigger, which simply defines AGI by monetary value: when the system can generate a certain magnitude of profit, for example around 100 billion US dollars, the exclusive rights will terminate.
But the conditions later changed. In October 2025, OpenAI completed its restructuring, with Microsoft obtaining approximately 27% of its shares, and an additional procedure added: even if OpenAI announces AGI, it must be verified by an independent panel of experts, and cannot be declared unilaterally. In February 2026, the two sides even issued a joint statement specifically confirming that the definition and certification process of AGI in the contract remains unchanged.
Then a huge turning point came. The latest revision in April 2026 removed all previous trigger conditions and replaced them with two dates: Microsoft's license will last until 2032, and its revenue share will last until 2030, neither of which will depend on whether OpenAI announces the achievement of AGI.
More than four months later, Brockman simply stated that AGI is now a spiritual concept.
I do not intend to infer any motives from this sequence of events. There are countless commercial reasons for the retention or removal of contract clauses, and Brockman most likely truly believes what he said.
On this timeline, the history of the term AGI went from being a name at the start, to a promise later, then to a contractual clause, and now to a slogan.
At least OpenAI still regards this term as its ultimate goal and measurement standard, but in fact there is no consensus even within OpenAI.
Because just one week before Brockman made that remark, in a TIME report, Sam Altman was still saying that OpenAI is "not fully at AGI yet", but he added that by the end of 2026, there will be an internal system that he is willing to call AGI.
Around the same time, Mark Chen, the company's Chief Research Officer, gave a number: We have completed 80% of the journey towards AGI.
Within ten days, three people at the same company gave three different answers.
Why does this question get so many different answers?
03
Four groups of people
A number of groups in the world have seriously answered the question "What exactly is AGI". These answers all seem to be talking about AGI on the surface, but they are not actually trying to solve the same problem at all.
The first group is OpenAI itself. The clause in its articles of association is its definition: outperforming humans at most economically valuable work. This is not the conclusion of an academic paper, it is the constitution of a company, and the source of that term in the contract. It defines intelligence through economics: whether it can replace human labor.
The leader of the second group is François Chollet. You may not have heard of him, but millions of developers around the world have used what he created: the deep learning framework Keras is his work. In 2019, he wrote an essay *On the Measure of Intelligence*, whose views were very unpopular at the time: No matter how large you make the model or how much data you feed it, that does not equal intelligence. Real intelligence is the ability to efficiently learn a task that you have never been taught before.
He did not just talk about it. He designed a set of tests that are interesting but effective: you are given a few examples, asked to deduce the rules, and then solve a completely new problem you have never seen before. These new problems can be solved by human children, but models struggle to get them right. This set of tests offered a bounty for five years without being cracked, and by the end of 2024 the best score only just exceeded 50%. It later became the most famous tough test in the industry, because it is specifically designed to distinguish between "truly solving the problem" and "memorizing the answer".
The third group is DeepMind, the company that beat Lee Sedol at Go back then, which is now owned by Google. Shane Legg — yes, the person who coined the name AGI in the mailing group in 2002 — is among the authors of their framework paper published in March this year.
Objectively speaking, unlike all others who define AGI with abstract concepts, this group has the most pragmatic attitude and the most quantifiable standards. Their point is: don't rush to ask whether we have reached AGI, first build all the test items we need. That paper breaks down cognition into ten categories — perception, generation, attention, learning, memory, reasoning, metacognition, executive function, problem solving, and social cognition — and compares each of them to the average human level. They also offered a 200,000 US dollar bounty for crowdsourcing, specifically calling for five incremental tests for which they believe tools and standards are most lacking.
They made a judgment on today's models: this is an uneven cognitive profile.
In other words, current cutting-edge models outperform most humans in mathematics, factual memory, and pattern recognition, but lag behind ordinary people in learning from new experiences, maintaining context in long conversations, and understanding social situations. If you are a rationalist, this set of frameworks they provide is probably the most illustrative of how far we have come towards AGI.
The fourth group is Anthropic, the parent company of Claude, whose founding team is a group of people who left OpenAI. CEO Dario Amodei simply does not use the term AGI, thinking it is too much like science fiction, and calls it "powerful AI" instead, describing it as "a nation of geniuses in a data center". But don't think he is being mystifying: he is the only CEO of all AI labs who has given an official timeline: this kind of system will appear between the end of 2026 and the beginning of 2027.
There is a more notable common point: all four definitions come entirely from the people who are building this technology. The people setting the standards are the same people submitting the papers.
This is not a conspiracy. The people building the technology are naturally the most qualified to describe it. But it means one thing: AGI has never been a discoverable fact, but a standard that needs to be agreed upon, and humanity has not yet reached that agreement.
04
What GPT-6 Astra has achieved
Let's lay out all the cards from the proponents first. The vast majority of people who are willing to mention the term AGI this time are not excited out of thin air.
Ethan Mollick, a professor at the Wharton School who has long studied AI's performance in real work, obtained early access to GPT-6. His evaluation is: this model can independently complete complex and meaningful work for him, and can work on a single task for several consecutive days.
The example he gave (in his own words, just a fun example) is a walk-through Library of Alexandria on a web page. The scene around 250 BC is based on historical records, the scrolls can be read, there are guided tours, and you can even switch between different historical hypotheses about how the library declined and was destroyed.
This achievement is far more important than a beautiful 3D scene, because it is not one single task, but five tasks combined: writing software, checking historical materials, designing interactions, organizing narratives, and arranging a large amount of content. These five tasks belong to five different job roles. Previous models could do any one of them well, but could not combine them into a finished product.
Another key point is that it has produced new outputs in mathematical research, and this time the outputs are verifiable.
On August 1, OpenAI announced that its internal version had solved ten open mathematical problems, one of which had been unsolved since 1999. They published a 249-page manuscript, as well as formal proofs that can be verified step by step by computers.
The value of this achievement does not lie in its difficulty, but in the three words "verifiable".
OpenAI once lost face on this matter. In October 2025, a vice president of OpenAI posted that GPT-5 had solved ten unsolved mathematical problems, but was immediately pointed out by the mathematician maintaining that problem list that it was a serious misreading — the model did not solve the problems on its own, but found existing literature that had the answers. OpenAI had to delete the post. Ten months later, when announcing mathematical research results again, they attached the proof process that the machine can verify step by step.
Next, a point that is easily overlooked: it can independently judge when to stop and ask humans for input.
OpenAI released a set of comparisons. For the same task, building a job search website for a person switching careers, the previous generation model worked silently for 13 minutes and 15 seconds before delivering the finished product. Astra stopped after 20 seconds and asked a question: Which industry do you want to switch to?
On the surface, this seems like a step backward.
This is mainly because for the past two years, the general standard for measuring agents has been "how long they can run continuously without human intervention", because they always break down halfway. But this metric has now reversed. Precisely because it can run for a long time now, the cost of going off track has become much higher. Working seriously for 13 minutes on a wrong premise is far worse than stopping to ask a question.
Being able to ask questions does not mean weaker capabilities, but that it knows which uncertainty will ruin the entire task. These two capabilities combined form a kind of modeling of humans, which is why many people believe that GPT-6 does have human-like intelligence.
Looking back at these three achievements, you will find a common point: none of them were done for the first time by Astra. Writing code, checking materials, designing, proving mathematics — the previous generation could do all of them. What changed is that for the first time, it can combine all these outputs into a deliverable finished product.
From "being able to do individual tasks" to "completing the whole task", this is the line that this generation of models has truly crossed.
But one point must be stated here, otherwise everything above is just promotion. Almost every positive example mentioned above comes from people who have obtained early test access or have a partnership with OpenAI. The footnotes of OpenAI's own demo page also clearly state that the displayed clips have been edited: