HomeArticle

GPT-6 scores 99.9%, the ARC AGI exam is forced to be remade, and the next level will test "invention"

新智元2026-09-16 11:19
GPT-6 has taken the industry by storm with its stunning performance on ARC-AGI-3, the next generation will focus on invention, and ARC-4 is scheduled to be launched next year.

AI scores nearly full marks on exam papers, what else can the next one test?

When GPT-6 Astra was released, it achieved a 99.9% score on the semi-private ARC-AGI-3 test set.

In OpenAI's own words: it has been completely "saturated".

Unexpectedly, its emergence directly invalidated the entire original plan for ARC-AGI.

The pre-planned ARC series now requires new questions to be developed.

Just today, François Chollet, founder of ARC Prize, officially unveiled the blueprint for the next-generation evaluation system:

ARC-AGI-4 is confirmed to be launched as scheduled in Q1 next year, while the truly disruptive ARC-AGI-5, even hailed as "the last exam paper for humanity", has been fully put on the agenda.

ARC 4 will focus on continuous learning and curriculum learning over longer time scales; ARC 5 is fully centered around the concept of "invention".

He even stated that "when there is no objectively measurable gap in learning efficiency between AI and humans, that will be the true AGI moment".

GPT-6 Crushes ARC-AGI-3, All Thanks to Engineering

GPT-6 Astra's performance on ARC-AGI-3 is the most dramatic event in the AI circle in 2026.

Under ARC Prize's "standard test framework", Astra scored 62.7%.

This score is of course the highest in history, but it is still far from full marks.

However, when OpenAI connected its own Provider Adapter, the score of the same model directly soared to 99.9%.

This is a context management adapter that can retain the model's hidden reasoning state between multiple calls, allowing AI to access capability plugins in a way closer to the production environment.

The result is: with the same set of weights and the same set of questions, the score difference is 36 percentage points.

The gold content of the ARC-AGI-3 leaderboard has therefore been re-examined, and the focus of the debate has shifted from "which model has stronger generalization ability" to "which harness is fairer".

This is not an isolated case. The same set of weights for Claude Opus was verified at 30% on the ARC-AGI-3 private set, and reached 95.5% after switching to a better harness.

NVIDIA's AVO even scored 100% on the public demonstration set.

The "saturation" of ARC-AGI-3 is more like an evaluation engineering event.

What ARC-AGI 3 Tests

The first two generations, ARC-AGI-1 and ARC-AGI-2, are essentially centered on "finding patterns from images".

The system provides several groups of small grid diagrams of "input → output", and the AI guesses the rules, then draws the output of a new group.

These are static questions with a unique correct answer, relying on the ability to "induce rules by looking at a few examples".

When it comes to ARC-AGI-3, the testing method has been completely changed.

It directly throws the AI into a game world, and lets it play, test, and find rules on its own across more than 1,000 levels.

Each game has its own internal logic, hidden rules and level clearance conditions.

There are no instruction manuals, no natural language prompts, no one will tell you that "collecting three red squares will let you pass the level".

The AI can only see the current screen, choose an action, observe the result, and then decide the next step.

It is like the blind man touching an elephant, testing step by step, and piecing together in its mind a model of "how this world probably works".

The ARC Prize team clearly listed four things this benchmark is designed to test: exploration; modeling; goal acquisition; planning and execution.

GPT-6 Astra completed tasks on 96% of the levels with fewer steps than the human median, using an average of 51.7% fewer actions per level.

ARC Prize calls this a "step change in capabilities", but at the same time emphasizes one point: benchmark saturation does not equal proof of AGI.

Because these environments are "bounded" and deterministic, not open-ended.

The Next Exam Paper Starts Forcing AI to "Invent"

According to the original plan, ARC 4 will carry forward the spirit of ARC 3, but will focus more on continuous learning and curriculum learning over longer time scales.

Each game has far more levels, and the levels are "compound interest" — every level requires reusing what was learned in previous levels.

Now, Chollet has incorporated "open-ended invention" into the shared R&D foundation of ARC 4 and ARC 5.

So, how exactly are the "invention-focused exam questions" designed?

ARC 1 to 3 all share a common premise: there exists a "correct" answer or a "winning" state, and humans can complete it on their first try, so that humans can be used as a measurement benchmark.

The entire set of human baselines for ARC 3 was measured through 486 people and 2893 attempts.

The definition of invention inherently includes "something that did not exist before", there are no standard correct answers, so there is no ready-made human baseline to divide against.

You can't say "how many steps humans took to invent this thing". So next, how exactly should this be measured?

A possible path is to return to Chollet's "definition" of intelligence: intelligence is not how many skills you master, but the efficiency of acquiring new skills.

If we follow this line of thinking, invention can be measured as — whether the new tool, new rule, or new notation created by the AI can make subsequent problems solved faster.

Astra created shorthand symbols for the game, which is actually the prototype of this action. It invented an intermediate representation, and used it to cut the number of steps in half.

For the next exam paper, ARC 5 will most likely turn this matter into "the exam question itself".

ARC 5 Goes Straight to the Ultimate Goal, How Many More Exam Papers Can Humans Produce?

It is worth mentioning that the exam questions of the ARC series will probably come to an end around the 6th or 7th generation.

Chollet once stated that as long as we can still propose tasks that "humans can do but AI cannot", we will continue to create new questions;

When we can no longer propose such tasks, when the gap between human learning efficiency and cutting-edge AI can no longer be measured, that moment will be AGI.

AI is advancing so fast that the time left for an evaluation to prove it is "difficult enough" is being compressed.

As soon as a difficult challenge is overcome, the question setter has to look for the next one.

After the scores approach the ceiling, it becomes increasingly difficult for the original exam papers to distinguish: how much stronger has the AI actually become, and where do humans still lead?

Following this trend, ARC 5's focus on "invention" is likely to be an early attempt to find an exam paper that is harder to be fully scored:

Let AI step out of pre-set solutions and create verifiable, reusable new things.

No one knows how long this question can hold up now.

But if even open-ended invention can no longer widen the gap between humans and AI, the phrase "no more exams to take" will turn from a sigh into an unavoidable issue in AGI discussions.

AI is still accelerating, how many more exam papers can humans produce?

References:

https://x.com/FakePsyho/status/2099086728963207537?s=20

This article is from the WeChat official account "AI Era", author: Taozi, published with authorization from 36Kr.