首页文章详情

Claude scores full marks, GPT scores 99 points, He Kaiming's team has completely aced the AGI exam.

新智元2026-10-04 20:00
Claude scores a perfect 100, GPT scores 99! The team led by Kaiming He has completely aced the AGI exam.

The new paper from He Kaiming's team has gone viral on X! 

The VISTA paper just uploaded to arXiv during the National Day holiday shows that in the past six months, all the world's top large models have been taken to the AGI examination room with their eyes covered. 

The same Claude Opus 5 only scored a meager 40 points in the official standard test. 

He Kaiming's team simply helped it remove the blindfold, let it see the game screen in person, and handed it an "album" that it can flip back to review at any time. 

As a result, Claude directly cleared all 25 public games of ARC-AGI-3 and got a perfect score of 100! 

And the number of steps it took is even more than 50% less than that of a human playing the game for the first time. 

For 18 never-seen-before games, no one gave it a single instruction, the AI figured out all the rules on its own and cleared all levels

Renowned AI blogger Mark Kretschmann wrote on X: "Giving an Agent the ability to 'take one more look' can make a world of difference." 

In the past few months, to clear ARC-AGI-3, top code experts from all walks of life have written thousands of lines of code to build a world simulator. 

He Kaiming's team concluded that the large models are already intelligent enough, what's holding them back are their "eyes" and "memory". 

Large models have been misled by a digital table for half a year 

ARC-AGI-3 is the AGI test paper released by François Chollet, the father of Keras, and the ARC Prize Foundation in March this year. It consists of a bunch of interactive pixel mini-games, with no instructions or rules at the start, and players have to explore and learn through trial and error entirely on their own. 

The scoring algorithm is extremely strict. The official team tested nearly 500 human beginners, and the AI can only get full marks if it not only clears the levels, but also takes no more steps than human players. 

However, what the official test platform gave the large model was only a 64×64 digital table. The little characters, mechanisms and routes that humans see are reduced to only 4096 numbers for the model. 

This is equivalent to asking a person to play Super Mario with their eyes closed, listening to others calling out coordinates.

For the same frame of the picture, the left side is the digital table given to the model by the official, and the right side is the image that VISTA allows the model to view directly

The first thing the VISTA team did was to replace the digital table with images. 

Without changing anything else, the score of GPT-5.6 Sol jumped from 13.33 points to 47.32 points. 

Viewing images also saves more computing power. One digital grid consumes about 4000 Tokens, while a 512×512 image only consumes 308 Tokens. 

After one full game run, the text version consumes 71.9 million Tokens, while the image version only consumes 30.7 million Tokens, with a higher score. 

Just by equipping the model with a pair of real "eyes", the large model regained 34 points. 

Only by replacing numbers with images, the score of GPT-5.6 Sol has more than tripled, and the number of games it can clear has increased from 1 to 9

Making AI never forget, earning an extra 24 points in one step 

Being able to see is not enough. A single game can have hundreds of steps, and a common problem of multi-modal models is that once they view an image, the image will be deleted once the context is full, or compressed into a few lines of text summaries. 

When the model wants to go back and check the details, that frame of the image has long been gone. 

The core trick of VISTA is a lossless visual memory "album". 

Every frame output by the game, including animation frames, is saved as it is by serial number. The model can call inspect to retrieve any frame at any time, compare several frames side by side, zoom in on the corners, and use read_pixels to read the exact color values. 

Most importantly, flipping through the album does not count as a step, only actual operations in the game count towards the step count. 

The team calls this mechanism "explicit attention", and the model decides for itself which frame to flip to and which part to look at. 

It also carries two notebooks with it, GUIDE.md records the game rules, and WORKING.md serves as a scratch pad. 

In one round of VISTA, the model views the screen, thinks about the rules, flips through the album when necessary, and finally takes one step

The paper records a very "human-like" operation moment. 

In level 3 of the BP35 game, an orange square appears on the screen. The model clicks on it, and the orange square turns into an X. 

It first pauses, retrieves the last frame of the previous round and the three animation frames just generated, and plays them back side by side. 

A few rounds later, it finds that clicking the X can make the orange square grow back, and immediately writes: "This is the tool I need, use it to build a ceiling to block the deadly rising obstacles." 

For this level, a human playing for the first time needs 44 steps, while it only takes 34 steps. 

After clicking the orange square, the model first retrieves the 4 animation frames before and after to compare frame by frame, and only proceeds after fully understanding the mechanism

In Pac-Man, the number of beans in the maze decreases by one after each consumption. The model even forcibly flips back to the first frame at the start of the game in the 13th and 36th rounds twice to memorize the maze map and find the way. 

Many people's first reaction is to give the model more context and more pixels. Tests in the paper show that neither of these two methods works. 

When the context is expanded from the default 200K to 780K, the number of Tokens per game rises from 30.7 million to 105.7 million, but the score drops from 99 points to 93.9 points. 

The same goes for images. When the image is zoomed in 16 times, the number of Tokens per frame rises to 1229, and the score drops to 88.3 points; when zoomed in only 4 times, the score reaches 99.7 points, which is higher than the default 8 times zoom. 

The paper explains that when the image is too large, the extra Tokens occupy the context for no reason, and the model does not see the content more clearly as a result. 

When the context is expanded from 200K to 780K and the image is zoomed from 8 times to 16 times, the score does not rise but falls

More is not always better, the entire VISTA design follows this idea, and the team mentioned on their blog that they deliberately pursue minimalism. 

There is no newly trained model in the system, the prompt for every game is the same, only four sentences in total, no word teaches it how to play the game specifically. 

All the prompts given to the model by VISTA

The code faction worked hard all summer, and VISTA caught up with one page of notes 

The high-score solutions that went viral on ARC-AGI-3 this summer, such as Tycho and Schema, all let the large model write a set of runnable programs for each game, and build a simulator from scratch for repeated trial and error. 

For the checkers game LF52 alone, Schema wrote a full 4000 lines of Python code. 

But the model in VISTA only used three sentences of natural language in its notes to figure out the same rule. 

For the same checkers rule, the left side is the code of the program solution (about 4000 lines in total), the right side is the three notes written by VISTA's model itself

According to the paper, VISTA is the first system that achieves full score or nearly full score on ARC-AGI-3 without writing programs. 

This point is even more important beyond games. In the real world, no one can write a 4000-line simulator for you in advance. 

Claude Opus 5 cleared all 183 levels of 25 games, got 100 points for each game, took a total of 7302 steps, which is only 0.43 times that of human players. 

GPT-5.6 Sol also cleared all levels and got 99 points. 

For the m0r0 game, humans need 1107 steps, while Claude only takes 219 steps. Both scores have scorecards on the official ARC Prize platform. 

The scorecard on the official ARC Prize platform, all 183 levels are cleared, all 25 games get full marks

Comparing 25 games one by one, the blue bars are all shorter than the gray bars. The blue ones represent Claude, and the gray ones represent humans playing for the first time

This purely mechanical, zero-training Harness is extremely versatile. 

When it is applied to 3D scenes with perspective, it still scores 84.12 points. 

When paired with GPT-5.6 Sol and applied to 34 web games (including Mario, 2048, Minecraft, etc.), the task success rate reaches 63.3%, which is higher than that of human beginners (