Kaiming He's team released the multimodal Harness framework
The vision-native Harness for multimodal models is here!
Recently, the team led by Kaiming He proposed a vision-native Harness in their latest paper VISTA: A Visual Harness for Reasoning in an Interactive World.
VISTA.
It enables existing multimodal models to directly observe the environment, save original frames, and review, zoom in and inspect details at any time during reasoning, without requiring full retraining from scratch.
Compared with the traditional approach that abstracts the environment into text or code, VISTA preserves raw visual information. The model can re-invoke previously viewed frames into the context based on the current problem to participate in ongoing judgment and reasoning.
Based on this method, visual experience becomes context that the model can call and review repeatedly at any time, and multimodal models are also equipped with a native scaffold built around visual memory.
In terms of evaluation, paired with VISTA, Claude Opus 5.0 completed all 25 public games of ARC-AGI-3, achieving a full score of 100 in the human-relative action efficiency metric, with 57.4% fewer game actions than the human baseline of first-time play.
When replaced with GPT-5.6 Sol, it also cleared all games and scored 99.
Furthermore, the team verified this method in browser games, mazes and connection puzzles, and listed embodied tasks that are closer to the physical world as the follow-up research direction.
How is this achieved?
Vision-Native Scaffold
From ResNet, Mask R-CNN to MAE, visual perception and representation learning has always been a main line of Kaiming He's research.
This time, VISTA targets a new problem: how to enable models to accumulate, store and utilize visual experience during continuous interaction?
The answer is Harness.
In November last year, Anthropic took long-duration programming tasks as an example to introduce how Harness helps models span multiple context windows and continuously advance tasks.
Specifically, Harness records pending items through a to-do list, and saves completed work through progress files and code records.
In this way, even when the model enters a new context window, it can understand what has been done before and what to do next by reading this information.
Complex tasks can thus be executed step by step, with results checked, and work continued across different context windows.
It can be said that this Harness, which mainly saves task status through text and code, has to a certain extent promoted the recent burst of Agent capabilities.
However, in real-world environments, relying solely on textual memory is obviously not enough.
Because once entering the real visual environment, the model not only needs to remember the position and orientation of objects, but also understand exactly what changes have occurred in the environment before and after each action.
Moreover, when the model sees an image for the first time, it does not necessarily know which detail will become important later.
At the same time, it is difficult to fully record all this visual information in a text note in advance, let alone ensure that the recorded information is still sufficient when the model executes subsequent tasks.
In benchmarks such as ARC-AGI-3, one existing idea is to first convert the screen into a text grid composed of numbers, and then let the model write programs to simulate environmental rules and verify action plans.
(Note: ARC-AGI-3 is an interactive visual reasoning benchmark that does not provide pre-defined game rules and goals. Agents need to explore how to clear the level on their own through observation and action)
However, the problem is that the more complex the environment is, the more difficult it is to completely convert the appearance, spatial relationship and dynamic changes of objects into text and code. Moreover, as the interaction history grows longer, early frames will gradually be moved out of the context or replaced by text summaries.
Therefore, the research question of VISTA becomes the following:
How to enable agents to re-inspect previously seen visual information when needed for reasoning, without additional fine-tuning of the base model?
Based on this, VISTA stores every frame returned by the environment as-is outside the context, including intermediate animation frames during actions, and retrieves them when the model needs them.
In the ARC-AGI-3 evaluation, it can not only compare frames from different moments, but also zoom in on local areas to inspect orientation markers on small squares.
Among them, text notes record the model's current understanding, and original frames retain evidence for it to re-judge.
The research calls this mechanism explicit attention to interaction history, where the model selects which visual information to re-enter the context based on the problem it is thinking about.
To make this selection possible, the system must save original frames and provide tools for retrieval and inspection at any time.
Specific Implementation
To enable the model to save and utilize past visual experience, VISTA is designed with three core components: visual observation, lossless visual memory and active visual inspection. These three components have their own divisions of responsibilities, which are responsible for enabling the model to see the frame clearly, save history, and review on demand respectively.
The first is visual observation, which is responsible for providing environmental frames to the model.
In the ARC-AGI-3 experiment, VISTA upscales the official 64×64 frame to a 512×512 PNG image, while preserving the color, appearance and spatial relationship of objects, so that the model can directly observe the environment.
The second is lossless visual memory, which is responsible for saving the frames seen by the model.
After each action is executed, VISTA completely saves all frames returned by the environment, including intermediate animation frames during the action, and establishes indexes by episode number and frame number.
In this way, the model does not need to judge in advance which information is worth remembering, but can first save the complete visual experience for subsequent use.
The last is active visual inspection, which is responsible for enabling the model to review historical frames on demand.
Through the inspect tool, the model can specify a historical episode to call up the corresponding frame, and can also crop and zoom in on local areas to inspect the details inside.
If you want to know exactly what changes an operation has made, the model can also call up multiple frames at once to compare the changes before and after the action.
The whole process is equivalent to equipping the model with a set of visual archives that can be accessed at any time. It can not only see the current environment, but also actively review past observations to provide a basis for subsequent actions.
Speaking of which, how do these three components cooperate specifically?
According to the execution flow of VISTA, at the beginning of each episode, the model first observes the current frame and available actions, then combines the previously accumulated experience to judge the state of the current environment, and puts forward hypotheses for the next action.
If the existing information is insufficient, the model can call tools to review historical frames, zoom in on local areas, and even read specific pixel information to further find evidence.
After finding the evidence, the model will not act directly, but first predict what changes this operation may bring.
After the action is completed, the actual result is compared with the prediction to check whether its judgment is correct, and the understanding of the game rules is revised accordingly.
Repeating this way, the model can explore the environment while accumulating experience in continuous interaction.
Of course, visual archives alone are not enough. To keep the model coherent in long-duration tasks, VISTA also introduces two text notes:
GUIDE.md: records rules and experience that can be reused across different levels.
WORKING.md: records the status, progress and follow-up plan of the current level.
When the context approaches the upper limit, the model will first organize a handover summary, and then enter a new context window to continue executing the task. Previous text notes, action history and visual archives will all be retained.
There is also a notable design here: although VISTA completely saves historical frames, it does not stuff all images into the model's context all at once.
In the complete solution, after each action is executed, the framework only shows the last frame to the model by default. As for the intermediate frames during the action, they are all stored in the visual archive and actively retrieved when the model needs them.
This not only retains complete visual information, but also prevents a large number of historical images from continuously occupying context space.
More importantly, VISTA does not require additional training of a new model for this purpose.
Environment understanding, action planning and reasoning are still completed by off-the-shelf multimodal models, while Harness is responsible for executing tool calls, saving visual history, and providing corresponding evidence when the model needs it.
In other words, what VISTA changes is not the model itself, but the way the model obtains, saves and uses visual information.
Experimental Validation
In addition to the experiments mentioned at the beginning, the team also tested VISTA on three benchmarks: GameWorld, AI GameStore and BabyVision, covering tasks such as browser games, mazes and connection puzzles.
Using the same GPT-5.6 Sol model, compared with the official base framework, VISTA increased the success rate from 40.0% to 63.3% in 170 tasks of GameWorld;
In 10 games of AI GameStore, the overall score rose from 47.3 to 140.3, where the human median was normalized to 100;
On the 39 maze and connection puzzles selected by BabyVision, the accuracy increased from 41.0% to 63.2%. For the last set of tasks with only static images, the model can also improve its judgment by zooming in on local areas and inspecting pixels.
These results show that saving original frames and allowing the model to review them on demand can further unlock the capabilities of existing multimodal models.
The same VISTA framework can be applied to different games and puzzles with minimal adaptation, which also demonstrates its application potential in more visual tasks.
In terms of conclusions, the research points the future verification direction to embodied tasks that are closer to the physical world: enabling the model to save, review and utilize its own visual experience in a continuously changing environment to determine the next action.
About the Authors
Finally, let's introduce the authors of this work.
The three co-first authors of the paper are Qiushi Han, Keya Hu and Linlu Qiu, and the other two authors are Cathy Wu and Kaiming He.
Qiushi Han (Josh Han) is currently a PhD student at the Operations Research Center of MIT, with Prof. Cathy Wu as his advisor.
Keya Hu is currently a PhD student in the Department of Electrical Engineering and Computer Science at MIT, co-advised by Kaiming He and Jacob Andreas.
She graduated from the ACM Class of Shanghai Jiao Tong University for her bachelor's degree, and her research interests focus on the intersection of language and vision, aiming to build agents with higher data efficiency and stronger generalization ability.
Linlu Qiu is currently a PhD student in the Department of Electrical Engineering and Computer Science and the Computer Science and Artificial Intelligence Laboratory at MIT, advised by Prof. Yoon Kim and Prof. Jacob Andreas.