H-JEPA, LeCun's world model startup has delivered its debut results, with all code and weights fully open-sourced.
LeCun's world model startup AMI (Advanced Machine Intelligence) has just quietly released its latest research results.
In the latest paper H-JEPA, researchers from AMI, New York University, INRIA Paris and Brown University jointly proposed a hierarchical world model oriented to visual planning.
H-JEPA stacks multiple JEPA layers by layer, enabling each layer to predict states after different time spans in its own representation space.
When planning, the high-level layer first formulates a rough plan, then takes the predicted intermediate state as a sub-goal and hands it over to the low-level layer. The low-level layer takes over and refines the plan step by step, finally generating executable actions.
Through this division of labor, the world model can not only plan long-term goals, but also handle immediate actions.
In the Visual AntMaze simulation maze task, the success rate of the three-layer H-JEPA reaches 73%, while that of the single-layer model is 18%.
Meanwhile, in multiple simulation tasks, the hierarchical model also achieves better performance with less planning computational cost.
In addition to the tangible test results, H-JEPA also follows the open-source route that LeCun emphasized at the beginning of his entrepreneurship.
The team has simultaneously made the code and pre-trained model weights public, so researchers can directly load the model to reproduce the planning experiments.
Why do world models need hierarchical planning?
LeCun has actually been talking about this for a long time.
As early as 2022, when he was still serving as the chief AI scientist at Meta, he put forward the concept of hierarchical JEPA:
Let the model make predictions at different abstraction levels and time scales, and then turn complex goals into actions through hierarchical planning.
This may sound a little abstract at first, but it is actually very easy to understand if you look at the examples LeCun often cites.
Suppose you are sitting in the office at New York University and are preparing to go to Paris. When planning the whole trip, you focus on the destination, flight and schedule.
After confirming the trip, you start to further arrange how to leave the office, go downstairs, and take a taxi to the airport.
When you actually walk out of the office, the information you pay attention to turns into the ground under your feet, the stairs, where to step next, and which foot to move first.
△
These planning levels cover different time spans and require different information. When planning the voyage, the height of a certain step under your feet is not important; when walking down the stairs, this detail directly determines how you should step.
And this is not a unique feature of human planning. For robots, the high-level layer needs to understand task goals, environmental relationships and future directions, while the low-level layer needs to master specific actions and body control details.
For example, a quadruped robot has reached the target position, but its leg posture is different from the target image. The model may still think that it is far away from the target because of these differences.
So at this point, the problem becomes:
How to make different layers learn their own suitable information representations respectively, and then gradually turn high-level plans into low-level actions?
How does H-JEPA work specifically?
Specifically, in order to solve the above problems, H-JEPA configures three components at each layer to process state, action and prediction respectively.
Among them, the state encoder is responsible for extracting the representation of the current environment. The lowest layer directly reads the image, while the higher layers further abstract based on the representation of the lower layer, retaining more important information at this time scale.
The action encoder is responsible for describing the changes brought by actions. The low-level layer represents specific actions, while the high-level layer compresses a series of continuous actions into coarser-grained changes.
The predictor combines the current state and action to predict the future state. Different layers cover different time spans. In the paper experiment, one step of the adjacent high-level layer corresponds to two prediction steps of the low-level layer.
In terms of model training, the above layers learn together through end-to-end training, and different layers are connected through sub-goals.
The intermediate state predicted by the high-level layer becomes the sub-goal of the low-level layer. The low-level layer will convert its predicted state into the high-level representation space, compare it with the sub-goal specified by the high-level layer, and then adjust the action accordingly.
However, when training the world model, there is still a classic problem to be solved: representation collapse.
If the encoder compresses all inputs into the same vector, and the predictor only outputs this vector, a very low prediction error can also be obtained, but the model does not actually learn any useful information.
To this end, H-JEPA follows the SIGReg previously proposed by LeCun's team in LeJEPA to constrain the representation distribution and prevent different states from being compressed into the same vector.
This year's LeWorldModel applies this method to world model training. H-JEPA takes it as the bottom layer, adds more JEPA upwards, forming a structure of co-training and top-down planning.
In the experiment, the researchers further observed the information trade-offs of different layers. In the maze task, the high-level layer gradually weakens the leg posture, but still retains the position information of the robot.
However, hierarchy also has applicable conditions. The two-layer model in the Push-T task performs better, and after increasing to three or four layers, the performance drops instead.
The paper believes that this may be because the shorter task trajectory limits the available training data for the high-level layer.
LeCun is still the same LeCun
In general, these designs of H-JEPA continue a basic proposition of JEPA: complete the prediction in the representation space.
The full name of JEPA is "Joint Embedding Predictive Architecture". LLMs usually learn to generate content by predicting the next token; visual JEPA first encodes images and videos into internal representations, and then predicts the representation of the target part without generating image pixels one by one.
In world models such as H-JEPA, the model will also combine actions to predict the representation of future states. The planner compares the consequences of different actions accordingly to find a path to the goal.
LeCun has always believed that understanding the world, predicting the consequences of actions and planning are the necessary capabilities to move towards human-level intelligence. His judgment on LLMs has always revolved around this point.
In September this year, at ECCV held in Malmö, Sweden, the familiar suggestion appeared again in his presentation slides:
If you are interested in human-level AI, don't work on LLMs.
After this slide was spread on social platforms, some people suspected that it was fake. LeCun replied:
It's 100% real.
He then added that LLM-based AI tools are very useful and everyone is using them. But he believes that LLMs alone cannot achieve human-level intelligence.
Nowadays, this set of research propositions has also become the entrepreneurial direction of AMI.
In March this year, AMI announced that it had completed a $1.03 billion financing with a pre-money valuation of $3.5 billion.
LeCun serves as the chairman of the company, Alexandre LeBrun as the CEO, Xie Saining as the chief scientist, Pascale Fung in charge of research and innovation, Michael Rabbat in charge of the world model, and Rabbat is also one of the authors of the H-JEPA paper.
However, to be fair, research results belong to research results, and another expectation of the outside world for AMI is still its products.
At least from the current official website, the company mainly displays its technical direction, team and recruitment information, and there is no product entrance for the public.
On the other side, World Labs, which also bets on world models, has ushered in a new turning point.
On September 28, AMD announced that it has signed an all-stock acquisition agreement of about 8.2 billion US dollars with World Labs founded by Li Feifei.
It can be said that one side is moving towards integration with chip companies, while the other side continues to move forward along its own research route.
But with the milestone progress of Li Feifei's world model startup, is the major breakthrough of Yann LeCun just a matter of time?
Reference links
[1]https://venturebeat.com/technology/h-jepa-teaches-world-models-to-plan-at-multiple-levels-of-abstraction
[2]https://arxiv.org/pdf/2610.06805
This article is from the WeChat official account "QbitAI", author: henry, published with authorization from 36Kr.