This year, world models have started taking real exam questions.
I did some checking:
At this year's WAIC, over 1,100 enterprises gathered, with more than 300 products making their global debut. The hottest term across the entire event is World Model.
01
How crowded is it? Robotics firms are launching new products, AI film and television companies are hosting forums, and even agricultural tech players have rolled out a plant growth World Model.
On July 19, Kunlun Wanwei dedicated session directly declared this year as the first year of World Model; you might, like me, first think: this new technology just emerged this year. But that's not true at all.
Count back three years:
At the start of 2024, Google DeepMind released Genie, marking the first time AI learned what "action" means independently from videos. In August 2025, Genie 3 achieved real-time interaction, allowing you to walk freely in a world generated live by AI.
In November 2025, Li Fei-Fei's World Labs turned World Model into a commercial product; three years saw four major leaps, with each stage hailed as a "breakthrough" by the industry.
Kunlun Wanwei entered the field long ago, and its Matrix-Game has iterated to the third generation, enjoying a rather special status in the industry.
After the second generation was open-sourced, the team led by Xie Saining, assistant professor at New York University and author of DiT, used it as the base to develop Solaris, the world's first multiplayer video World Model; the third generation, Light Interaction jointly released by NVIDIA and Zhejiang University, was also developed based on it.
What's more interesting is that when international giants like NVIDIA released SANA-WM and Adobe released RELIC, they all used Matrix-Game as a benchmark for comparison.
What does this mean? When others take exams, they use your standards as their passing score line.
An open-source model that can reach such a status—what does that indicate? It was already usable last year. The technology has long been ready, yet the "first year" is counted from this year. What's the difference?
It's the exam paper. In the robotics field, there's a universal benchmark called LIBERO that all models must test on.
Robots are assigned a set of standard tasks to measure their completion rate. After three years of testing, the results this year are rather awkward: the scores of over a dozen mainstream models all cluster between 96 and 99 points, with the top ones differing by less than 1 point, and nearly everyone approaching full marks.
The total score gap among the top 10 performers in the provincial college entrance exam is smaller than the points of a single multiple-choice question. Can this paper still rank contestants? Yes, but even the paper's designers no longer trust the rankings it produces.
What does a test paper where everyone scores near full marks indicate?
The paper is obsolete, and it can no longer effectively differentiate performance. Your ranking depends on who you include in your competitor list; swap out a few opponents, and the top spot goes to someone new. In recent years, the phrase "refresh SOTA" in product launches has become increasingly meaningless.
The root cause lies here: everyone is grinding the same test paper that everyone can almost ace.
The other half of the World Model field faces the same issue; what do we use to evaluate video generation? Image stability, scene realism; all these metrics share one common trait: they rely entirely on human eyes.
Judges score by looking at screens, and audiences applaud at demos; the human eye has been the only examiner for three years, and this examiner is almost fully satisfied.
This is the situation in 2026: old benchmarks can no longer differentiate performance, and the old examiner is reaching the limit of satisfaction. When a technology reaches this stage, there are only two paths: continue competing for decimal points on the old test paper, or switch to a new test paper that shows no favoritism.
Therefore, when looking at World Model developments this year, don't count how many products each player releases; at this WAIC, Kunlun Wanwei is the one that laid out three new test papers at once.
02
The first new test paper features a robot that scored zero.
The protagonist is Riemann-1.0, the first assignment submitted by Kunlun Wanwei's subsidiary Riemann Dynamics: a "general brain" for robots.
Let me explain how this zero score came about.
Riemann's technical materials include a set of controlled experiments: a model trained from scratch without pre-training, installed on a dual-arm robot, to perform four types of household chores: folding clothes, clearing tableware, tidying desktops, and stacking blocks.
The success rate was 0%—not because the robot didn't move at all.
The robot extended its hand, touched the objects, and executed roughly 20% of each action correctly; but none of the four tasks were fully completed. That's the nature of this new test paper.
Models in simulation can score 95 points because simulation gives "process points": you get points for reaching your hand correctly, and lose points if your direction is slightly off. The physical world doesn't do that.
Diagram explanation: Riemann-1.0 technical architecture diagram: images, robot status, and actions form a causal chain within the same model; it first decides how to move, then predicts how the world will change.
Why do top performers in simulation turn in blank papers when transferred to real robots?
Because simulation is a lenient examiner: the lighting is always even, the table never shakes, and the folds of clothes are all pre-calculated.
The real world shows no mercy: fabric slips, light reflects, and if your grip strength is slightly off when you pick up a corner of clothing, the entire subsequent process fails.
This gap is known in the industry as the "Sim-to-Real Gap," where countless robotics teams have failed.
A piece of clothing is either folded properly or not—there is no "almost done" score tier, you have to be honest; look, the glory of two decimal places on the old test paper is completely reset here.
What separates 0 points from 85 points?
232,000 hours of video recordings.
86% of the training data fed to this robot brain by Riemann is first-person daily life videos of humans: how people fold clothes, tidy tables, and insert pens into pen holders.
Simply put, this robot brain "grows up" by watching humans live their daily lives. The remaining small portion of data is the robot's own operation records.
Can the skills learned from watching videos be successfully transferred to real robots? On July 19, Kunlun submitted its test paper; in the official real-machine evaluation, Riemann-1.0 achieved an average success rate of 85% across four types of household chores, 15 percentage points higher than the strongest open-source competitor.
There's an even larger test field: RoboCasa-365.
If you're not familiar with it, just remember it's the benchmark that most closely resembles real household scenarios; it covers 365 types of long-process kitchen tasks, from fetching ingredients to washing dishes, where points are only awarded after the full sequence of actions is completed. Riemann-1.0 scored 62.6%, while the second-place performer got 54.2%.
Notice the pattern of these two sets of numbers: it's a clear lead that creates significant differentiation. As soon as the new test paper is adopted, the ability to distinguish performance returns.
Two more figures explain the situation better than scores.
First: the entire robot brain runs on a single RTX 4090—yes, the same graphics card you use for gaming. The hardware threshold for a robot "general brain" is only over 10,000 RMB.
Second: in one evaluation task, the robot was asked to put a Rubik's Cube into a storage box; the Rubik's Cube and the specific box were never shown in the training data, yet it succeeded 10 out of 10 times.
It can handle tasks it was never taught—this is what the 230,000 hours of video truly brought: a model that has "seen the world."
This new test paper comes at a cost: it's so expensive that every participant has to trade something else to take the exam. The video generation model we mentioned earlier has just paid its tuition.
03
What's the other tuition fee? The answer is: speed.
This model is called Matrix-3.5, the latest generation of the Matrix-Game series that was previously used as a base by Xie Saining's team.
In its release materials, there's a number that looks like a typo in any product launch: the previous generation ran at 40 frames per second, while this generation runs at 20 frames per second.
It's 50% slower.
The unwritten rule of this industry's product launches is only one direction: "faster"—doubled frame rates, halved latency, improved responsiveness, which audiences understand immediately and judges reward with points. You can't find a second company that voluntarily says "we've slowed down."
Therefore, there's only one explanation for this number: it traded half of its speed for something it considers more valuable—memory retention and computing power lightweighting.
Previous AI-generated worlds had a flaw, like goldfish with short-term memory: when you walk forward, buildings and trees are there; turn around and look back, the original building is gone, replaced by a new one—it has forgotten the content it generated just one second ago.
For three years, World Model players have been shouting about "building virtual worlds," but the bottleneck has always been this: making the world able to remember itself.
In short, what Matrix-3.5 does can be summed up in one sentence:
Label every frame the AI generates with 3D coordinates and store it in an archive; when you turn around, it retrieves the original scene from the archive.
When you turn around, the building is still there; walk away for a minute and come back, the building remains in its original place with the same windows—this one minute of "persistence" is what those 20 frames bought.
To clarify: the 50% slowdown comes from switching from multi-GPU to single-GPU setup, and the team chose this path before working tirelessly to recover speed.
To prevent the 20 frames from dropping further to 10 frames, the engineering team compressed the generation steps from dozens to 3, cut the decoder size by three-quarters, prioritized memory retention, then gradually recovered speed, finally settling at the 20-fps sweet spot where both functions remain usable.
This is a clearly calculated trade-off.
By the way, regarding the "building remains there after turning around" task: previously, Google's Genie 3 was the best in the world, but it's closed-source and only accessible through demos. Matrix's solution is open-source.
You may notice that this choice is the same as the robotics path: the robot surrenders scoring rights to the physical world, taking a no-nonsense test paper beyond the screen.
Matrix-3.5 gave up "faster" for "better memory retention," voluntarily picking a harder problem within the virtual world; speed pleases the eyes, while memory retention pleases those who treat the virtual world as a real, persistent space.
One test paper beyond the screen, one within the screen. There's one last test field that will never have standard answers.
04
For example: music. Music can never be quantified by rules. No universal law in the universe can judge whether a song sounds good or bad; the only examiner for this task is the human ear, which you cannot replace.
So in a test field where the examiner cannot be replaced, how do you establish the "first year" concept?
The third model Kunlun released at this forum, the AI music model Mureka, gives an answer: if you can't replace the examiner, be completely honest with them.
In its technical materials, there's a number that rarely appears in product launch slides because it's unflattering:
Let the new model generate 100 songs, and invite professional judges to score them using consistent criteria; only 28 songs passed all three checks: listenable, usable, and on-topic. The comprehensive usability rate in this round of evaluation is 28%.
Does this number look impressive? No, it doesn't.
Don't feel sorry for it yet. Think about yourself: have you ever used AI to generate images? Write copy? Out of 10 attempts, how many times can you use the output directly without any edits?
Yes, that's exactly the feeling; the 28% figure matches the real experience of every person who has used AI products.
Moreover, think in reverse: a company that dares to print 28% must have another internal calculation in mind.
This 28% is the yield of products that pass all three checks, adhering to the "directly usable" standard; if you loosen one check and only count "listenable," the number will jump significantly. But they didn't loosen the criteria.
Because it clearly understands that user experience cannot be hidden: if you claim 90%, users will see through it after three attempts; if you claim 28%, users will feel you are being honest after using it three times.
The phrase "matches real experience" is far more scarce in this year's product launch season than any impressive number; the entire field is flooded with 99% claims and "far ahead" slogans. How long has it been since you saw a company print its less-than-30% yield rate in official release materials?
Not just this number, the same document acknowledges another trade-off:
To make the music sound "more human and less AI-like," the team compromised on the richness of arrangement, which is not the highest in peer tests.
This trade-off has a living witness at the forum scene.
At the roundtable, Huang Xiaoming talked about his experience recording a new song: The composition was written by humans, but all demo vocals were AI-generated. His first reaction was "it sounds so good, I wish I could sing that well."
Then the producer poured cold water on him: precisely because it was AI, it was too perfect, to the point where there was no sense of breathing, no human touch at all.
As a singer with decades of experience, he was struck by the "over-perfection" of AI; too much perfection is exactly where the artificial "AI flavor" comes from, while breath and imperfections are the source of human touch.
Mureka gave up the richness of arrangement in exchange for "less AI flavor," betting on this direction.
Diagram explanation: WAIC Kunlun Wanwei World Model and Multimodal Paradigm Revolution Special Forum Session
Look, they wrote their own shortcomings openly. This is how you take exams in a field with no standard answers: if you can't give the impartial scores from the physical world, give a score that's 100% honest.
By the way, Huang Xiaoming has actually used Mureka.
He asked Mureka to generate a lyrical song in Aska Yang's style, and his evaluation was "Mureka is indeed pretty good," adding specifically: the lyrics are quite touching.
The first two test