Kunlun Wanwei's Fang Han Unveils Three Major Industrial Predictions at WAIC: World Models Will Step Beyond Screens, Ushering in a "Paradigm Shift" in the Gaming and Music Industries
On July 19, during the 2026 World Artificial Intelligence Conference (WAIC), Kunlun Wanwei's special forum "World Models and the Paradigm Shift of Multimodality" was grandly held at the West Bund International Expo Center in Shanghai. As a leading enterprise deeply rooted in the AI technology sector, Kunlun Wanwei announced a major upgrade of its core models at the forum — the Matrix-Game 3.5 World Model, Mureka v9.5, and O3 Music Model, which comprehensively cover cutting-edge tracks such as world models and AIGC creation. Fang Han, Chairman and CEO of Kunlun Wanwei, announced at the forum that 2026 is the "First Year of the World Model", marking a critical transition for AI from content generation to understanding and interacting with the physical world.
The scale of this WAIC is unprecedented. Xi Jinping, President of the People's Republic of China, attended the opening ceremony and delivered a keynote speech. This is the first time since the launch of WAIC that a top-level national leader has been present. The conference is co-hosted by more than ten ministries and commissions including the Ministry of Foreign Affairs, the National Development and Reform Commission, and the Ministry of Industry and Information Technology, together with the Shanghai Municipal Government. The total exhibition area has exceeded 100,000 square meters for the first time, with more than 1,100 participating enterprises, over 3,000 exhibits on display, and over 300 products making their global debut. The two core tracks of intelligent computing and embodied intelligence each gather more than 200 enterprises. Over 140 forums were held intensively, and 9 Turing Award and Nobel laureates came to share their insights. This is a truly national-level, world-top-tier AI event.
Fang Han: 2026 is the "First Year of the World Model", with three industrial predictions anchoring the future direction
During the keynote speech session of the forum, Fang Han, Chairman and CEO of Kunlun Wanwei, officially declared 2026 as the "First Year of the World Model". He pointed out that the core proposition of the AI industry in the past two years has been "generation" — generating text, images, videos, and music. Starting from 2026, the core proposition of AI will shift to "understanding" and "interaction" — enabling AI to truly understand the operating laws of the physical world and engage in closed-loop interactions with the real environment. The world model is exactly the key technical foundation to achieve this leap.
Fang Han, Chairman and CEO of Kunlun Wanwei
In his speech, Fang Han systematically elaborated on three industrial predictions:
First, world models will step out of the screen and into the physical world. We have observed that embodied intelligence is moving from laboratories to real-world environments. This trend requires robots to perform environmental reasoning before taking actions — predicting the consequences of actions, evaluating environmental changes, and adjusting response strategies in a timely manner. The breakthrough of this capability will shift the focus of competition from single tasks to the general model capability in unfamiliar environments. Whoever can train a model that can still "understand correctly and act properly" in complex scenarios will get the admission ticket to the physical intelligence era, which is also the underlying capability urgently needed by China's intelligent manufacturing and robotics industries.
Second, the gaming industry will be the first to be transformed by world models. Open-world games will no longer rely on hundreds of team members spending years on manual construction. Matrix-Game 3.5 allows the virtual world to grow and evolve in real time as players take actions. In the next three to five years, world models will become the infrastructure of the gaming industry, completely revolutionizing the content production method, and domestic models have already taken the lead globally in this track.
Third, tools such as AI music and AI video will evolve from technical experiments to mass creation tools, promoting the in-depth development of AIGC. As the strongest music generation model in China and one of the top two globally, Mureka is the first to eliminate the "AI flavor", allowing ordinary people to turn vague inspirations into emotional and warm works. This is not to replace musicians, but to significantly lower the creation threshold, enabling more people to express themselves freely, thus fundamentally changing the creative ecosystem of the music industry and even the entire AIGC industry.
Behind the three industrial predictions lies the same direction: enabling AI to move from "being able to generate" to "understanding the world" and "being capable of acting". Fang Han stated that this is not only the answer from Kunlun Wanwei alone, but also a microcosm of China's AI moving from R&D to industrial applications.
Global upgrade of AI models, sparking a revolution in full-modality content generation
At this forum, Kunlun Wanwei completed a major global upgrade of its core models in a concentrated manner.
2.1 Matrix-Game 3.5: Patch-level Memory Injection, Redefining Interactive World Models
As the latest iteration of Kunlun Wanwei's world model series, Matrix-Game 3.5 brings revolutionary technical upgrades at this forum.
Cheng Yu, Chief Scientist of Skywork, officially released the Matrix-Game 3.5 World Model at the forum. Matrix-Game 3.5 focuses on interactive world model technology, achieving Patch-level memory injection and system-level performance optimization. The 5B model can achieve real-time generation at 20FPS on a single GPU under 720p resolution, with 1-minute memory capability.
Cheng Yu, Chief Scientist of Skywork
Cheng Yu pointed out that the world model is the "infinite data engine" for embodied intelligence. Traditional data collection faces pain points such as high labor costs, time-consuming and labor-intensive real-machine collection, and extremely low efficiency, which cannot meet the large-scale demands of large models. Matrix-Game 3.5 builds an infinite data production system supporting the next-generation world model, breaks through traditional bottlenecks, creates three automated data production pipelines, and outputs high-quality training data in the format of Video+Pose+Action+Language. Currently, it has accumulated more than 5 million high-quality video clips, over 10,000 effective training hours, and covers more than 1,200 game scenarios.
In terms of technical architecture, Matrix-Game 3.5 innovatively splits historical frames into Patch Memory with 3D coordinates (lifted to the world coordinate system via Depth+Pose). During inference, it retrieves visible Patches based on the current camera frustum and reprojects them to the current perspective to form a Mosaic, which is then injected into DiT as Memory Token to achieve long-term consistent Patch-level spatial memory. This design upgrades Frame-level FOV Memory to Patch-level FOV Memory, bringing four core advantages: more accurate spatial memory retrieval and cross-frame consistency, more stable camera motion generation, In-context Learning that maximally preserves the dynamic generation capability of the base model, and controllable memory that allows users to freely edit memory blocks.
In addition, Matrix-Game 3.5 has achieved full-link optimization of DiT inference and VAE decoding at the inference acceleration level. Through three measures: "reducing model throughput + DiT model optimization + VAE pruning", it has reached the industry-leading performance of 20FPS 720p real-time inference on a single GPU, enabling high-quality world models to truly have real-time interaction capabilities.
From version 1.0 to 3.5, Matrix-Game has always adhered to the open-source path. The 3.5 version announces that its core architecture is open-source, attracting global developers to jointly build the ecosystem. Matrix-Game 2.0 is the first open-source implementation in the real-time interactive world model technology paradigm, and Matrix-Game 3.0 is the first to systematically incorporate the memory problem into an open-source solution.
Previously, the team of Xie Saining, Assistant Professor at New York University and author of DiT, based on the open-source base of Matrix-Game 2.0, released Solaris, the world's first multi-person video world model, which also confirmed the base model value of this open-source solution from the practical use in the academic circle. The work Light Interaction jointly released by NVIDIA and Zhejiang University was developed based on the Matrix-Game 3.0 model. In addition, world models from top international manufacturers such as NVIDIA's SANA-WM and Adobe's RELIC all use Matrix-Game-related models as comparison benchmarks.
Matrix-Game 3.5 Open Source Address
https://github.com/Riemann-Dynamics/Matrix-Game-3.5
2.2 Mureka v9.5 and O3: The AI Music Model with the Least "AI Flavor", Leading Music into the Era of "Versioned Production"
The release of the Mureka v9.5&O3 large music model has become one of the most emotionally impactful moments of this forum.
From the internal test of SkyMusic 1.0 to Mureka v9.5&O3, Kunlun Wanwei's music model has gone through an evolution path from "valid works" to "controllable creation" and then to "multimodal creation". The v9.5 version first put forward the slogan of "the AI music model with the least 'AI flavor'" — it no longer relies on superposition of sound tracks and complex arrangements to create a powerful first impression. Instead, in a real music framework, it handles arrangements, vocals, emotions, and Prompt intentions in a more restrained manner, making musical expression more natural and sincere. O3 is built on the new V9.5 version and extends the inference process of MusiCoT through test-time scaling. It does not simply "generate more" or "think longer", but allows the model to continuously examine the current expression during the generation process: whether the local melody still serves the overall goal of the whole song, whether the arrangement overshadows the vocals, whether the emotion is overdrawn in advance, and whether the structure is deviating from the original intention. The extra inference space is used to review deviations and perform self-correction, enabling the music CoT to converge gradually instead of proceeding all the way after a single decision.
Subjective scoring of songs
From the evaluation results, V9.5 has completed a more solid base upgrade along the MusiCoT path: melody, vocals, sound quality, and arrangement have been comprehensively improved, and instruction accuracy has increased from 6.92 to 7.62. It no longer takes "fuller and more complex" as the only standard, but handles arrangements, vocals, emotions, and Prompt intentions in a more restrained manner within a real music framework, making songs more natural and sincere. O3 is built on V9.5, and through continuous reasoning, reviewing deviations, and correcting subsequent decisions, it further enhances melodic development, arrangement structure, and instruction understanding, making the song structure more complete, emotions more coherent, and paragraph relationships more reasonable. Simply put, V9.5 is responsible for making songs more natural, and O3 is responsible for enabling the model to review, think more, and revise more before delivery.
Nowadays, a song is often consumed together with videos. Short videos, games, film and television trailers, brand videos, and virtual human performances all require sound and images to express together.
This major upgrade of Mureka also integrates music and video MV into the same creation pipeline. Users can start from an image, a video, or a moodboard, let the system understand the rhythm, characters, space, and emotions of the picture, and then generate matching music; of course, they can also start from a song, let the Agent understand the lyrical narrative, beat arrangement, and emotional fluctuations, and then continue to generate visual solutions for characters and MV. Lyrics determine the narrative, beats determine editing, melody determines memory points, and images determine communication scenarios. What the Agent needs to do is to make this information continuously transmitted under the same creation goal, instead of simply splicing several generation tools together.
More importantly, Mureka has evolved from a single music generation tool to a complete "versioned production" platform. Its differences come not only from the model effect, but also from the systematic combination of "model training — optimization selection during inference — versioned creation tools — platform distribution". Multiple versions of the same idea can be quickly generated, supporting partial retention and replacement in dimensions such as melody, vocals, and structure; Studio transforms "operating DAW" into "directing creation", lowering the threshold while raising the upper limit. At the same time, Mureka has launched the Character function — a piece of sound or a photo can generate an exclusive singer role that runs through the entire pipeline of songs and MVs; AI soundtrack can automatically analyze the scenes, actions, and emotions of video frames to create music that better fits the content; Mureka Agent can call all capabilities such as full-song generation, single-track generation, Remix, extension, MIDI export, and voice cloning through a single natural language dialogue, without switching interfaces throughout the process.
As the vision proposed by Kunlun Wanwei: AI music is no longer just simple consumption content, but has been upgraded to a language for self-expression. Users are not only passive listeners, but also active expressers and participants. Mureka is redefining the music creation ecosystem in the AI era with a complete closed loop of model capabilities, creation tools, and platform distribution, allowing every ordinary person's inspiration to be realized with warmth, from "an idea" to the complete creation pipeline of "lyrics — vocals — characters — songs — MV", fully releasing the in-depth creativity of AIGC.
Academicians and industry leaders gather to discuss the future of world models
This forum also invited many heavyweight guests. Zhou Zhihua, Academician of the Chinese Academy of Sciences, Professor and Vice President of Nanjing University was invited to deliver a keynote speech, sharing in-depth insights on the cutting-edge directions of world models. In addition, we invited Professor Liu Yang, Founder of Riemann Dynamics Robotics, and Alberto Taiuti, CEO and Co-Founder of Reactor Technologies, Inc. to respectively introduce the Riemann Robot Model 1.0 and Reactor's experience in accelerating the global implementation of world models.