HomeArticle

After Agentic AI, is the next paradigm Physical AI?

36氪的朋友们2026-07-20 09:42
An increasing number of practitioners have begun to ask questions.

When the competition among large language models evolves into a "super capital-intensive" arms race, an increasing number of practitioners are beginning to ask the same question: Where is the next paradigm of AI?

The answer is moving from "inside the screen" to "outside the screen".

On the evening of July 17, during the 2026 World Artificial Intelligence Conference (WAIC), the "Tencent WAIC Night", hosted by Tencent News and Tencent Technology, with special support from Tencent East China Headquarters and Tencent Cloud Intelligence, was held in Shanghai.

At the event, Chen Yu, Managing Partner of Qiming Venture Partners, Huang Xiaohuang, Co-founder and Chairman of Coohom, Luo Yihang, Co-founder and CEO of ShensTech, and Zhang Zhizheng, Co-founder and LLM Lead of Beijing Galaxy Universal Robotics, jointly launched a roundtable discussion on spatial intelligence, world models and Physical AI, centered on the topic "After Agentic AI, What Is the Next Paradigm?".

From left to right: Chen Yu from Qiming Venture Partners, Huang Xiaohuang from Coohom, Luo Yihang from ShensTech, Zhang Zhizheng from Galaxy Universal Robotics

Huang Xiaohuang put forward a viewpoint at the opening: he believes that simple Agents with pure software and no industry accumulation can hardly maintain their advantages, because large models can basically replicate them directly. In the future, tools, hardware, data and large models need to collaborate to continuously amplify the value of AI products. Coohom's focus is on "letting large models generate an interactive world", a direction beyond the "range of large language models" — Physical AI.

At CES this January, Jensen Huang, founder of NVIDIA, repeatedly emphasized that AI will evolve from perception, generation, agents, and eventually to "Physical AI" that can understand the physical world, and asserted that "the ChatGPT moment of Physical AI" is coming soon.

Following the thread of the "physical world", Luo Yihang summarized it as a "general model for the physical world": just as the Scaling Law of language models has verified generalization, the physical world will eventually have its own "ChatGPT moment". He judges that the generalization of embodied intelligence may arrive in 3 to 5 years.

Zhang Zhizheng's statement is the most concise — from Digital AI to Physical AI, the learning paradigm will shift from offline learning to online learning, "learn to interact first, then keep learning from interactions", where interaction is no longer the purpose, but a data flywheel that drives capability growth.

The shift from Agentic AI to Physical AI is being revalued by capital and policies.

The World Economic Forum judged in early 2026 that AI is "beginning to operate in the real economy as a physical system"; Deloitte's "2026 Technology Trends" report stated directly that Physical AI is ready for mainstream deployment. Data also supports this: the global embodied intelligence market size was about 4.44 billion US dollars in 2025, and is expected to reach 23 billion US dollars by 2030, while China's industrial scale is expected to exceed one trillion yuan by 2035.

Although all parties have a common understanding of the direction, there are obvious differences in the paths — the focus is precisely on the "world model".

Huang Xiaohuang believes that the world is driven by invisible laws such as gravity and friction, and vision is only an appearance, so he is more optimistic about returning to physical simulation. At present, Coohom's exploration of world models has gradually converged to spatial intelligence. On the other side of spatial intelligence, Li Feifei presented a different vision with Marble, the first commercial world model from World Labs, whose valuation soared from 1 billion US dollars to 5 billion US dollars in more than a year. Luo Yihang advocates integration, arguing that every single path will eventually hit a bottleneck; Zhang Zhizheng breaks down the "world model" into three indispensable capabilities: modeling the change of state caused by actions, using modeling to promote policy learning, and having general evaluation capabilities.

The three also have different focuses on the implementation pace: Huang Xiaohuang regards "digitizing physical world information in a low-cost and large-scale manner" as the primary challenge; Luo Yihang admits that current models and data are "not good enough", but has already seen the "precursor" similar to the evolution from GPT-2 to GPT-3.5; Zhang Zhizheng's judgment is the most radical — the logic of Physical AI is to "pursue full generalization within a limited scope", so the arrival of Physical AI Agents may greatly shorten the cycle compared to Digital AI.

The answer to paradigm shift may not come from a pre-determined technical route, but gradually emerge in this continuous debate.

The following is the transcribed text of this panel, with adjustments and deletions made without changing the original meaning.

01

The Next Paradigm of AI

Chen Yu: Thank you for Tencent's invitation. Over the past year, the development of the entire AI industry can be described as "advancing at a tremendous pace" — whether it is the leap in AI capabilities itself, or the performance of several large model companies in the capital market. But AI is no longer limited to language models, it has expanded to embodied intelligence, video generation, and spatial intelligence.

Today we are honored to invite the CEOs and chairmen of companies that are either pre-IPO or listed, focusing on the three directions of spatial intelligence, video generation and embodied intelligence, to discuss together — after Agentic AI, what is the next AI paradigm? First, please give a brief self-introduction from left to right, and answer in one sentence: what do you think the next paradigm of AI will be?

Huang Xiaohuang: Hello everyone, I am Huang Xiaohuang, Co-founder and Chairman of Coohom. Our company is positioned in the field of spatial intelligence. When I first started my business, I used GPU clusters to do spatial rendering. Later, with the development of the times, I began to do "inverse operations" — inferring structured data from rendering results or real-world results, and thus "stumbled into" this track.

In recent years, I have been thinking about one thing: large language models have entered a super capital-intensive competition, so I often wonder which models are outside the range of large language models?

After observing the development of large models for more than a year, my feeling is that pure Agents that only focus on pure software and have no industry accumulation will have increasingly weak moats. Because large models can basically replicate software or Agents directly, it is difficult to establish long-term competitive advantages relying solely on tools. My preliminary judgment is that in the future, tools, hardware, data and large models need to be deeply integrated.

So where should we go in this era?

I spent two or three years thinking and exploring, and came to the conclusion that large models must be developed, but they must be built outside the range of large language models. So we initially trained four or five different models, and finally converged to spatial intelligence — how to generate an interactive world that is difficult to describe precisely with language. It is a bit like Li Feifei's Marble, but we think that pure static scenes are not enough, and certain interactions are also required.

In short, it is to let large models generate an "interactive world", which is the direction we chose, at least it is outside the range of large language models.

Luo Yihang: Hello everyone, I am Luo Yihang, Co-founder and CEO of ShensTech. ShensTech focuses on multimodal generation and world models. What is the next generation paradigm? I think it must be a point that no one can think of now, that has not converged, and has no consensus. My conclusion is: a general model for the physical world.

Why? First, the core paradigm of this generation of models is the Scaling Law. Language models have verified their versatility and generalization. I believe that a model similar to the language model paradigm will definitely appear in the physical world to solve the interaction and generalization problems of the physical world. It needs to achieve generalization and logic: first, there will definitely be architectural innovations, just like Transformer for language models, we are also exploring hybrid architectures such as Diffusion Transformer or MoE to build the architecture of world models; second, there will definitely be innovations in data. I don't think relying solely on current embodied ontology data is a good path.

In the future, we can imagine that various robots with different ontologies and architectures in the physical world are just like Agents and WorkBuddy in the digital world today — they can plan independently, collaborate, and autonomously complete some long-term tasks, which are extensions of Agents in the physical world. The ontologies of these robots must be diverse, but their intelligence must be relatively general.

We firmly believe that world models will have their own "ChatGPT moment", giving birth to such versatility and generalization; when the embodied world model reaches a certain level of intelligence, combined with the development of ontology and cerebellum, I think that in 3 to 5 years, the moment of embodied generalization may come.

Zhang Zhizheng: Hello everyone, I am Zhang Zhizheng from Galaxy Universal Robotics, co-founder and large model lead of the company. If we talk about the next change of AI, I think it is very clear: the shift from Digital AI to Physical AI — from building an AI model for the digital world to building a model that can interact with the physical world. This is also one of the reasons why embodied intelligence is attracting so much attention in academia and business today.

In this process, the learning paradigm will undergo a fundamental change. In the past, when we built GPT and large models for the digital world, many of them adopted offline learning; but if we want to build a model that can reliably interact with the physical world, we must first change offline learning to online learning — learn to interact first, then keep learning from interactions.

Under this paradigm, interaction is not the ultimate goal of learning, but a method that turns our AI into a flywheel of capabilities and data, continuously moving forward.

Driven by this data flywheel, I believe we will move from the widely used digital agents today to physical world agents that can help us handle various tasks in the physical world. This will bring greater changes to the entire industry — whether in R&D or application, as well as our lifestyles and work patterns.

02

"World Model": The Debate Between 3D Simulation and Pure Video

Chen Yu: All of you just mentioned that Physical AI may be the next-generation AI development paradigm, and you also mentioned world models. I'll just advance the topic of world models a little later.

There are different opinions on world models now. I may receive five BP (business plans) about world models a day, but everyone has different understandings and definitions of world models. How do you understand world models? What exactly are they?

Huang Xiaohuang: Let me speak first. Our team has actually been researching world models for a long time. In summary, there are basically two main routes: one is to follow traditional 3D simulation, using structured data to carry out simulation; the other is to follow the route of pure video training.

The former believes that the operation of this world exists beyond vision — the movement of objects is formed by the interaction of various forces such as gravity and friction, or the interaction of certain "causes". Therefore, the world is composed of invisible factors, and they think that the pure visual approach is unreliable, and vision only "deceives humans".

The latter believes that as long as the data volume is sufficient and the training algorithm is upgraded, just like humans mainly rely on vision to understand the world, the pure video model approach should work.

Our past accumulation is all in parametric models and structured data.

I can't prove which direction is definitely right or wrong now, we think both are possible, but we are more optimistic about the former — we believe that the world is composed of countless physical laws and constraints. In the end, the operation of the world still needs to return to physical simulation. Only after "guessing" all the physical and 3D characteristics of the world can we accurately predict the world.

So we take a route that leans towards physical simulation, but we do not deny the feasibility of the pure video model. This is our viewpoint during the research process.

Chen Yu: But an important reason why people use video models is that video data is relatively easy to obtain. Do you think the former route is more difficult to implement technically?

Huang Xiaohuang: A lot of our underlying data also comes from videos. But in the final calculation, whether it is based on the video model or based on physical values is an essential difference. It doesn't mean that we can't get information from video data. I think 80% of the work in these two methods may be similar, and the last 20% will diverge.

Even within our company, there are ongoing debates between the two ideas: because everyone sees that large language models, as long as the Scaling Law works and the data volume is infinite, why can't they simulate that gravity is 9.8? Some people think that gravity of 9.8 is something that can be solved with one physical parameter, but you may need millions of videos to train it out.

So there are different viewpoints here. In addition, I personally have a judgment — the pure video model may be the "dish" of large companies, and we don't dare to do it easily.

Chen Yu: So what do you think, Mr. Luo? In a sense, your company and ByteDance are competitors.

Luo Yihang: Our idea is: first figure out what the end point is, then reverse deduce what kind of route we need. The first key point is that this model needs to form a closed loop from perception, understanding, prediction to action. Second, I don't think a single route can work through — just like in language models there are not only LLMs, in video generation there are not only DiT, there are also autoregressive routes now, and they are actually integrating with each other.

So I think world models will also follow an integrated route, because every single route will have bottlenecks.

Our own Motubrain follows a unified integrated route, which includes not only video data, but also some physical reinforcement data, ontology data, Ego data, and so on.

We think it needs to meet two conditions: first, it should be like a human, not taught purely step by step, but need to "watch", then "think" after watching — for example, when we pick up an object, there is already a world model in our mind after watching, which needs to predict, imagine, and then act. This is what we think is the essential logic; on the other hand, I think the architecture will continue to evolve, and it will not converge right now.

Chen Yu: But there is a problem here: many things cannot be seen from videos. For example, when I use my hand to pick up a cup, how much force I need to apply in the middle is very difficult to get from watching videos.

Luo Yihang: In the data for pre-training, mid-training or fine-tuning, those physical data and tactile data are also a type of data; including the parts that RL (reinforcement learning) needs to align later, these data are also included, which are part of the input signals.

Chen Yu: Okay, Mr. Zhang, what's your opinion?

Zhang Zhizheng: World models can be said to be the most frequently asked topic to me this year.

Now all walks of life are talking about world models, and people researching different directions have different definitions and understandings of it: people who used to work on content generation and video generation say that video generation models are world models; but people working on embodied intelligence will say that models that can predict the future and generate actions are world models.

In fact, I think the current public understanding of world models is incomplete. If we return to the essence, what kind of model can truly be called a world model? I summarize it as three capabilities, all of which are indispensable.

The first capability is modeling: given the current state and an action you take, what change will this action bring to the state? Whether you predict the future in pixel space or latent space, you must be able to model the changes your actions bring to the environment.

The second capability is that after you model this change, you can use the modeling of the state to promote policy learning for a given task — learn a skill faster, and output a more reliable action.