Zhu Songchun: China's AI Industry Should Not Be Misled by the "Musk Belief"
The last time ZHU Songchun appeared at the World Artificial Intelligence Conference (WAIC) was back in 2019. For the following seven years, he was almost entirely absent from China's largest AI industry showcase.
That absence itself was a statement of stance.
Over the past few years, WAIC's center stage has belonged to large language models, the Scaling Law, and the belief in "miracles born of sheer scale." ZHU Songchun, President of the Beijing Institute for General Artificial Intelligence (BIGAI), has been one of the most consistent and persistent critics of this dominant narrative.
Ever since ChatGPT ignited the large model wave at the end of 2022, he has repeatedly emphasized that large language models cannot naturally lead to Artificial General Intelligence (AGI), and that simply scaling up parameters, data, and computing power will not automatically bring about AGI emergence. In his view, the capital-driven marketing and international geopolitical narrative behind this boom are far more powerful than genuine technological breakthroughs.
On July 19, 2026, ZHU Songchun made his return to WAIC.
During his one-hour speech at the "Thinkers Forum," he summarized the past several years of the AI boom as a capital and geopolitical narrative jointly promoted by the "Silicon Valley-Wall Street-Washington trio." China's technology community has long been trapped in problems defined by the United States: "Information flows from the U.S. to us, then we invite American experts to restate it, creating a cycle of reverse information import and amplification."
He even singled out a phenomenon he called the "Musk Faith": investors' blind adherence to grand tech leader narratives, noting that "80% of what Musk promises never gets delivered." Projects like the Hyperloop have been shut down, full self-driving has been delayed more than a dozen times, and brain-computer interfaces have fallen far short of expectations — yet the grander the promises and the slower the delivery, the more fanatical the market reaction becomes.
From June 2026 onward, global tech stocks have experienced dramatic volatility. U.S. tech giants keep expanding their AI capital expenditures, putting free cash flow under pressure; SpaceX's stock fell below its IPO price after going public; Meta has also begun exploring commercializing its surplus computing power externally. ZHU Songchun's judgments are partially coming to pass. "When the bubble is stretched to its limit and shows signs of starting to burst, people might be more willing to listen to what I have to say." That is part of the reason ZHU Songchun returned to WAIC this year.
Yet the mainstream consensus in the industry remains on a different track. OpenAI, Google DeepMind, and Anthropic are still doubling down on large models, arguing that there is far more untapped potential in architectural improvements, reasoning enhancements, and multimodal fusion. Leading domestic AI companies generally hold that large models are not the end point of AGI, but they are currently the most efficient technical tool and commercialization path — and the ecological advantages already established should not be easily abandoned. Both sides have their own evidence, and their own vested interests.
ZHU Songchun's core judgments have never changed over the past three years. In early 2024, he used the metaphor of "climbing Mount Everest vs. landing on the moon" to describe the distance between large models and general artificial intelligence. In his 2025 book *Giving Minds to Machines*, he characterized large models as "brains in vats": they can speak, but they do not understand the real world, lacking a genuine connection between "words" and "reality." By 2026, he reviewed four successive boom-and-bust cycles: the AI "Four Little Dragons" era in computer vision, the metaverse craze, the "hundred-model battle," and embodied intelligence. His conclusion is that the faith in computing power, the faith in big data, and the faith in large models are all being dismantled one by one by real-world outcomes.
Technically, ZHU Songchun has taken a path distinct from the Transformer paradigm: centered on causal reasoning and intrinsic value-driven mechanisms, to build general agents that can autonomously generate tasks, learn, and act.
Along this path, his team has proposed the CUV general intelligence framework. Here, C represents the agent's cognitive architecture, U covers perception, cognition, and action capabilities, and V stands for the intrinsic motivation and value system. ZHU Songchun has also divided AGI innovation into five layers: philosophy, mathematical and physical frameworks, models, algorithms, and implementation. In his view, most current industrial innovations are still concentrated on the algorithm and implementation layers, and truly original breakthroughs require tracing back further to the philosophical and fundamental theory layers.
Built around this system, the Beijing Institute for General Artificial Intelligence has developed the TongOS general artificial intelligence operating system and the TongPL programming language, and launched the general intelligent agent "Tongtong." Unlike large model products that are primarily driven by external instructions, "Tongtong" is designed as a prototype of a general agent driven by an intrinsic value system, capable of autonomously understanding the environment, generating tasks, and planning actions. According to the team's evaluation standards, some of its cognitive and behavioral capabilities have reached the level of a five- to six-year-old child.
In March 2026, "Tongtong" 3.0 made its debut at the Zhongguancun Forum, with upgrades in spatial intelligence, cognitive intelligence, and social intelligence. The concurrently released "Tong Brain" was positioned as the core embodied intelligence engine connecting general agents to physical robots.
In 2025, *Science* consecutively reported on ZHU Songchun's team's research around the general intelligent agent "Tongtong" and their roadmap from individual AGI agents to AGI societies, focusing on their efforts to extend general intelligence from individual cognition to multi-agent interaction, social simulation, and artificial civilization research.
At this WAIC appearance, ZHU Songchun formally brought "social intelligence" to the forefront: "understanding responsibilities, rights, and interests, inferring intentions, reading subtext, and participating in complex social collaboration." He argues this is the final barrier to AGI, and the dimension where all current large models are still handing in blank papers.
For the industry, however, the more practical question remains: before "Tongtong" completes a viable commercialization loop, does the path ZHU Songchun has chosen represent a more forward-looking technical judgment, or an academic bet yet to be proven?
During WAIC, Q had an in-depth conversation with ZHU Songchun. He further elaborated on his views around topics such as whether large models can lead to AGI, whether bubbles exist in the current AI boom, whether China truly lacks original technical paths, and how cognitive architectures can move toward industrial implementation.
01 Scorching Weather, Hottest Large Models, and Industrial Bubbles
Q: This year's WAIC is still buzzing with activity, with notable new progress in areas like robots and intelligent agents. What's your overall impression of the current AI industry?
ZHU Songchun: There are huge crowds and a lively atmosphere, and robotic technology has indeed made noticeable progress compared to previous years. Yet a fundamental question still hangs over the entire industry: can the technical path we are currently taking ultimately deliver the general artificial intelligence everyone talks about?
Back when the large model boom first started from late 2022 to early 2023, I pointed out that large language models cannot naturally lead to AGI. Over three years later, large models have improved in language generation, dialogue, and tool use, but there remains an essential gap between them and true general intelligence. The most prominent problem right now is equating incremental improvements in partial capabilities with the imminent arrival of AGI, which in turn creates overinflated technical expectations and capital valuations.
Q: Is the question of whether large models can lead to AGI a separate matter from whether large models can generate industrial value?
ZHU Songchun: Absolutely two different things. As a tool, large models can play a role in scenarios like chat, content generation, and code assistance. But having practical application value does not make a system AGI, nor does it prove that continuing to scale parameters, data, and computing power along the same path will necessarily achieve AGI.
Scenarios that have truly formed stable commercial closed loops are still limited. The "Agent" concept is very hot, and there are many products, but moving from demo to real production environments requires solving a whole series of issues around accuracy, long-term planning, result verification, cost, and accountability boundaries. A model that can chat and call a few tools is still very far from a general agent that can autonomously understand the world, set goals, and complete open-ended tasks.
Q: Why do you frame this round of AI boom within the context of international politics and capital narratives?
ZHU Songchun: Any major technological wave is never just a technical line — behind it lies a narrative co-constructed by politics, capital, and communication. Around 2014 to 2016, when DeepMind was acquired by Google, OpenAI was founded, and AlphaGo entered the public consciousness, the globalist narrative was also shifting toward the "America First" paradigm. Artificial intelligence, and AGI in particular, gradually became a key strategic anchor for the U.S. to reshape its technological competitive advantage.
This narrative combines policy needs, Silicon Valley companies' fundraising needs, and Wall Street's capital needs, drawing global capital, talent, and infrastructure investment to concentrate in the U.S. Even if the bubble bursts in the future, the U.S. will already have secured tangible benefits such as data centers, electricity, chips, talent, and tax revenues. So to understand AI, we cannot only look at model leaderboards — we also need to see who is defining the problems, setting the standards, and organizing the resources.
Q: Are chips and computing power the most critical bottleneck for China's AI development?
ZHU Songchun: Chips are certainly important, and China must build self-reliant and controllable software and hardware foundations. But we cannot reduce every problem to "not having enough GPUs." Computing power only generates value when combined with correct technical architectures, clear tasks, and real-world demands. Massive chip purchases and intelligent computing center construction can still lead to resource waste if there is a lack of effective utilization and industrial closed loops.
The Scaling Law portrays the continuous expansion of data, parameters, and computing power as a sufficient condition for reaching AGI, but I believe this logic has fundamental flaws. It's like a person who never builds a knowledge structure or thinking ability, but only hoards piles of exam papers — they will never develop true general competence. AI's real bottlenecks also include cognitive architectures and value systems.
Q: Do you think the AI industry has already formed a bubble? What is the essence of this bubble?
ZHU Songchun: I believe there is an obvious bubble right now. The essence of the bubble is prematurely and excessively discounting the future value of an immature technology into the present. Capital markets rely on grand narratives to sustain expectations, companies keep switching buzzwords to secure financing and valuations, while real technological maturity, user demand, and business revenues fail to keep pace.
From large models and Agents to embodied intelligence and world models, hot topics keep rotating — there is genuine progress, but also a huge amount of labeling and homogenization. Judging bubbles should not only look at stock price movements, but also check whether input and output match, whether the technology solves real problems, and whether the business model is sustainable. When training and operating costs keep rising, models need rapid iteration, but revenues cannot cover investments, this structure is extremely fragile.
Q: How do you expect this bubble to adjust? Can bubbles also have positive effects?
ZHU Songchun: Bubbles are not entirely without positive effects. High-density capital investment accelerates infrastructure construction, attracts talent into the field, and may drive breakthroughs in some key technologies. After the dot-com bubble burst, companies with real core technologies and viable business models survived — the AI industry may go through a similar process.
The problem is that the costs of bubbles are equally massive, including capital losses, talent misallocation, distorted public expectations, and damage to the long-term scientific research ecosystem. A huge number of teams chase short-term fads, investment institutions only back projects that can be quickly packaged and exited, while truly original research requiring 5 or 10 years of accumulation struggles to get support. Only when the bubble recedes can the industry have the opportunity to refocus on the essence of technology and real-world scenarios.
Figure: ZHU Songchun speaking at the Thinkers Forum
02 Stop Misusing AGI as a Marketing Buzzword
Q: There is no unified definition of AGI in the industry right now. How do you define true general artificial intelligence?
ZHU Songchun: AGI is not a marketing term that can be arbitrarily interpreted. Strictly speaking, general artificial intelligence should possess physical common sense and social cognition, be capable of autonomously generating tasks and completing infinite open-world tasks, and its behaviors should be driven by intrinsic values aligned with human values and social norms.
I believe there are at least three fundamental characteristics: first, infinite task generalization, breaking through predefined tasks and closed scenarios; second, autonomous task generation, being able to form goals, plan actions, and adjust based on feedback; third, value-driven, knowing why to act, what should be done, and what should not be done. A system that only answers questions based on external instructions, or performs pattern matching within existing data distributions, cannot be called true AGI.
Q: Why do you think large language models cannot cover all the capabilities required by AGI?
ZHU Songchun: Large language models primarily process statistical correlations in language sequences. Language ability is important, but AGI also encompasses many domains including vision, robotics, cognitive reasoning, multi-agent collaboration, social intelligence, and social minds. A system that can fluently generate text does not truly understand the physical world, causal relationships, other people's intentions, or social norms.
The current mainstream path tends to mistake external performance for internal intelligence: when a model scores high on certain exams, or completes a sequence of tool calls, people interpret it as "emergent" general intelligence. But without stable world representations, causal mechanisms, goal generation, and value judgment capabilities, these performances can hardly generalize consistently and reliably in open environments.
Q: Why do you consider "cognitive architecture" and "value" more important than model parameters?
ZHU Songchun: Architecture determines what kinds of capabilities an agent can develop. Different biological species have different cognitive architectures — no matter how much data and training you give them, they will never naturally gain capabilities beyond the boundaries of their architecture. Today's many large models focus heavily on parameters, data, and computing power, yet lack stable world models, causal mechanisms, goal systems, and value structures.
We proposed the CUV framework to uniformly describe agents through the cognitive architecture C, the capability or potential function U, and the value function V. The cognitive architecture organizes information and makes decisions; the capability system supports perception, language, movement, and learning; the value system drives goal selection and behavioral direction. True general intelligence requires the three to co-evolve. The idea of "giving minds to machines" means endowing agents with internal goals, value judgments, and structures for understanding others.
Q: You divided AGI innovation into five layers: philosophy, mathematical and physical frameworks, models, algorithms, and implementation. Why start from the philosophical layer?
ZHU Songchun: Most current industrial competition is concentrated on the model, algorithm, and implementation layers — for example, parameter scaling, training tricks, inference speed, and engineering deployment. These efforts are certainly important, but if the underlying assumptions about "what intelligence is" never change, all innovations are just optimizations within the existing paradigm.
The philosophical layer determines how we understand the subject, goals, and values of intelligence; the mathematical and physical framework layer translates these fundamental understandings into computable, verifiable forms; only then come models, algorithms, and engineering implementation. Truly original breakthroughs must be able to redefine problems, not just run faster on someone else's track.
Q: How can we prove that an original architecture is closer to AGI than the Transformer path?
ZHU Songchun: Originality cannot rely solely on conceptual claims — it must ultimately be tested through a set of open, systematic standards. AGI testing should not only focus on language Q&A or single benchmarks, but cover capabilities such as perception, reasoning, causal understanding,