HomeArticle

The top 10 trends I saw at WAIC 2026

量子位2026-07-22 11:55
Humans are the most valuable resource for the development of AI.

There were simply too many people.

Extreme heat and thunderstorms could not stop the crowds; scalper tickets were driven sky-high yet still impossible to secure; every exhibition hall and forum was packed to the brim, yet no one’s enthusiasm dimmed.

If you were on site at World Artificial Intelligence Conference (WAIC) 2026 in Shanghai, you would likely share this exact feeling.

We delved deep into the exhibition halls and forums to sift out the ten most pivotal core trends.

If you could not attend in person, or did not have time to catch every launch and exhibit, this concise and to-the-point summary will help you quickly grasp the key trends showcased at this year’s WAIC.

Ten Core Trends:

WAIC has become the China AI World Cup, gathering models, products, and partnerships under one roof

AI software and hardware have both matured: booths are no longer limited to one-way demos, with more to see, more to play with, greater interactivity, and significantly extended visitor dwell time

AI Chat has become commonplace; the new gold standard is capability for real work and deliverables — Show me the Agent!

Foundation models have entered the semi-final stage: six leading model players have emerged, each with distinct strengths

World models are red-hot, the top track under the physical AI wave

The embodied intelligence battle has spread upstream to the supply chain, as robot mass production enters a dividend period for data and components

Robots advance by leaps and bounds each year: high-productivity players do not fixate on humanoid forms, while more human-like designs lean heavily into emotional value

Domestic computing power has moved beyond the single-chip protagonist era, with supernodes, computing networks, and ecosystems now determining competitive outcomes

The "shovel sellers" of the gold rush have finally stepped into the spotlight, as embodied infrastructure sees a collective boom

400,000 attendees flooded WAIC, creating a globally unrivaled AI user ecosystem advantage

Trend 1: WAIC Becomes the China AI World Cup, Uniting Models, Products, and Partnerships

At this year’s WAIC, the exhibition’s "launch cycle" is continuously lengthening and deeply integrated with the core exhibition period.

In previous years, WAIC was largely seen as a centralized showcase for cutting-edge technologies; this year, booths are universally synchronized with product launches, strategic signings, and new feature rollouts. Leading AI companies are actively anchoring their core business timelines to the July event in Shanghai.

Around WAIC, Moonshot AI released Kimi K3 with 2.8 trillion parameters, SenseTime debuted the multimodal foundation model SenseNova-U1 Pro, while Stepfun showcased its full terminal ecosystem matrix including Step AOS, STEPX Neo, and "Super Eva"...

While each company’s product entry point differs, their core priorities share similarities: model iterations, the "first mile" of commercial implementation, and ecosystem partner support all converge to break out at the WAIC milestone.

This change stems from the evolution of the AI industry’s implementation logic.

Previously, large model updates relied mostly on one-way online releases, driving discussions through social media and live streams; today, for models to truly enter the industry, developers, hardware terminals, cloud service providers, and vertical scenario customers must collaborate in the same physical space.

The core value of the conference has shifted from pure technology display to facilitating tangible partnerships and cross-industry testing.

Therefore, when experiencing this year’s event, we should not only count new models — the real change is that China’s AI now has its own annual event calendar.

Whoever reveals their full hand here implicitly accepts the collective scrutiny of the on-site audience, the industry, and the market.

Trend 2: AI Software and Hardware Are Both Matured — Booths Bid Farewell to One-Way Demos, With More to See, Play, and Interact With, and Significantly Longer Dwell Times

If last year’s AI exhibition was most defined by "onlookers," this year, nearly all booths that retain visitors let audiences participate at any moment.

Whether it is interactive robot control, edge device experiences, or creative application trials, the old one-way technical demos are now broken down into small, specific tasks that users can operate and experience hands-on.

This trend is especially prominent in AI cultural and entertainment scenarios.

However, the cultural and entertainment production chain is long and tools are scattered — what form should AI capabilities take to land in practice?

Many booths on site tell us that hardware is an absolutely indispensable link.

Take first-time exhibitor Tantan Technology as an example: its booth unifies AI music, AI dubbing, and AI 3D into a continuous closed loop from inspiration to distribution.

It demonstrates a clear software-hardware integration approach: the self-developed multimodal music generation foundation model "Tianpule" extends into the conversational creation agent Tunee, finally landing as a physical carrier — the world’s first generative AI guitar, connecting the full "model-application-hardware" chain.

On-site experiences are built entirely around this system.

Users can generate songs and MV by chatting with Tunee, and make continuous edits during the conversation; pick up the AI guitar to generate exclusive music directly, or hum a melody that the model turns into a playable track in minutes, with the guitar’s "light-tracking" system lighting up in sync, and gesture changes triggering instant audio feedback.

Software lowers the barrier to creation, while hardware lets AI truly step out of the screen into daily life, forming sustainable usage habits.

From another perspective, while these booths may look like amusement parks, they reveal the most critical product proposition on the AI application side:

User retention directly depends on whether the system can deliver a sense of participation, achievement, and a takeawayable work in a very short time.

Across the venue, today, the competitive logic for AI ToC products has long moved beyond the initial stage of traffic grabbing.

At the same time, no matter how strong the underlying model capabilities are, if limited to parameter comparisons and one-way Q&A, they remain distant from ordinary people’s daily lives.

Both booths and the market are teaching AI products how to retain users. Whether a product can extend experience duration, drive social distribution, even support multilingual localization for overseas monetization, and build a habit of opening the app the next day, determines whether it can truly integrate into life scenarios.

Visitors are voting for products with their real patience and hands-on operations.

Trend 3: AI Chat Has Become Commonplace — Capability for Real Work and Deliverables Is the New Gold Standard, Show Me the Agent!

"How many parameters does this model have?"

At this year’s WAIC, this question is still asked, but it is clearly no longer the core focus of attention.

More "me-relevant," practical, and implementation-focused questions have become the top priority this year — for example, what exactly can it do for me? How long does it take to complete this task? Can the final deliverable be used directly?

In the past two years, large model companies have mostly focused on demonstrating parameter scales, training methods, and various benchmark scores.

But walking into Hall H1 this year, Agents have become the default topic that nearly 70% of exhibitors jointly showcase and discuss.

Baidu brought its general-purpose agent "Baidu Buddy" to the site, Tencent showcased a wide range of agents for mobility, office, and content scenarios, while Alibaba’s discussions centered on the Agentic era.

In addition, Kingsoft Office launched the Lingxi Professional Edition for individuals and WPS Comate for organizations, while Stepfun introduced agent-native operating systems to mobile and automotive scenarios.

Among the ten selected "Hallmark Exhibits," agent-related applications already occupy four spots.

Why?

The traditional chat box excels at answering questions, but real task execution requires calling tools, remembering context, continuously checking intermediate results, and finally completing the entire process smoothly.

The former demonstrates that the model "can talk," while the latter requires tangible proof that it "got the job done." It can be said that the ability to deliver effective outcomes is becoming the most practical, most substantial standard for measuring large model capabilities.

The traditional chat box will not disappear entirely, but gradually retreat to the position of a task entry point. Users only need to state their goal, and the system will independently break it down, call resources, execute the task, and finally return the finished product to the user.

"Show me your Agent" can be said to be the clearest unspoken line at this year’s WAIC site.

Trend 4: Foundation Models Enter the Semi-Finals — Six Leading Model Players Emerge, Each With Distinct Strengths

The six rising stars of large models are the first wave of influential enterprises in the AI 2.0 era.

By 2026, the landscape of core players remains, but each has carved its own path, with strategic directions clearly diverging.

MiniMax and Kimi continue to compete head-to-head in cutting-edge technology breakthroughs and open-source influence, forming the most concentrated competitive tier of domestic foundation models alongside Zhipu AI and DeepSeek.

They have all collectively entered the foundation model playoffs alongside internet giants ByteDance, Tencent, and Alibaba.

On the other side, the other rising stars are precisely defining their own business boundaries.

Stepfun bets on end-to-end terminal closed loops, integrating models, operating systems, and smart terminal devices into a unified narrative framework; Baichuan Intelligence enters the market with a medical-enhanced large model, deeply cultivating the vertical medical scenario; 01.AI shifts to enterprise-level decision intelligence, calling itself the "leopard" in the "five tigers and one leopard" lineup.

By the way, Dr. LI Kaifu, founder and CEO of 01.AI, is the only top leader among the six rising stars who stationed himself at the exhibition booth.

This proactive stance of top leaders showing up in person is just one side of the exhibition’s competitive landscape. But physical presence is not the only standard for measuring an enterprise’s industry influence.

The absence or low profile of Zhipu AI, DeepSeek, ByteDance, and others has not diminished audience attention to their industry impact.

Their names are frequently heard in various technical seminars across the exhibition.

Overall, the large model track has moved past the homogeneous construction phase, and the industry focus has shifted to more practical supply-demand matching — identifying which scenarios each enterprise is best suited to empower and which core business problems it can solve.

The competitive logic during this reshuffle period is especially strict —

Maintaining industry standing still requires solid foundation capabilities, but the ability to build unique business moats and offensive positions will directly determine tier differentiation in the next round of competition.

Trend 5: World Models Are Red-Hot, the Top Track Under the Physical AI Wave

As the LLM paradigm gradually converges, physical AI has officially taken the baton, becoming the next surging, chaotic frontier.

Among them, World Models have made frequent appearances in major forums and achievement releases, arguably becoming the next contested high ground after VLA — and the most promising brain architecture leading to physical AGI at this stage.

Suddenly, everyone is showing their cards:

RoboSense: AWE 3.5

Ant Group Lingbo: LingBot-VA 2.0

GigaVision: GigaWorld

……

At this year’s WAIC, the "Six Dragons of World Models" summit forum was the most highly anticipated roundtable, reflecting the real picture of China’s current world model research — a hundred schools of thought contending, with no unified conclusion.

Every industry is working on world models, and each company insists that theirs is the native world model. However, to date, no implemented result or demo has proven to be a clear game-winner, and no one can convince the others.

In LI Feifei’s words: this is one of the most important, yet most overused terms in today’s AI field.

At the WAIC site, three main camps have roughly formed, centered on the definition of "world" —

1. The Video Camp: The world is pixels

Following a renderer-heavy route, outputs are oriented to the human eye, with the core evaluation standard being visual fidelity. This line is relatively lively at WAIC — demos are intuitive, and draw large crowds of onlookers.

But it also faces the most intense criticism: critics argue that this is just riding the physical AI narrative, treating pixels as the next predicted token without touching the essence of the physical world.

2. The Abstraction Camp: The world is latent space

This route is more practical than the pixel camp, and is the absolute main force at WAIC, as well as the camp with the most "chaos," since each company defines latent space in completely different ways.

Another problem is that latent space is a black box, with no pixel output for intuitive evaluation. As a result, each company builds its own evaluation system, with no unified benchmarks.

WorldArena, RoboTwin, Libero... each company tops their own respective leaderboards, and everyone has claimed first place.

3. The Spatial Intelligence Camp: The world is 3D

Named by LI Feifei’s World Labs, this camp is most closely integrated with the transformation of AI applications. Numerous companies deeply engaged in 3D asset generation are accelerating their shift in this direction.

Who exactly is building the real world model?

QbitAI has tirelessly asked this question over and over, but no clear answer has emerged so far.

Perhaps the final answer will only come from real, on-scene battles that produce true champions.

Trend 6: The Embodied Intelligence Battle Spreads Upstream to the Supply Chain — Robot Mass Production En