HomeArticle

Actual Measurement of "WeChat Micro": When Tencent Starts to Unleash the AI Potential of Its National App

极客公园2026-06-24 07:10
Is This Time Really Different?

WeChat AI has finally arrived.

In the past two years, the AI industry has been seeking an answer to a question: Where is the entry point for the next - generation Agent? An independent app? A browser? Or an operating system?

WeChat has provided another answer. A few days ago, WeChat's native AI assistant, "Xiaowei," started a small - scale gray - scale test. In the latest version of WeChat (8.0.75), users can find the "Xiaowei" entry in the upper - left corner of the main interface. Clicking on it will lead to the AI assistant. Meanwhile, swiping the main interface to the right can also quickly evoke the interaction with "Xiaowei."

It can be seen that Xiaowei doesn't appear as a new app, nor does it attempt to become a new chat entry. Instead, it's like a layer of AI capabilities embedded within WeChat, appearing in multiple scenarios such as chat boxes, official account and video account content, Moments, collections, and mini - programs, waiting to be called by users when needed.

Xiaowei is connected to the relationships, content, and services that WeChat has accumulated over the past decade. However, it doesn't rush to complete all operations for users. In terms of product form, it's closer to the AI middle - layer that WeChat is trying to build.

Behind this choice is also WeChat's exploration of what role an Agent should play in a super - platform.

Is Xiaowei not a new entry but WeChat's "AI Middle - layer"?

After entering Xiaowei, the most intuitive feeling is that it doesn't change the original way of using WeChat. Instead, it adds a layer of understanding ability to the existing scenarios.

Currently, Xiaowei's entry points are scattered in multiple locations in WeChat. The main entry can be accessed in the upper - left corner of the home page. It can also be called in the "+" menu of the chat window. There are also corresponding entry points on the content pages of official accounts and video accounts. Different entry points correspond to different capabilities.

In the chat scenario, Xiaowei can help users summarize chat content, extract key points based on the group chat context, and even assist in generating replies.

For the content of official accounts or video accounts, users can directly ask questions about articles and videos:

In the Moments scenario, Xiaowei can help users quickly understand the recent dynamics of their friends:

For the files in the collection, it can also assist in organizing information. However, this organization is still in its infancy. Currently, it can only send information to Xiaowei, and Xiaowei will automatically extract key information and organize it into notes and save them in WeChat's collection. It can't achieve the ideal goal of helping to organize the entire collection folder.

These functions are not unfamiliar when viewed individually, but they all point to a change: AI is no longer a new tool that users need to actively open. Instead, it has begun to become a basic ability within WeChat.

However, Xiaowei's design is also very restrained. It has a strong information - understanding ability, but its execution permissions are not fully open.

Xiaowei can read information within the scope of user authorization, help summarize chats and analyze content. However, when it comes to social and asset - related operations such as sending messages and sending red envelopes, it will require user confirmation. Even when helping users enter mini - programs and recommend services, it mainly focuses on "finding for you, organizing for you, and guiding you there" rather than directly completing transactions for you. The limited function of organizing the collection folder is also a design consideration for privacy protection.

WeLM: What WeChat needs is not the largest model but the model that best understands the scenarios

Behind Xiaowei is WeLM, a large - scale model independently developed by WeChat. It doesn't use Tencent Group's Hunyuan large - scale model. Instead, it has brought an internal model that has been iterated quietly for many years to the forefront.

WeLM was actually released in 2022, around the same time when ChatGPT became popular. After that, it has had little external publicity, and was even once considered by the outside world as a shelved project.

But in fact, it has been iterating quietly. Now, it's the WeLM - V4 version that supports Xiaowei. In the past few years, the large - scale model industry has been competing in terms of parameters and general - ability rankings. WeLM has never joined this fray. Perhaps from the day it was born, its goal was not to be the "strongest general - purpose model" but the "special - purpose model most suitable for the WeChat ecosystem."

Behind this positioning may be the inevitable real - world constraints of a super - app: With a user base in the billions, it can't use the ideas of general large - scale models for implementation.

For independent AI products, a higher inference cost and a longer delay are just issues of experience quality. But for WeChat, every group - chat summary, every content Q&A, and every service scheduling involves hundreds of millions of daily calls. A slight deviation in cost and delay will result in astronomical expenses and unacceptable experience degradation after full - scale implementation.

Therefore, WeLM has chosen a highly sparse MoE route in terms of architecture. By activating only part of the parameters in a single inference, it controls the amount of computation. Coupled with a series of structural optimizations such as KV - Mirror, grouped query attention, and multi - Token prediction, its core goal is only one: to keep the single - round inference cost within a range that can serve one billion users while ensuring the ability meets the standard.

Besides cost, another hard constraint is response speed.

For Agents in professional scenarios, such as writing code and making plans, users can tolerate a "thinking time" of dozens of seconds or even minutes. But the AI interactions in WeChat are all fragmented and immediate needs - asking about the key points of a group chat, checking a chat record, summarizing an article. Users only have a few hundred milliseconds of patience.

In actual tests, Xiaowei can almost respond instantly. Behind this is the special design of WeLM for low - latency, such as the Hidden Decoding technology: All complex inference processes are completed within the model, and the long - winded generation process is not shown to the user, taking into account both inference quality and interaction efficiency, so that users don't have to pay for "AI's thinking."

And the most unique advantage of the WeChat ecosystem has in turn become a natural strength of WeLM: long - context.

For general large - scale models to obtain users' personal context, users need to actively upload files, authorize data, and connect to third - party tools. But WeChat itself is a container for continuously precipitating context - chat records, social dynamics, collection content, browsing history, and service consumption records. All personal data naturally exists in this ecosystem.

WeLM expanded the context window to the 128K level during the pre - training phase. It's not to boost the ranking data but to natively adapt to the multi - dimensional and continuously accumulated personal data environment in WeChat. This is why, when it comes to content summarization and information organization, Xiaowei's experience is more in line with users' real context than general AI.

Of course, WeChat doesn't follow a completely self - developed and closed - off route. Currently, Xiaowei adopts a multi - model strategy of "self - development as the mainstay, external supplementation": WeLM takes the lead in scenario scheduling and daily interactions within the ecosystem to ensure cost and efficiency. When it comes to long - tail needs such as complex inferences and professional knowledge, external models such as DeepSeek are called to provide support. This combination not only ensures the controllability of the core experience but also doesn't have to tackle all technical dead - ends. It's a typical pragmatic approach.

Ultimately, WeLM's approach is also a microcosm of WeChat's underlying logic for creating an Agent: It doesn't pursue the ceiling of parameter scale, nor does it compete for rankings in general - ability. All technical choices revolve around the four words "ecosystem adaptation."

After Xiaowei is fully launched, WeLM will most likely become one of the large - scale models with the highest call volume in China. It may not be the "smartest," but it must be the one that best understands the WeChat scenarios and is closest to ordinary people's daily digital lives.

Should the real national - level Agent appear in a super - app?

One of the most imaginative functions in Xiaowei may be mini - programs or small tools. With just a natural - language description, a fully - formed lightweight application can be generated.

I gave Xiaowei an instruction to "create a fragmented note." It almost immediately generated a fragmented note mini - program covering work, inspiration, and life. After creation, it can be continuously used in the small tools section of the settings.

For ordinary users, this means that the process of creating a mini - program, which originally required learning code, designing pages, and developing and launching, has been compressed into a natural - language conversation.

Of course, the current small tools are still in the early stage. The generated tools can only be used by the creator and cannot be shared or distributed. They also don't support complex capabilities such as payment and multi - user collaboration. They are more like personal efficiency tools.

But the signal it sends is very important. In the past, the supply of mini - programs mainly came from developers: someone proposed a demand, a team was responsible for development, and then it was distributed to users through the platform.

And AI is changing this chain. In the future, an ordinary user may also generate a small tool that meets their own scenario by describing the demand. The supply of the WeChat ecosystem may no longer only come from professional developers but from the joint creation of users and AI.

But small tools are just the beginning. WeChat's real advantage is not just allowing users to generate tools. It already has a context that is difficult for other products to replicate.

WeChat is no longer just a chat tool. Users' social relationships, work communications, content consumption, payments, and local - life services are all continuously precipitated in this ecosystem.

Therefore, when AI starts to understand and call this information, it changes not only the way of answering questions but also re - activates the data accumulated in the past. For example, summarizing Moments may seem like a simple function, but it corresponds to a real demand: users don't have time to browse all the dynamics but hope to know about their friends' recent situations. For some industry users, Moments are even an important channel for obtaining industry information.

The value of AI is to reorganize these scattered information. This is also the biggest difference between Xiaowei and other AI assistants.

Currently, the capabilities provided by Xiaowei don't completely go beyond the existing scope of the industry. Products such as Meituan, Xiaohongshu, and Douyin have also added AI assistant capabilities.

But the uniqueness of WeChat lies in its long - term accumulated personal context. It is naturally closer to a person's digital life. In the future, Xiaowei may further evolve into a personal assistant.

Users don't need to remember where a certain function is or what a certain mini - program is called. They only need to express their needs, and AI can understand the intention and call the underlying services to complete the task.

This is actually a change from GUI to LUI. In the past, WeChat continuously added functions, but the entry points that users could see and use were always limited. Many services are not without value, but are hidden in the complex product structure.

AI has become this new scheduling system. In the past, it was "users looking for services." Users needed to search, click, and learn about entry points. In the future, it may become "AI looking for services." Users only need to describe their needs.

For example, instead of actively opening the food - delivery mini - program, searching for products, and filling in information, users can tell AI to "arrange a dinner suitable for three people." Instead of rummaging through chat records, users can let AI help organize information about a certain project.

This is also the biggest imaginative space for WeChat to create an Agent. In the past few years, the competition among AI products has always revolved around a question: Who can become the next super - entry.

But WeChat's answer may be different. It doesn't create a new entry but turns an existing super - app into a smart entry.

Of course, this path is not easy. When AI starts to understand user relationships, call services, and participate in transactions, issues such as privacy boundaries, user authorization, and the interest balance between the platform and ecosystem partners will become new challenges.

But if the end - goal of an Agent is to become everyone's digital assistant, then the most likely place for it to grow may not be a new AI app but a super - app like WeChat that already carries people's digital lives.

*Source of the header image: AI - generated

This article is from the WeChat official account "GeekPark" (ID: geekpark), written by Lian Ran, and is published by 36Kr with authorization.