HomeArticle

Office Agent has gained explosive popularity, and the model battle has entered the next stage.

极客公园2026-08-05 08:00
The battle of models is, after all, a battle over entry points and ecosystems.

The battle of models is ultimately a battle of entry points and ecosystems.

From June to early August this year, in less than 60 days, one thing happened almost simultaneously among the new BAT giants.

First in June, Alibaba took the lead in launching a major integration of internal Agent product lines and capabilities. On August 3, while officially announcing the Qwen3.8 model with significantly improved capabilities in Coding and professional office (Cowork) scenarios, it also formally merged the three internally incubated Agent products QoderWork, Wukong and MuleRun into a unified Agent product "Qianwen Workspace".

Shortly afterwards, on July 20, Tencent launched business integration, and the relevant business and team of QClaw were merged into the WorkBuddy system.

10 days later on July 30, ByteDance announced that the entire Feishu product team would be merged into Doubao, with Xie Xin, head of Feishu, reporting to Zhao Qi, head of Doubao.

With different organizational structures, different product accumulations and advantageous markets, the three top Chinese Internet enterprises made similar integration choices at nearly the same time. This is hardly a coincidence.

A more reasonable judgment is that this kind of integration is an inevitable choice against the background that the AI industry has gone through the exploration period and entered the stage of entry point competition: in the past, large tech companies needed multiple teams to verify different paths such as desktop operation, cloud execution and enterprise collaboration; now, the main product forms have gradually become clear, and the competition has also shifted from whether Agent can be developed to who can become the unified entry point for users and enterprises.

As the latest product launched in this round of convergence, Qianwen Workspace is also the most suitable sample to observe the changes in the industry.

As Alibaba's phased capability restructuring, it integrates the previous accumulation scattered in models, desktop terminals, cloud terminals and enterprise collaboration scenarios into a new independent product.

At the same time, it also embodies several most critical issues in the competition of enterprise Agents: why the internal horse race mechanism ends at this time, why enterprise IM and cloud have become the underlying assets of Agents, and how models, products and real tasks form a co-evolution closed loop.

Understanding it means understanding the real match point of enterprise-level Agents.

01 Why are large tech companies converging their paths at the same time

Why did large tech companies start integrating Agent products almost at the same time?

Take Alibaba as an example. At the beginning of this year, it was reasonable to maintain QoderWork, Wukong and MuleRun at the same time.

The three products represent three different judgments.

QoderWork runs on users' computers and is good at reading and processing local files; Wukong tries to access DingTalk's enterprise accounts, permissions and application system; MuleRun has run on the cloud from the very beginning, can perform tasks for a long time, and entered the overseas market earlier. As of May this year, it has served enterprises and users in 43 countries and regions, among which users paying more than 200 US dollars per month account for 34%.

When the market is not yet finalized, three teams testing local execution, enterprise collaboration and cloud operation respectively can help Alibaba find the direction faster.

The problem is that the trial-and-error period will not last forever. In the first half of this year, the industry's understanding of general Agents has undergone obvious changes.

Anthropic found that many non-technical employees would bypass the ordinary chat interface and directly use Claude Code to organize files, process tables and complete multi-step knowledge work. Therefore, Cowork was developed later: it retains the task execution capability of Claude Code, but adopts an interaction mode more suitable for ordinary knowledge workers.

A similar change also took place at its old rival OpenAI. At OpenAI, some users handle programming problems in ChatGPT, while other users complete non-programming tasks such as report, image and data processing in Codex. The originally clear product boundaries between chat, programming and office began to be actively broken by users. So in late July, OpenAI merged Codex, which had operated independently for nearly a year, into the ChatGPT desktop client, and so far the three ends of Chat, Work and Codex have been integrated.

As for this move of Qianwen Workspace, it is also a recombination of capabilities verified by Alibaba's three experiments: the desktop end inherits the capabilities of local file and computer operation, the cloud end undertakes long tasks, resource scheduling and multimodal generation, and the enterprise end continues to access organizational identities, permissions and business systems. Concentrate the originally scattered product capabilities and feedback data into one entry point, one set of engineering system and one iteration closed loop.

Actual measurement of Qianwen Workspace shows that in addition to quickly producing professional reports according to requirements, it can also directly connect to DingTalk and send reports to the chat interface

Users are voting with their feet spontaneously, and the market has not left much time for large tech companies to continue scattered experiments.

The second-quarter report released by Analysys shows that in June 2026, the total visits of 17 mainstream desktop AI office agents have exceeded 60 million times. The total visits of Tencent, ByteDance and Alibaba's products are about 56.22 million, leaving less than 5 million market share for other players.

The technical route is gradually determined, and the head effect has also begun to appear. At this time, decentralization will become a burden on resource allocation.

The same is true for the engineering layer. In the past, every time the model was upgraded, the three teams had to re-adapt to task planning, context management, tool invocation, failure retry and permission control, which was essentially unnecessary repetitive labor.

"Qianwen Workspace" is the product born under this background of every second counts.

02 The watershed of enterprise Agents

Most of the general office Agents on the market currently still solve personal productivity problems first. They can read files in users' computers, analyze personal documents, understand historical conversations, organize personal schedules, and call limited external tools according to the permissions owned by users themselves.

But these are mainly personal contexts. The superposition of personal contexts is not equal to the enterprise context. The sum of the efficiency of all employees in an enterprise is not equal to the enterprise efficiency either. Equipping each employee with a smarter personal assistant can improve local output, but may not necessarily improve the overall information flow, task collaboration, and the efficiency of converting personal results into organizational actions and outcomes.

Enterprises are faced with another type of problems.

An enterprise is not a simple collection of individual accounts. The sum of all employees' files, conversations and tasks will not automatically become the operation mode of the enterprise. For the same file, a personal Agent may only regard it as a material to be summarized; for an enterprise Agent, it also involves who can read it, who needs to confirm it, which departments the conclusion should be synchronized to, and who must approve the follow-up tasks.

This is the most fundamental difference between enterprise-level Agents and personal office Agents. It is also a cognitive basis for the obvious difference in the behavior logic of Chinese and American large tech companies under the similar integration background.

OpenAI and Anthropic are mainly merging the entry points of chat, programming and desktop execution. The scope of integration of Chinese large tech companies includes enterprise collaboration software such as DingTalk and Feishu, and at the same time makes the cloud part of the underlying capabilities of Agents.

Reflected in the internal product collaboration of enterprises, while the Doubao product line integrates Feishu, it also deeply combines with Volcano Engine. Qianwen Workspace merged Wukong, which previously focused on the DingTalk scenario, and also inherited Alibaba Cloud's all-round resources in the enterprise-level service market in the past, including industry cognition, data management and resource scheduling.

In other words, Agents are starting to absorb IM and cloud, turning them into the underlying infra.

The importance of cloud as the core infra of Agents goes without saying. Here we mainly look at the integration narrative of IM.

Traditional office software divides the world by functions: email is responsible for communication, documents for writing, spreadsheets for data... But in the real world, when users want to achieve a goal, what they need is the combination of multiple tool capabilities, and the same is true for Agents.

That is also why, over the past decade or more, DingTalk-style IM, which integrates communication, employee identities, organizational relationships, collaboration networks, files, meetings, approvals, permissions and connections to various enterprise applications, as the software with the most functions and complete organizational context, will inevitably become the most underlying capability support in the Agent era.

I simulated a three-person topic discussion in my self-built DingTalk organization group, and asked Qianwen Workspace to summarize the key points of the discussion and extract to-do items. It not only identified five to-do items from the conversation, assigned the responsible person to each one correctly, marked the urgent one with "high" priority and "upcoming deadline", and directly wrote them into the DingTalk to-do system.

Of course, for an external third-party Agent, it is not difficult to read group messages through a connector, then generate a summary, or complete investment research, web page production or spreadsheet processing. The real difficulty is writing back: creating to-do items, adjusting schedules, sending emails, submitting approvals, or assigning tasks to the right people.

Reading only requires interfaces, while writing back requires higher trust. Enterprises must determine what the Agent can see, who it can operate on behalf of, which steps need to be approved, and how to trace and revoke after an error occurs. From this perspective, over the past decade, what DingTalk has accumulated for Agents is not only tools, but also the moat of trust and enterprise-level context.

03 Models and Agents become an intelligent flywheel

If we further focus on the first-tier players in China, it is not difficult to find that Alibaba has another rare condition in this competition. It has first-tier large models, enterprise collaboration platforms and cloud computing infrastructure at the same time.

However, having both models and applications in hand will not automatically translate into advantages. This combination is truly valuable only when models and Agents can provide feedback to each other.

Therefore, we see that on the same day that Qianwen Workspace was launched, Alibaba quietly released Qwen3.8 with greatly improved capabilities in Coding and professional office (Cowork) scenarios.

The logic behind this is that in the production environment, Agents can provide models with more valuable feedback than chat records, and models can solve the underlying problems in Agents once and for all.

For example, many problems in many Agent products that seem to occur at the application layer actually have their roots at the model layer. In the past, if the model selected the wrong tool, the product team could add an invocation rule to solve it; if the model forgot the initial requirements in a long task, the team could add an extra workflow outside; these patches on the application can objectively reduce the error rate, but cannot improve the model's own judgment and cognitive ability.

Therefore, we can see that when OpenAI launched Codex, it did not just connect the general model to the code editor, but trained codex-1 for software engineering tasks. It uses real programming tasks for reinforcement learning, can read and modify files, run tests, and continue to adjust according to the test results until it gets a passing result. Its appearance has solved a large number of problems for developers in programming tasks at one stroke, and has also become the key for OpenAI to achieve a counterattack in the programming field.

The model solves the underlying capability problem of Agents, and Agents can also bring more complete feedback to the model than chat records.

For OpenAI and the Qwen series of models, the complete action trajectory left by Agents can tell the model how it understands the goal, what plan it has made, which tools it has invoked, at which step it deviates from the requirements, what the user has modified, and whether the final product is really used. Especially in code tasks, whether the code can run, whether the test passes, and whether the file is modified correctly can usually be verified by machines. And the failure trajectories in these real enterprise tasks can be fed back to the model and become the driving force for model iteration.

Finally, Agents provide real task feedback for models, and models provide stronger judgment capabilities for Agents. The combination of model + Agent forms a flywheel system that keeps running and evolving.

04 Conclusion

Of course, objectively speaking, the completion of this integration of Qianwen Workspace does not mean that Alibaba's exploration of Agents can rest easy from now on. But it at least shows that Alibaba has found its own main line.

When more and more Agents can make PPTs, analyze spreadsheets and control browsers, the functional differences between various To C-style Agents will shrink rapidly. The real gap in the future will come from the system capabilities behind the products: how much enterprise context it masters, how much execution authority it can obtain, and whether it can turn every real Agent task into the basis for the next model upgrade.

In this process, the context of IM, the underlying capability construction of infra, and the flywheel of model and Agent are all indispensable.

Qianwen Workspace has already got half of the admission ticket.

This article is from the WeChat official account "GeekPark" (ID: geekpark), written by Cynthia, edited by Jing Yu, and published with authorization from 36Kr.