HomeArticle

Doubao, the terminator of AI model horse racing, has started to enable agents to "work independently"

字母榜2026-08-25 15:57
Mixing chat scenarios and work scenarios together will hinder the self-evolution of the model.

In the office Agent track, Alibaba, Tencent and Baidu have already engaged in fierce head-to-head competition, and now ByteDance finally can't hold back anymore.

According to reports, ByteDance has fully integrated the TRAE and Coze teams into the Doubao system, while separating the "work" function from Doubao independently, and will launch the independent AI office product "Doubao Work" as soon as this week. The report says the specific launch date is set on August 29.

In fact, this adjustment has long been traceable.

One month ago, on July 30, the Feishu product team was just merged into Doubao. Xie Xin, head of Feishu, reports to Zhao Qi, head of Doubao, and Feishu was downgraded from one of ByteDance's six independent BUs to a capability module behind Doubao.

But this time, ByteDance is taking a much bigger step.

Doubao has carried out several updates in the past week, including remote control of computers via mobile phones, Windows virtual desktops, skill stores, connectors, and work partners, launching more than 200 skills at one go. Now it seems that this is the general mobilization before the war.

That said, why does the work mode of Doubao have to be split out independently?

There are many reasons. On the product side, the current Doubao is too bloated: from generating images and videos to making phone calls, cloud disk services, and built-in browsers... The originally simple UI design is no longer concise, which dilutes the core attributes pursued in the "office" scenario.

On the other hand, as early as June, the official blog post of Seed 2.1 revealed that Seed for Seed is the development direction of the development team.

At the current stage, only by splitting Doubao Work out and combining it with TRAE and Coze can this technical direction be realized.

1

When you open Doubao, the first thing you see is always the familiar main dialogue interface. To access the "Work" section, you have to find the entry from the sidebar. In Doubao's product architecture, "office" has always been an accessory, not the main scenario.

From the users' perspective, Doubao is first and foremost a chat AI, and secondly a tool for getting things done.

Although Doubao is highly user-friendly, it still feels disconnected when used in office scenarios. What office users want is to "open the app and start working directly", not "chat for a few words first, then manually switch to work mode".

Even if the work mode is upgraded to a first-level entry, it cannot avoid another more fundamental problem: Doubao is too "fat".

This app is packed with functions such as dialogue, PPT generation, video generation, image design, data analysis... The function density is extremely high. When an office user wants to do a proper job, he has to rummage through a whole universal toolbox to find the screwdriver he needs.

What office users want is a focused workflow, not a general store stuffed with everything.

Competitors have long seen this clearly.

Tencent's WorkBuddy has been an independent product and independent brand from the first day it was launched. Users know clearly that they come to handle office work as soon as they open it. It launched the desktop version on March 9, and its monthly PC visits reached 8.85 million in the same month, exceeding 20.97 million by June, a 2.4-fold increase in three months.

Alibaba's Qianwen Work was also launched for public beta on August 30, with the local desktop client released first, followed by the web version and the DingTalk built-in version later.

2026 is called "the first year of Office Agent" by the industry, and the three major manufacturers have already engaged in fierce competition at the entry level. When all competitors have shown their independent weapons, ByteDance cannot leave its main functional equipment hidden in a chat App.

After Feishu was merged into Doubao on July 30, it gave Doubao Work more confidence. Now when you open Feishu, Doubao is placed in the most prominent position, and users are willing to use Doubao in Feishu, since Feishu is originally oriented to office scenarios.

The original "Aily Smart Partner" of Feishu has also been replaced by Doubao.

In the past, Doubao's office mode and chat mode shared the same dialogue window, and functions such as video generation and image generation were all crowded here. But now, when Doubao enters the work mode, only projects, risk operation confirmation, skills and connectors are left.

On August 17, the function of mobile remote control of computers was launched, allowing users to remotely access files via mobile phones. It does not matter if files are stored on two separate computers in the office and at home, users can retrieve them respectively and then merge them.

On August 18, the Windows virtual desktop function was launched subsequently. The AI runs tasks in an isolated virtual desktop, and users can use their own computers as usual without mutual interference. Data in the local mode is stored on the local machine, and the cloud computer mode keeps persistent online, so tasks can still run even if the computer is shut down.

On August 20, the sidebar workbench was launched, with a row of multi-tabs on the right side of the dialogue page, which can open local files, Feishu documents, online web pages and codes. AI dialogue and practical operation are carried out on the same screen, and the content generated in the document can be written back directly, eliminating the need for copying and pasting. The official description of this function is "real-time synchronization and retention".

On August 21, the three sets of functions: skill store, connector and work partner were all launched at one go. More than 200 skills and connectors have been launched on the platform, and users can also precipitate commonly used work steps, delivery standards and reference templates into custom skills for repeated reuse.

Within a week, there was almost one update per day.

As mentioned at the beginning, this week's high-frequency launch is not so much a product iteration as a general mobilization before the war.

All the keys are ready. The door is about to open.

2

The confidence for Doubao Work to be split out independently does not come from the product, but from the model.

On June 23, at the Volcano Engine 2026 Summer FORCE Conference, the Doubao large model 2.1 series (Doubao-Seed-2.1) was officially released, divided into two versions: Pro and Turbo. It was fully launched on Doubao and TRAE on the same day, and the API was simultaneously launched on Volcano Ark.

The official positioning of it is "a brand-new agent oriented to real productivity scenarios". The most critical change lies in the evaluation criteria: starting from version 2.1, it no longer focuses on the common static benchmarks in the past, but targets the "performance in actual workflows".

Tan Dai, President of Volcano Engine, said that only when the model capability crosses the "qualitative change point" can it truly meet the needs of production scenarios.

Back to the capability itself, let's look at the measurable indicators first. In the Agent direction, on GDPval which measures the economic value of real work tasks, Seed 2.1 Pro scored 87.9 points. For comparison, GPT-5.5 scored 84.9 points, and Claude Opus 4.7 scored 82.7 points; it scored 83.8 points on MCP-Atlas which inspects MCP tool calls, also higher than the latter two.

In the mobile terminal scenario, it scored 73.1 points on MobileWorld for mobile GUI tasks, which is the highest score among the participating models; on Agents' Last Exam, its full pass rate was 19.5% and the comprehensive score was 41.9%, higher than 40.5% of Claude Opus 4.7 but lower than 47.9% of GPT-5.5; it scored 0.788 on OSWorld, ranking first among the 20 models included in llm-stats, and it also reduced the average number of steps to complete tasks by 16% through reinforcement learning.

In the programming direction, it scored 71.0 points on Terminal Bench 2.1, close to 71.7 points of Claude Opus 4.7; it scored 59.8 points on SciCode scientific code, exceeding 58.4 points of GPT-5.5; it scored 47.0 points on NL2Repo-Bench repository-level code generation, also higher than 45.1 points of GPT-5.5.

In terms of multi-modal understanding, it scored 86.4 points on CharXiv-RQ for complex document understanding, 82.7 points on MMMU-Pro, and 89.2 points on VideoMME for video understanding.

But here comes the problem: if all these capabilities are buried in a chat App packed with all kinds of functions, it is actually a kind of damage to the model.

In the technical stack of Agent, there is a concept called Intent, which is roughly the same as the intent in public cognition, and it is the anchor point of the entire execution chain.

When a user says a sentence, the first thing the model does is not to rush to reply, but to determine the intent: "What exactly do you want me to do". Only after determining the intent, can it be split into sub-tasks, select tools, execute step by step, and finally verify and deliver. The entire operation of the Agent is centered on the intent.

If the intent is wrong, every subsequent step will be in vain; if the intent is vague, the model has to guess repeatedly, wasting computing power on "guessing what you want to do", while the computing power for each task is limited.

The main Doubao App, just stretches the intent space too much.

Behind one input box, there are functions of dialogue, video generation, image generation, phone call, cloud disk, browser and so on.

When the model receives a sentence "Help me make this into a PPT", it has to exclude a lot of intents first: is it generation? Is it editing? Or is the user just chatting about PPT?

Every guess will consume the reasoning budget that should have been left for "getting work done". The more bloated the functions are, the harder it is to determine the intent. The more energy the model spends on "guessing", the less computing power will be left for "doing".

What's more troublesome is that chat does not retain intent.

Traditional chat follows the question-and-answer mode, each round is independent, and the context scrolls with the window. The model does not need to take continuous responsibility for the same thing. However, office work is a long-term task. For example, to write a report, users need to consult materials, read documents, modify charts, and iterate repeatedly. The intent must run through the whole process, and the model must remember "why I am here" at every step.

This relies on task state persistence, which means to wrap the entire process in an independent execution environment. But chat App is inherently designed with the rule of "finish replying and then forget", and the intent cannot survive more than one round.

Looking further at the underlying level, it is a problem of alignment direction. The direction in which the model's capability grows depends on what scenarios it is evaluated and trained in.

If the model keeps answering questions in the chat App every day, its capability will develop towards "being good at chatting". Only when the evaluation criteria becomes "Is the work done?" and "Is the work done well?", will it develop towards being a "capable Agent that can get work done", which is why all the benchmarks released with Seed 2.1 in June are related to Office Agent.

Therefore, splitting the work function out is inevitable at the technical level. The product shell locks the intent for the model first: when the user enters the "work mode", the intent is "getting work done" from the very beginning, and the model no longer needs to guess. The independent execution environment is responsible for supporting the intent all the way to the end of the task. Evaluation and training can finally focus on the real workflow.

The ultimate direction of all this is the sentence in the official blog of Seed 2.1: Seed for Seed.

The core logic can be summed up in one sentence: let the Seed model participate in the R&D process of the Seed model itself as an Agent, that is, self-evolution.

The traditional large model R&D process is "people discover problems → people write code → people run experiments → people adjust parameters", and the whole cycle is very long. Seed for Seed allows the model to participate in diagnosis, data synthesis, framework optimization and experimental verification, and many links can be run in parallel and automatically, which shortens the R&D cycle.

The R&D process of Seed for Seed itself is an extremely complex Office Agent scenario, which is essentially the same as the scenario where users use Doubao Work to make research reports, write plans, and process spreadsheets, except that the complexity is higher.

In other words, ByteDance has launched the large model version of "Doubao Work" in its internal most difficult knowledge work scenario. When ByteDance productizes this set of capabilities and delivers it to ordinary users, the current "Doubao Work" comes into being.

Office Agent is not a product developed from scratch, it is the export of the Seed for Seed technical capability for downward compatibility.

Conversely, the massive real user task data and feedback accumulated by Office Agent can be used as training and verification materials for Seed for Seed.

This forms a two-way cycle: the internal R&D scenario explores upward, and the external office scenario lands downward, both of which share the same set of Agent base.

3

All actions at the product and technical level must finally be implemented on people.

At the all-staff mid-year meeting in August, Liang Rubo summarized ByteDance's business strategic principles into 11 characters: "High priority, thick main trunk, optimize for the long term".

This is ByteDance's second all-staff meeting within the year, breaking the previous convention of holding only one all-staff meeting a year. Only three core businesses are retained: AI, information platform, and transaction services.

"Thick main trunk" means to cut off all the trivial branches and pour resources into one core channel. Douyin is the existing main trunk, and Doubao is to be cultivated as the next main trunk.

The integration of TRAE and Coze into Doubao, and the componentization of Feishu are all moves under this logic. ByteDance no longer wants to support a bunch of AI teams exploring separately. It wants a unified AI entry, a unified office brand, and a unified team.

Standing at the center of this integration is Zhao Qi.

After TRAE and Coze are merged, the product and operation teams report to Zhao Qi uniformly. In less than a year, Zhao Qi has grown from the product head of Doubao to the overall head of ByteDance's entire AI office product line (Doubao + Feishu + TRAE + Coze).

TRAE is originally an AI programming IDE, and Coze is a low-code Agent platform, which have been exploring in the Agent direction respectively.

After the adjustment, the departments have transformed from "technical exploration" type to "product implementation" type. This is exactly the organizational meaning of "thick main trunk": reduce the middle layer, and let resources be directly poured into one core channel.

TRAE is split into two parts: the general office part of TRAE Work is integrated into Doubao Work, and TRAE IDE and CLI are retained as the programming product line under the Doubao brand.

The Agent platform capability of Coze is absorbed by Doubao, and the skill store and connector system are directly reused, but the prospect of Coze as an independent brand has become unclear. ByteDance's official statement says "the rights and interests of existing users will not be affected".

However, this is a double-edged sword.

When TRAE and Coze were under the engineering architecture department, they were essentially a dual-track experiment. Each of them ran its own line, and the assessment focused on technical indicators, developer ecosystem and technical influence, with little intersection between each other.

Now that they are all under the management of the product head, it means that the conclusion of the experiment has been reached. ByteDance no longer needs two technical teams to prove their respective directions, but to pour all capabilities into one product to compete head-on with WorkBuddy and Qianwen Work for the desktop market.

Technical ideals give way to product KPIs, and some people may not adapt to this change; but on the other hand, the capabilities that have been polished for two years finally have a channel to reach users on a large scale, and there is no need to consume energy in internal competition any more.

This article is from the WeChat official account