HomeArticle

After spending 41 minutes, what kind of Agent future is Qianwen Office betting on?

奇点研究社2026-08-06 11:23
What you buy is not Token, but the AI's ability to complete tasks.

When Token becomes the new business language of the AI era, will enterprises actively push Agents to become increasingly "heavy"?

What can you do in 41 minutes? Finish a meeting, drink a cup of coffee, or wait for Qwen Office to deliver an industry weekly report.

On August 3, Qwen3.8-Max and Qwen Office were unveiled together. We got hands-on with them immediately, and let three AI office products, namely Qwen Office, Workbuddy and TRAE Work, run the same research task. The results showed that Qwen Office took 41 minutes, TRAE Work took 16 minutes, and WorkBuddy took only 7 minutes.

Judging only by speed, this test has no suspense at all, and Qwen Office lost completely.

But when we spread the three reports out side by side, things started to get interesting. Behind the three time-consuming figures lie three completely different Agent routes, and a more thought-provoking question: When Token becomes the new business language of the AI era, will companies selling Token have the incentive to make Agents increasingly "heavy"?

One Weekly Report Reveals Three Agent Routes

The task of this test is not complicated: generate a weekly report on the embodied intelligence industry, with a statistical cycle of one week, covering financing events, enterprise dynamics, technological progress and other contents.

Apparently, this is a task very suitable for Agents to play their strengths, because it is not just simple information search, but also includes finding information, judging value, screening events, verifying sources, and finally forming an industry report available for reading and decision-making reference.

However, the answers delivered by the three AI office products present three completely different product philosophies.

Let's look at WorkBuddy first. It took 7 minutes, the fastest among the three products to complete the task. From the final result, the content delivered by WorkBuddy is closest to a "work draft that can be directly modified".

It has a complete structure and clear titles, covering industry dynamics, corporate news and financing events. For ordinary office scenarios, if users only want AI to quickly sort out materials needed for an afternoon meeting, such results are already sufficient.

Most office tasks do not require AI to be an industry expert. What users need more is to complete a task quickly and reduce the cost of sorting out materials from scratch.

WorkBuddy chooses the strategy of "finish first, then optimize", but the problem also arises here: its judgment chain is relatively short.

In the weekly report, WorkBuddy judged that there were no major public dynamics retrieved for Leju Robotics and Tushi Intelligent Navigation this week. But further verification found that both companies actually had public events happening.

Leju Robotics' IPO application was accepted by the Shenzhen Stock Exchange this week, and it also participated in relevant competition events; the co-founder of Tushi Intelligent Navigation was selected into the MIT TR35 list.

The problem is not that the information does not exist, but that Workbuddy deduces "no information found" as "this event did not happen".

There is a critical gap between these two judgments. In ordinary office scenarios, this deviation may have limited impact, but when the AI output is an industry report used to judge market trends, competition patterns and even investment directions, wrong judgments will affect subsequent decisions.

If WorkBuddy reflects the priority of efficiency, then TRAE Work is more inclined to information coverage. Its process is closer to searching keywords, obtaining information, summarizing and sorting, and generating reports.

From the surface results, TRAE Work has more information volume than WorkBuddy. It lists the dynamics of ten companies in the report, six of which are marked as "no verifiable new dynamics retrieved this week", including Qianxun Intelligence, Leju Robotics, Digua Robotics, Tushi Intelligent Navigation, Extreme Dynamics and Zhongke Diwuji.

However, after checking one by one, it is found that at least five of these companies have public dynamics. This week, the founder of Qianxun Intelligence gave a speech at the Industrial Internet Conference, and there are also relevant in-depth reports;

Leju Robotics not only has IPO-related progress, but also other public activities; Digua Robotics has reports on supply chain ecological cooperation and computing power platforms; Extreme Dynamics displayed a full-size humanoid robot at ChinaJoy; Zhongke Diwuji also had patent disclosures and forum activities.

This shows that the problem of TRAE Work is not simply insufficient information. To be more precise, it emphasized information acquisition rather than in-depth judgment in this task.

Another noteworthy point is its source structure. A large number of links in the report generated by TRAE Work come from content aggregation platforms such as Toutiao, rather than the original sources of events. For ordinary information sorting tasks, this method does not have major problems. But for industry research, finding information is only the first step. A piece of news about the same company being reposted by ten media outlets does not mean that ten pieces of independent information have appeared.

If the Agent cannot identify the relationship between information, what it obtains may only be more links, rather than higher information density.

In contrast, Qwen Office spent 41 minutes completing this weekly report, and it added more judgment steps during the task.

For example, it will actively mark risks instead of forcing a definite answer.

When mentioning a piece of information about the US FCC restricting the import of Chinese humanoid robots in the weekly report, Qwen Office did not write it directly as a confirmed fact, but marked "Currently there is only a single source, further verification is recommended."

This seems to be just a detail, but it reflects a distinctive processing method. Many Agents, when faced with uncertain information, tend to give a complete answer. Because it looks "smarter" for being complete, but for research tasks, knowing where the uncertainty lies is also a kind of capability.

Secondly, it will also distinguish between the occurrence time of events and the publication time of reports. For example, the cooperation between Guanglun Intelligence and Siemens. This event itself took place during WAIC on July 20, which has exceeded the statistical window of the weekly report. But Qwen Office did not simply delete it, but marked "This is an event before the statistical window, but relevant reports appeared within the window, so it is included".

In addition, it will also explain its own trade-off logic. When counting the news of Extreme Dynamics, Qwen Office wrote "It is in a relatively silent period after the Pre-IPO round within the window, so one piece of news is included". In addition to telling users the result, it also explains "why".

Finally, it will actively explain the research methods. A complete "Report Description" is added at the end of the report, including statistical caliber, source format, reliability description and coverage scope.

This is also the biggest difference between Qwen Office and the other two products: what it outputs is not just a report, but also a demonstration of a set of research processes.

In another set of tests, we asked Qwen Office to read the meeting transcript of the AI hardware A1 recording card in DingTalk. It did not pretend to read successfully without obtaining authorization, but clearly prompted that it could not access the data currently.

After re-authorization, it not only completed the extraction of transcript content, but also actively identified speech recognition errors in it. For example, VLA was mistakenly recognized as "VOA", and CVPR was recognized as "cpr", which it corrected according to the context.

Faced with unconfirmed company names, it did not force a guess, but marked them as suspected transcription errors and left them to users to judge.

The two test scenarios are completely different, but they reflect the same capability behind them: Faced with uncertain information, Qwen Office does not rush to give answers, but knows when verification is needed and when the boundary needs to be explained.

But it should be emphasized that this does not mean that WorkBuddy and TRAE Work have no research capabilities. After all, occasional test results cannot define the capability boundary of a product. The final performance comes from the combined effect of many factors such as model capability, Agent process, tool chain design and resource input.

To be more precise, the three AI office products show different ways of investing intelligence.

The Watershed of Agents Is Not Speed, But Judgment Cost

In the past, when judging the capability of AI products, users were used to comparing response speed and answer smoothness. But when AI really enters the office scenario, the core of competition shifts from "who responds faster" to "who delivers usable work", and from "output capability" to "execution credit", that is, whether users believe that it can stably deliver reliable results in complex tasks.

In this light, the differences between the three AI office products actually come from their different understandings of the role of Agents.

WorkBuddy is more like an efficient executor. Once the task is in hand, it disassembles, executes and submits the work in one go, compressing the information sorting that originally takes half an hour to a few minutes. This logic perfectly fits the creed of traditional office software: tools are responsible for speeding up, people are responsible for making decisions, pursuing results first then optimization, and "good enough".

TRAE Work is more like an information assistant, whose strength is to cast a wide net, expand the search radius as much as possible, and push more relevant materials to users. In the early stage of AI, this capability is very scarce. After all, the biggest pain point of using search engines in the past was "not knowing where to look". But when we enter the deep water zone of complex tasks, the bottleneck is no longer finding information, but judging information. Especially in research-oriented tasks, the quantity of information never equals the value of information.

Qwen Office has chosen a heavier path, it tries to move the Agent from "helping users do things" to "taking over judgment for users".

Behind those 41 minutes are more intermediate steps: cross-checking sources, verifying credibility, locking event time, eliminating duplicate noise, marking risks, and explaining trade-offs. In other words, it makes the work process that was originally hidden in the minds of human researchers explicit.

This is also the biggest difference between Agents and traditional AI assistants. In the past, the core value of software tools was to improve efficiency: Office helps people write documents, Excel helps people calculate data, and search engines help people find information. Tools are responsible for execution, and people are responsible for judgment.

After the emergence of Agents, this relationship loosened. It began to participate in the judgment process. It needs to know which information is worthy of trust, which places need verification, and when it cannot give a definite answer.

The ability to build user trust in complex tasks will become the core of future Agent competition.

Token Is Becoming the New Business Language of the AI Era

Why are people willing to invest more intelligent costs to build this trust? The answer points to Token.

Business in the Internet era revolves around user scale, traffic is the core asset, and DAU, MAU and user duration are the scales to measure business value.

In the AI era, the number of users is still important, but what is more important is how much AI capability users have actually called. Because for AI products, opening an App once generates little commercial value. What really generates value is how many times the model is called, how many tasks are processed, and how many Tokens are consumed.

Token has evolved from a technical indicator to a new business measure in the AI era.

This is also the reason why Alibaba reorganized its AI business around Token. In March this year, Alibaba established the ATH Business Group, which is in parallel with Alibaba Cloud, e-commerce and other businesses, and has become one of Alibaba's core businesses, personally led by group CEO Wu Yongming. The core direction of ATH is summarized as "create Token, deliver Token, apply Token". The implication behind this sentence is that in the AI era, competition is not only