After a model evaluation, I began to re-understand WorkBuddy
What a model can ultimately deliver depends not only on its inherent capabilities, but also on the working environment it is placed in.
Flash models are increasingly becoming the "basic staple" in the AI world.
They are fast, low-cost, and more than capable of handling daily tasks. As a result, a fixed routine has emerged for horizontal model evaluation: assign the same task to several models, compare their outputs to see which one performs better, and then rank them.
This round of Flash model testing initially followed this established routine.
We ran two tasks using DeepSeek-V4.1-Flash, Yunzhisheng U2-Flash, and Qwen-3.8-Flash: an in-depth research report on the mobile phone industry, and the development of a match-3 mini-game. U2Flash and DeepSeekFlash stood out in terms of rigor and insight respectively, while Qwen showed obvious shortcomings in factual verification for the research report and game delivery.
However, when we reviewed the entire testing process, we suddenly realized that this test was unfair from the very beginning: the three models did not run in the same working environment.
DeepSeek Flash and U2Flash completed the whole process in WorkBuddy, with access to internet search, file reading and writing, and tool invocation; Qwen3.8 Flash, on the other hand, was tested on the web experience version of the Qwen AI platform, with nothing but an isolated dialog box throughout the process.
Since we could not directly call the Qwen3.8 Flash model in the Qwen Office desktop client during the initial test, we conducted an additional round of re-testing with strictly controlled variables, using exactly the same tasks and materials to re-run the whole process in WorkBuddy.
The result was unexpected. Qwen3.8 Flash, which performed poorly in the initial test, saw a significant improvement in output quality in the re-test: a large number of data errors from the first round were re-verified and corrected; the previously stuck match-3 mini-game also completed the full closed loop of locating bugs, modifying codes, and conducting practical verification.
The same model can produce such vastly different delivery results when switched to a different working environment, which makes us wonder: what exactly are we testing in the so-called model evaluation? Is it the model itself, or the actual deliverables the model can produce after being integrated into the workflow?
01
Qwen Model Runs "Bare" on Its Own Platform
The starting point of this evaluation stemmed from a failed integration attempt.
We originally downloaded the Qwen Office desktop version, intending to call the Qwen 3.8 Flash model directly inside it, but soon ran into obstacles. Qwen Office explained that the specific model to be used is determined by the session configuration of the product, and users cannot manually switch or specify the underlying model during conversations. Whether a model is accessible also depends on the specific entry point, account, and version.
Later, we settled for the next best option, turned to the web experience version of the Qwen AI platform, found the 3.8Flash entry, and then officially started running the tasks.
At that time, it seemed that we had just changed an operation entry, but during post-test review, we realized that this was the most important hidden variable in the entire test.
Let's look at the first round first.
The research report task had quite a few hidden pitfalls. The five material packages and eight task lists we provided contained misattributed data, unit errors, confusing time calibers, and unsubstantiated assertions. The game task required the model to build a fully functional match-3 mini-game.
DeepSeek Flash
U2Flash
U2Flash handled the data traps most meticulously. After finding that the figure of "OPPO 68.72 billion yuan" conflicted with the data from Transsion Holdings' annual report, it traced all the way to the business entity and statistical caliber. DeepSeek Flash's strength lies in independent judgment, as it can put forward sharp viewpoints such as "chip forking".
The problems of Qwen3.8Flash were concentrated in the section "Panorama of New Phones".
The real new phones listed in the materials were not all available yet, but it submitted the report in advance, using old product lines from its training data to fill the 2026 new phone list: vivo was written as the X200 series, while Huawei had models such as Pura 80 and Mate 70. The revenue, net profit and terminal business data in Huawei's 2025 annual report were not accurately captured; a set of OPPO data was also misjudged.
The shortcomings exposed by the game task were even more serious. It only generated a piece of plain text code that needed to be manually copied to a notepad or editor to save and run. After opening it in the browser, the deliverable still stayed at a very primitive web demo stage: it could be launched and played, but the scoring logic was completely stuck, like a functional but soulless empty shell.
Qwen3.8flash web version
These performances certainly have factors related to the model itself, such as using old historical knowledge to fill current information gaps, and the lack of a self-check mechanism after generation. But in the first round of testing, Qwen3.8Flash did not have access to the "weapons" that the other two models had.
DeepSeek Flash and U2Flash run in AI office products like WorkBuddy. When consulting Huawei's annual reports, IDC shipment data, and industrial and commercial data, they can search online; when encountering data traps in materials, they can initiate cross-verification; after the report is written, they can directly save it to local files, organize it by chapters, and package the output.
At this point, the situation became awkward. We thought we were comparing the technical capabilities of the three Flash models horizontally, but what we actually compared was three completely different sets of model invocation and delivery methods.
On one side, the model is connected to a workbench with internet access, tool invocation, local reading and writing, and self-check capabilities; on the other side, the model's capabilities are limited to a single web dialog box, running "bare". The delivery gap obviously cannot be simply attributed to "insufficient model intelligence".
So we reset the variables, kept the tasks, materials and model unchanged, only moved Qwen3.8 Flash into the same working environment WorkBuddy, and re-ran the whole process.
The result was unexpected. In the research report task, Qwen Flash almost reproduced the engineering standard of U2Flash from before. Right at the beginning, it output a list of credentials containing seven error checks, and each of them was cross-verified.
Qwen Flash retest version
It also established a three-level source system of "official, institutional, and third-party estimation", and marked each entry in the text. In the initial test, Qwen made a large number of factual errors in the "Panorama of New Phones" section; after the retest, it not only completed the eight-chapter report, but also proactively conducted a round of data error checking.
In the game development task, the difference was even more significant. On the web version, Qwen could only provide a static piece of text code; while in the WorkBuddy environment, when there were rendering bugs or stuck scoring logic, Qwen Flash could re-read the error logs through the workbench's execution environment, complete the closed loop of "locating errors, modifying codes, locally recompiling and verifying", and finally deliver a mini-game with decent playability.
The reason why the model showed such completely different upper limits of delivery is that it was given the conditions to re-verify information, re-read contextual materials, call environmental tools and actively correct errors. The capabilities that were previously trapped in the dialog box were maximally released after it was equipped with a complete working link.
DeepSeek Flash
U2 Flash
This also makes us re-understand the current path divergence of AI office products.
Qwen Office tends to integrate model capabilities, application tools and interactive interfaces into a highly standardized delivery solution. This "out-of-the-box" experience greatly lowers the usage threshold, but also restricts users' right to customize the underlying model. In order to pursue the "absolute unity" of the front-end experience, the model's potential for adaptation in complex and dynamic scenarios is invisibly "sealed".
In contrast, WorkBuddy does not forcibly bind a single model ecosystem. It allows users to freely mount DeepSeek, U2 or Qwen in the same workspace according to task requirements, treats the model as a "pluggable computing power engine", and completely returns the right of choice to the task itself. Here, the model is more like a replaceable computing resource, and the workbench is responsible for organizing tasks, contexts, tools and models.
Under this path, the core value of AI office products no longer depends on "how powerful my model is", but on "whether I can accurately and seamlessly embed the most suitable model into complex real workflows".
02
Agent Competition Has Escalated
From "Model Entry" to "Task Entry"
Looking back at this test, the three experiences correspond to three completely different AI usage methods, and also reflect three stages of the evolution of AI applications.
The first stage is conversational AI assistants, which are single-model chat tools mainly for personal scenarios. Users enter the preset suite provided by the vendor, where models, tools and interaction logic are all fully packaged. Its essence is an auxiliary tool that improves personal efficiency. Its advantage is simplicity, at the cost of completely giving up the right of choice.
The second stage is a schedulable "Agent Workbench", which