Xiaolong bypasses the office agent
Office Agent has become a new arena for tech giants.
Yet amid this boom, second-tier startups including Zhipu AI, Moonshot AI and MiniMax have stayed fairly calm, making little noticeable noise.
In fact, these startups have also rolled out relevant products: Kimi launched Kimi Work in June, Zhipu AI released AutoClaw in March, and DeepSeek open-sourced DeepSeek Harness in August.
But obviously, these are not real "Office Agents", merely personal desktop tools.
It is not that these startups are unwilling to follow the trend. The Office Agent sector targets enterprise procurement, which brings high and stable revenue. Especially at the stage of preparing for IPO, several such contracts can greatly polish the financial reports.
The reality is that these startups cannot keep up. Their inherent genes and scale naturally determine that they are not suitable for this track. Once they force their way in, it will not only lead to losses, but also possibly delay the company's core AGI development roadmap.
1
Tech giants dare to go all-in on Office Agents not because their models are exceptionally powerful. On the contrary, the startups actually take the lead in this field.
Taking the Agent benchmark test as an example, Kimi K3 scored 42 points in the SWE Marathon long-horizon engineering test, higher than Claude Opus 4.8 and GPT-5.6 Sol; it reached 91.2% in BrowseComp long-cycle information retrieval, ranking first globally; it also scored 88.3 points in Terminal-Bench and 84.8 points in OSWorld-Verified.
The GLM-5.3 released by Zhipu AI on August 14 performs even better: its Terminal-Bench 3.0 score surged from 4.6 of version 5.2 to 28.3, the DeepSWE v1.1 score rose from 46.2 to 66.9, close to Claude Fable 5's 69.7; its Agents' Last Exam score increased from 23.8 to 28.5, surpassing Kimi K3's 27.6 and Opus 4.8's 25.7. It took the first place among all open-source projects in three Agent-related benchmarks.
So in fact, the tech giants are the "weaker side" in terms of model capability.
Then where do the giants win? They win in the understanding of customers' organizational structures.
What the giants do is embed Agents into customers' organizational relationships, which involves more permissions, approval processes and workflows, rather than the mathematics, coding and reasoning that the startups are good at.
If code repositories and documents are the context of our daily office work, then the above elements are the enterprise context, which is the real moat of the tech giants.
These data only exist on the servers of DingTalk, WeCom and Lark, which the startups cannot access, let alone use to train their models.
What the giants need to do is to leverage the enterprise context to turn Agents into "employees" of the organization, so that they can work for humans under the existing permission system.
Even MyContext, the personal work context project recently open-sourced by Qianwen Office, distills a "second version of you" from chats, documents and meeting records in Lark and DingTalk. It aims to convert the work traces precipitated in the organization into a way for Agents to recognize each individual user.
Data is only one aspect. On the other hand, at the product and technical level, the competition between Office Agents lies in Runtime.
Large language models are brains, but they have no hands, feet or memory. When you tell it "help me sort out the meeting minutes from last week, generate a weekly report and send it to the boss", it can understand the meaning of the sentence, but it cannot execute the instruction. It is a pure text system that "only thinks but cannot act".
Runtime is the body and nervous system equipped for this brain. It is a set of program frameworks responsible for converting the model's "ideas" into "actions". When the model says "I want to read this file", Runtime will open the file and read the content inside. The model itself cannot touch the real world, and Runtime is the only interface between it and the real physical world.
Agents need to work continuously. The model first thinks about what to do next, Runtime executes the task, feeds the result back to the model after execution, and the model decides the next step according to the result until the task is completed. This loop is the core of Runtime.
The Harness we often talk about is the upper layer of Runtime. Harness is responsible for managing thinking and calling logic, while Runtime is responsible for the final execution.
No matter it is WorkBuddy, Qianwen Office or Doubao Work, their core is to turn the long-accumulated enterprise operation logic into Runtime.
We can split the term Office Agent: the model determines the "Agent" part, while Runtime determines the "Office" part.
The Runtime of tech giants is an extension of the organizational system, while the Runtime of startups is an extension of personal devices.
The Runtime of tech giants is rooted in the organizational structure. The Runtime of Qianwen Office is the enterprise capability layer of DingTalk; the Runtime of WorkBuddy is the Agent stack of CodeBuddy, wrapped with an additional Tencent security gateway; the Runtime of Doubao Work is directly a cloud computer.
The users of their Runtime are not "you", but "your organization". Permissions, approval, audit and data are all designed around organizational relationships.
The Runtime of startups is rooted in you as an individual.
The Runtime of Kimi Work is the local execution kernel Kimi Code, and the Runtime of AutoClaw is connected to IM workbenches like Lark through OpenClaw. The default users of their Runtime are "you", centered on your personal devices, personal browsers and so on.
In other words, the Runtime of startups is your personal employee, while the Runtime of giants is the enterprise's employee. This does not mean they narrow down the scope, but attack from a completely different direction.
Obviously, the startups and the giants are not at the same starting point. The giants follow the Runtime logic and find the Office Agent track, while the startups cannot force their way into this field.
2
According to Analysys data, in June, the total monthly visits of 17 mainstream domestic desktop AI-native office agents exceeded 60 million, while the number was only 20 million in March. It tripled in just three months.
Tencent's WorkBuddy topped the list with 20.97 million monthly visits, exceeding the sum of the second and third place. The "2026 White Paper on China's Enterprise-level AI Agent Industry" released by institutions under the China Academy of Information and Communications Technology shows that the enterprise-level AI agent market reached 21.2 billion yuan in 2025, and is expected to increase to 44.9 billion yuan in 2026, almost doubling, and is expected to exceed 332 billion yuan in 2029. Gartner predicts that by the end of 2026, 40% of global enterprise applications will have built-in task-based AI agents.
This is the reason why the giants quickly adjusted their organizational structures and gathered resources to develop Office Agents.
The entry point is shifting from "applications" to "Agents". Once users get used to talking directly to Agents instead of opening office software, the organizational assets accumulated by DingTalk, Lark and WeCom over the past decade may depreciate within a single product cycle.
The Office Agents of tech giants are both offensive and defensive moves.
Qianwen Office defends DingTalk's core B-end market, WorkBuddy defends WeCom, and Doubao Work defends Lark. They are even willing to burn money for the entry point: Doubao Work gives 30-day subscriptions to users upon login; Qianwen Office is free during public beta and charges only after the official launch; WorkBuddy gives free tokens for using Hy3. For the giants, Office Agent is the cost center of their ecological strategy, an insurance to prevent the depreciation of their core business.
The startups themselves do not have similar assets to defend.
Thus there is a saying: Kimi and Zhipu AI are independent model companies, and each product line needs to be self-sufficient, which cannot use subsidies to seize the market like the giants.
This statement is half correct. Kimi and Zhipu AI may also subsidize products for the entry point, but the duration and scale of subsidies they can afford are far less than large tech companies. The real difference lies in the "timing".
Although enterprise-level Office Agents have high unit prices, they require heavy operation and heavy customization, involving integration of ERP, CRM, OA and HR systems, permission sorting, private deployment of data, and the contracts need to go through layers of checks of information security, procurement and compliance.
The general process is: the business department raises the demand first, the IT department makes technical selection, then works with the vendor for POC verification (ranging from two weeks to one month). After the verification is passed, it enters the bidding process. The procurement department compares prices and completes supplier access, then the procurement committee approves the contract.
The legal department reviews contract terms, intellectual property rights and data responsibility division, the information security department conducts penetration testing, equal protection assessment and data review, the compliance department confirms whether it meets industry regulatory requirements, and finally the vice president in charge or even the CEO needs to sign off.
In the procurement process of a large client, you need to contact more than a dozen to 20 key stakeholders. If any link gets stuck, the entire order will be suspended.
From contacting the client to actually signing the contract, it often takes a quarter or even half a year, and the situations vary greatly among different clients.
It is not a once-and-for-all deal after acceptance. What enterprise clients buy is not software, but services. It requires exclusive customer success managers to make regular return visits, 7x24 hour technical support response, SLA guarantee, regular security inspection and version upgrade, and continuous iteration for customized demands raised by clients.
What is more troublesome is that the organizational structure and business process of the enterprise will also change. For example, when the department is adjusted, the approval flow is modified, or a new system is launched, the corresponding Office Agent configuration needs to be modified accordingly, which requires the vendor's team to follow up continuously.
Behind a large client, there is usually an exclusive service team of 2 to 5 people.
For startups that have not yet formed stable cash flow, this is obviously quite difficult.
3
In the Office Agent track, DeepSeek is an alternative. It never planned to make a complete office software, and its official definition is only one sentence: everything is a plugin.
Tools, files, conversations, sandboxes, task loops, and even the interface itself, all are made into replaceable plugins.
After its release on August 13, the GitHub repository of DSH exceeded 20,000 stars in 1 hour, 92,000 stars in 28 hours, more than 180,000 stars in 9 days, and now it is close to 200,000 stars. This is the fastest growing open source project in GitHub's history.
More than 10,000 plugins have emerged in the community soon: desktop clients, memory systems, workflow tools, multi-model switchers. Even the deepseek-harness-desktop project that converts it from a highly threshold Web UI to an executable file has already gained 13,000 stars.
In the past, Claude Code and Codex made Harness a black-box product, but DeepSeek chose to open source Harness itself, which is equivalent to distributing the "capability to build Office Agents" to everyone.
Thus, a brand new participation mode appears in the Office Agent war. DeepSeek only distributes tools, letting everyone build what they need by themselves. If you want it to understand images, just download the corresponding plugin from the GitHub plugin market and install it. DSH only provides the environment, not the implementation method.
DeepSeek does not want applications or entry points, what it wants is the "standard", making DSH the public base of Office Agents, so that developers get used to using its framework and process.
Just like OpenClaw and Claude Code, the product logic of all Agent products comes from these two products.
The advantage of this approach is that the Agent is strongly related to the task: you only get what you need. If you cannot implement a certain function by yourself, you can turn to the community for help. If you do not need a function, it will not be there, unlike other Office Agents that come with pre-installed functions you may never use for your whole life.
There is only one problem: DSH uses the MIT license, which means anyone can modify its code. Both the giants and the startups have their corresponding products, and they can also carry out secondary development based on DSH.
This is very similar to Android, but Android is backed by Google's search and advertising revenue, while DeepSeek only has API as its only charging method. The premise of betting on the standard is to bet that others are willing to stay in your system for a long time.
No matter how prosperous the community project is, it is difficult to convert it into DeepSeek's own revenue. The only revenue DeepSeek can actually get is the API fee for calling DeepSeek model through DSH by default.
The plugins of DSH can connect to any model. It runs DeepSeek today, and may switch to GLM, Kimi or GPT tomorrow. The underlying framework may be emptied at any time. The more open the ecosystem is, the more likely DeepSeek will become an infrastructure, but the harder it is to monopolize the industrial chain.
Although the Office Agent track has just started, one direction is certain: the future office will definitely focus on a certain Office Agent, instead of piling various software, data and reports on the desktop as it is now.
Agents will become the new entry point, and the form of organizational software will be rewritten.
This article is from WeChat Official Account "LetterRank" (ID: wujicaijing), written by Miao Zheng, authorized to be released by 36Kr.