Today, Gemini has fully plunged into the fiercely contested AI office space.
Reported by Zhidx on October 9, in the early hours of this morning, at the Gemini at Work 2026 conference, Google Cloud officially launched Gemini Agent, a universal enterprise agent, aiming to assign tasks that were previously scattered in office software, data platforms and development tools to an agent that can independently plan, call tools and execute continuously.
The capabilities of Gemini Agent are no longer limited to answering questions or generating a piece of text. Users can directly describe the final goal in natural language, allowing the agent to decompose tasks on its own, call enterprise data, write and run code, generate documents and multimedia content, and continuously advance complex work in the cloud.
When facing large-scale tasks, it can also organize multiple sub-agents to collaborate for execution, and even participate in team collaboration as a digital colleague with an independent enterprise identity.
Taking preparing PPT materials as an example, users can directly ask Gemini Agent to make PPT, build a revenue forecasting model, and generate supporting promotional videos. The agent needs to understand the goal on its own, call corresponding skills, enterprise data and generative models, complete all links, and then deliver the results to users.
In this process, users can still check the progress, adjust requirements and decide whether to authorize key operations, but they do not have to stay in front of the computer all the time to guide each execution step step by step.
To enable agents to integrate into daily enterprise work, Google Cloud has also supplemented capabilities including four-tier memory mechanism, cross-platform collaboration, enterprise data access, security isolation and cost control around Gemini Agent, forming a complete agent system that can push enterprise tasks from instructions to execution results.
This product form of Gemini Agent reminds people of domestic products such as Doubao Work, Tencent WorkBuddy and Alibaba's Qwen Office. They are all office agent platforms, and all focus on "independently decomposing tasks, calling tools, and continuously advancing work".
However, it is worth noting that Gemini Agent is an enterprise-level product. The specific official commercial release date and pricing plan have not yet been announced, and it is only open to some enterprise users for preview.
01.
From Conversation to Execution
Gemini Agent Aims to Become a Universal Colleague for Enterprises
In the past, AI capabilities in enterprises were usually scattered in different products: some people used chatbots to write copy, some used programming assistants to process data, and others used AI features in office software to create presentations.
Gemini Agent hopes to integrate these capabilities into a unified agent, so that the same agent can handle different tasks such as knowledge Q&A, document generation, code execution and multimedia creation.
In terms of access entrances, Google Cloud extends it to Web, mobile, desktop and command line environments, promotes the collaboration between the agent and work tools such as Google Workspace, Microsoft 365 and Slack, and can also access other services through MCP. Users can call the agent in their daily work environment without having to explain the background from scratch every time.
Gemini Agent also provides specialized versions for the financial services and legal industries to better complete tasks in related fields.
Long-span tasks are another threshold that Gemini Agent tries to break through.
Ordinary chatbots usually work around the current conversation. Once a task involves multiple systems, long execution time or repeated verification, users need to follow up continuously. Gemini Agent relies on the cloud operating environment to maintain the task status, so that part of complex work can continue to advance after the user leaves the computer.
For example, a data analysis task can go through multiple stages such as data extraction, code generation, model training, result check and report organization. The agent does not have to wait for the user to issue new instructions after completing each step, but can continue to execute according to the established goal, and seek user intervention when encountering links that require judgment or authorization.
Google Cloud has also introduced a dynamic multi-agent collaboration mechanism for Gemini Agent. When the task scale is large, the system can split the work to multiple sub-agents, which respectively process links such as data, code or content generation, and then summarize the execution results.
In addition to tasks initiated actively by employees, Google Cloud also proposed the concept of Coworker Agent.
This type of agent can have an independent enterprise email, exclusive storage space and company directory identity, and participate in email and document collaboration. With the corresponding authorization mechanism, it can undertake continuous tasks as the work identity of a team member, instead of only starting to work after waiting for employees to open the chat window.
For example, enterprises can assign certain periodic work such as data sorting, information summarization or email preparation to the agent, so that it can continue to advance according to the established rules.
02.
Four-Tier Memory Enables Agents to Understand Enterprises
Supports Models Such as Gemini and Claude
To enable an agent to participate in enterprise work for a long time, it is not enough to only understand the current instructions. It also needs to remember what tasks the user is processing, how the enterprise defines business indicators internally, and how the past workflows are usually executed. Focusing on this problem, Google Cloud proposes four types of memory mechanisms, and further decouples model selection from specific tasks.
The four types of memory solve different problems respectively:
The first type is Session Memory, which is used to retain the context of the current task or conversation, so as to prevent users from explaining the same thing repeatedly.
The second type is Semantic Memory, which is used to organize enterprise knowledge and business concepts, helping agents understand internal company terms, data definitions and knowledge relationships.
The third type is Procedural Memory, which is used to precipitate workflows and execution methods, so that agents can reuse defined task steps.
The fourth type is Episodic Memory, which is used to record experiences and results related to past tasks, providing references for subsequent work.
These four types of memory do not simply store all historical conversations in a larger context window, but try to distinguish different types of information, so that agents can call corresponding background when needed.
For example, when an employee requests to generate a sales analysis report, the agent not only needs to understand the current instruction, but also needs to know the company's definition of indicators such as sales revenue and profit margin, understand the format that the report should adopt, and refer to the execution experience in related historical tasks.
For enterprises, the value of such capabilities lies in reducing the cost of repeatedly explaining backgrounds. However, whether the memory can play an accurate and continuous role still depends on enterprise knowledge quality, permission settings and information update mechanism.
At the model level, Google Cloud hopes to avoid enterprise agents being completely bound to a single model.
Gemini Agent supports selection between Google's own models and third-party models such as Anthropic Claude, and distributes different tasks to suitable models through intelligent routing. Relatively simple work can be handled by models with lower cost, while complex reasoning tasks can call more powerful models.
Google Cloud stated that after adopting intelligent routing, the comprehensive operating cost of related tasks can be reduced to 1/3 of that when all cutting-edge models are used, and the execution speed is increased by 40%.
In terms of underlying models, Google also mentioned Gemini 4 Argon released last week. This model is oriented to complex reasoning and long-span software engineering tasks, and has made progress in related benchmark tests and Google's internal engineering applications, providing model capability support for agents to perform complex work.
03.
Enterprise Data Is No Longer Just a Stored Asset
Google Cloud Completes the Data Foundation for Agents
What affects whether an agent can be put into the production environment is also whether it can accurately call data, continuously execute processes, and comply with enterprise rules throughout the whole process.
Google Cloud has built multiple Skills for specific scenarios for Gemini Agent. The Operation Report Skill allows users to understand enterprise operation data through simple conversations, while the Machine Learning Skill can play a role in AI development scenarios.
If an agent needs to perform tasks independently, it must be able to understand and use enterprise data. However, a large amount of enterprise information is still scattered in databases, cloud storage, office documents and business systems. Even if the model itself is powerful enough, it is difficult to complete end-to-end business tasks if it cannot understand the meaning of data or obtain appropriate access permissions.
Google Cloud has integrated capabilities such as Knowledge Catalog, Smart Storage and Borderless Lakehouse for Gemini Agent this time, trying to lower the threshold for agents to use enterprise data.
The focus of Knowledge Catalog is to help agents understand the business semantics behind the data.
The same indicator in an enterprise may adopt different field names or calculation methods in different systems. For example, when an employee proposes to query "net profit margin", the agent not only needs to find relevant data, but also needs to know the specific definition of this indicator in the enterprise and where the corresponding data comes from.
Through Knowledge Catalog, Google Cloud hopes to associate business concepts with underlying data structures, and is compatible with enterprise data environments such as Databricks, dbt, Looker modeling system and SAP. In this way, when performing analysis tasks, agents can locate data according to the existing business definitions of the enterprise, instead of completely relying on field names to guess the meaning.
In the conference demo, the data science agent extracted data, generated code and trained the XGBoost model according to natural language instructions, demonstrating the automated process from business problems to machine learning analysis results.
Smart Storage extends AI capabilities to unstructured files. Images, PDFs and scanned documents usually lack structured information that can be directly used for retrieval. Even if enterprises store a large number of materials, they may not be able to quickly find the really useful content.
Google Cloud hopes to use Gemini to conduct contextual understanding and labeling of these files, so that agents can extract relevant information from the materials in storage.
For example, Snap connects the Prism agent with storage archives for technical diagnosis and troubleshooting. Google Cloud disclosed that this solution shortened the relevant troubleshooting time from 30 minutes to 30 seconds. Hitachi uses the HMAX agent to compare the wiring photos before and after the process at the production site, helping frontline personnel identify differences, achieving a 30% productivity improvement.
Another capability, Borderless Lakehouse, targets the problem of scattered cross-cloud data in enterprises. Google Cloud hopes