Not waiting for Gemini 4, Google has launched an office Agent that supports calling Claude.
Alright, Google has also joined the fierce scrum of office AI agents, and the platform that emphasizes multi-model invocation has even arranged for its competitor Claude to be available for use!
In the early hours of this morning, Google Cloud released a long article titled Welcome to Gemini at Work 2026: Introducing the Gemini agent, signed personally by Google Cloud CEO Thomas Kurian, which focuses on introducing a series of updates of Gemini in the enterprise agent field.
Among them, the most groundbreaking launch is the brand new Gemini Agent.
According to Google's vision, this will be a general-purpose office agent that covers all types of work tasks and can run continuously for a long time.
It can help you look up information, write emails, make PPTs, analyze data, write and run code, call tools across applications, plan tasks independently, and gather multiple sub-agents to work together when necessary.
There are also two points particularly worthy of attention.
First, it can have its own enterprise identity and become an AI colleague with an email address, a calendar, and an independent account.
Second, it can automatically select the underlying model according to tasks, and even call Claude from Anthropic.
Wow, Google doesn't even have to insist on using its own Gemini exclusively?
What's more interesting is that all the functions launched this time are clearly targeted at the recently booming office agent market.
Mark Zuckerberg would probably have the following inner monologue after hearing the news:
Dude, you stopped obsessing over your underperforming Gemini model and turned around to compete with me????
Gemini agent
The working principle Google sets for it is as follows:
You give it objectives, not instructions.
It is assigned goals instead of step-by-step commands.
That means if you want the Gemini agent to help with work, you don't need to tell it step by step what to do first and then what to do next.
You only need to state the final expected result, and leave the rest to it to arrange on its own.
Its core capabilities include a unified entry, persistent cloud operation, cross-platform invocation, multi-agent collaboration, event triggering and scheduled tasks, etc.
The example given in the article is to ask it to help complete a market analysis report.
It can search and organize information on its own, analyze relevant data, build financial models in Google Sheets, and then call Google Slides to generate a presentation.
Throughout the whole process, it can select tools and call different applications according to the needs of the task, and finally return the completed work to you.
You may be full of questions — similar cross-application operations are nothing new, what are you trying to do that's different?
Most netizens are not optimistic about this, and the comments under the official announcement tweet are mostly full of playful teasing:
However, what Google emphasizes when launching Gemini Agent this time is that it tries to integrate all kinds of work capabilities into one single agent.
What makes it different?
Under the general framework of Gemini Agent, Google also proposed the Coworker Agent, which literally translates to "colleague agent".
Google now allows enterprises to create a long-existing identity with clear job responsibilities for Gemini, so that it can participate in daily collaboration just like other members in the team.
In addition, the AI colleague is configured with an independent Google Workspace account, enterprise email, calendar, Google Drive storage space... and it can even appear in the company's address book.
First of all, the threshold for getting started is almost zero. Secondly, after the creation is completed, team members can invite it to join Google Chat group chats or @ it in documents just like dealing with ordinary colleagues.
The AI colleague uses its own identity, and the operations it performs will leave its own records, which is convenient for enterprises to manage and track exactly what it has done.
Up to this point, it still looks very normal...
What is relatively different is that apart from the independent identity, Gemini Agent is equipped with a more complete memory mechanism.
According to the official introduction, it has a total of four types of memory.
The first one is Session Memory.
It is mainly used to save the context of the current task.
Even if a task lasts for several days, Gemini Agent can remember which steps have been completed and what work remains to be done next.
The second one is Semantic Memory.
Gemini Agent will accumulate knowledge from the documents it has read, communications with people, and collaborations with other agents, and organize it into a structured knowledge base.
For example, it can remember what products the company has, what each team member is responsible for, and what a certain business term specifically refers to within the company.
The third one is Procedural Memory.
It is mainly used to record how a task should be completed.
For example, what data is needed for a certain type of report, what process to follow for analysis, and what format to use for final delivery.
There is also a detail here: Gemini Agent can even write Skills on its own, saving the methods learned in the work process for use in subsequent tasks.
The fourth one is Episodic Memory.
It is used to record what work has been completed in the past and the corresponding execution experience.
With the combination of the four types of memory, Gemini Agent will gradually have the opportunity to get familiar with users' work habits, company business and team collaboration methods.
For example, the first time you ask it to make a weekly project report, you may need to explain the report structure, data sources and work process.
By the second and third time, it can reuse the previously accumulated knowledge and methods, reducing the content that users need to explain again.
Google even used a metaphor specifically, saying that "Gemini will be like a new employee, gradually getting familiar with you, your tools and your team before officially starting work".
Wow.
Now it can even complete the onboarding training all by itself.
The only problem is that this "employee" seems to never take the initiative to resign (smiley face here).
In addition, this tool can independently select the large model it uses for work.
Google has made it clear this time that the underlying model of Gemini Agent can be selected separately.
At this stage, it can dynamically choose between Google's own Gemini series and Anthropic's Claude series of models, and it plans to support more closed-source and open-source models in the future.
As for which one to use specifically, Gemini will make a trade-off between effect, speed and cost according to task requirements.
For relatively simple work, models with lower cost can be used; for complex tasks, models with stronger capabilities can be called.
Multiple models can also be combined in one project to complete different links.
Google calls this set of capabilities multi-model orchestration, and it is equipped with the Smart Routing automatic routing capability.
It is worth noting that Google also emphasized one point —
The underlying model can be changed, but the context, Skills and enterprise data accumulated by the agent can be retained.
The main purpose is to avoid letting it start understanding the work from scratch every time, and to allow it to continue the previous tasks, knowledge and working methods.
This means that even if the ranking of model capabilities changes in the future, enterprises can adjust the underlying model without having to completely overthrow and restart all the established workflows.
Google, Meta, OpenAI, the three-way battle for overseas daily agents?
Google's entry into this market this time comes at a very subtle point in time.
Just on September 29, OpenAI just launched Dots that can work 7×24 hours.
Earlier than that, Meta also launched the personal AI agent Muse.
Now Google joins the battle with Gemini agent, and the three giants have officially met on the agent track.
Let's take a look at what paths the other two have taken.
Meta's Muse is more focused on personal life scenarios, which can help users shop, book trips, send emails, and even complete payments.
For example, if you take a fancy to a product, Muse can help you find goods, compare options, operate on the website, and finally complete the purchase.
Compared with the complex workflow inside enterprises, Muse pays more attention to those trivial and time-consuming things in people's daily life.
This is also related to Meta's existing social products and huge user base.
Letting ordinary users get used to handing over demands such as shopping and travel to AI is obviously also an opportunity Meta wants to seize.
OpenAI's Dots has more similarities with Gemini Agent.
Dots has an independent cloud computer, supports 7×24 hours of operation, can remember users' long-term goals and work habits, and can connect to more than 4000 applications.
For example, when a developer receives new user feedback, Dots can actively analyze the problem, modify the code, and prepare a PR for review.
For content creators, it can also sort out materials and make content drafts after new interview records come in.
However, Dots currently puts more emphasis on working continuously around individual users.
It will gradually understand users' preferences, goals and judgment standards, and actively look for items that can be moved forward.
OpenAI is also exploring professional Dots for enterprise positions, allowing agents to have independent identities and responsibilities, but this part is currently mainly in the early stage of enterprise pilot programs.
In contrast, what Google focuses on demonstrating this time is to let agents further integrate into the existing organizational and office systems of enterprises.
Gemini Agent can use the business context in applications such as Gmail, Docs, Sheets, and Calendar to understand what employees are doing.
Enterprises can also create Coworker agents with independent emails, calendars, accounts and permissions, so that they can participate in team collaboration with their own identities.
At the same time, Google has also incorporated model selection, enterprise data connection, security audit and cost control into this system.
Even its competitor Claude can become the underlying model that Gemini agent calls to complete tasks.
Looking at it this way, the focus of the three companies is quite clear.
One More Thing
By the way, Google's long article also released other related content. Interested readers can go to check it out on their own.
- Gemini is fully integrated into Google Workspace, supporting cross-application work completion
- Open access to Skills, Tools and enterprise connectors
- Launch professional agents for data analysis, finance and legal fields
- Enhance agent security governance and cost control
Original Google article:
https://cloud.google.com/blog/products/ai-machine-learning/welcome-to-gemini-at-work-2026/
This article is from the WeChat official account "QbitAI", author: Heng Yu, authorized for release by 36Kr