Can OS3 do what the smash-hit Muse and Instinct cannot?
After Mark Zuckerberg unveiled the Muse Charm, the Rabbit r1 was brought back into the spotlight.
When many people see this new device, they can't help but have a sense of déjà vu, recalling the small orange square box from over two years ago. On X, some users posted side-by-side photos of Zuckerberg holding the Charm and Lu Cheng holding the r1 to make memes, and many foreign media also mentioned this first-generation AI hardware when covering the Charm.
This small, portable device can understand user requests through voice and have AI handle tasks on behalf of people. This scenario does feel very familiar.
It's not just the hardware that feels familiar.
In 2024, at the r1 launch event, Rabbit used scenarios like hailing a ride, booking tickets, and ordering food delivery to explain to the public what a Personal Agent can actually do. Since then, these few tasks have repeatedly appeared in demos of successive generations of Agent products, gradually forming a set of familiar "standard packages" for the audience. The review site Assistant Benchmark also includes booking flights, reserving restaurants, and shopping in the test dimensions for Personal Agents.
The reason why this set of scenarios is popular is easy to understand: the delivery results are clear, such as whether the food delivery has arrived, whether the ticket is booked, and whether money is saved during shopping. Compared with explaining how many points the model has improved, these scenarios make it easier for people to understand exactly what AI can do for them.
But are the tasks that are easiest to demonstrate and spread really the most valuable use cases for Personal Agents?
Even large companies face obstacles when getting AI to place orders
After Muse went viral, it soon encountered resistance from platforms.
On September 20 local time, Amazon stated that it had blocked Muse from shopping on its website on behalf of users. Amazon claimed that Meta did not inform it in advance that Muse would access its site, and that Muse did not identify itself as an agent while browsing, raising privacy and security concerns about the way it handles user credentials.
From a technical perspective, Personal Agent solutions using cloud virtual machines have already faced many cases of being denied access. Lu Cheng also mentioned in an interview with tech media WIRED: "We went through all of this a year and a half ago."
Whether AI can stably complete operations and place orders on behalf of people is a separate issue from whether it is authorized to perform such operations. Sometimes even if the Agent knows how to operate, the platform may not grant access, and commercial access and the interest distribution behind it can also block orders.
What platforms are worried about, apart from account security and transaction liability, is their own business.
In the past, users searched, compared products, and viewed recommendations on platforms, from which platforms earned advertising fees and transaction commissions. If all these processes are handed over to AI, users may only remember their own Personal Agent and no longer care which platform the order comes from. Platforms may lose the opportunity to influence user choices, and their bargaining power for advertising and commissions will also be challenged.
After Muse gained attention, the stock prices of travel companies such as Airbnb, Booking, and Expedia came under pressure, reflecting the market's concerns about this change: Will the future transaction entry be controlled by platforms, or by Personal Agents?
Muse's solution is to leverage the resources and influence of large companies to build an Agent-friendly business ecosystem.
After being rejected by Amazon, Muse announced a partnership with US grocery delivery platform Instacart. Travel technology company Duffel also announced that US Muse users can now query real-time flight inventory and prices, and handle bookings, cancellations, and subsequent itinerary arrangements.
On the user side, Muse provides free usage quotas to attract more users, letting them experience the results of Personal Agents first before gradually cultivating usage habits.
Getting more platforms to accept Agents requires handling commercial access, interest distribution, and transaction responsibilities, as well as continuous investment. Large companies like Meta have the resources to attract users and the ability to influence the ecosystem. That's the significance of Muse's rise: promoting more services to integrate, so that Personal Agents have more available scenarios. If these collaborations lead to broader openness, they will also create opportunities for other teams.
However, startup teams with limited resources can hardly copy the investment model of large companies. They need to rely more on technological and product innovation to secure space for survival and development.
Solutions from startups
Lu Cheng once mentioned in an interview that Rabbit, a team of only 15 people, will improve the capabilities of Personal Agents through multiple iterations of the Action Model.
In January 2024, with the release of the r1, Rabbit's first-generation LAM went online, which could only be used via the r1 at that time; in October, Rabbit launched the standalone LAM playground, which can be used on the web, allowing the Agent to browse web pages, find information and perform operations according to user requirements.
According to the technical specifications at the time, it combines page screenshots and web page structures to determine clickable positions, and adjusts subsequent actions as the page changes. When encountering an unfamiliar website, the Agent will find the entry on its own and break down the goal into operational steps to achieve it — this is also the technical route that most Personal Agents still adopt today.
In February 2025, Rabbit released the Android Agent research preview, expanding AI's operational scope to Android applications and system functions of devices. A few months later, some AI phones that followed this technical route were blocked from access by major apps.
Until earlier this year, Rabbit upgraded LAM to DLAM, allowing Agents to access users' own Windows and Mac computers, and operate browsers and desktop software after obtaining authorization.
The effect of this solution is that executing tasks through real devices can effectively reduce the restrictions encountered when cloud proxies access services. The execution environment also extends from cloud virtual machines to users' own computers, where pre-installed software, saved files, and ongoing projects can all be directly accessed and processed by the Agent.
Judging from these release timelines, Rabbit has been leading the industry in exploring Personal Agents. With OS3, Rabbit has started integrating these capabilities to explore what other problems the next-generation Personal Agent needs to solve.
New explorations of OS3
The day before the Muse Charm was unveiled, Rabbit released OS3.
Users don't need to buy an r1 first. By opening the web page, connecting to the model and computer of their choice, they can let OS3 use existing files, software, and tools to handle tasks. One account can connect up to five devices, and users can also interact via iMessage, Telegram, or the r1.
A key focus of this launch demo is to let users assign multiple tasks at once, to see if the Agent can handle them properly.
Lu Cheng put forward several requirements in one go: find a philosophy book, make a reading guide, and send it to himself via Slack; research Liverpool Football Club and generate an interactive web page; make an Agent competitive analysis report, and by the way write a user manual for OS3.
After finishing speaking, he asked OS3 to split the tasks and start executing, while he continued to ask questions in the same conversation.
An intuitive interaction change is that the web main interface of OS3 adopts continuous conversations. Users can assign multiple things in the same conversation, and multiple tasks can be executed in the background simultaneously without interfering with each other, while users can continue to ask questions, or supplement and modify previous requirements.
The arrangements that users no longer need to make become problems that the system needs to solve: which sentence is a new task, which sentence is modifying an old requirement, which file does "the document I mentioned earlier" refer to, and what existing information needs to be retrieved. The significance of continuous conversation is to hand over part of the task organization and context management to the system.
Rabbit's bet is: personal assistants should be able to handle instructions that appear at any time and keep changing, so that users spend less time organizing prompts and managing conversations.
Other products are also exploring this direction. The new version of Projects that Claude Code recently began testing also allows users to put forward requirements in one main conversation, with AI splitting and coordinating background tasks.
OS3 also extends this arrangement to the devices connected by users. After understanding the requirements and splitting the tasks, it also needs to judge: what files and tools are needed, and which device should be used to complete the work.
This involves Rabbit's second technical bet: integrating into users' existing work environments and intelligently coordinating multiple devices.
Cloud virtual machines have a very direct benefit: users can turn off their computers, and tasks will continue to run in the cloud. But apart from platform access restrictions, there is another experience cost: users need to move their work into this new environment. To get the Agent to start working, users need to re-login to accounts, upload materials, and configure tools in the virtual machine.
But in users' own computers, all these things already exist: half-finished projects, installed software, files that have been revised several times, and long-accumulated resources.
Accessing users' existing devices allows Agents to use ready-made files and software, reducing the cost of re-preparing the work environment.
However, accessing users' computers also means that tasks are restricted by the device status: if the computer goes to sleep or loses network connection, the tasks running on that computer will stop.
The solution adopted by OS3 is that it supports connecting up to five devices, including local computers, office machines, and users' own cloud virtual machines.
If a task requires software on a certain computer, it will be operated there; for tasks that need to run for a long time, they can be arranged on devices that stay online continuously. Users can specify the execution location, or the system can intelligently select a device according to the task, or relay the task between multiple powered-on devices to complete it.
For example, OS3 can automatically find the video on Computer A, then use Computer B which has professional editing software to complete the editing, and then use Computer C with the logged-in account to finish the upload. The capability to assign tasks across multiple devices and execute them in relay does not seem to be implemented by any other Agent at present.
In actual use, we tried to get OS3 to complete a data collection task: research the Muse Charm and organize it into a report, then send a message to the mobile phone after completion. About 15 minutes later, the mobile phone received a completion notification from OS3 via iMessage, which listed the research content and the save location of the report in iCloud.
Muse is driven by Meta's own Muse Spark, while OS3 supports integration with different model services, and can also use local models. Models are updated very quickly, and users may switch to another service for reasons of capability, speed, or price.
Several domestic models are also supported for integration. We used Alibaba's Qwen during the test, and it performed very well.
According to Rabbit's design, after changing the model, the previously accumulated memories, configured skills, and connected devices can be retained. OS3 itself is free to use. To use cloud models, users need to prepare the API keys of supported services and pay according to usage.
Rabbit hopes that OS3 can take over this part of accumulation: the underlying model can be updated, and the established working relationship between users and Personal Agents can continue to be retained. A Personal Agent that gradually knows what projects you are working on, how you are used to processing files, and what things have been discussed before should not have all these accumulations invalidated when the model is replaced.
We learned from Rabbit's Community that within three days of the launch of Rabbit OS3, new users have successfully completed more than 10,000 tasks in total, and the total daily token consumption of users using their own API keys is nearly 1 billion, with the per capita level far higher than most Agent products.
A long-term usable Personal Agent also needs to be able to verify task results, and apply user feedback to the next work, so that mistakes corrected last time can be avoided next time. Only when users no longer need to repeatedly assign tasks, check results, and rework, can Agents truly save time for people.
Model progress can improve understanding and execution capabilities, and platform collaborations can bring more available services, but how to make good use of these capabilities still requires continuous exploration of products.
The next-generation Personal Agent not only needs the ecosystem to open up more possibilities for it, but also requires specific innovations one by one to make it truly easy to use.
This article is from the WeChat official account "Letter AI", author: Yuan Xinyue, published with authorization from 36Kr.