HomeArticle

From Siri to Muse, how do we hand over our lives to AI step by step?

窄播2026-10-08 10:23
A relay.

Siri is catching up with the do engine vision it originally initiated. The newly emerging Muse products also need to go through a continuous process of improving reliability, security and cost-effectiveness.

Tech video creator Matt Robb listed a keyboard on Facebook Marketplace, and let Muse handle communications with potential buyers on his behalf. Then Muse independently accepted a lower offer, sent the buyer his home address, and arranged a time for the buyer to come to his apartment. It was not until a stranger showed up at his apartment door that Robb realized Muse had completed a whole series of decisions for him.

This is an AI-led operation that breaks through the screen and enters the real world in the form of a stranger standing at the user's doorstep. While this operation has sparked concerns over the loss of control of Muse's permissions, it also demonstrates the most appealing side of this type of product.

Less than a month after its release, some users are already using Muse to negotiate prices with telecom operators, cancel broadband services and claim flight compensation, while others are letting Muse purchase daily necessities and book cleaning services. Some specific items on personal to-do lists can now be directly completed by Muse and crossed off the list. This type of experience has attracted a large number of users for Muse. Sensor Tower estimates that by the end of September 2026, Muse's downloads have exceeded 5 million.

This way of gaining popularity is very similar to that of OpenClaw at the beginning of this year. Both are Agent products that bring shocking experience, and both have attracted widespread attention from all parties in a short period of time. The difference is that OpenClaw shows people that an Agent with memory, file permissions and message access can work continuously like a digital employee; while Muse demonstrates to the public a machine butler that can better integrate into shopping, travel, social interaction and family life.

Creating such a machine butler is a tech fable that runs through the development history of the Internet and AI. The independent Siri product launched in 2010 proposed to become a "do engine". After that, generations of personal assistant products have successively supplemented the capabilities of entry point, initiative, memory, understanding, reasoning, action and continuous operation under different technical conditions, aiming to bring the machine butler into reality.

It was not until the emergence of Muse that these capabilities were truly integrated, forming a consumer-grade product paradigm that can be promoted to the general public.

01

Do engine, not a chatbot

The idea and attempt to let computers accept entrustment and complete tasks like human beings appeared even earlier than the launch of the independent Siri product.

In 1987, Apple depicted a digital assistant in its concept video "Knowledge Navigator". The assistant had its own image, could find papers for professors, remind users to return calls, and continue to retrieve information during calls.

A few years later, this assumption was turned into a concrete product. In 1994, the personal assistant Wildfire could handle voice messages, dial numbers and manage voicemail on the phone; Portico, launched in 1998, already allowed mobile phones to connect to the Internet, and complete tasks such as reading emails and sending short messages through voice operation.

However, none of these products gained widespread popularity. The reason was that the technical conditions at that time could not support the emergence of a truly reliable assistant. At the same stage, Microsoft also launched Clippy, an assistant with a paperclip appearance, trying to provide active help for users. But because it always brought unnecessary disturbances to users, it was listed by Time magazine as one of the 50 worst inventions in history.

It was not until Apple officially released the App Store and SDK to the public in 2008 that the three founders of Siri saw an opportunity to try to create a machine butler again. Job applicants at Siri were required to read Michael Dertouzos' book "What Will Be: How the World Will Change When We Live in the Computer", because Dertouzos believed that devices should serve people, instead of people serving devices. People who did not agree with Dertouzos' views were considered unsuitable to work at Siri.

In 2010, the independent Siri product was launched. What fascinated Steve Jobs about it was not only that it was good at chatting, but also that it could complete specific tasks on behalf of users. Users could directly say to Siri "Help me find a romantic Italian restaurant near my office", then Siri would integrate information from Citysearch, Gayot, Yelp, Yahoo! Local, AllMenus.com, Google Maps, BooRah and OpenTable to give the answer.

At that time, ordinary users regarded this as another type of search engine service. But both the founders of Siri and Steve Jobs firmly believed that Siri was a do engine rather than a search engine.

Adam Cheyer, co-founder of Siri, later summarized the core capabilities of the product as: "Text understanding and information integration." He also recalled that when Siri was first launched, the team made several magazine cover-style images, one of which showed Siri squeezing and deforming Google's logo, implying that it would defeat Google in the future.

When asked whether the acquisition of Siri was to enter the search business, Steve Jobs clearly replied: "They are not a search company. They are an AI company. We have no plans to enter the search field." That is to say, Apple was not buying another search box, but an intelligent layer that can understand users and schedule the capabilities of mobile phones.

In 2011, the acquired Siri became a system function with the launch of iPhone 4S. Many ordinary users experienced for the first time that they could directly say a task goal to their mobile phone to complete the operation, instead of opening the app first, looking for buttons, and then completing the operation step by step. Its personalized tone also made the communication between humans and machines less rigid.

On Siri in 2010 and 2011, we could already see the embryonic form of a machine butler. It could understand natural language, call third-party services, control hardware devices, and had a personalized tone. The founders of Siri even envisioned a business model where it could charge service fees by selecting restaurants for users. This is not just an interaction method, but a potential business entry point.

02

Become an entry point, but trapped in a pre-set world

Steve Jobs might have seen the value of Siri as an entry point, but he passed away before he had time to realize it. After that, Siri experienced a strategic deviation inside Apple. Bill Stasior, who took over the management of Siri in 2012, came from Amazon's search department, and tended to build Siri into a search engine to aggregate Apple's search capabilities with Siri. The third-party ecosystem construction that Steve Jobs promised when acquiring Siri did not surface until 2016.

At this time, Amazon's Alexa had risen rapidly relying on the Echo smart speaker, bringing personal assistants into the living room and to the position of business entry point.

When Amazon launched Echo in 2014, what it really bet on was not a speaker, but to make Alexa eventually become a unified entry point in the home through the speaker. Far-field microphones, wake words and always-on status allowed users to summon Alexa without picking up their mobile phones. Then Amazon also opened the Skills platform, integrating music, shopping, calendar, smart home and third-party services into it.

Users could directly say a sentence to Alexa to play music, control home appliances, check the weather, or call a certain life service. This was originally Siri's own vision for itself, but Amazon turned it into a mature product experience faster.

Subsequently, Google followed up and launched Assistant, aiming to provide users with an "ambient experience that exists across devices and scenarios". Google Assistant can appear on mobile phones, speakers and chat applications, understand the world where users are located with the help of search, maps and knowledge graph, and even call restaurants and barbershops, talk with real people, handle schedule conflicts and complete reservations.

However, Apple's Siri team was not required to cooperate closely with the Beats team to integrate the smart voice assistant service into HomePod until a meeting in 2015. In the same period, with the ecological advantages of Apple, Siri quickly accessed Mac, Apple Watch and Apple TV, trying to become the "voice operating system" of Apple's hardware. But at this time, Siri had lost the openness it had in its startup period.

Domestic teams also learned from Amazon to build new entry points — if voice can replace menus and app icons, personal assistants may restructure the distribution methods of content, home furnishing and services. Therefore, Xiaomi connected mobile phones, TVs, speakers and IoT devices through Xiao Ai; Tmall Genie tried to bring e-commerce, payment and content into families; Xiaodu competed for users with low-cost speakers and screen-equipped devices.

The increasing number of peers does not mean that personal assistants have become mature. The early Siri relied on a large number of vocabulary libraries and manual rules, adding new expressions might take weeks, and complex functions even required a development cycle of nearly one year. In order to let Alexa complete food ordering, ride-hailing or device control, the platform needed to develop a separate Skill for each service. Every time a new service is accessed, engineers have to design processes, handle authorization and exceptions.

More and more pre-settings were added, but the flexibility and freedom of the assistant were still not unlocked. After leaving Apple, Siri founders Adam Cheyer and Dag Kittlaus founded Viv Labs, trying to combine services temporarily according to the questions raised by users through "dynamic program generation", instead of writing pre-set processes for each demand. This was already close to people's vision of Agent today, but limited by technology, it did not achieve the ideal effect.

At the same time, technical limitations also reduced the permissions that platforms were willing to open. The consequences of playing the wrong song are limited, but booking the wrong flight ticket, sending the wrong message or completing the wrong payment will generate real responsibilities. The weaker the understanding ability, the more cautious manufacturers are to open permissions; the more restricted the permissions, the harder it is for assistants to obtain real ability to handle affairs.

Personal assistants at this stage are in a contradictory development state: they are more active than mobile apps, but more mechanical than real machine butlers; they occupy more and more entry points, but can only act in the pre-set world.

03

Generative AI breaks the pre-set bottleneck

At the end of 2022, ChatGPT broke the development bottleneck of personal assistants.

It did not take over mobile phones or smart home devices, but for the first time let the public see that AI can continue to talk around an open question, accept follow-up questions, explain reasons and generate new content.

Old assistants needed to put users' expressions into pre-set categories first, while large models can understand goals, constraints and context, identify intentions, and have capabilities closer to "understanding". Users no longer need to use fixed sentence patterns familiar to the system, nor do they need to split the problem into several standard steps first. The machine can first understand what people really want to do, and then decide how to respond.

After that, OpenAI demonstrated two attempts to evolve towards personal assistants: one is that GPT-4o further reduces voice latency, supports interruption, visual recognition and real-time feedback, turning voice from an input method into an interface that shares the on-site environment with users; the other is Plugins that connects the model to search and third-party services, enabling the model to perform limited real-world actions after being authorized.

The presentation of this possibility also accelerated the exit or iteration of the previous generation of personal assistants.

In 2023, Microsoft stopped supporting the independent Cortana application, and directed users to Copilot on the page. This is not just a product offline, but a paradigm shift: Cortana was built around wake words and fixed skills; Copilot takes the large model as the understanding layer and enters Windows and Microsoft 365.

Apple also tried to rebuild Siri. The new Apple Intelligence proposes screen perception, personal context and cross-app operation, hoping that Siri can find information from emails, photos and messages, and then complete actions on behalf of users. However, the delayed launch of capabilities exposed Apple's insufficient capability reserve. Understanding complex requests in the Demo and executing requests safely and stably on hundreds of millions of devices are not the same problem.

In China, products such as Doubao, Qwen and Kimi first completed the popularization of generative conversation, and mobile phone manufacturers are also continuously upgrading their own assistant products. They brought search, writing and multimodal interaction into the mass market, but at this stage they mainly generate responses, without entering into specific tasks.

The launch of DeepSeek R1 strengthened the product imagination of "reason first, then act", and quickly made reasoning capabilities penetrate into personal assistants. For personal assistants to handle affairs, it is inherently necessary to understand a sentence, sort out constraints, disassemble steps, compare solutions before taking action, and then adjust the plan according to environmental changes. The penetration of reasoning capabilities provides personal assistants with capabilities close to "thinking".

When both understanding and thinking capabilities are available, various companies began to find the specific landing point of the new generation of do engine along the resources they have mastered.

Alibaba connected Qwen to services such as Taobao, Alipay, Fliggy and Amap; Tencent is exploring the combination of Agent and super app entry point through WeChat Xiaowei; Doubao has developed new features such as outfit recommendation and exhibition tour explanation based on video call capabilities. Alexa+ recombines generative conversation, memory, family calendar and smart home, and connects services such as OpenTable, Uber Eats and Spotify.

However, whether it is Plugins or Alexa+, actions still mainly occur in the already connected services. Generative AI can discuss almost any task with people, but it does not yet have a pair of hands that can enter the real digital world.

04

Agent makes personal assistants start to take action

The next problem becomes specific: personal assistants can already discuss and think about a task with the help of large models, so how to make them really operate software, use accounts, continue to work after users leave, and get rid of pre-settings for actions.

A series of attempts centered on execution capabilities began to appear intensively.

Anthropic's Computer Use and Doubao Mobile Assistant no longer wait for third parties to provide APIs, but let the model understand the interface through screenshots, then click, scroll and input like a human. It is equivalent to letting machines enter the digital world along the path that humans have already taken. OpenAI Operator allows the model to complete form filling, product purchase and service booking relying on the browser. Manus further lets tasks run in an independent cloud environment.

With these explorations, the execution unit of personal assistants has begun to change from "calling a function" to "moving between multiple tools around a goal", turning the question-and-answer entry point into a work entrustment entry point.

OpenClaw pushed the execution paradigm to the stage of truly executable tasks. Users can hand over files, browsers, emails, calendars and message entry points to a continuously running Agent, letting it receive tasks in the background, save memories and call tools repeatedly. It proves that people are willing to entrust more things to Agent, but it still lacks the basic guarantee for popularization among the public.

The opportunity to encapsulate this capability into mass products falls on two paths. One path is to create Agent products in productivity scenarios, enabling them to write articles, make PPTs and write codes, including Codex, Qwen Office, WorkBuddy and Doubao Work that have chosen this path. The other path is the personal assistant path that Muse has taken, which is more focused on life and consumption scenarios.