HomeArticle

AI hardware is waiting for the "unsung heroes behind the scenes"

具身研习社2026-10-09 17:52
Nobody cares which socket the electricity enters the room through.

AI hardware is rewriting its own competitive unit.

Glasses, pendants, earphones, rings, desktop robots, AI toys, vehicles, speakers, and even desk lamps have all started to access intelligence. Their forms keep expanding outward, yet the "next iPhone" that has been lingering over the industry for the past two years has gradually faded into the background.

Recent new moves from Apple have made this change more noticeable. The new smart home hub has been brought to the forefront, while devices such as door locks, doorbells, thermostats, and cameras continue to spread across every corner of the room. A company that is best at integrating software and hardware into a single product has not tried to make one machine handle all tasks this time.

Devices keep expanding outward, and the center is beginning to shift from the devices themselves to the layer above them.

This perfectly reflects the current state of AI hardware: no one knows what the end point looks like, so people first let the answers diverge.

This may not be a well-considered plan, but it is at least more honest than pretending to know the answer.

The proposition that used to be pinned on one "ultimate device" has been split into many lighter, more specific rounds of trial and error.

However, this dispersion also exposes new gaps. While each piece of hardware can independently gain perception, memory and action capabilities, an individual's identity, context and permissions are still fragmented across different systems. Every new touchpoint may generate an additional record of the same person, a new set of accounts, and a new ecological wall.

The entry point is spreading from a single device into a network. Perhaps what was really miscalculated in the past was not the device, but the hierarchy.

What really needs to be centralized is not necessarily a certain form of hardware, but the layer above the devices. It does not matter whether it is finally called a hub, a super app, or something else.

What matters is that this layer has not truly taken shape yet.

The net has been cast, and the next round of competition is starting at the end where people pull the net back.

Hello, this is AI Hardware Review.

When everyone is chasing the answer, we will first see the problem clearly.

When They Choose "Spread-the-Net Betting"

Dispersion is not a pre-designed route, but the result of a collective shift in direction.

Before the shift, the industry actually shared the same imagination: a new computing revolution should be wrapped up by a new personal terminal.

This imagination largely comes from the successful experience left by smartphones. Over the past two decades, the iPhone has successively integrated communication, maps, payment, entertainment and internet services into one screen, leaving a deep mindset in the hardware industry: there should always be a central entry point in an era, and the entry point is preferably a device.

Therefore, when the wave of large models came, almost everyone's first reaction was the same question: what will the iPhone of the AI era look like?

Humane AI Pin and Rabbit r1 are the two most serious respondents to this question. Both set out with the ambition to "reinvent the personal terminal", and both paid a high price. Humane has exited the market, while Rabbit completed a meaningful transformation in September this year: OS3 has been moved to the cloud, allowing users to access the same account via Web, Telegram and the r1 device, and it can also connect to multiple computers to perform tasks.

Rabbit's evolution from r1 to OS3

A piece of hardware that once tried to be the answer has now become part of the answer.

Their experiences have made the entire industry see one thing clearly: new hardware itself will not automatically become the new computing center just because it is connected to AI.

At least up to now, no one has solved the problem of a single entry point.

But the industry has not abandoned hardware because of this.

Meta is still betting heavily on smart glasses, and OpenAI is still looking for new personal devices. What has really changed is that the tight binding between hardware and intelligence has begun to loosen.

Muse can continue to access new devices from mobile phones, computers and the Web. Qualcomm has directly proposed that the future will be more "agent-centric, rather than device-centric", and Alexa+ has also started to appear in TVs, cars and speakers. More and more companies are beginning to accept that AI can have many landing points, without having to pin all its value on a single shell.

The pace of launching new hardware has not stopped, but the mindset of betting on hardware has changed.

It is precisely at this moment that trying a new position is becoming cheaper and cheaper.

Capabilities of models, voice, multimodality and Agents are becoming infrastructure that can be purchased and combined, and upstream chips and platforms are also covering more and more terminals. Volcano Engine disclosed that as of June 2025, the shipment volume of AIoT products connected to the Doubao large model has exceeded 1 million units; JD's JoyInside has also connected to more than 40 toy brands.

When "installing AI on hardware" is no longer a R&D capability but a purchasable and combinable infrastructure, the strategy naturally changes from "think clearly before acting" to "try everything out".

As a result, the industry's method of finding answers has changed.

The allure of the "next iPhone" has not disappeared, but it is no longer the only way to bet.

Almost all major manufacturers now adopt the standard strategy: one hand bets on possible central devices, and the other hand spreads touchpoints to more real-world locations.

The gambling table is still there, but no one goes all in on the same card anymore.

Every Piece of Hardware Is a Questionnaire

If you look at these scattered pieces of hardware as products, most of them still look rough. But if you treat them as questionnaires, they become quite interesting.

When the ultimate form is temporarily unsolvable, these products begin to take on another role: splitting the big question of "how humans and AI should get along" into dozens of small experiments, and putting them into real life for testing. Each product is asking a question on behalf of the industry, and the reviewers are the users.

The first thing that has been tested out is the boundary of perception.

If AI wants to get more context, the hardware needs to be closer to people. Note-taking ornaments, voice recording pendants, and smart glasses are essentially fighting for the part of life that machines could not access in the past. But the problem quickly shifts from "can it record" to "where should it record".

Limitless added Consent Mode to its recording function, and DayoTag emphasizes local storage. While Meta continues to make glasses with cameras, it also launched the Ray-Ban Meta Audio that completely removes the camera, leaving only sound and AI. Even though technical capabilities can move forward further, products have taken the initiative to step back.

This kind of concession is more worth noticing than adding a new sensor.

The first-person perspective is certainly precious, and continuous recording can also allow the model to get more context. But after a piece of hardware truly integrates into life, users are not the only participants. Strangers on the street, colleagues in the office, and friends sitting opposite may all be accidentally included in the data scope of the device. How much AI can see and hear will ultimately not only be determined by sensors, but also by how much people and the surrounding environment are willing to open up.

Therefore, the stronger the perception, the more important the boundary becomes.

Moving forward, the question becomes: why can an AI hardware stay in people's lives for a long time?

Companion products most often take "how smart the model is" as their selling point, but actual use is constantly pushing the evaluation criteria to other directions. Ropet disclosed that its 90-day retention rate reaches 80% to 90%; LOVOT's long-term user data spans up to three years. Both of them remind people of one thing: the real difficulty for companion products is not to make people feel interesting for the first time, but to remain worth existing after the novelty fades.

This also explains why some products have taken the initiative to reduce conversations.

LOVOT does not rely on complex language to maintain relationships. bibo gives more interactions to eyes, postures and touches, and YareLampGo expresses its status by turning its head and changing lights. They are testing a very specific thing: the AI interface does not necessarily have to be a chat box.

Official websites of LOVOT and Ropet

Companionship in real life does not always require high information density. One action, one response, or even keeping quiet at the right time may be more suitable for long-term coexistence than saying more paragraphs. For this type of product, intelligence is a capability, while low burden is closer to the condition for the product to be viable.

This is also an easily overlooked change in this round of exploration: AI hardware is not only looking for new device forms, but also re-searching for the interaction language of AI.

Going further, the question shifts from coexistence to action.

Google's new Nest camera can already answer "what time did the garbage truck come on Thursday", Alexa+ has been integrated into BMW vehicles, and Muse is preparing to shop and book tickets for you and earn commissions from it.

If you answer the previous questions wrong, you lose part of the experience. But if you answer this question wrong, you hand over the keys.

None of these questionnaires have been completed, but the outline of the boundary has been tested out. Real scenarios let wearing duration, usage frequency, retention rate and continuous authorization filter the answers for the industry.

And the exploration itself thus has another layer of value.

Many products today may not eventually become a stable category, but what they have tested out may not disappear together with them. How continuous memory should be collected and called, how to let people perceive the status of the Agent, whether actions can become the interface, and at which step permissions should be returned to humans, these experiences will be separated from specific products and gradually become reusable capabilities for later products.

The early days of the mobile internet also went through a similar process. A large number of apps disappeared, but some interaction methods that originally belonged to individual products remained, and finally became the language that everyone understands by default.

AI hardware is probably at such a stage now.

On the surface, everyone is making products, but in fact, everyone is looking for a set of interaction primitives in the real world for the yet-to-appear AI hub system.

When memory, perception, expression and action are respectively placed in different devices, the questions that a single product can answer will reach their limit here.

Who Will Become the "Super App" Behind All AI Hardware

This dispersion does not mean that no one has tried to bring everything together.

Within the walled garden of a single manufacturer, "the same AI working across devices" has already happened. Siri's conversation history can be continued across different Apple devices via iCloud; Meta is also allowing Muse to continue to access new hardware entry points from mobile phones, computers and the Web. The device can be changed, while accounts, sessions and part of the context start to stay at a higher level.

Inside the wall, the connections have started to be established, and the real breakpoint appears after stepping out of the wall.

A person may soon have multiple AIs: one in the glasses, one in the earphone, one in the watch, one in the car, and one at home. They are all accumulating information about you, but each holds its own version.

Every time you change a device, you have to introduce yourself all over again.

They all know you separately, but none of them know each other.

The missing layer here can actually find a reference from the super apps in the mobile phone era. WeChat does not care what brand of mobile phone you use, it cares whether your chat, payment and travel all happen inside it.

Now, similar logic is stepping out of the screen. The so-called "hub" of AI hardware is closer to this kind of continuity: identity, memory and permissions stay at the upper layer, and different devices and Skills are called according to tasks.

In the past, super apps aggregated digital services. Personal AI goes one step further, and it needs to organize the information, services and action capabilities scattered in different scenarios.

Bill Gates made the same judgment three years ago: Android, iOS and Windows are all platforms, and Agent will be the next platform. When he said that, he was still talking about software. Today, this judgment is further extending into the hardware field.

It is precisely here that the problem gets stuck.

Matter solves the problem of "devices can be connected", and MCP starts to solve the problem of "models can call tools". But at the higher layer, there is still no public answer to questions like which device has the right to know what at the moment, which memory should be inherited, and how far AI can act on behalf of users.

What is more troublesome is that this gap is not unseen by everyone, but everyone wants to fill it in their own way.

Whoever masters the continuously growing user memory will master the starting point of the next service. In other words, it is very likely that whoever masters this hub layer will hold the entry point of the next era.

Few manufacturers are willing to easily hand over such assets to a public platform. So everyone chooses to build their wall higher first, and leave interoperability to the future.

As a result, "small-scale hubs" appear earlier.

The home is a very typical space. There are enough dense devices, and the boundaries of accounts and space are relatively stable.

Xiaomi disclosed that the number of its global IoT connected devices has exceeded 1.1 billion. Miloco 2.0 starts to organize these devices at a higher level: using cameras as environmental input, MiMo is responsible for understanding, then the Agent calls Xiaomi Smart Home devices, while adding identity recognition, family memory and long-term tasks.

It is no longer just "controlling a home appliance with one sentence", but trying to let the system understand the goal, and then organize a set of devices to complete it.

Lu Weibing, President of Xiaomi, recently talked about Xiaomi Smart Home and said: "No matter how responsive it is, it is essentially a 'remote control'. The next step is to move towards active intelligence."

Xiaomi is far from building a universal Personal AI hub, but the home scenario at least points out a possibility: the hub does not necessarily grow out of a portable device, and it may first grow out of a space.

Xiaomi Miloco 2.0 

This is the real cost of dispersion. The more dispersed the forms are, the easier the context will be broken, and the more expensive the absence of the hub will be.

The truly scarce thing in the industry next will also shift its position. In addition to stronger models and newer hardware forms, the continuity of context and the action authorization built on this continuity will become more and more important.

The business logic will also be rewritten by this change.

In the past, hardware was a one-off deal: after the device was sold, the relationship ended. Now the device is only the beginning of the relationship, and the real value-added part is the increasingly accumulated memory in the cloud. Hardware makes money once, while memory makes money in the long run.

What will really widen the gap in the next stage of AI hardware is no longer just who can make the next hit hardware, but who can let an AI continuously live across different hardware.

Looking back at the history of hardware, humans have always been adapting to objects: learning keyboards, learning touch screens, learning to translate intentions into buttons.

This round of dispersion is the dust