HomeArticle

From the black tape test to cross-device intelligent agents, two veterans in the voice industry have been waiting for ten years.

尔山2026-08-06 17:35
How can AI build full-scenario connections with users and continuously obtain context outside the chat window?

Years ago, the central control screen of an XPeng vehicle was tightly sealed off with black adhesive tape.

During that media test, the tester could neither see the interface nor touch the screen, and could only complete in-vehicle operations through voice commands. The test was ultimately a success. Zhao Hengyi, who was the head of XPeng's voice business at the time, later recalled: "In our understanding, the screen is only an auxiliary tool. The core is how to help users truly get things done through effective communication."

Back then, in-vehicle voice technology was still focused on solving basic problems such as wake-up, recognition, and fixed command execution. The black tape test reflected a touch of stubbornness from the product team, and in effect pushed a longer-term issue to the forefront: When the screen moves to the background, how do machines understand human intentions, and how do they respond to users' next needs?

Years later, large language models have brought this question back to the surface. Machines have begun to understand semantics, accept interruptions, and call tools, and AI is expanding from chat boxes into the real world. In 2025, OpenAI integrated the io team founded by Jony Ive and others into the company; Amazon and Meta later acquired ambient AI device company Bee and wearable device firm Limitless respectively.

The goal behind these tech giants' scramble for hardware is crystal clear — How can AI build full-scenario connections with users and continuously obtain context beyond the chat box?

"The connection between AI and people should cover all scenarios," Zhao Hengyi told 36Kr. When a person walks from their car to the office and then back home, the devices around them keep changing, but the same agent should still recognize them, retain previous memories, and call the capabilities of the devices in front of them.

The invisible and intangible voice has taken on the core role of connecting interactions in this process. It can move along with users, has fast input speed, and carries hesitation, certainty, pauses and emotions. Zhao Hengyi describes voice as "so natural like air and water", so it has long been underrated.

Hojo, founded at the end of 2024, is entering the market from this overlooked niche, positioning itself as an AI Stack for smart hardware.

Zhao Hengyi, Founder and CEO of Hojo, has previously worked on voice platforms, robots, smart cabins and AgentOS at LeEco, AISpeech, XPeng, NIO and SenseTime successively. Su Dan, Co-founder and Chief Scientist, has worked his way up from Baidu and Didi to Tencent AI Lab, experiencing the entire cycle of voice recognition evolving from near-field mobile phone use to far-field sound pickup, multimodal interaction and large voice models.

Zhao Hengyi and Su Dan have known each other in the voice industry for nearly a decade. The former has advanced along product, hardware and supply chain tracks, while the latter has delved deep into models, acoustics and engineering. It was not until the advent of the large model era that their career paths finally converged at the intersection of Hojo.

Zhao Hengyi on the left, Su Dan on the right

They want to jointly answer one question: After AI leaves the chat box, how can it enter real hardware and continuously understand a person there?

01 Two voice veterans finally meet in the large model era

Zhao Hengyi has been working in the voice field for nearly ten years.

Over the past decade, from the mobile internet and IoT era's smart speakers, to automotive intelligence, and then to today's large models, technological waves have changed several times. Surprisingly, his career path has never deviated from the voice track.

At AISpeech, he was in charge of the DUI platform and launched the conference business in 2018. When he moved to XPeng, he integrated voice capabilities into vehicles and robots. NIO's NOMI also made him realize that users would develop emotional attachment to an interactive object with voice, expression and personality.

Zhao Hengyi believes that the value of voice never lies only in input speed. Hesitation, pauses, certainty and emotions are all hidden in sound, and this information can hardly be fully preserved by a well-edited text.

This firm belief in voice originated from a sentence during his campus days — "Without voice, even the best drama cannot be presented perfectly." He loves music and Hi-Fi, so he gave up the visual research direction and entered the laboratory to study voice and audio. To some extent, this also laid the foundation for his later product philosophy: when AI enters the real world, the most natural and information-rich communication method between humans and machines is still voice.

Su Dan's career trajectory is more like an epitome of voice technology development.

He is a typical technical expert from top tech companies. When he was pursuing his doctorate at Peking University, he studied voice recognition. When he joined Baidu, smartphones had just brought voice search and input methods to real users. Later at Didi and Tencent, his work expanded from voice recognition to multimodal voice interaction and large voice models, focusing on solving problems such as multilingual integration, vertical domain optimization, and streaming full-duplex experience.

Voice recognition is one of the first fields where deep learning achieved large-scale industrial application. Now, large models have re-integrated the previously scattered capabilities of recognition, synthesis and conversation. This round of changes has impacted technical personnel at large companies far more fiercely than expected.

In fact, large models are changing the way these technical experts view organizational structures.

Previously separated technical directions are rapidly converging, and models, data, computing power and business need to be re-coordinated, but the resources of large companies are difficult to move synchronously with the technical route. Su Dan hopes to "deploy resources and advance work in a more flexible organization" and hopes that technology can access real applications earlier. Zhao Hengyi, who has also worked at several automakers, faced troubles at the organizational process level: when promoting cross-departmental projects in large companies, a lot of time was spent on explanation, coordination and resource acquisition, and some decisions took months to implement.

Their career paths intersected long ago, but they never actually worked together. After the rise of large models, models, data, products and hardware were pushed into the same system, and the two reached a clear consensus to start a business together: Integrate voice models, product experience and hardware engineering into one company, and build a complete AI Stack for smart hardware.

At the end of 2024, Hojo was registered and established. The division of labor after starting the business followed their respective strengths: Zhao Hengyi is closer to users, products, customers and the supply chain, while Su Dan is responsible for breaking down requirements into acoustics, model and engineering links. Su Dan commented on Zhao Hengyi's product intuition: "He will first clarify the needs of users and products, and then figure out what parts the technology needs to supplement."

This is also the first time the two have fully explained the name Hojo to the public: it comes from the first letters of the two founders' English names. The team later endowed HOJO with the meaning of Holistic Omni Joint Operator, which represents full-domain, full-modality, cross-device and interactive system. The name connects the division of labor of the two founders and the goal of the company.

Just like the division of labor between the two in this interview, Zhao Hengyi answers more questions about users, products, customers and the supply chain; when the topic turns to acoustic front-ends, model boundaries or hardware lifespan, Su Dan takes over. The two rarely express superfluous emotions, and usually directly break down problems into technical conditions and business constraints.

At the beginning of starting the business, the two decided to invest in ASR, small-size TTS and native full-duplex voice models at the same time. Su Dan attributed these three routes to "requirements forced out by the product side": devices need a set of stable, low-latency and cost-controllable model solutions. Among them, the native full-duplex voice emotional dialogue large model was first proposed and promoted by him. He believes that the next generation of voice interaction cannot stay in the turn-based mode of "you finish speaking - I recognize - I answer", but should allow users to interrupt at any time, so that the model can listen, think and speak at the same time.

The three model lines are not isolated from each other. They share the same data cleaning and augmentation pipeline, training framework and evaluation system, and ASR labeling and TTS synthetic data will also feed back to full-duplex training. Su Dan said, "The voice pipeline as an independent industry is disappearing, but voice capability, as the underlying organ of agents, will only become more important, especially in the two main directions of human-machine dialogue and content creation."

In other words, general models can absorb basic capabilities, while problems such as streaming transformation, vertical domain adaptation, large model lightweighting, and inference acceleration still need to be solved on real devices. Memory, task execution, apps and subscriptions will not appear automatically just by accessing an API.

For a startup, betting on three model lines at the same time is a heavy investment. But Su Dan's judgment at that time was that the stronger the general model, the narrower the space for selling ASR or TTS capabilities separately; what Hojo really needs to master is the part that is strongly coupled with hardware and product experience, as well as the long distance after the model enters the device.

In June 2026, Hojo open-sourced Hojo-ASR-V1, which ranked second in the actual measurement list of Hugging Face Open ASR Leaderboard, second only to Microsoft. Its small-size Hojo-TTS-Light-40M has only 40 million parameters, and the first sound response of the full-duplex model reaches the 100-millisecond level. These indicators only cover part of the technical chain, but what Hojo wants to solve is whether a piece of sound can be stably understood after entering the device, converted into actions, and precipitated into memory that can be called for the next interaction.

02 From one vehicle to more devices

Zhao Hengyi's vision for cross-device agents did not come out of nowhere, but evolved from his several experiences at automakers.

XPeng made him seriously think for the first time about how voice can enter the physical world. Both vehicles and robots require the system to understand users outside the screen, and the black tape test back then verified exactly whether voice could complete tasks independently.

At NIO, NOMI made him see that users will develop emotional attachment to an interactive object with voice, expression and personality. NIO was also developing mobile phones at that time, and Zhao Hengyi hoped that NOMI could appear on multiple terminals, allowing users to continue interacting wherever they went.

However, the capability boundary of an automaker can hardly extend to all life scenarios outside the vehicle. "If I have the opportunity to start a company, I should not be limited to the in-vehicle scenario," Zhao Hengyi recalled. It was at that time that he began to envision extending the AI-human connection from the cabin to broader life scenarios, which later became the starting point for the establishment of Hojo.

Later, when working on AgentOS, he further pushed this line of thinking to cross-device scenarios: The same Agent should know what the user has done on different terminals, and also understand whether the device in front of it can record, broadcast, or display text. After Hojo was registered in 2024, the team did not expand all scenarios immediately, but first returned to the most familiar automotive cabin to promote the first core project.

Cars are similar to a high-pressure laboratory for voice technology. There are multiple people talking, wind noise, tire noise and weak network in the car, and the system also needs to meet low latency and high stability requirements. Such an environment makes the capabilities of the two complement each other — Su Dan has long been dealing with far-field voice and complex acoustics, familiar with microphone arrays, front-end enhancement and robustness in noisy environments; Zhao Hengyi is more familiar with product definition, customer requirements and mass production delivery.

According to Hojo's disclosure, the company has delivered some functions to a large automotive customer, and some capabilities have been deployed in vehicles. Compared with quickly replicating more large customer projects, Zhao Hengyi prefers to make one key customer a benchmark first. If projects expand too fast, engineers will be split into different requirements, and R&D resources will be fragmented. For a startup with limited resources, making one core project a benchmark can test capabilities better than simply increasing the number of contracts.

Zhao Hengyi accumulated experience in integrating software and hardware in the consumer electronics field in his early years, and automotive projects further strengthened his capabilities in complex acoustics, weak network, low latency and mass production. Now these capabilities are being transferred to headphones, recording devices and portable hardware. The acoustic structure, Bluetooth link, firmware and App of various devices are different, but far-field sound pickup, low-latency interaction, cloud-edge collaboration and mass production experience can be reused.

The difficulty Hojo wants to overcome is how the audio collected by the microphone can be stably streamed to the App, whether users can start recording without taking out their mobile phones, and how the device can degrade when the network is weak or even disconnected. Zhao Hengyi said that accessing a large model API cannot solve these problems automatically.

In a sense, automobiles continue to serve as a high-standard verification field, while consumer electronics undertake the task of large-scale promotion. The latter has a broader market. IDC predicts that global shipments of wearable audio devices will reach 407.6 million units in 2026, and shipments of smart glasses will be about 13.6 million units, a year-on-year increase of 41.4%. Compared with the highly concentrated smartphone market, these categories are occupied by more brands. Many manufacturers are good at industrial design, supply chain and channels, but do not have the ability to independently build models, Agents, Apps and subscription systems.

Although mobile phones are the largest smart terminals, Hojo has strategically given up this market. Global smartphone shipments exceeded 1.2 billion units in 2025, and leading manufacturers control the operating system entry, accounts and user context, and most of their core AI capabilities are kept in-house. Hojo voluntarily gave up the entire mobile phone ecosystem, and focused on the scattered categories of wearable audio, recording and portable devices with a large number of brands.

03 Hundreds of millions of devices, only pick the most "high-value" voice

The hardware market can be very broad, but what Hojo wants to do is to narrow down its scenarios instead.

For scenarios, Zhao Hengyi repeatedly mentions the term "information density". The company did not expand product categories along the idea of "connecting everything to AI", but first calculated two accounts: How much effective information does one interaction transmit? Can the value obtained by users cover the cost of continuous reasoning and service? In other words, the form of devices is determined by brands and the supply chain, and Hojo cares more about what the microphone finally picks up, and whether this content is worth users paying for in the long run.

Smart home scenarios were the first to be excluded. Saying "adjust to 26 degrees Celsius" to an air conditioner only changes a switch or a value, which can be done by mobile phones, remote controls and automation rules. "If you only complete control through voice, it cannot bring high value to users, and it is essentially a one-time task," Zhao Hengyi said.

Cars also have vehicle control functions, but in Zhao Hengyi's view, the automotive cabin is an independent digital space, where music, search, content and services continuously generate context; touch screen operation is inconvenient during driving, and using mobile phones is even more difficult, so voice has more opportunities to become a high-frequency interaction entry.

The problem with AI toys lies in their life cycle. Children may lose interest after a few months, and the one-time selling price can cover the cost of this period; but model calls, storage and maintenance will continue. Zhao Hengyi said: "If the product is sold once, but services still need to be provided continuously afterwards, all costs will become operating costs; if users are not willing to pay, you will eventually face the harsh commercial reality."

This is exactly the fate that many AI hardware cannot get rid of. Hardware revenue is confirmed at the time of sale, but cloud reasoning costs accompany the entire life cycle. Humane AI Pin stopped selling less than a year after its launch. After the cloud service was shut down, core functions such as calls, messages and AI queries became invalid, exposing the life cycle risks that AI hardware highly dependent on cloud services may face.

For devices with no other source of income, subscription is a relatively clear cost recovery path. But the premise of all this is that the product can continuously provide clear value.

AI recording device company Plaud provides a referable market sample. The company disclosed in June 2026 that its services cover more than 2 million users, and its software subscription annual recurring revenue reached 100 million US dollars. This shows that in high information density scenarios such as meetings, interviews and knowledge work, hardware can become an entry point for continuous services. Users are willing to pay continuously, because after recording, there are transcription, retrieval, summarization, task disassembly and knowledge precipitation functions.

The scenario Hojo wants to anchor is exactly like this. Zhao Hengyi took this interview as an example