From the black tape test to cross-device intelligent agents, two voice industry veterans have waited for a decade.
Years ago, the central control screen of an XPeng vehicle was tightly sealed with black adhesive tape.
During that media test, the participant could not see the interface or touch the screen, and could only complete in-vehicle operations through voice commands. The test was successfully passed in the end. Zhao Hengyi, who was the head of XPeng's voice business at the time, later recalled: "In our understanding, the screen is only an auxiliary tool. The key is how to help users get things done effectively through smooth communication."
Back then, in-vehicle voice technology was still focused on solving problems such as wake-up, recognition and fixed command execution. The black tape test, which carried a hint of stubbornness from the product manager, actually pushed a longer-term issue to the forefront: When the screen is moved to the background, how do machines understand humans and respond to their next needs?
Years later, large language models have brought this question back to the surface. Machines have started to understand semantics, accept interruptions, and call tools, and AI is expanding from chat boxes to the real world. In 2025, OpenAI integrated the io team founded by Jony Ive and others into the company; Amazon and Meta later acquired ambient AI device company Bee and wearable device company Limitless respectively.
The purpose behind these giants competing for hardware is crystal clear — How can AI build full-scenario connections with users and continuously obtain context beyond the chat box?
"The connection between AI and humans should cover all scenarios," Zhao Hengyi told 36Kr. When a person walks from a car to the office and then back home, the devices around them keep changing, but the same agent should still recognize them, retain previous memories, and call the capabilities of the devices in front of them.
The invisible and intangible voice acts as the core of the interactive link in this process. It can move with people, features fast input speed, and carries hesitation, certainty, pauses and emotions. Zhao Hengyi describes voice as "so natural like air and water", so it has long been underestimated.
Hojo, founded at the end of 2024, is cutting into this inconspicuous niche, positioning itself as an AI Stack for smart hardware.
Founder and CEO Zhao Hengyi has previously worked on voice platforms, robots, smart cockpits and AgentOS at LeEco, AISpeech, XPeng, NIO and SenseTime in succession. Co-founder and Chief Scientist Su Dan has worked at Baidu, Didi and then Tencent AI Lab, experiencing the entire cycle of voice recognition development from near-field mobile phone use to far-field sound pickup, multimodal interaction and large voice models.
Zhao Hengyi and Su Dan have known each other in the voice industry for nearly a decade. The former has been advancing along product, hardware and supply chain tracks, while the latter has been delving into models, acoustics and engineering. Until the advent of the large model era, their career paths finally converged at the intersection of Hojo.
Zhao Hengyi on the left, Su Dan on the right
Together, they want to answer a question: After AI leaves the chat box, how can it enter real hardware and continuously understand a person there?
01 Two voice veterans finally meet in the large model era
Zhao Hengyi has been working on voice technology for nearly a decade.
In the past ten years, from the mobile Internet and smart speakers in the IoT era, to automotive intelligence, and then to today's large models, technological waves have changed several times. Surprisingly, his career path has never deviated from the voice track.
At AISpeech, he was in charge of the DUI platform and launched the conference business in 2018. At XPeng, he integrated voice technology into vehicles and robots. NIO NOMI also made him realize that users would form emotional stickiness to an interactive object with voice, expression and personality.
Zhao Hengyi believes that the value of voice never lies only in input speed. Hesitation, pauses, certainty and emotions are all hidden in the voice, and this information can hardly be fully preserved in a piece of organized text.
This certainty about voice comes from a sentence he heard in his campus days — "Without voice, even the best play cannot be presented perfectly." He loves music and Hi-Fi, so he gave up the visual direction and entered the laboratory to study voice and audio. To some extent, this has also laid the foundation for his later product philosophy: when AI enters the real world, the most natural and information-rich way of communication between humans and machines is still voice.
Su Dan's career trajectory is more like a microcosm of the voice technology industry.
He is a typical technical expert from large tech companies. When he was pursuing his doctorate at Peking University, he studied voice recognition. When he joined Baidu, smartphones had just brought voice search and input methods to real users. Later, at Didi and Tencent, his work expanded from voice recognition to multimodal voice interaction and large voice models, focusing on solving problems such as multilingual integration, vertical field optimization, and streaming and full-duplex experience.
Voice recognition is one of the first fields where deep learning has achieved large-scale industrial applications. Today, large models have re-integrated the originally scattered capabilities of recognition, synthesis and dialogue. This round of changes has impacted technical personnel in large enterprises far more severely than expected.
In fact, large models are changing the way these technical experts view organizational structures.
Previously separated technical directions are converging rapidly, and models, data, computing power and business need to be re-coordinated, but the resources of large companies can hardly move synchronously with the technical route. Su Dan hopes to "allocate resources and advance work in a more flexible organization" and allow technology to access real applications earlier. Zhao Hengyi, who has also worked at several automakers, is troubled by organizational processes: when promoting cross-departmental projects in large companies, a lot of time is spent on explanation, coordination and resource acquisition, and some decisions are implemented only after several months.
Their career paths have intersected long before, but they have never really worked together. After the rise of large models, models, data, products and hardware are pushed into the same system, and the two reached a clear consensus to start a business together: Integrate voice models, product experience and hardware engineering into one company, and build a complete AI Stack for smart hardware.
At the end of 2024, Hojo was registered and established. The division of labor after starting the business follows the areas that the two are good at. Zhao Hengyi is closer to users, products, customers and the supply chain, while Su Dan is responsible for breaking down requirements into acoustics, model and engineering links. Su Dan commented on Zhao Hengyi's product intuition: "He will first figure out the users and products, and then reverse deduce what parts the technology needs to supplement."
This is also the first time the two have fully explained the name Hojo to the public: it comes from the initials of the two founders' English names. The team later gave HOJO the meaning of Holistic Omni Joint Operator, which represents full coverage, full modality, cross-device and interactive system. The name connects the division of labor of the two and the goals of the company.
Just like the division of labor between the two in this interview, Zhao Hengyi answers more questions about users, products, customers and the supply chain; when the topic turns to the acoustic front end, model boundaries or hardware lifespan, Su Dan takes over. The two rarely express excessive emotions, and usually directly break down problems into technical conditions and business constraints.
At the beginning of starting the business, the two decided to invest in ASR, small-size TTS and native full-duplex voice models at the same time. Su Dan attributed these three routes to "requirements forced by the product side": the device needs a set of stable, low-latency and cost-controllable model solutions. Among them, the native full-duplex voice emotional dialogue large model was first proposed and promoted by him. He believes that the next generation of voice interaction cannot stay in the cycle of "you finish speaking - I recognize - I answer", but should allow users to interrupt at any time, so that the model can listen, think and speak at the same time.
The three model lines are not isolated from each other. They share the same data cleaning and augmentation pipeline, training framework and evaluation system, and ASR annotation and TTS synthetic data will also feed back to full-duplex training. Su Dan said that the "voice pipeline" as an independent industry is disappearing, but voice capability, as the underlying organ of agents, will only become more important, especially in the two main directions of human-machine dialogue and content creation.
In other words, general models can absorb basic capabilities, but problems such as streaming transformation, vertical domain adaptation, lightweighting of large models, and inference acceleration still need to be solved in real devices. Memory, task execution, apps and subscriptions will not appear automatically with a single API access.
For a startup, betting on three model lines at the same time is a heavy investment. But Su Dan's judgment at that time was that the stronger the general model, the narrower the space for selling ASR or TTS capabilities separately; what Hojo really needs to master is the part that is strongly coupled with hardware and product experience, as well as the long process after the model enters the device.
In June 2026, Hojo open-sourced Hojo-ASR-V1, ranking second in the actual measurement of the Hugging Face Open ASR Leaderboard, second only to Microsoft; its small-size Hojo-TTS-Light-40M has only 40 million parameters, and the first sound response of the full-duplex model reaches the 100-millisecond level. These indicators only cover part of the technical chain, but what Hojo wants to solve is whether a piece of voice can be stably understood after entering the device, converted into actions, and precipitated into memory that can be called for the next interaction.
02 From one car to more devices
Zhao Hengyi's vision for cross-device agents is not a sudden inspiration, but a continuous evolution from his experiences at several automakers.
XPeng made him think seriously for the first time about how voice can enter the physical world. Both cars and robots require the system to understand humans beyond the screen, and the black tape test back then verified whether voice could independently complete tasks.
At NIO, NOMI made him realize that users will build emotional stickiness to an interactive object with voice, expression and personality. NIO was also developing mobile phones at that time. Zhao Hengyi hoped that NOMI could appear on multiple terminals, allowing users to continue interacting wherever they went.
However, the capability boundary of an automaker can hardly extend to all life scenarios outside the vehicle. "If I have the opportunity to start a company, it should not be limited to the car interior," Zhao Hengyi recalled. It was at that time that he began to imagine extending the connection between AI and humans from the cockpit to a broader life scenario, which later became the starting point for the establishment of Hojo.
Later, when working on AgentOS, he further pushed this line of thinking to cross-device scenarios: The same Agent should know what the user has done on different terminals, and also understand whether the device in front of it can record, broadcast, or display text. After Hojo was registered in 2024, the team did not immediately expand to all scenarios, but first returned to the most familiar automotive cockpit to promote the first core project.
Cars are similar to a high-pressure laboratory for voice technology. There are multiple people talking, wind noise, tire noise and weak network in the car, and the system also needs to meet low latency and high stability. Such an environment allows the capabilities of the two to couple with each other — Su Dan has long been dealing with far-field voice and complex acoustics, and is familiar with microphone arrays, front-end enhancement and robustness in noisy environments; Zhao Hengyi is more familiar with product definition, customer needs and mass production delivery.
According to Hojo's disclosure, the company has delivered part of the functions to a large automotive customer, and some capabilities have been deployed on vehicles. Compared with quickly replicating more large customer projects, Zhao Hengyi prefers to make one key customer a benchmark first. If projects are expanded too fast, engineers will be split into different requirements, and R&D capabilities will be fragmented. For a startup with limited resources, making one core project a benchmark can verify capabilities better than simply increasing the number of contracts.
Zhao Hengyi has accumulated software and hardware integration experience in the consumer electronics field in his early years, and automotive projects have further strengthened his capabilities in complex acoustics, weak networks, low latency and mass production. Today, these capabilities are being migrated to headphones, recording devices and portable hardware. The acoustic structure, Bluetooth link, firmware and App of various devices are different, but far-field sound pickup, low-latency interaction, cloud-edge collaboration and mass production experience can be reused.
The difficulty that Hojo wants to overcome is how to stably push the audio collected by the microphone to the App, whether users can start recording without taking out their mobile phones, and how the device degrades when the network is weak or even disconnected. Zhao Hengyi said that accessing a large model API cannot solve these problems automatically.
In a sense, cars continue to serve as a high-standard verification field, while consumer electronics undertake the task of large-scale promotion. The latter has a broader market. IDC predicts that global shipments of ear-worn devices will reach 407.6 million units in 2026, and shipments of smart glasses will be about 13.6 million units, a year-on-year increase of 41.4%. Compared with the highly concentrated smartphone market, these categories are occupied by more brands. Many manufacturers are good at industrial design, supply chain and channels, but do not have the ability to independently build models, Agents, Apps and subscription systems.
Although smartphones are the largest smart terminals, Hojo has chosen to strategically abandon this market. Global smartphone shipments in 2025 exceeded 1.2 billion units, and leading manufacturers control the operating system entry, accounts and user context, and most of the core AI capabilities will remain internal. Hojo voluntarily gave up the entire mobile phone ecosystem, and focused on ear-worn, recording and portable devices with scattered categories and numerous brands.
03 Hundreds of millions of devices, only pick the most "valuable" voice
The hardware market can be very broad, but what Hojo wants to do is to narrow down the scenarios instead.
For scenarios, a word Zhao Hengyi repeatedly mentions is "information density". The company did not expand categories along the idea of "connecting everything to AI", but first calculated two accounts: How much effective information does one interaction transmit? Can the value obtained by users cover the continuous reasoning and service costs? That is to say, the form factor of devices is determined by brands and the supply chain, and Hojo cares more about what the microphone finally picks up, and whether this content is worth users paying for in the long run.
Smart home scenarios were the first to be excluded. Saying "adjust to 26 degrees Celsius" to an air conditioner only changes a switch or value, which can be done through mobile phones, remote controls and automation rules. "If you only complete control through voice, it cannot bring high value to users, which is essentially a one-time thing," Zhao Hengyi said.
Cars also have vehicle control functions, but in Zhao Hengyi's view, the automotive cockpit is an independent digital space, where music, search, content and services continuously generate context; touch screen operation is inconvenient during driving, and using mobile phones is also more difficult, so voice has more opportunities to become a high-frequency interaction entry.
The problem with AI toys lies in their life cycle. Children may lose interest after a few months, and the one-time selling price can cover the cost of this period; but model calls, storage and maintenance will continue to happen. Zhao Hengyi said: "If the product is sold once, but you still need to provide continuous services afterwards, all costs will become operating costs; users are not willing to pay, and in the end you will face the cold commercial reality."
This is the fate that many AI hardware can hardly get rid of. Hardware revenue is confirmed at the time of sale, but cloud reasoning costs accompany the entire life cycle. Humane AI Pin stopped selling less than a year after its launch. After the cloud service was shut down, core functions such as calls, messages and AI queries became invalid, exposing the life cycle risk that AI hardware highly dependent on cloud services may face.
For devices with no other source of income, subscription is a relatively clear cost recovery path. But the premise of all this is that the product can continuously provide clear value.
AI recording device company Plaud provides a referential market sample. The company disclosed in June 2026 that its services cover more than 2 million users, and the annual recurring revenue from software subscriptions has reached 100 million US dollars. This shows that in high information density scenarios such as meetings, interviews and knowledge work, hardware can become the entry point for continuous services. Users are willing to pay continuously, because after recording, there are transcription, retrieval, summarization, task disassembly and knowledge precipitation.
The scenario Hojo wants to anchor is exactly like this. Zhao Hengyi took this interview as an example: the information transmitted