Developing voice interaction solutions tailored exclusively for Asian working professionals, ElevTalk has closed its Pre-Seed round of financing.
ElevTalk, a company focused on multilingual AI voice interaction, recently completed its Pre-Seed round of financing, with the financing amount undisclosed. Investors in this round include Patriot Fund (jointly operated by Bass Ventures and EO Studio), Xiaoxiao Fund, and a founding member of Baidu. The funds raised in this round will be mainly used for the iteration of core multilingual AI voice technologies, product polishing, and market expansion in North America and Asia.
ElevTalk is mainly targeted at professionals who work in multiple languages, with its initial focus on Asian professionals. Taking voice input as the entry point, the product integrates original language transcription, multilingual mixed input, expression optimization and cross-language translation into a single workflow, and has been further optimized for languages including English, Chinese, Japanese, Korean and Vietnamese.
Different from traditional voice input products, what ElevTalk aims to solve is not only "converting voice to text", but also enabling non-native English users to complete communication directly with their most natural language and expression patterns.
AI voice enters a new round of growth, while gaps still exist in the multilingual workplace scenario
With the development of speech recognition, large language models and real-time voice technologies, voice is evolving from a traditional dictation tool to a more important human-computer interaction entry in the AI era. The capital market has also begun to verify the value of voice input as an independent product form. Public reports show that the US AI voice input product Wispr Flow has a valuation of nearly 2 billion US dollars.
However, the ElevTalk team has observed that at this stage, the world's mainstream voice products are still mostly developed around English usage scenarios, while a large number of professionals in the real world are not native English speakers. Especially for Asian users, the language state in daily work is often not a single language.
For example, when a Chinese entrepreneur replies to an American investor, they may use Chinese to organize the main content, interspersed with English company names, product names and industry terms; even if they use English throughout the whole process, Asian accents, personal names and technical terms are prone to frequent errors in speech recognition. Similar scenarios are also widely found in South Korea, Japan and other non-English speaking markets.
ElevTalk's two co-founders Dai Zheng and Fu Boming, as well as many entrepreneurs and investors around them, are heavy users of voice input, and this usage experience became the starting point for the team to first identify the problem.
Before founding ElevTalk, Dai Zheng was engaged in product strategy work at a leading AI voice company, and had access to the voice models, products and supplier ecosystems in different regions of Asia and other parts of the world. The team observed that the capabilities of globally universal voice models are improving rapidly, but the technical supply for different Asian languages is still relatively scattered. When users enter scenarios such as multilingual mixed expression, accented English and cross-language office work, products often need to handle multiple links including recognition, translation, expression optimization and context understanding at the same time.
Therefore, the team formed a judgment: the competition of next-generation voice products is not only about "who can dictate English more accurately", but also about who can understand global non-native English users more naturally. ElevTalk chooses to cut into the market from the group of Asian professionals, who have concentrated demands and high usage frequency.
Starting from the workflow, hundreds of testers were acquired within 3 days at the Demo stage
Focusing on the above scenarios, ElevTalk currently provides capabilities such as voice input, expression optimization and cross-language translation, and supports multilingual mixed input.
Users can either directly use English for voice input, express in their mother tongue such as Chinese and Korean, or mix multiple languages in one sentence. After recognition, ElevTalk can also perform expression optimization or cross-language conversion according to user needs, making the results suitable for emails, instant messaging, social media and other office scenarios. As an alternative to native keyboard input, users can get the desired input content in any App without switching interfaces.
For example, Chinese users can directly speak a piece of content in Chinese mixed with English company names and technical terms, and then ElevTalk generates expressions suitable for English work scenarios, without requiring users to translate by themselves first and then use AI tools for secondary polishing.
For Asian users who use English directly, ElevTalk focuses on optimizing scenarios where traditional voice input tools are prone to errors, such as Asian accents, personal names, company names and technical terms.
At the technical level, ElevTalk is currently focusing on investment in Asian language speech recognition, multilingual mixed input, cross-language intention understanding and low-latency interaction.
According to the company's internal unified test, the Character Error Rate (CER) of ElevTalk in Korean and Chinese speech recognition tasks is 1.55% and 2.25% respectively; under the same test set and test conditions, its Korean and Chinese recognition indicators are lower than those of the world's leading voice input products such as Wispr Flow.
ElevTalk's earliest demand verification even took place before the formal product was formed. After the first version of the Demo was completed, the team released product information on social media. At that time, the product had not yet reached the state of a complete commercial product, but within 3 days after its release, hundreds of users took the initiative to contact the team and hoped to participate in product trials.
These early users share relatively concentrated common characteristics: many of them are Chinese and Korean entrepreneurs, investors and technology practitioners working in the United States, who frequently need to switch between English and their mother tongue. This feedback prompted ElevTalk to further focus on multilingual professionals, rather than developing a general dictation product for all users.
At present, without paid market promotion, ElevTalk has accumulated more than 1,000 early users. According to data provided by the company, core users are actively using it for 5 to 7 days on average per week, and the most active users input more than 6,000 words by voice in a single day; the week-over-week retention rate of weekly active users is about 50%.
For the team, a more important indicator than the number of registered users at this stage is whether users continue to use ElevTalk in real work processes such as emails, instant messaging, content creation and AI interaction, and gradually take voice as their default input method.
Asian users are just the starting point, targeting the global non-native English-speaking market
In terms of business model, ElevTalk currently takes personal subscription as the main commercialization direction, and first serves professionals who frequently use voice input and have cross-language communication needs.
The team's long-term judgment on the voice market does not stop at "improving typing efficiency". As a human-computer input method, the keyboard has been used for more than a hundred years, but from the perspective of human communication habits, speaking is always a more natural way of expressing information. In the past, restricted by recognition accuracy, latency and context understanding capabilities, voice has always been more of an auxiliary input tool.
With the development of large language models and new-generation voice models, voice recognition can now further complete intention understanding, translation, expression optimization and even task execution. As a result, voice has the opportunity to evolve from a pure input tool to a more complete human-computer interaction interface.
ElevTalk chooses Asian professionals as its first user group, on the one hand because this group has more concentrated multilingual demands, and on the other hand it is related to the team's own capability portfolio.
ElevTalk has Dai Zheng as CEO and Fu Boming as CTO. Both of them graduated from the School of Mathematical Sciences of Peking University with bachelor's degrees.
Dai Zheng is a Schwarzman Scholar at Tsinghua University, holding a master's degree in management. He used to work at McKinsey and has experience in AI voice product strategy. Previously, Dai Zheng founded an edtech company with cumulative revenue exceeding seven-digit RMB. At present, he is mainly responsible for ElevTalk's product strategy, market entry and global user expansion.
Fu Boming has 7 years of experience in voice and multimodal model training. He once served as a partner at an AI startup supported by frontline VC, and has work experience in large technology companies, having participated in the development of products with millions of users. At present, he is mainly responsible for the construction of ElevTalk's model, engineering and overall technical system.
Product insights, industry and market judgment, Go-to-Market experience, as well as long-term multimodal technology accumulation, constitute the current capability portfolio of ElevTalk's founding team.
Dai Zheng said that Asian users are the starting point of ElevTalk, not the end point. The team's long-term goal is to solve the problem of how global non-native English users can interact with computers more naturally through voice, and gradually expand from voice input to a more complete AI workflow.
Hyungjun Yang, Director of Bass Ventures, and Taeyong Kim, CEO of EO Studio, investors of this round from Patriot Fund, said: "Voice is becoming a new interaction entry in the AI era, but the existing mainstream products are still mainly built around native English users. A large number of professionals around the world switch between English and their mother tongue every day, and there are still obvious product gaps in multilingual mixed input, accent recognition and cross-language expression. The ElevTalk team not only has a personal understanding of this problem, but also demonstrates fast product development and iteration speed. Asian users are just a starting point, and the team is actually facing the huge global non-native English-speaking population. We expect ElevTalk to start with voice input and gradually grow into the Business OS in the Voice era."
After the completion of this round of financing, ElevTalk will continue to invest in product and core technology iteration, and further expand the markets in the United States and Asia.