Three "voice-first" educational AI companies have successively secured financing. Why have they all started to let AI "teach by speaking out"?
Recently, an interesting common trend has emerged in the educational AI field: a growing number of startups have begun to emphasize "voice-first", putting "speaking" at the core of AI tutoring products.
In late September, Eevi, an AI language learning company based in Amsterdam, the Netherlands, completed an undisclosed Pre-seed funding round led by Rockstart. Targeting international working professionals, it positions itself as a "voice-first" AI language tutor, aiming to solve a typical problem in language learning: users may have mastered a large number of words and finished plenty of exercises, but still get stuck when they actually try to speak out.
Eevi refers to this phenomenon as "Fluency Illusion", namely the illusion of fluency — seeming to have mastered a language does not mean being able to actually use it. Its product allows users to have real-time voice conversations with AI from the very start, instead of making users constantly click on multiple-choice questions, memorize words or input text.
Coincidentally, almost at the same time, two other AI tutoring products are also following a similar path.
Aristotle from the United States also announced in September this year that it had raised $5 million in a seed round of funding led by True Ventures. Targeting students aged 13 to 18, it defines itself as a voice-first AI tutoring platform. Before its official launch, Aristotle has already allowed more than 1,000 students to complete over 1,500 hours of one-on-one AI tutoring.
Indian AI education company YoLearn.ai secured an undisclosed seed round of financing in early September this year, which was led by ABP Education, an educational arm of the well-known Indian media group ABP Group. It combines real-time voice conversations with interactive whiteboards, enabling AI teachers not only to "speak" but also to draw and work through problems on the screen. At present, it covers Indian students as well as exam scenarios such as JEE and NEET, and supports 22 languages including Hindi and Hinglish.
Three products, one for language learning, one for K12 tutoring in India, and one for youth education in the United States, all put "voice-first" at the forefront of product design.
This may be an interaction transformation taking place in AI education.
General large language models may also have voice features, but they only add the "voice" function as an extra. Products like Eevi, YoLearn and Aristotle, by contrast, design their entire learning process around voice interaction from the very beginning.
TCOH, a practitioner in the edtech industry, shared with Duozhi: "Voice interaction is the most commonly used and most efficient way of interaction between people, and it is also the most frequent and natural interaction method in educational scenarios."
Real teaching is never a simple one-question-one-answer mode. Teachers will ask follow-up questions, interrupt, and judge whether students have really understood the knowledge according to the tone of their speech; students will also expose their knowledge gaps through speaking. In many cases, being able to "write something out" and being able to "say it out" are two completely different abilities.
This is especially true for language learning.
A student may know words like "restaurant", "reservation" and "available", and be able to arrange them into correct sentences in the app, but when he actually walks into a restaurant and needs to say "I'd like to make a reservation" on the spot, he may still fall into a pause.
This is also the product opportunity behind Eevi's so-called "Fluency Illusion".
Therefore, the real space opened up by voice interaction lies in enabling AI to participate in the process of "retrieving and applying knowledge", rather than only staying at the level of knowledge Q&A.
This is exactly the most noteworthy part of voice as an educational interaction method.
For language learning, voice is almost the target ability itself by nature: end users need to understand others, organize language and respond instantly.
For subjects such as mathematics and science, the situation is slightly different.
YoLearn's approach is very typical: the AI communicates with students through voice, while using an interactive whiteboard to display formulas, graphics and derivation processes. Voice undertakes real-time interaction, and the visual interface undertakes knowledge presentation, so voice and vision together form the interaction interface of AI teaching.
Aristotle follows a similar idea. It positions itself as an AI tutor, focusing on dynamically adjusting explanations, pace and support methods according to students' understanding level, so that voice conversations can truly serve one-on-one teaching.
This means that what is really worth paying attention to behind "voice-first" is: AI education products are evolving from "content products" to "interaction products".
Today's real-time voice models allow AI to keep listening and speaking, and conduct multi-turn communication based on context. For educational products, this means that for the first time, AI is increasingly approaching the form of "one-on-one practice partner".
This is also why early-stage capital is starting to bet on such products.
However, it is worth noting that "voice-first" is still more like a product hypothesis worthy of verification, rather than a proven educational conclusion.
A preprint study published in September this year conducted a randomized experiment in a corporate finance course of Boston University's Online MBA program. The study found that in the voice mode, the conversation density of students is about 1.8 times that of the text mode, and the number of questions asked is about 2.4 times that of the text mode; students also prefer voice. But in this five-week experiment, there was no statistically significant difference in weekly learning mastery between the voice and text modes. And the cost of voice is about 2.8 times that of text.
(Data shows that students prefer voice interaction)
(Data shows that voice interaction and text interaction have almost no difference in their impact on students' academic performance)
This result is actually very important. It can be seen that voice has significantly changed "how students learn with AI", but there is currently no evidence proving that it has significantly changed "how much students have learned".
There is still a complete set of pedagogical design between an AI that can chat with you smoothly and an AI that can really help you learn things.
Therefore, from this perspective, what Eevi, YoLearn and Aristotle really need to verify now is not "whether AI can chat with students". This problem has been basically solved.
They need to answer three more difficult questions.
First, can voice practice partners bring about ability improvement?
For example, Eevi cannot just tell users "you have spoken for 20 minutes today", but needs to prove that users have indeed improved their oral fluency, vocabulary use, listening comprehension or real-scene communication ability.
Second, can AI actually complete teaching in the process of conversation?
A real human teacher will judge when to correct, when not to interrupt; when to continue asking follow-up questions, and when to reduce the difficulty. If AI only converts the text answers of ChatGPT into voice, it is essentially still just a chatbot.
Third, is voice the final form of educational AI?
This is probably the most noteworthy question. It is very likely that the answer is not that "all educational AI in the future will only rely on speaking".
Future educational AI will increasingly be like a multimodal private tutor: communicating through voice, presenting content through vision, working through problems on the whiteboard, helping review with text, and dynamically adjusting courses according to students' performance.
The products of Eevi, YoLearn and Aristotle have been developing in this direction.
In China, "voice-first" has not yet become a widely emphasized label for educational AI products, but voice has been integrated into the interaction system of mainstream AI teaching products. Whether it is the AI tutor in learning machines or the AI tutor of Doubao Aixue, they are all interacting with students in real time through multiple methods such as voice, text and images.
In any case, voice is only the entry point, and the end point is still "teaching and learning".
What is really worth paying attention to is that educational AI is moving towards real-time teaching interaction: AI hears students' answers, understands their feedback, and asks follow-up questions, corrects mistakes and adjusts teaching pace at any time. Voice is just the most natural implementation method at present, and better interaction forms may emerge in the future.
For Eevi, YoLearn and Aristotle, what needs to be verified now is whether more natural and higher-frequency real-time interaction can eventually be transformed into quantifiable learning effects.
References:
When AI Tutors Speak: Evidence from a Randomized Field Experiment https://arxiv.org/abs/2609.23958?utm_source=chatgpt.com
This article is from the WeChat official account "Duozhiwang" (ID: duozhiwang), written by Wang Shang, and authorized for release by 36Kr.