Just now, AI has broken through the "video Turing Test", with 17 million netizens watching the live broadcast online.
In the past, video calls were often regarded as a safety measure for identity verification. Text can be written by others, and voice can be synthesized, but asking the other party to turn on the camera, chat for a few words and observe their expressions always seemed to add an extra layer of certainty.
Today, even this certainty of "seeing is believing" is beginning to be challenged.
Recently, the AI research team Tavus released Griffin, a real-time video interaction model, and announced the results of an experiment. After a one-minute video call with its research preview version Griffin-Lite, 26 out of 54 participants, accounting for about 48%, believed that there was a real person sitting on the other side of the screen.
▲ As of press time, nearly 17.36 million netizens have viewed this content
As a control, Tavus's previous generation system, under the same experimental procedure, only 1 out of 41 people was mistaken for a real person, a proportion of 2.4%.
Tavus thus claims that Griffin is the first model to pass the real-time "video Turing test", and classifies it into a new category called Human Interaction Model, abbreviated as HIM.
However, the "first pass" is still Tavus's definition of the experimental result for now. At the same time, since the preview has not been fully opened, the suspicion of a demo that only exists on paper cannot be ruled out. To fully understand this release, three issues need to be clarified at the same time: how the participants were convinced, how the model completes the interaction, and how much the one-minute performance can prove.
The convincing details are hidden beyond speech
When chatting with AI, the moment that most easily breaks the immersion may not be the absurdity of the answer, but sometimes just the fact that it is too bad at picking up the conversation.
You pause to organize your thoughts, and it rushes to answer immediately; you tell a sad story, but the face on the screen still wears a smile; you interrupt it, and it continues speaking at its original pace. The answers may not be wrong, but the communication always feels a bit awkward.
What Griffin tries to handle are these details that happen beyond language. According to Tavus's introduction, it can see, listen and speak at the same time, and continuously decide what to do next based on the other party's expressions, line of sight, tone and pauses.
In the official demonstration, the model showed the ability to change emotions, tones and movements along with the conversation, and also participated in scenarios such as imitation instruction games, Rubik's Cube interaction, and observing users welding circuit boards. Tavus hopes to demonstrate that the model can participate in communication by combining visual and temporal information.
For example, for the same period of silence, some people are thinking, some have finished speaking, and some may want the other party to continue explaining. The system needs to judge by combining context and non-verbal signals, avoiding treating every pause in sound as the end of a speech.
The scope of visual generation has also been expanded. Tavus states that Griffin can generate the entire frame in real time starting from a single reference image, including the person's face, arms, fingers, as well as chair movements, shadows and background changes.
Therefore, the character on the screen can also nod, adjust their line of sight or change their expression while listening to others, instead of suddenly moving only when their own lines start.
If AI wants to learn to chat, it has to get rid of the turn-based answering rhythm
Behind the above capabilities is the change of interaction mode.
Common cascading systems will complete speech recognition, large language model response, speech synthesis and digital human driving in sequence, with multiple links connected one after another. After the user finishes a sentence, the system processes it, and then passes the result to the next link.
Griffin adopts a full-duplex video interaction scheme. The so-called full-duplex means that while outputting audio and video, the system still continuously receives and processes the user's voice and images.
▲ The left is the original image, the right is AI-assisted translation, for reference only
It consists of two main parts.
The continuous dialogue modeling engine is responsible for perceiving the scene, deciding what to say and when to respond, and outputting control signals such as emotions, expressions and movements at the same time; the audio and video generation engine converts these signals into continuous sound and images.
The dialogue engine re-evaluates the communication state at intervals of less than one second, so the model can decide to pause, continue or give up the floor to the user based on the user's reaction even in the middle of a sentence.
Achieving this coordination also requires the generation speed to keep up.
The speech module of Griffin-Lite adopts streaming generation, which can output sound gradually as control signals arrive, and supports cloning voices using about 10 seconds of reference audio.
The video module generates 720p, 25fps video in real time in segments of 320 milliseconds each. Through model distillation and autoregressive training, the team compressed the original multi-step generation process into fewer steps, while improving the picture stability during long-time output.
In the H100 test announced by Tavus, the average delay between the arrival of audio and the appearance of its corresponding effect on the screen is 0.43 seconds, which is about half of the next fastest method among the comparison objects. This figure measures the generation delay from audio to video, and cannot be directly regarded as the time from the user's question to receiving the complete answer.
In terms of image quality, Griffin-Lite achieved the best results in three indicators: DOVER, FID and THEval in the comparison with four published streaming diffusion models; its lip sync indicator LSE-C scored 7.27, ranking second.
Its ranking is close to humans, but the one-minute experiment still has limitations
In addition to internal tests, Griffin-Lite also appears on Nvidia's public VideoFDB leaderboard. This benchmark evaluates whether the model can understand non-verbal signals in communication, and whether it can generate situation-consistent sounds, expressions and movements respectively.
In the generation track, Griffin-Lite scored 3.83 in total, the human reference score is 3.92, and the next-ranked Gemini 2.5 plus Anam combination scored 2.80. In the perception track, Griffin-Lite scored 3.73, the human reference score is 4.20, and the highest score of other public baselines is 3.44.
Both scores are ahead of other AI systems on the leaderboard, but they are scores under specific benchmarks. In particular, it should be noted that some comparison objects only receive audio, while some receive both audio and video, and the input conditions are different, so the ranking cannot be directly extended to a comparison conclusion of all capabilities.
As for the most concerned 48% result, it comes from another experiment organized by Tavus, where participants were recruited through an independent research platform.
Participants were told in advance that they would have a one-minute video call with another participant to discuss the most anticipated thing of the year. What actually appeared on the screen was an AI character whose face, voice and responses were generated in real time by Griffin-Lite.
After the call, participants first evaluated the naturalness, credibility and communication experience of the other party, and were not asked whether they ever suspected that the other party was not a real person until the end of the questionnaire. At the end of the study, all participants were informed of the AI identity of the character.
Finally, 26 people judged that the other party was a real person, and the average confidence level of this group of participants in their judgment was 79%; the participants who thought the other party was AI had an average confidence level of 81%. Those who had doubts usually noticed it within 20 seconds after the call started.
The model has not yet eliminated all communication barriers. In the 7-point evaluation, the naturalness score was 5.4, the credibility score was 5.6, the willingness to communicate again score was 5.8, and the dialogue fluency score was 4.9, which was the lowest among the five evaluations.
The experiment shows that the real-time synthesized face can already make some people believe it is real under specific conditions. However, the sample size was only 54 people, the call lasted only one minute, and the participants were told in advance that the other party was another person, so the results still need to be verified on a larger scale, for longer periods and in more scenarios.
The more realistic it is, the more it needs to clarify its identity
In Tavus's vision for Griffin, you will never think about the "mechanism" of communication. You won't think about "what words should I use to make it understand", you just express yourself naturally, and let the conversation flow like water.
Imagine that future online tutoring will no longer be cold PPT recorded broadcasts. The AI teacher can see from your furrowed brows that you don't understand, and then voluntarily switch to a more popular metaphor;
For future customer service, you don't need to rack your brains to describe the model of a part, you just need to hold it up to the camera, and the AI can work with you to figure out how to repair it.
The machine finally bends down and steps into the human context.
This is also why, while releasing the Griffin-Lite preview, Tavus announced that it will only open the research preview to some trusted testers, and it is not yet available for ordinary customers.
The team said that it is developing an AI identity disclosure function, and carrying out further safety and alignment assessments, and the wider release will depend on the progress of handling relevant issues.
In the past, the challenge for digital humans was how to make people willing to communicate with them; as real-time generation becomes more and more natural, products must also answer how to let people know who they are communicating with. Making AI learn to read social cues may lower the usage threshold, and letting users see its identity clearly will determine whether this convenience can be built on trust.
This article is from the WeChat official account "APPSO", author: Discover Tomorrow's Products, published with authorization from 36Kr.