HomeArticle

Seedance 2.5 Review | Seven-dimensional Hands-on Test: How Much Has AI Video Improved From Short Clips to Full Narrative?

36氪AI测评2026-08-12 18:45
The first half of the AI video track is over, which is our biggest takeaway after we completed the testing of Seedance 2.5.

What does the first half focus on? It is a competition for cleaner datasets, more coherent video actions, and higher resolution. But when applied to real creation scenarios, you will find that the real bottleneck of AI video goes far beyond the competition of these parameters.

 

On July 31, Seedance 2.5 was officially released and launched.

To this end, the 36Kr AI Evaluation Team and 27 36Kr AI Evaluators spent a week restoring AI content based on part of the plot of *Journey to the West*, and conducted an in-depth evaluation of the Jimeng Seedance 2.5 model from seven dimensions.

36Kr AI Evaluators: Covering AIGC content creators, heads of AI manhua drama teams, creative directors of 4A advertising agencies, CG R&D personnel, technology industry observers, top Bilibili UP owners, and practical experts in AI video production and implementation across various industries.

 

This time we chose *Journey to the West* as the evaluation carrier, precisely because the story has high public awareness, which can avoid subjective deviations caused by plot understanding and allow us to directly examine the strengths and weaknesses of the model itself.

Standing at the turning point of AI video iteration, what real production pain points have been solved by the update of Seedance 2.5, and what unsolved problems are still left?

 

Let's start with the conclusion:

1. This upgrade has three hard-core features that directly change the creation method: the duration is doubled from 15 seconds to 30 seconds; the upper limit of reference materials is increased from 15 to 50 full-modal; a new local editing function is added.

2. The improvement on the picture level is also obvious. The camera movement has a sense of design, the impact sense of action scenes is close to game CG, the character expressions are delicate, and the overall texture is close to real shooting.

3. But the problems are also very obvious:

 · The consistency of characters in long shots is still the biggest pain point.

 · The model's ability to understand prompts cannot keep up with its generation ability, and its understanding of prompts with literary descriptions still needs to be improved.

 · The usage cost is relatively high, and the threshold for a complete finished video starts at hundreds of yuan.

 

The following is the detailed evaluation of the seven dimensions.

 

01

Spatial Structure: A Sharp Tool for Professional Users

Seedance 2.5 supports white model reference, which is particularly important.

Senior evaluators generally gave high evaluations:

A film crew can first place cubes, cylinders, character positions and motion paths in a simple 3D space, and then let the model quickly generate a preview close to the final visual effect based on this spatial information. It understands the positional relationship of characters in space very accurately, who is in front, who is behind, who is facing whom, and there is almost no obvious deviation.

Compared with version 2.0, Seedance 2.5 has strengthened its spatial geometry understanding ability. The model no longer only recognizes the pixel information in the reference picture, but can extract depth information from the 3D reference and reconstruct the 3D spatial relationship in the scene. This ability was only in its embryonic form in the previous generation of products, and it has gradually matured in version 2.5.

 

However, for ordinary users who have no 3D production experience, the experience is completely different.

A problem repeatedly appeared in the evaluation process: When the reference pictures are insufficient, the model is still prone to deviations in judging the orientation of characters.

For example: if you want to generate a scene where Sun Wukong is facing away from the audience and facing the monster, the final result often shows that Sun Wukong is facing the audience and facing away from the enemy, and the character relationship is opposite to expectations.

The reason for this problem is: Pure text description itself lacks spatial depth information. The model can only infer the character orientation based on semantics. When the description is not specific enough, it will preferentially generate the most common frontal character composition. Compared with version 2.0, this problem has been improved, but it has not been completely solved, so ordinary users still need to rely on more accurate prompts or continuously adjust the results through multiple generations.

At the same time, Seedance 2.5 performs more stably in picture spatial hierarchy. The relationship between the foreground, middle ground and background is clearly divided, the depth of field is natural, and the overall picture has a good sense of spatial depth, and it is no longer easy for all elements to be stacked on the same plane.

We used *The Daughter Kingdom* as the test scene. In the final generated result, the spatial relationship is clear, the background blur transition is natural, there is no obvious plot hole or position error, and the overall spatial expression is close to the effect of a real photographic lens.

 

36Kr AI Evaluation View

The spatial understanding ability of Seedance 2.5 has become one of its most competitive advantages, but this ability currently still relies more on reference samples. For professional creators, it can significantly improve the production efficiency of complex scenes; for ordinary users, character orientation is still the most in-demand experience link to optimize.

 

02

Camera Language: The Camera Movement Finally Has a Sense of Rhythm

If the spatial structure is "static space", then the camera language is "dynamic space".

The change of camera movement in the new version is holistic.

The movement of the lens has rhythm, a sense of breathing, and the logic of cinematic camera movement: starting from a close-up shot, the speed will change, slow at the beginning, fast in the middle, and slow down at the end, just like a real photographer operating it.

This change comes from the model's deeper learning of camera language data. In the 2.0 era, the model learned the form of camera movement, while in 2.5 it learned the purpose of camera movement. It begins to understand the emotions and narrative functions corresponding to different camera movement methods, so the movement is well-organized.

The most obvious improvement is the transition logic.

The coherence of the video shots generated by the new version is significantly improved, and the audience will not feel obvious sense of fragmentation when watching.

An evaluator who has long been producing AI manhua dramas commented directly:

"The storyboards of version 2.0 are more like AI splicing shots, while the storyboards of version 2.5 are extremely natural."

Storyboard design is one of the core abilities of directors. If AI can independently create well-designed storyboards and camera movements, then what it touches is not only the execution layer, but also begins to enter the creation layer.

The transition of the part of Red Boy is a good example: the lens transitions from the flame special effect of Red Boy to the next scene, using the burning flame as the transition element, and the whole transition is very natural. This has gone beyond the category of simple lens movement, and contains narrative thinking inside.

Although it has not reached the level of a professional director, the direction of understanding is correct, and the progress speed is very fast.

 

03

Motion Performance: Full of Impact Sense, Walking Is Still "Unpredictable"

All evaluators have highly consistent conclusions on motion performance: the impact sense and interaction performance are very excellent.

A professional evaluator's comment is very realistic:

"If this was three years ago, the cost would be at least 100,000 yuan."

The impact sense sounds very abstract, and you will understand it at a glance: when you punch out, you can feel the power, see the tensed muscles, the fluttering of clothes, the force reaction, and the changes in the environment caused by it. The performance of 2.5 in this aspect is already close to the level of some game CG.

Being able to achieve this level is directly related to the training data: The Seedance team has specially carried out enhanced training for action film and television materials, so that the model can learn the physical laws of human movement and force feedback. This shows that the 2.0 model only learned the appearance of actions, while 2.5 learned the mechanical logic of actions.

The stability of continuous actions has also been greatly improved. The common problems of stuttering, frame dropping, character disappearance, and props flying around in version 2.0 are significantly reduced in 2.5. Characters will no longer have weird actions of swinging their fists randomly, and every attack has a clear target and trajectory.

In the test of the fight between Sun Wukong and Erlang Shen, the model also added special effects independently. Sparks, sparkles, energy fluctuations, matched with appropriate sound effects and impact rhythm. It even completed the filter effect of black and white highlight scenes, and the visual experience exceeded expectations.

 

04

Character Performance: More Delicate Inner Drama, But There Is Still Deviation in Prompt Understanding

Character performance has always been an important indicator to measure the intelligence level of AI video,