Keling Observation ② | Recreating *Farewell My Concubine* with Keling: The cinematic quality is fully achieved, so how to make the complex narrative more stable?
High-quality visuals, cinematic shots, advertisements and commercial short films are the application scenarios most frequently mentioned by users.
This time, we further expanded our testing: if cinematic quality is the most anticipated feature of Kling, can it support the creation of complete short films that include characters, actions and emotions?
Centered on the theme of *Farewell My Concubine*, we originally created the test script *Farewell My Concubine: Past and Present*, and used the first-to-last frame image-to-video and subject binding features of Kling 3.0 to complete a round of stress testing. This test was carried out around four dimensions: character consistency, composition and spatial hierarchy, light and shadow and atmosphere shaping, action and camera movement design, which not only retained shots with high completion, but also fully recorded the trial-and-error process in complex tasks.
First, the core conclusion: Kling 3.0 is more suitable for the creation of key shots.
The model performs outstandingly in composition, light and shadow, emotional atmosphere and character subject consistency, and can stably output high-quality visuals for short action shots with clear goals. If you need to use Kling to process complex narrative films, a more reliable idea is to split long actions into short shots with single goals, and then connect them into complete segments through storyboard design and post-editing.
When the task involves cross-space movement, multi-prop and multi-character fight scenes at the same time, the final visual effect highly depends on prompts, reference images, shot duration and splitting methods.
The purpose of this stress test is to find a more stable creation method, not to draw a final conclusion on the model's capabilities.
[Figure 1 | Reference image used in this test]
Title
Cinematic quality is first implemented in key shots
Judging from the test results, the core advantage of Kling 3.0 is concentrated in the visual completeness of key shots.
When the shot has a clear goal, simple character relationship and short action duration, the model can generate high-quality visuals. The performance of light and shadow, composition, characters and emotional atmosphere is all outstanding.
For example, the scene of the opera performer preparing in front of the mirror most intuitively reflects the visual expressive power of Kling.
[Figure 2 | Opera performer preparing in front of the mirror]
Kling does not only perform simple dynamic processing on this scene. The mirror frame limits the character to the visual center, the beads and opera costume on the left form the foreground, the character and bronze mirror are placed in the middle ground, and the backstage corridor and warm light act as the background, forming a standard three-layer depth structure of foreground, middle ground and background.
The picture gets rid of the flat sense and presents a distinct cinematic spatial texture.
The core value of Kling in this type of shot lies in generating a complete visual picture with composition, light and character state at one time.
This also indirectly confirms the core reason why users generally associate Kling with "cinematic quality" and "high-quality shots" in the first review.
Title
Short action shots can stably output high-quality dynamic effects
Cinematic quality does not only come from static composition. Kling can also complete impactful cinematic dynamic shots.
For the scene of Xiang Yu fighting on the battlefield, we obtained a short shot with high completion after multiple rounds of testing.
[Video 1 | Xiang Yu breaking through the encirclement]
Visual focus: the momentum of the character, the sense of oppression on the battlefield and the rationality of the action direction in a short period of time.
In this clip, Xiang Yu's breakout route is clear, and the character's movement conveys the tension of killing. The effects of yellow sand battlefield, clashing weapons, falling human bodies and diffuse dust jointly build a sense of oppression in the melee.
The core premise for Kling to generate long-shot action scenes is to have a clear target shot that only carries one main action. The clearer the goal, the higher the stability of the output result.
Title
Atmosphere shaping and character consistency: provide anchors for narrative
1. Atmosphere shaping
The emotional shots before and after Xiang Yu's suicide by the Wu River are another set of materials with high completion in this round of testing.
[Video 2 | Xiang Yu committing suicide by the Wu River]
Visual focus: light and shadow and dark area details.
The overall brightness of this scene is low. In the close-up shot of Xiang Yu, the sword body shows a faint reflection, and the cold gray sky light paired with the dirt, blood stains and tears on his face jointly set off the tragic emotion of the hero at the end of his road.
2. Character consistency
The ability of subject binding also provides support for the finished film. Xiang Yu, Consort Yu and the opera performer maintain the consistency of faces, costumes and character identities in different shots without deviation.
Subject binding solves the basic problem of narrative: ensuring the recognizability of characters, so that the audience can continuously identify the same character.
After the characters are stabilized, shots of different locations, shot sizes and emotions have the basis to be edited into the same narrative.
Title
Complex narrative: need to split into a standard shot workflow first
After this round of testing, we further clarified the reasonable creation method of Kling 3.0. A complete narrative is not impossible to achieve, but all creation tasks cannot be concentrated on a single long shot.
1. A single shot carries a single task
If a shot contains elements such as position movement, turning around, sword dancing, pursuers, horses and camera circling at the same time, the model needs to process too much information at the same time. A more reliable way is to shorten the shot duration so that it only completes one clear action, avoiding shape deviation during long-term movement.
2. Battle scenes are realized through storyboard and editing
The complete action chain of "charging, swinging the sword, hitting, the enemy being impacted, falling to the ground, and the pursuers surrounding" does not need to be fully presented by a single shot. It can be split into independent shots such as charging, swinging the sword, the enemy retreating and battlefield reactions, and then through sound effects and editing, let the audience perceive the complete battle process.
3. Split the spatial path into nodes in advance
The movement path of the character walking from the backstage to the side stage and then entering the stage itself contains multiple spatial nodes. Setting each node as an independent shot, or setting the first and last frames for key transitions respectively, is easier to control the effect than just emphasizing "entering the stage from the side" in the prompt.
This creation method does not bypass the model's capabilities, but organically organizes the key shots that the model is good at into a complete narrative. AI is responsible for shot generation, and creators are responsible for task splitting, continuity construction and final effect judgment.
Title
Stress test verification: complex space and physical interaction still require enhanced control
After clarifying the advantages and creation workflow, let's look at the most difficult scenes in this round of stress test.
The test results are affected by multiple factors such as prompts, first and last frame design, subject reference images, shot duration and generation randomness. The conclusion is only applicable to the parameters of this round of test, and does not represent the effect of all scenarios.
1. Precise spatial routes require more refined shot constraints
We tested the continuous scene scheduling of the opera performer walking from the backstage to the stage for preparation. This task requires the model to understand the spatial relationship between the backstage, side stage and stage, not just complete the action of "the character stepping onto the stage".
The results of 5 rounds of tests show that the character can stably complete the semantic result of "stepping onto the stage", but it is difficult to continuously follow the specific movement path of "entering the stage from the side".
[Video 3 | Spatial connection of the opera performer from backstage to stage]
Visual focus: whether the character enters the stage from the side, and the spatial connection logic between the backstage and the stage.
The tone, character state and stage atmosphere of this scene all meet expectations, and the deviation mainly occurs at the path connection around the 7th second. Kling generates a visually smooth movement effect, but cannot stably reproduce the real theater's stage entry movement path.
In actual creation, a more reasonable processing method is to split the backstage, side stage and stage into two shots, and build spatial continuity through character orientation, line of sight guidance and editing techniques.
2. Multi-prop and multi-person battles are more sensitive to generation conditions
Opera and battlefield shots include double swords, water sleeves, armor, spears, hand movements, body center of gravity and camera movement at the same time, which is the test task with the highest information density in this round.
In some test versions, problems such as the change of the number of double swords, the deformation of water sleeves, and the mismatch of feedback after killing characters appeared. Prompts can provide direction constraints, but cannot completely eliminate such deviations in long continuous actions.