HomeArticle

What are video generation models competing for?

新剧观察2026-08-10 11:44
Ecosystem and business efficiency

If the key keyword of 2023 was large language models, then from 2024 to 2026, the most intuitive competition of generative AI has taken place in the video domain.

In the past, when people discussed AI video, they focused on "whether it can make an image move"; today, the questions have evolved into: whether it can maintain character consistency, understand complex actions, automatically complete storyboarding, synchronously generate dialogue, music and ambient sound, and whether the generated results can be directly integrated into advertising, short drama, game and film & television production workflows.

As of August 2026, Chinese models including ByteDance Seedance, Kuaishou Keling, Alibaba Wanxiang, Tencent Hunyuan, MiniMax Hailuo, and Shengshu Technology Vidu have formed a fairly dense competitive landscape. They are no longer just catching up with overseas models, but are competing head-to-head in multimodal input, native audio, generation speed, product pricing and commercialization.

However, behind the bustling scene, a more critical question remains: what stage have China's large video generation models actually reached? Are they only bringing an upgrade to content production tools, or may they restructure the entire video industry?

From "Can It Generate" to "Can It Produce"

Video generation is not simply playing image generation technology in sequence.

An image only needs to look reasonable at a single moment, while video must process space, time and causal relationships simultaneously. When the same character walks through several shots, their face cannot change suddenly, and their clothes cannot disappear out of thin air; a car turning must conform to the laws of motion, and a cup falling to the ground cannot pass through the tabletop; after the shot switches, the characters, environment and lighting must also remain continuous.

Therefore, the difficulties faced by video models are not only about "whether the drawing looks realistic", but more importantly, whether they can understand a dynamic world.

Early AI videos often gave people the impression of "stunning at first glance, breaking immersion at the second glance". The picture might look beautiful, but fingers would deform, objects would drift, and characters would change faces halfway through walking. It was suitable for demonstrating technology, but hardly qualified for real production processes. What creators got was not a finished piece, but a segment of material that required repeated attempts, selection and retouching.

Over the past two years, the industry's progress has mainly occurred in four directions.

The first direction is subject consistency. In the past, the biggest problem in generating multi-shot videos was that "the same person does not look like the same person". Now, models can lock in characters, costumes, scenes and visual styles as much as possible through character images, video clips, voices and style references. Vidu takes subject consistency as a core capability, allowing users to establish visual constraints for characters or objects through multiple reference images; Alibaba Wanxiang also supports multimodal reference inputs including text, images, video and audio.

The second direction is controllability. Real commercial creation cannot rely solely on unexpected surprises, it needs certainty. Advertising clients will not accept random changes in product colors, and directors cannot tolerate the protagonist's actions being completely determined by the model. Functions such as first and last frame control, action migration, camera movement, local editing, and character replacement are essentially transforming AI video from a "random generator" into a "controllable production tool".

The third direction is integrated audio and video. Early models were only responsible for generating silent pictures, and users had to dub, find music, make sound effects and adjust lip sync separately. Today, native audio has become the focus of competition for leading models. Keling 3.0 can handle multimodal input and output including text, images, audio and video, and supports multiple languages, dialects and accents; Vidu Q3 can synchronously generate dialogue, narration, sound effects and music; the new generation of Seedance model is also enhancing the joint generation of audio and video and long narrative capabilities.

The fourth direction is moving from single generation to unified workflow. In the past, text-to-video, image-to-video, video editing and action control were isolated tools; now, leading manufacturers are trying to use one model to complete understanding, generation, modification and continuation. Users no longer need to re-render the entire video, but can use natural language to ask the model to "change the character into another piece of clothing", "keep the action but change the scene" or "extend the last shot by five seconds".

This is the real industry inflection point for video generation.

To judge whether a technology is mature, we cannot only look at how stunning its best works are, but also how many qualified results ordinary users can get after using it ten times in a row. Video models are shifting from pursuing the "upper limit of capability" to raising the "lower limit of results": they should not only occasionally generate a blockbuster clip, but also deliver usable materials stably, quickly and at low cost.

In other words, the competition for AI video has shifted from "can it generate" to "can it produce".

Not Just a Model Competition,

But Also a Competition for Ecosystem and Commercial Efficiency

China's large video generation models have not followed a completely identical development path.

Kuaishou and ByteDance have huge content platforms and creator ecosystems. For them, video models are not isolated technical products, but a natural extension of the original content system.

Kuaishou can observe how real creators produce short videos, which functions are used most frequently, and what kind of content is easier to spread, then apply these feedbacks to model and product iteration. As a result, Keling has formed a closed loop from models and creation tools to content distribution. It is worth noting that this business has already achieved verifiable commercialization through financial data: Keling AI generated 340 million RMB in revenue in the fourth quarter of 2025, with monthly revenue exceeding 20 million USD in December 2025, corresponding to an annualized revenue run rate of 240 million USD.

The significance of this set of data is not only the revenue scale, but also proves that video generation no longer relies entirely on capital investment and technical narratives, and there are indeed people willing to pay for it.

ByteDance's advantage lies in its more complete content production tool chain. From the models to Dreamina, Jianying, CapCut, and then to Douyin and TikTok, ByteDance has the opportunity to connect scripts, images, videos, editing, dubbing and distribution. For ordinary users, they may not care which underlying model is used, but only care whether they can complete a video with lower thresholds in familiar tools. What ultimately determines the outcome is not necessarily leading by a few points on a certain ranking list, but who can turn model capabilities into functions that users use every day.

Alibaba and Tencent represent another path: combining models with cloud services.

On the one hand, Alibaba Wanxiang builds a developer ecosystem through open source, and on the other hand, provides API, deployment and fine-tuning capabilities to enterprises via Alibaba Cloud. Its value is not only about generating a video for individual users, but about enabling e-commerce merchants, advertising agencies and enterprise clients to produce content in batches.

Tencent can embed video models into advertising, games, social networking and cloud services. For this type of comprehensive technology company, even if video generation does not independently become a huge consumer product, it can serve as a basic capability to reduce the content production cost of the entire business system.

Startups such as MiniMax and Shengshu Technology face a more complex situation. They do not have super platforms to provide natural traffic, so they must form differentiation in technical features, overseas markets or vertical scenarios. Vidu continuously emphasizes subject consistency, generation speed and anime content; MiniMax enters the global creator market through Hailuo AI. In July 2026, MiniMax released the H3 video model, which supports text, image, video and audio input, can generate videos of up to 15 seconds in 2K resolution with native stereo audio, and announced the open availability of model weights.

From this perspective, manufacturers' advantages are not only "larger model parameters" or "better single generation effect", but the positive cycle formed between iteration speed, engineering capabilities and industry scenarios.

China has a huge market for short videos, e-commerce, games, online literature, animation and short dramas. These industries generate a large number of low-cost, high-frequency content demands that require rapid testing every day. The denser the demands, the faster the model gets real feedback; the cheaper the model, the easier it is to create new demands.

However, the market competition is also extremely fierce.

When all models can generate high-definition pictures, maintain character consistency and output audio, basic capabilities will quickly converge. The lead formed solely by model performance may only last for a few months. Price wars will also compress the gross profit margin of generation services, forcing manufacturers to continuously bear the costs of training, inference and product iteration.

Therefore, this competition will most likely not produce only one winner. The model layer may gradually concentrate on a small number of companies with capital, computing power and data capabilities; the application layer may remain decentralized, with a large number of professional tools emerging in scenarios such as advertising, e-commerce, animation, games, short dramas and education.

The real moat will shift from "I can generate video" to "I master which type of users, which kind of workflow and which set of commercial data".

Will AI Video Replace the Traditional Film and Television Industry?

Whenever a new video model is released, the most frequent question people raise is: will actors lose their jobs? Will directors be replaced? In the future, can one person make a movie alone?

These questions capture the impact of the technology, but tend to overestimate short-term substitution and underestimate long-term restructuring.

A mature film and television work is not a simple splicing of several beautiful shots. It requires stable character relationships, coherent narrative logic, precise emotional expression and a large amount of cross-shot scheduling. Even if the model can generate videos of tens of seconds or even dozens of seconds, it does not mean that it has understood the character growth and story structure in a 90-minute movie.

At present, the easiest entry points for AI video are not long film and television works, but industries with short production duration, high production frequency and tolerance for rapid trial and error.

For example, e-commerce merchants can convert product images into display videos in batches; advertising teams can generate storyboards and samples before official shooting, and quickly produce multiple versions of materials for different audiences; game companies can use AI to create concept animations, character previews and promotional videos; short drama and animation teams can reduce production costs for some scenes, special effects and transition shots.

The common feature of these scenarios is that they were not impossible to achieve in the past, but too costly and took too long to complete. What AI changes first is not "whether it can be done", but "whether it is worth doing".

When the production cost of an advertising material drops significantly, enterprises may not necessarily reduce their total budget accordingly, but instead produce more versions to conduct more granular audience testing. In the past, a brand could only produce five advertisements, but in the future, it may generate five hundred at the same time, and then select them based on click-through rate and conversion rate. This means that video generation not only brings cost reduction, but may also create new content demands.

Therefore, the impact of AI on employment will not be simply replacing humans with machines. What is more likely to happen is the redistribution of the value chain.

Repetitive execution tasks will be reduced, while the importance of creative planning, aesthetic judgment, material management, model control and result selection will rise. Future video creators may no longer complete every shot by hand, but need to manage multiple models just like a director manages a film crew. The production capacity of one person will expand significantly, but people who can stably produce high-quality content still need to understand narrative, rhythm, composition and user psychology.

Before large-scale commercialization truly kicks off, the industry still has to cross several thresholds.

The first threshold is stability. Commercial clients do not care about the best performance the model can achieve, but whether the promised effect can be repeatedly realized. As long as generation still relies heavily on random attempts, the seemingly low price may be offset by the time cost of repeated trials.

The second threshold is cost. Video generation needs to process a large amount of spatio-temporal information, so its inference cost is naturally higher than that of text and images. In the future, models will not only compete on performance, but also on generation speed, success rate and cost per second under the same quality standard.

The third threshold is copyright and security. Whether the model's training data is legal, whether the generated content imitates specific actors and film & television characters, and how the uploaded faces, voices and commercial materials of users are protected, all may affect the industrial boundary. Current Chinese regulations require generative AI services to respect intellectual property rights, portrait rights and personal information rights; the "Measures for the Identification of AI-Generated Synthetic Content" released in 2025 also requires explicit or implicit identification of relevant content.