HomeArticle

MiniMax and Alibaba are simultaneously deploying their video Agent businesses. After finishing the competition on models, are they starting to compete on workflows?

AI价值官2026-09-03 13:33
The AI video industry is competing on production tools, with numerous new players entering the market and further intensifying the market competition.

Over the past half year, multi-Agent collaboration has almost become the standard configuration for AI video platforms. Platforms including LibTV and Flova have quickly achieved rapid growth in user base and business scale by encapsulating model APIs into complete workflows that cover the whole process from script to finished video.

Nowadays, upstream model vendors have begun to enter the market collectively, shifting the focus of industry competition from underlying model capabilities to the competition of production tools and production experience.

On August 31, Qwen Creation's Agent Teams was officially opened to all users — five Agents including screenwriter, director, artist, cinematographer and editor are connected to form a complete pipeline from creative idea to finished video.

Just 10 days ago, MiniMax launched its official creation workspace MiniMax Design, which integrates full-link creation capabilities such as generation, editing and clipping into the official workspace. ByteDance also launched the internal test of "Manhua Creation Tool" and Seedance Studio in the same period: one focuses on team-based mass production of long-form manhua dramas, and the other targets professional film and television level creation.

The three players have made intensive moves within one month, but their identities and strategies are not equal: Alibaba and MiniMax, as chasers, are impacting the market with a two-pronged approach of "official tools + third-party channels"; ByteDance, as the market leader, chooses to maintain its advantages and consolidate its ecosystem through closed-loop tools and gradient pricing; while the third-party aggregation platforms in the middle are under the pressure of simultaneous tightening from both upstream and downstream.

The consensus across the entire industry is getting clearer: the arms race of model capabilities has reached a new equilibrium point, and the next battlefield is about who can turn models into competitive industrialized production tools.

Expand coverage through channels, deepen capabilities with Agent platforms: Alibaba and MiniMax's two-pronged breakout strategy

When their model capabilities catch up with the first tier of the industry, Alibaba and MiniMax almost simultaneously moved in the same direction: instead of only exporting underlying model capabilities, they take the official Agent workspace as the carrier, encapsulate the models into complete production tools, and then distribute them through third-party channels, advancing on two fronts to impact the market.

Alibaba's action rhythm best illustrates this logic. Only one week after the official release of Wan 3.0, Qwen Creation launched the custom-tailored Agent Teams feature for Wan 3.0, which is a production layer deployment fully supporting the new model.

Wan 3.0, launched on August 24, can generate 30-second videos at a time with native 4K resolution. It supports direct input of multi-format documents for the first time, with significantly improved character consistency and picture stability. During the public beta period, it was evaluated by creators as "stable, realistic and high-quality".

Although Wan 3.0 has been connected to platforms such as Flova and LibTV immediately, the problem also arises here: the production process and data are deposited on third-party platforms. As the model provider, Alibaba can neither accumulate its own creator ecosystem, nor package the model capabilities into a recognizable production experience. For practitioners of manhua dramas and short dramas, Wan 3.0 is still just one of the many optional models at present.

Qwen Creation's Agent Teams is designed to fill exactly this gap. It integrates the full-power version of Wan 3.0 Prime and Qwen-Image 3.0 Pro into a multi-Agent system designed according to the division of labor of a film and television crew: users put forward an idea, the screenwriter produces the script, the director makes the storyboard, the artist sets the style, the cinematographer generates the footage, and the editor produces the finished video, with the whole process closed-loop on the same platform.

Every link from creative idea to finished video remains in Alibaba's own ecosystem. As a result, Wan 3.0's model capabilities are transformed into production capabilities that practitioners can call directly.

MiniMax follows the same path, but its open-source background makes its demand for application layer even more urgent. The H3 video model open-sourced in July is priced at less than one third of the mainstream price in the industry, and performs well in scenarios such as TVC advertisements and e-commerce videos.

However, open source is a double-edged sword: while penetrating the market rapidly, anyone can deploy and fine-tune it. Relying purely on model capability output without application layer support, creators can hardly perceive the substantial difference between H3 and competing products, which will eventually lead to a price competition quagmire.

MiniMax Design launched on August 20 is a direct response to this situation. It starts four Agents for copywriting, image, video and audio in parallel, directly encapsulating H3's capabilities in editing, typesetting, transition and soundtrack into out-of-the-box workflows.

It is worth mentioning that the preset templates provided by MiniMax Design are not simply applied, but have been style-calibrated for H3's attention mechanism, making the finished video more consistent with the output texture of professional editing software, so that the advantage of "the native model has a more thorough understanding of content" can be truly implemented in actual production scenarios.

On the very day when Qwen Creation's Agent Teams went online, MiniMax announced that H3 Max was officially connected to MiniMax Design. This version can generate a complete 5-second 768P audio and video in less than 3 seconds, with the generation speed faster than the playback speed. This breakthrough makes it possible for the model to access continuous live streaming, extending the service scenarios from single video production to continuous frame output, and MiniMax Design is exactly the first entry for creators to experience this capability.

Alibaba and MiniMax have both chosen the combined strategy of "expanding coverage through channels and deepening capabilities with tools": they quickly reach users through third-party platforms to expand their scale, and then use their own workspaces to precipitate users and production data. The ultimate goal is to capture more share from the market dominated by Seedance.

Promote creation tools on one hand, launch gradient pricing on the other: how ByteDance maintains its leading advantage

For ByteDance, which occupies a dominant position in the market, competition is no longer limited to the comparison of model parameters, but has spread to the competition for production tools and creator ecosystem.

On one side, there is the cost-effectiveness involution of model vendors. The capabilities of Wan 3.0 and MiniMax H3 have been improved successively, and H3 has entered the market with extremely low pricing. As competing products have gradually crossed the commercial qualification line, Seedance's pricing advantages and capability barriers are being compressed.

On the other side, there is the risk of the production link and creation ecosystem "falling into other hands". According to the statistics of daily average computing power consumption, Seedance accounts for more than 80% of the domestic AI video market in China, with a penetration rate as high as 95% in the short drama industry. However, as the model provider, ByteDance can only obtain the call volume. Those truly valuable vertical training materials, including how creators disassemble scripts, adjust storyboards and reuse assets, are all deposited inside third-party platforms, which is exactly the core fuel for the iteration of multimodal models.

What's more tricky is that after MiniMax H3 went open source, some leading aggregation platforms have begun to contact computing power vendors, trying to fine-tune and launch their own brand of video generation models based on H3 — third-party platforms are no longer just simple agency channels, and begin to have the possibility of reducing upstream dependence.

ByteDance's response is to take actions at both the tool layer and the pricing layer, with the core goal of opening up the full-link closed loop from model, production to distribution.

The manhua creation tool in internal test on August 18 is a direct response to this situation: it covers the whole process from script upload to video synthesis, supports one-click generation of finished video for scripts up to 200,000 words, allows 30 people to be online simultaneously in a single team space, and can also be directly bound to the Douyin Short Drama Creator Center account, so that the data of the production end and the distribution end are connected.

Just one day later, Seedance Studio was also exposed to enter a small-scale invitation test, adopting project-based management logic, covering the full link from demand understanding to finished video promotion, targeting the professional demand of "making a whole film" rather than "generating a short video clip".

The strategic fulcrum of these two new products lies in that they truly string ByteDance's three existing advantages: model, tool and distribution, into a closed loop. The manhua creation tool has been embedded into ByteDance's ecosystem since its design: the upstream directly calls the massive IP of Fanqie Novel, the middle stream completes the full-link AI collaboration from script disassembly to video synthesis within the tool, and the downstream synchronizes to the Douyin Short Drama Creator Center with one click, realizing full connectivity of listing, review and revenue sharing.

Seedance Studio extends this logic to professional film and television creation scenarios. The character assets, storyboard history and project iteration records accumulated by creators in the tool are naturally bound to ByteDance's distribution and revenue system.

At the same time, ByteDance has also launched a dynamic pricing combination strategy: using high-end new models to set quality benchmarks, and reducing the price of old versions to maintain the cost-effectiveness market base, to cope with the competition from Wan 3.0 and H3.

While Seedance 2.5 is pushing the upper limit of high-end image quality, ByteDance has simultaneously lowered the prices of the mini and fast versions of the 2.0 series to a more competitive price range. The entry-level 2.0 mini 720P version is as low as 0.2 yuan per second for a limited time, focusing on low-cost batch trial and error; the mid-tier 2.0 fast 720P version is as low as 0.6 yuan per second for a limited time, striking a balance between generation speed and image quality, directly targeting the pricing range of MiniMax H3.

High-end products are responsible for building brand reputation, mid-tier products compete for market share, and entry-level products expand the user scale. Through this gradient strategy, ByteDance tries to hedge the cost-effectiveness pressure brought by competing products while stabilizing its existing market base.

Scenario expansion drives mixed use of multiple models, aggregation platforms remain irreplaceable

When model vendors enter the market collectively, the most direct pressure falls on third-party aggregation platforms such as LibTV and Flova.

The price war in August has reached an abnormal level. Taking Seedance 2.5 as an example, even considering the difference in resolution, the price gap between different platforms at the same gear can be several times. This reflects a real industry dilemma: when the core capabilities of all platforms come from the same batch of upstream models, it is difficult for the products themselves to form differences, and competition naturally narrows down to the competition of access speed and subsidy intensity. However, aggregation platforms do have living space.

They still have two values that cannot be ignored.

The first is that they are the most important distribution channels for model vendors. Wan 3.0 was connected to Flova, LibTV and RunningHub on the day of its launch, and MiniMax H3 chose LibTV for its debut, which shows that the channel value and creator ecosystem of these platforms are still valued by model vendors. Even ByteDance will not completely bypass them, after all, aggregation platforms contribute a considerable proportion of Seedance's API call volume.

The second is that the demand for mixed use of multiple models provides a living soil for aggregation platforms. This can be confirmed from the actual production habits of creators: the mixed use of models such as Seedance, MiniMax and Keling has almost become a consensus among AI short drama practitioners.

As AI video expands from short dramas to advertisements, cultural tourism videos, mid-length dramas and online movies, creators will choose different models in different projects and creation stages according to different considerations such as cost, effect and application scenario, so as to achieve the optimal cost-effectiveness.

In contrast, the official Agents of model vendors naturally tend to schedule their own models, which is a constraint in complex projects that require cross-model combination.

Moreover, aggregation platforms are rapidly building their moats: LibTV has launched the "Skill Super Creation Incentive Plan", investing 10 million yuan to support Skill creators, precipitating the workflows of top directors into SkillHub and converting them into replicable production templates; OiiOii has launched a preset style gallery with 149 styles and a global asset locking system, fundamentally avoiding the problem of style drift; TapNow uses an open-source workflow library to target distributed teams that have strict requirements for version control; Flova promotes the Skill production system, encouraging creators to encapsulate their methodologies into Skills to promote sharing and reuse.

These platforms hope to use the small flywheel built by creators and Skills to outperform the big flywheel of leading model iterations in user mindshare.

However, the time window for this race is narrowing. When ByteDance holds three cards of model, tool and distribution at the same time, and when MiniMax and Qwen directly cover the target customer group with official Agents, aggregation platforms are facing the pincer attack of simultaneous tightening from both upstream and downstream. Platforms with differentiated assets and vertical depth still have a foothold, but the window period for pure tool vendors that have neither ecosystem binding nor irreplaceable capabilities is shrinking rapidly.

The core of competition for AI video tools has long shifted from model capabilities to