Flova and LibTV are still bickering with each other, while ByteDance's MiniMax has already joined the fray.
Time is running out for video model aggregation platforms.
On August 20, MiniMax officially launched its multi-modal creation Agent workspace MiniMax Design, which features one-click video generation and native model advantages. The platform directly integrates the generation, editing, and clipping functions of third-party platforms, and solidifies developer experience directly into Skills.
Two days earlier, ByteDance's other creation platform, the manhua drama creation tool beyond Skylark, was exposed for internal testing, which supports one-click video generation for scripts of up to 100,000 words. An AI creation product matrix covering short dramas and manhua dramas is taking shape.
On the other side, aggregation platforms are still waging fierce price wars to compete for users.
"At this price, it's impossible to get it even at the procurement cost."
This comment comes from the comment section of the Flova discount topic. On August 18, Flova launched the 818 Limited-Time Carnival Month event, announcing the annual bottom price for popular models, reducing the video generation price of 480P Seedance2.5 to 0.175 yuan per second.
According to the official pricing, the price of Seedance2.5 video of the same quality is about 0.67 yuan per second.
This is not the first time Flova has challenged its peers on model pricing. Back on August 7, the day Seedance2.5 was officially launched, LibTV announced a sharp price cut, setting the price of 720P Seedance2.5, which originally cost 1.5-2.1 yuan per second, at 0.4 yuan per second.
Subsequently, FlovaAI quickly followed up, offering an ultra-low price of 0.23 yuan per second through a 53% discount for 10 days and a membership points buy-one-get-one-free deal.
RunningHub's 0.21 yuan per second and SkyProduction's 0.3 yuan per second even cut the price of Seedance2.5 to a fraction of the original overnight. Higgsfield even offers free trials directly.
From the official price of more than 1 yuan to the platform price of 20 cents, the price difference is more than five times.
Apart from pricing, there is even more heated narrative competition.
On July 14, LibTV launched the Agent function, announcing that it already has more than 100 director-level Skills encapsulating professional creation methods, aiming to build the world's largest professional video Skill Hub in one go.
Three days later, Flova posted a taunt: "We have noticed that more and more products in the AI video field have begun to learn from Flova's product form and creation paradigm. We are glad that the video Agent creation mode pioneered by Flova has attracted the attention of the industry."
Two platforms in the same track fought across the air over the almost unprovable proposition of who invented the video Agent first. What they are competing for behind the scenes is the same group of AI video creators.
But the underlying logic of this scuffle has quietly changed.
On the supply side, new players such as OiiOii, SenseTime Seko, and Pexo are entering the market rapidly. The most scarce resource has changed from computing power to users. In the past, the track competed for who had a stronger model, but now the main models of all players are Seedance2.5 and MiniMax H3, so they can only push the price to the extreme.
At the same time, a new crisis is coming. When model manufacturers themselves start to compete for users directly, how many stories can the middlemen who started by encapsulating APIs tell?
Middlemen have to burn money to grab business
The business of AI video aggregation platforms is essentially a transaction with a threshold so low that it is almost non-existent.
The Agent dispute between LibTV and Flova is a representative case.
Both platforms have a background related to Jianying. Chen Mian, founder of LibTV, once served as the head of global commercialization of Jianying and CapCut, while Guo Lie, founder of Flova, is the founder of FaceMoji. After the team was acquired by ByteDance, it became an important early team of Jianying.
Both products adopt the product philosophy of Jianying, that is, low-threshold creation to replace professional workflows. LibTV promotes that its Agent mode can generate complete videos through one-sentence prompts, and Flova promotes that users can generate excellent videos just by chatting.
But if you delve into the technology, the two sides are actually fighting a mirror war.
AI video aggregation platforms do not develop underlying large models, do not invest in core algorithms and computing power, but encapsulate the APIs of upstream manufacturers, and then wrap a layer of better user experience.
There are open source projects and ready-made workflow codes everywhere on GitHub, and a decent platform can be built with slight modifications. That's why so many players have crowded into this track in just one year.
All players provide one-stop services from script, storyboard to finished video output, and the core selling points are all workflows and skills. Tracing back to the source, their methodologies can all find traces from the rise of Jianying.
At the beginning of its launch, the core benefit that LibTV promoted was that it could integrate the entire creation process of scripts, storyboards, shots, and editing with one canvas, and was equipped with exclusive functions such as three-view character design, character subject library, plot deduction, and picture deduction.
The seemingly complex exclusive functions are essentially the efficient invocation and encapsulation of model capabilities by the platform, using a set of precise instructions to command the model to complete tasks, and organizing mature creation processes into automated workflows. Although the production is cumbersome, the difficulty lies in the product/function design level, not the algorithm itself.
Flova is similar, except that the platform goes a step further, emphasizing its self-positioning as an Agent and putting more emphasis on intelligent agent capabilities, but the core remains the same. Both are trying to connect the steps of script writing, storyboarding, generation, and editing into an automated pipeline.
But when all players' capabilities come from the same batch of upstream models, competition will inevitably degenerate into a race of access speed and subsidy intensity. Whoever can access new models earlier and offer more free quotas will gain an advantage in the user base.
A relevant person in charge of RunningHub once publicly stated that top productivity tools should be a sharp weapon for creators to make money, not a burden for creation. But the result is that the product itself cannot form differentiation, and the only sharp weapon left for platforms to make money is price.
Public reports show that the official API cost of Jimeng 3.0 Pro 1080P is about 1 yuan per second, while the price of Jimeng premium membership is about 1.3 yuan for 15 seconds. The platform has to use current limiting and queuing methods to reduce losses. The official side is already in such a situation, so the money burning speed of third-party platforms whose prices have been cut to 20-30% of the official price can be imagined.
Users attracted by subsidies will also be taken away by the next cheaper and faster-access platform. It is still unknown what kind of user loyalty the platforms have bought with real money.
Model manufacturers do not want to be just water sellers
"The Fired Girl" was once a representative AI short drama work of LibTV. Director Feifeifei created an 11-episode work with the help of the LibTV platform, which gained over 220 million native views on Douyin, with stable likes of more than 300,000 per episode. The cost of each episode was only about 10,000 yuan, and the production cycle was no more than 3 days.
But in a recent public interview, the creative team revealed that the second season will continue to be produced on ByteDance's Skylark platform.
The reason is that Skylark provides the production team with a series of incentives, such as being selected into the Skylark Creator Program, Skylark Premium IP Support Program, Douyin AI Creation Competition, and even directly signing directors to join the Skylark Super Creator Support Program, providing points and technical support, which is equivalent to directly binding them as platform partners.
This is a competitive method that cash-strapped third-party platforms cannot easily achieve. The official side directly gets involved in content production, and indirectly influences the creator group by binding the head exclusive content ecosystem.
Third-party platforms are fully capable of providing technical support, points, and revenue sharing, but when it comes to the content platform side, ByteDance's advantages with Jimeng and Douyin come to the fore. Creators ultimately need to monetize on short video platforms with Douyin as the core, and ByteDance's incentive policy toolbox is always larger than that of third-party platforms.
On the tool side, ByteDance's layout is even more intensive. Jimeng AI is deeply connected with Doubao, Jianying and Douyin, forming an integrated closed loop from creativity to distribution. At the same time, ByteDance is also testing the manhua drama creation tool and the professional AI film and television work platform Seedance Studio, targeting the two hottest vertical scenarios of AI videos: short dramas and manhua dramas.
The ways third-party platforms rely on to retain users are long-term memberships on the one hand, and the fact that creators' workflows and data are deposited inside the platform, leading to high migration costs. But when ByteDance, which controls the entire process from production to distribution, enters the market, the process assets held by third-party platforms will rapidly depreciate.
ByteDance has not even used its "nuclear weapon" yet. Once Jianying and CapCut integrate AI video generation capabilities, they can instantly retain a huge number of ordinary users, tightening the three links of tools, content and distribution at the same time.
What's more subtle is that MiniMax, another upstream player, has also taken action.
MiniMax Design is entering the market. MiniMax H3 itself is known for being cheap and open source. Coupled with the product positioning of the multi-modal creation Agent workspace, it supports users to describe requirements, disassemble tasks and generate works in natural language. The official target user group highly overlaps with Flova and LibTV.
Platforms like LibTV are facing a pincer attack where both upstream and downstream are tightening. ByteDance's official ecosystem is competing for creators and content ecosystems; upstream video model manufacturers take the initiative to enter the market, directly covering the target customer group with native Agent workspaces. In addition to price, aggregation platforms need more strategies to participate in competition.
Is it worth being twice as expensive as Meitu?
According to a Bloomberg report on August 6, Yanyu Technology, the parent company of products such as LiblibAI, LibTV and Xingliu, has started preparations for a Hong Kong IPO, and is about to complete a new round of financing with a valuation of 3 billion US dollars.
In terms of performance, as of May 2026, the company's annual recurring revenue has exceeded 300 million US dollars, a year-on-year increase of over 3000%. LibTV was launched in March, with a single-day revenue of over 1 million US dollars. By May, its revenue had reached more than 13 times that of the first month. Tencent, Granite Asia, Shunwei Capital, Ant Group, HSG and other well-known institutions are all its investors.
Judging from the PS valuation statistics, Yanyu Technology with a valuation of 3 billion US dollars corresponds to 300 million US dollars in annual recurring revenue, with a price-to-sales ratio of about 10 times. In the same creation tool track, Meitu's PS valuation is about 4.7 times.
In terms of performance, this is a counter-intuitive phenomenon. Meitu has hundreds of millions of users, mature subscription revenue and positive cash flow, while LibTV is still relying on subsidies to drive growth and burning money to expand its scale.
The high valuation stems from its ultra-high growth rate. The performance growth rate of over 3000% plus the current premium for the popular AI application track gives LibTV higher imagination space. In the eyes of investors, LibTV is expected to sprint to the ecological niche of Jianying and replace Adobe to become the national creation tool of the next era.
But from the perspective of cost structure, compared with Jianying's heavy asset amortization logic, LibTV is still in the retail distribution logic.
The core assets of Jianying are self-developed AI models and massive user data. The marginal cost in the later stage can be amortized by its huge user base and internal ecosystem, and it has control over market pricing power.
The core assets of LibTV are product experience and workflow rather than underlying models. As a middleman, it needs to continuously purchase computing power and model invocation permissions from the upstream Volcano Engine at retail prices, and then sell them to consumers at prices far below the procurement cost.
The bigger problem lies in the moat.
At present, the differentiation that various platforms focus on can be roughly divided into two routes. LibTV takes the route of professional controllability + Agent invocation, using an infinite canvas to allow users to precisely control every shot like a professional director, and at the same time encapsulating more than 100 style templates for invocation.
Flova takes the route of natural language orchestration + low-threshold creation, using sub-Agents to lower the creation threshold to a one-sentence inspiration. TapNow focuses on node-based infinite canvas, and RunningHub's RHTV provides intelligent planning workflows.
But both of these two routes have their own risks. The more control you have, the higher the learning cost; the more automated you are, the more unpredictable the results will be. Closed Agents will weaken the sense of control of professional creators, and complex canvases will discourage ordinary users. The industry as a whole is leaning towards micro-innovation.
Once a function is verified to be effective, other platforms can quickly follow up and replicate it. A single update at the model level may smooth out all the workflow advantages of third-party platforms.
A typical case is the release of MiniMax H3.
It integrates editing logic, font typesetting, transition design, rhythm control,