HomeArticle

AI Video Generation Showdown in July: Over 200 Billion Yuan of Capital Pours Into the Track, Who Has Made It to the Final Round?

IT桔子2026-07-28 15:20
July is not even over yet, and the AI video generation track has already become scorching hot.

Before July even ended, the AI video generation track had already reached a boiling point.

From July 3 to July 23, over a mere 20 days, 5 companies successively announced large-scale financing deals. Kling AI secured $3 billion, Shengshu Technology raised $500 million, Aishi Technology completed a multi-hundred-million-dollar Series C+ round, Zhixiang Future closed a 1.5 billion RMB Series C round, and even FlovaAI, a new firm founded less than a year ago, landed an $80 million angel round.

The combined financing of the top 4 unicorn-level companies exceeded 26.2 billion RMB. In other words, the capital influx in July 2026 alone surpassed the total funding the entire track attracted over the previous two years (2024-2025).

Why did all these financings cluster in July? What exactly happened? This is certainly no coincidence.

Kling AI: Historic Joint Investment from BAT, ARR Quadruples in One Year

On July 3, Kling AI announced a Series A financing with a $3 billion maximum capital increase limit, led by CPE Yuanfeng and Guofang Innovation, with 34 institutions participating. The investor list features a rare simultaneous appearance of Tencent, Alibaba Cloud, and Baidu. Lighthouse Capital served as the exclusive financial advisor.

The $3 billion Series A capital increase cap corresponds to a post-money valuation of approximately $18 billion, with approximately 19 billion RMB already signed for the committed portion.

The reason investors dare to assign this valuation lies in the market confidence boosted by Kuaishou's 2026 Q1 earnings report.

Kuaishou's 2026 Q1 financial disclosure shows Kling AI's quarterly revenue exceeded 650 million RMB, growing over 300% year-over-year. Its Annual Recurring Revenue (ARR) skyrocketed from $100 million in March 2025 to nearly $500 million in March 2026, a 4x year-over-year increase.

This figure stands out as a clear outlier in the AI video track — apart from ByteDance's Seedance, no other domestic player has reached this scale.

Kling's revenue is driven by two pillars: API calls from B2B enterprise clients and paid consumer subscriptions. As of June 2026, Kling AI has surpassed 100 million global users and nearly 50,000 enterprise clients. For example, Kling was deeply involved in creating virtual scenes and special effects for the hit Chinese historical drama "The Peaceful Era", and supported the generation of hundreds of high-quality shots for the Hollywood series "The House of David".

In terms of products, the company launched its 3.0 series of models this February, enabling up to 15-second video generation, storyboard control, subject consistency, and synchronized audio-video output.

Two noteworthy details stand out: First, the valuation is slightly lower than the rumored $20 billion target from May, with a more pragmatic pricing strategy that actually accelerated transaction execution. Second, the announcement includes a hidden performance clause — if Beijing Kling fails to complete an IPO before October 2031, investors have the right to demand share repurchases at principal plus 8% annual simple interest.

The combination of corporate spin-off, financing, equity incentives, and repurchase clauses sends a clear signal: Kling is paving the way for an independent public listing.

Shengshu Technology: $500 Million Financing on the Eve of IPO, Evolving from Video Models to "World Models"

On July 6, Shengshu Technology announced a $500 million Series B+ financing round.

Looking at the full timeline, the company secured 3 funding rounds within 5 months: over 600 million RMB in Series A+ (led by Zhongguancun Science City and Xinglian Capital) in February, 2 billion RMB in Series B (led by Alibaba Cloud) in April. Adding this July's $500 million round, total cumulative financing has exceeded 5 billion RMB.

Three rounds in 5 months — this pace indicates that capital is rushing to get on board. Why the urgency?

By the end of March this year, Shengshu Technology had completed shareholding restructuring — a standard pre-IPO procedure. Market sources report it could launch its Hong Kong IPO process as early as the first half of the year, explaining why investors are scrambling to enter ahead of the public listing.

Shengshu Technology's core product Vidu launched globally in July 2024, and has iterated to the Vidu S1 version three years later — a real-time interactive model that supports voice-controlled framing and unlimited duration generation (from 1 minute to 2 hours).

In 2025, Shengshu Technology achieved over 10x growth in both users and revenue, generating more than 400 million videos cumulatively. Its B2B client roster is particularly impressive: in advertising and e-commerce, partners include JD.com, Alibaba 1688, Amazon, Meituan, Focus Media, Blue Focus, L'Oréal, and Anta. In film and animation, it serves Tencent Animation, China Literature, CCTV Animation, iQiyi, and Mango TV. In gaming, its clients include Lilith Games and 37 Interactive Entertainment.

Not content with only video generation, Shengshu Technology began expanding into embodied intelligence this year.

In April 2026, the company officially released its general-purpose world action model Motubrain, adopting the World Action Model (WAM) technical route that unifies perception, prediction, and action modeling. This endows robots with capabilities for "unified foresight" and "unified multi-functionality", while the commercial version of MotuBrain has entered industrial-level validation.

This Tsinghua University-founded company, only 3 years old, is rapidly evolving from a "video model company" into a "general-purpose world model platform". After the $500 million funding arrives, the capital will be primarily allocated to R&D for the "general-purpose world model". This is the core narrative that truly attracts investors.

Aishi Technology: A Unicorn in Three Years with 150 Million Global Users

On July 14, Aishi Technology's Series C+ round was led by Alibaba. Combined with the $300 million Series C round completed in March (led by CDH Investments, with nearly 20 participants including China Ruyi and 37 Interactive Entertainment), the total Series C financing reached 2.98 billion RMB, pushing post-money valuation past $2 billion.

Alibaba previously led Aishi's $60 million Series B round in September 2025, and is increasing its stake again in this Series C+ round. Ant Group also led a Series A2 round of over 100 million RMB back in April 2024.

Alibaba Group's continuous investment in Aishi demonstrates that it views the company as a strategic layout point in the AI video track.

As of March 2026, Aishi Technology has exceeded 150 million global users, with 15 million monthly active users and an Annual Recurring Revenue (ARR) of $40 million.

Notably, co-founder Xie Xuzhang revealed that Aishi's training cost for developing comparable models is only about 10% of its peers. This cost advantage is achieved through key measures: filtering high-quality data to reduce ineffective training iterations, leveraging an optimized Diffusion Transformer architecture to improve resource utilization and lower per-training costs, and reusing engineering experience across projects.

This January, Aishi released PixVerse R1, claiming it as the world's first general-purpose real-time world model supporting 1080P resolution. Unlike traditional "pre-recorded" video generation methods, R1 enables "real-time dynamic generation" — users can input new instructions at any point during video playback, and the visuals achieve natural, smooth transitions in lighting and camera angles in approximately 0.5 seconds.

Aishi takes a different path from other players: while competitors focus on generation quality, it prioritizes interactive capabilities. Just one week after R1's launch, China Ruyi announced a strategic partnership with Aishi and invested $14.2 million.

The latest Series C+ financing explicitly states its use of funds will target basic video generation models, real-time world models, and global business expansion.

Zhixiang Future: A New Unicorn with the Alternative Strategy of "Model + Agent + Hardware"

Zhixiang Future is the "new unicorn" on this list — after its 1.5 billion RMB Series C round closed, its post-money valuation surpassed $1 billion. Founder Mei Tao is a Foreign Fellow of the Canadian Academy of Engineering and former Vice President of JD.com.

The company intensively completed three rounds of financing in nearly three months, raising over 2.1 billion RMB cumulatively.

On July 23, Zhixiang Future announced its 1.5 billion RMB Series C round, led by the National Social Security Fund, Sichuan Industrial Revitalization Fund, and ICBC Investment, with participation from film industry capital including SMG New Vision Fund and Huace Film & TV, plus 18 other institutions. This combination of national long-term capital, local state-owned assets, and industrial capital is rarely seen in the AI video track.

Zhixiang launched the world's first publicly accessible video generation model based on the DiT architecture back in May 2024. At the recently concluded WAIC 2026, it released its multimodal creative agent vivago R1, focused on unlimited-duration video generation and editing.

Currently, its products cover over 100 countries worldwide, serving more than 50 million users and 40,000 enterprise clients. Its full-year 2025 revenue exceeded 100 million RMB, and 2026 Q1 revenue already surpassed the entire previous year. During the 2026 Chain Expo, Mei Tao stated that due to the high pre-training costs of large models, the company expects to achieve monthly break-even in 2029.

Mei Tao put it plainly: "We won't compete with ByteDance on underlying capabilities. We focus on commercial marketing and professional film and television collaboration."

Zhixiang's path is clear: avoid direct confrontation with large tech firms, and deeply cultivate vertical scenarios. It first builds its user base through image models, then evolves toward video and world models.

FlovaAI: $80 Million Angel Round, Guo Lie's Third Entrepreneurial Venture

Guo Lie, founder of FlovaAI, is a prominent figure in China's mobile internet history. In 2013, he developed FaceQ, followed by FaceU. In 2018, his team was acquired by ByteDance for $300 million. After that, he participated in incubating BeautyCam, Xingtu, and Jianying at ByteDance — yes, Jianying was developed with his involvement.

In 2025, Guo Lie embarked on his new venture, founding Shenzhen Yuzhou Technology, and launched Flova.ai in October. HSG, IDG, and Yunji Capital collectively invested over $80 million.

Investors dared to commit $80 million at the angel round not just for the product, but for Guo Lie himself. Every consumer-facing tool product he previously developed became a phenomenal hit.

Flova is an AI-native video creation Agent platform. Users express their creative needs through natural conversation, and the Agent completes the entire workflow from scriptwriting, character design, and storyboarding, to image/video/audio generation and timeline editing assembly.

Different from the mainstream canvas-based tools on the market, Flova features a unique "storyboard" workspace. The Agent receives instructions in the chat dialog, displays its progress on the left-side storyboard, and users view results directly in the central preview area, with the ability to double-click to intervene and make edits at any time if unsatisfied.

Multiple early users report that Flova's standout feature is its "memory" — after creating more than a dozen episodes, it can still remember a specific storyboard from the very first episode. This long-context memory capability is highly distinctive among similar products.

At this stage, Flova adopts a zero-gross-margin strategy. Guo Lie said: "During the development phase, we will strive to achieve zero gross margin." The team is not in a rush to pursue short-term revenue, but instead focuses on effective consumption ratios: what percentage of Token usage is perceived as valuable by users, and what portion represents ineffective waste caused by product imperfections.

Why All the Rush in July? Plausible Explanations

Returning to the core question: why did these 5 companies coincidentally complete their financings intensively in July? Piecing together all the information, four factors create a perfect resonance.

First, ByteDance Seedance's "siphon effect" forced everyone to accelerate.

Since ByteDance Seedance 2.0 launched this February, its monthly revenue has exceeded 1 billion RMB, with approximately 95% penetration in the short drama industry, and holding over 80% of the market share by daily Token consumption volume.

Volcano Engine raised its 2026 MaaS revenue target from last year's 1.5 billion RMB to 15 billion RMB — a 10x increase almost entirely driven by the Seedance model. Moreover, Seedance 2.1 is soon to be released, with expected 20% performance improvements and a low-tier version priced as low as 0.5 yuan per second.

As ByteDance executes a dual-pronged attack of "disproportionate model capability lead + price war", other players that don't quickly raise capital, stockpile computing power, and accelerate iteration may lose their very right to stay in the game.

Everyone is seeking a fulcrum to counter ByteDance in the long run, but with completely different strategies: Kling takes a platform-based full-modal path, Shengshu pursues a tech-nerd route, Aishi focuses on differentiated real-time interaction, Zhixiang deeply cultivates vertical scenarios, and FlovaAI adopts an Agent workflow approach.

Investors are not betting on a single narrative, but placing wagers on different technical routes and business models.

Second, commercialization has been validated, and capital sees clear exit paths.

Last year, the industry was still debating "whether AI video can be profitable", but this year's data speaks for itself: Kling's ARR is nearly $500 million, Shengshu achieved over 10x revenue growth in 2025, Aishi's ARR stands at $40 million, and Seedance's monthly revenue exceeds 1 billion RMB.

Paid scenarios for AI video across advertising, e-commerce, short dramas, gaming, and film have all been verified one by one. This state of "visible cash flow" is the direct trigger for concentrated capital entry.

Third, the track is upgrading from "video generation" to "world models", elevating the industry narrative.

In the first half of this year, the competitive landscape of foundational large models became increasingly concentrated, and primary market funds began searching for the next growth direction, with multimodal generation and world models emerging as the consensus breakthrough.

Notice this trend: none of the 4 companies that secured massive financing are positioning themselves as mere "video generation" players. Kling is promoting an