With revenue exceeding 100 million yuan, this dark horse in the multimodal generative AI industry embarks on a new journey.
Nearly two years have passed since OpenAI released its text-to-video model Sora, but AIGC enterprises in China and the United States have shown completely different development trajectories: on one side, there is Sora 2 with persistently high costs that can never be scaled up, and the Sora App with almost zero user retention rate; on the other side, Chinese enterprises rooted in the vast application soil are getting better and better, ushering in an overall commercial explosion.
Intelligent Emergence recently learned that ZX AI, a generative AI startup focused on visual multimodality, has achieved an annual revenue of over 100 million yuan in 2025. Its C-end product vivago.ai also hit its peak download volume recently, with nearly 10 million new users added in January, ranking top 10 in the "Video Players & Editors" category on Google Play in more than 100 countries and regions around the world, showing huge commercial development potential.
Since its founding, ZX AI has successively released the HiDream-I1 image generation large model and the HiDream-E1 interactive editing model, which were fully open-sourced in April 2025, and topped the internationally authoritative AI evaluation list Artificial Analysis within 24 hours of open-sourcing.
This Hefei-based company has found a perfect balance between generation quality and efficiency through its self-developed large model with over 100 billion parameters and the world's first diffusion autoregressive architecture. At this stage, its products have been widely applied in cultural creativity, film and television, advertising and other fields.
Intelligent Emergence exclusively learned that the financing process of ZX AI is further accelerating: the Series B financing has entered the delivery stage, and the term sheet for the next round has been signed in advance. A close insider of the company revealed that both rounds of financing are worth hundreds of millions of yuan. Against the backdrop of intensified competition in the AI visual generation track, ZX AI has continuously obtained heavy investment from leading capitals with its hardcore technical strength and clear commercialization path.
What other commercial potential does ZX AI have at the moment when the implementation of multimodal applications is accelerating?
The most industrialized scientists, the most down-to-earth romance
From the very beginning of its establishment, ZX AI has found a pragmatic romance. Its founder Mei Tao is a foreign academician of the Canadian Academy of Engineering, who previously worked at Microsoft for 12 years. He has published more than 300 papers in the fields of multimedia analysis and computer vision, and won the best international paper award 15 times.
But Mei Tao's experience goes far beyond academia. In 2018, Mei Tao joined JD and served as the vice president of JD Exploration Research Institute. This career experience made him see the path from technology to commercial implementation.
When Mei Tao decided to found ZX AI, he had a clear vision. On one hand, multimodality is the most likely path to achieve general AGI, a view that later became the industry consensus. At the same time, in terms of commercial prospects, multimodality has a broader space than pure language models. "Currently, 50%-60% of global AIGC revenue comes from image and video related applications, which is higher than that of pure text models. When we made the entrepreneurial decision in 2023, multimodal companies like Midjourney had proved strong commercialization capabilities through SaaS tools, and clearly verified the product's market fit." Mei Tao told 36Kr in mid-2025.
And this is exactly Mei Tao's main battlefield, with profound accumulation in the fields of computer vision (CV) and multimodality.
However, for Chinese innovative enterprises at that time, Sora was a huge obstacle standing in front of them when they first entered the market. Considering its restoration of the physical world and amazing implementation effects, the industry at that time was quite expecting whether Chinese startups could produce generation results comparable to Sora.
A competition thus kicked off. After the release of Sora, it only took ZX AI half a year to launch its self-developed multimodal large model. In April 2025, ZX AI even open-sourced the image generation large model HiDream-I1 and the interactive editing model HiDream-E1 in one go, opening up the closed loop from dialogue to image creation. HiDream-I1 topped the authoritative list Artificial Analysis within 24 hours, becoming the first Chinese self-developed generative AI model to enter the global first echelon, and refreshed the industry records in the three dimensions of image quality, semantic understanding and artistic expression.
However, many entrepreneurs later reviewed and concluded that Sora was actually somewhat backward in terms of architectural innovation. Mei Tao also felt at that time that the overall functions of Sora were close to expectations. Only half a year later, with the entry of startups like ZX AI, OpenAI no longer has much advantage in the current video generation field. Especially from the perspective of product implementation, other products both overseas and in China are actually almost at the same level.
At the same time, in exploring the multimodal architecture paradigm, ZX AI is even at the forefront. The company first developed dual models for generation and understanding, and then planned the integration of understanding and generation, which is regarded as the best path to the physical world.
ZX AI has also been on the way to break through industry difficulties. In 2025, with the open sourcing of the latest model and the release of products such as vivago 2.0, Mei Tao also told 36Kr that the DiT (editor's note: Diffusion Transformer) architecture uses the powerful capabilities of Transformer to process video data, enabling AI models to efficiently model spatio-temporal relationships and flexibly generate videos of different resolutions, which is an important progress. However, for the entire generative AI industry, the realistic restoration of complex physical phenomena remains an unsolved problem: dynamic details that can be intuitively perceived by humans, such as the trajectory of splashing water droplets and the mechanical feedback of object collisions, are still in the exploration stage of "similar in form but different in spirit", and visual incongruity often appears in related scenarios.
ZX AI found an excellent balance between generation effect and running speed through the Sparse DiT architecture. Then through adversarial distillation technology, while increasing the reasoning efficiency, it greatly enhanced the details and aesthetics of the picture. This finally contributed to a number of creative achievements of ZX AI's HiDream-I1 model.
Take a unique path in algorithms to solve the last-mile problem
Different from the logic of large manufacturers that focus on basic models and expand parameters, small manufacturers pay more attention to innovation and implementation. In Mei Tao's view, this is also the value of ZX AI, which is to solve the last-mile implementation problem of AI.
He once told 36Kr, "From the first day of our entrepreneurship, we have a strong sense of crisis, thinking about how to find PMF. We took a relatively early and fast step in commercialization. Although we did not raise the most money, every penny we spent and every employee we recruited was carefully considered."
In the early stage of its establishment, ZX AI formed a "1+3+N" layout, that is, one core multimodal large model, driving three major products: creation tool platform, interactive marketing content tool and one-stop video creation Agent. Up to now, its services have covered more than 20 million individual users and more than 40,000 enterprise users around the world.
After positioning correctly, the core is to do a good job in delivery, serve customers well, and let AI generate real value.
Mei Tao told 36Kr that ZX AI has the most complete multimodal copyright corpus, hundreds of thousands of hours of copyright video materials and tens of thousands of authorized IPs in China. It not only covers 70% of domestic film and television data, but also has formed hundreds of millions of AIGC secondary creation materials, which are now widely used in film and television, cultural tourism, marketing and other scenarios.
"At Microsoft Research, we often say that it may take a hundred engineers to turn a technology into a product; to sell the product well, it may take another hundred solution experts or BDs, which shows how big the gap in the middle is. At that time, I thought I must find a place to open up the whole chain."
It is precisely this full-chain capability from technology to implementation that has made ZX AI highly favored by capitals since its birth.
In 2024, ZX AI completed hundreds of millions of yuan in Series A financing, led by Hefei Industrial Investment Group, with participation from institutions such as Anhui Provincial Artificial Mother Fund. At the end of 2025, JD Group increased its investment in ZX AI as a strategic investor. Its huge business scenarios behind it, including logistics, retail, health care and industry, are exactly the best implementation test fields and application fertile ground for multimodal AI technology.
Subsequently, insiders revealed that ZX AI had stepped up preparations for Series B financing and planned to complete the delivery in early 2026.
36Kr recently learned that ZX AI has successfully obtained the term sheet for the next round, in which old shareholders continued to increase their holdings, and new shareholders include industrial capitals, listed companies with deep business cooperation and well-known investment institutions. At present, the amount of Series B financing has reached hundreds of millions of yuan.
Yuan Guoliang, CEO of Shanghai Dunhong Asset, commented on ZX AI: "We firmly believe that video generation technology, as a new generation of productivity tools, will fully empower thousands of industries. Especially in the e-commerce field, video has become the core medium connecting commodities and consumers. HiDream has initially verified its application value and commercial potential in e-commerce scenarios through products, reflecting that the team not only understands technology, but also understands the industry. At the same time, we believe that its technical architecture and evolution direction have the possibility of expanding to a world model with more versatility and cognitive depth, which is a leap of underlying capability. We look forward to exploring the long-term path of technology and industry integration with the team, and helping promote multimodal generation to become a universal and intelligent industry infrastructure."
The best target with both commercial strength and architectural innovation
2025 is the first year of explosion for China's multimodal generative AI. With the maturity of AIGC technology, productivity and creativity have been significantly improved, driving the application market to show an explosive growth trend. According to IDC data, the compound annual growth rate of the global generative AI market is expected to reach 63.8% in the next five years, and will reach 284.2 billion US dollars by 2028, accounting for 35% of total AI investment. ZX AI has become a beneficiary of this with its extremely strong technical strength and industrial implementation thinking. The company's commercialization process is rapid, and 36Kr learned that ZX AI's annual revenue in 2025 has exceeded 100 million yuan.
The ability to quickly achieve such results in the highly competitive multimodal generation field benefits from ZX AI's unique business model thinking and strong underlying innovation capabilities. It can be said that ZX AI is one of the few enterprises in the industry that pays equal attention to both commercialization and technological innovation.
During the three years since its establishment, ZX AI has experienced different business models. The model in 2023 was MaaS, selling models and APIs, similar to the PaaS model of cloud computing. The model in 2024 was SaaS, mainly selling tools to allow users to use tools on ZX AI's platform to produce content.
Nowadays, it has upgraded its model and officially transformed to RaaS, that is, a result-delivered, user-value-oriented business model, including tools, content materials, and limited video production/delivery only charge a small basic fee, and the main income comes from commission sharing after the increase of customers' GMV. According to Mei Tao, he believes that such customer value is relatively clear, which can basically realize zero-risk investment and share incremental benefits.
As entrepreneurship is getting better, Mei Tao also said that he has found a balance between commercial returns and capability improvement. On one hand, he will continue to raise the bar and do a good job in the research of vertical basic models, and a more powerful underlying architecture with more advanced methods will surely lay a better foundation for the model's capabilities. In addition to independent research and development, ZX AI also embraces a broader ecosystem through open source to increase the possibility of success. On the other hand, it still focuses on solving the last-mile problem, digging into the actual scenario needs of users, opening up more vertical data in industries such as education, e-commerce and cultural tourism, doing fine-tuning, and truly solving industry problems.
Intelligent Emergence also learned that ZX AI is currently developing a new generation of multimodal generation architecture with multimodal reasoning drive and unlimited memory, which will greatly improve the model's reasoning capability and achieve a higher level of horizontal scaling up between multiple tasks.
Nowadays, with the resonance of technology, market and policy, the industry is also realizing that AI video is no longer a geek's toy, but a productivity tool that can directly generate cash flow. Since last year, popular AIGC videos generated by AI such as "Cat and Dog Sports Meeting" and "Knife Cutting Glass and Fruit" have become popular on social platforms, attracting more and more creators to join in. It is a common choice for leading players and ordinary C-end users, which ultimately accelerates the commercialization process of the video generation track.
According to data from international research institution Fortune Business Insights, the global scale of AI video generation in 2024 was about 620 million US dollars, and it is expected to reach 2.56 billion US dollars in 2032, with a compound annual growth rate of 20% from 2025 to 2032.
At this stage, AIGC has been the mainstream choice in marketing and specific content fields. A more promising prospect is that when the model can stably solve the problems of character consistency and long-term sequence coherence, AIGC will detonate the market in high-end applications such as film and television and games. When the model breaks through the problem of understanding and generation consistency, it can truly understand the physical world and generate more realistic and controllable content and details. At that time, it will be the real explosion moment of the video generation track. In this race, ZX AI has been at the forefront.