Tsinghua-backed AI video unicorn plans to seek a Hong Kong IPO
Shengshu Technology, a generative AI company founded by Zhu Jun, a professor at Tsinghua University, and heavily backed by Alibaba, is reportedly preparing for a Hong Kong stock listing.
On August 18, multiple media outlets cited sources familiar with the matter saying that Shengshu Technology is considering an initial public offering in Hong Kong, with a planned fundraising scale of over 500 million US dollars (about 3.9 billion Hong Kong dollars). It has already carried out listing-related cooperation with China International Capital Corporation and CITIC Securities, and is expected to officially submit the application as early as next year.
In response to the above rumors, Shengshu Technology, CICC and CITIC Securities all stated that they would not comment.
Signals of listing preparation have been released long ago. On March 30 this year, the main entity of the company completed the shareholding system reform, changing from "Beijing Shengshu Technology Co., Ltd." to "Beijing Shengshu Technology Co., Ltd. (unlisted)"; as early as April, market news spread that the company would launch the Hong Kong stock IPO process as soon as the first half of 2026.
01
Founded in March 2023, Shengshu Technology's founder Zhu Jun is a professor in the Department of Computer Science and Technology of Tsinghua University, IEEE Fellow, and also serves as the Vice President of Tsinghua University Artificial Intelligence Research Institute and the Director of the State Key Laboratory of Intelligent Technology and Systems. Its core founding team is one of the world's earliest research groups deeply engaged in deep probabilistic generative models, laying a distinct Tsinghua-origin technology foundation for the company.
In the three years since its establishment, the company's financing pace has continued to accelerate, and it has completed three rounds of heavyweight financing only since 2026.
In February, it completed the over 600 million yuan A+ round financing; in April, it secured the nearly 2 billion yuan B round financing led by Alibaba Cloud, with participation from institutions including China Investment Corporation, Jiu'an Haitang, and TAL Education Group. The post-investment valuation exceeded 2 billion US dollars, with total fundraising of over 2.6 billion yuan within two months. In July, it completed another new round of financing of 500 million US dollars, setting a new record for the largest single financing in China's general world model field.
In terms of shareholder lineup, Alibaba's affiliated entities did not enter in the later stage — just 3 months after the company was founded in 2023, Ant Group led the angel round investment, with Baidu Ventures and Qiming Venture Partners as early continuous follow-on investors. It also introduced state-owned industrial capital such as Zhongguancun Science City, with a shareholder structure covering industrial, financial and state-owned backgrounds.
On the technology and product side, Shengshu Technology started with video generation, and gradually extended to the world model and embodied intelligence sectors.
In April 2024, the company released its first-generation large video model Vidu, becoming the first text-to-video product in China that fully benchmarks against Sora. The Vidu Q1 iteratively launched in April 2025 surpassed models such as Runway, OpenAI Sora, and Kuaishou Ling on both the authoritative video generation evaluation benchmarks VBench-1.0 and VBench-2.0, taking the first place on both lists in the text-to-video track.
In terms of commercialization, 8 months after its launch, Vidu's annual recurring revenue (ARR) exceeded 20 million US dollars (about 140 million yuan). Its users cover more than 200 countries and regions around the world, and both the user scale and revenue achieved more than 10 times growth in 2025.
At present, Vidu has been deployed in multiple industries including advertising, e-commerce, animation, cultural tourism, radio and television, and education, and has reached in-depth cooperation with enterprise-level platforms such as Feishu. With the generation cost as low as 1.34 yuan for 1080P 5-second videos, its cost-performance advantage has become the core support for the rapid growth of commercialization.
In July 2026, the company released the Vidu S1 model for real-time interactive scenarios, supporting real-time video calls and voice-controlled video progression. Users can manipulate the behavior of digital humans through voice commands to achieve continuous interaction of unlimited duration.
02
Apart from video generation, Shengshu Technology has extended its technical base to the embodied intelligence track in the physical world.
In December 2025, the company jointly released the general base world model Motus with Tsinghua University, positioned as the "brain" for embodied intelligence in the real world. Motubrain launched in April 2026 won the first place in two international authoritative benchmarks, WorldArena (world modeling capability) and RoboTwin 2.0 (task execution capability), refreshing the SOTA with an average success rate of 95.8%. It is the only model in the list with an average score of over 95 in the random interference environment, and has the generalization capability across different robot entities and task scenarios.
With the homologous deep probabilistic generative technology base, Shengshu Technology has become one of the few domestic manufacturers that lay out both AI video generation and embodied intelligence world models, with significantly higher technology reusability than companies in a single track.
03
Behind the rapid growth, track competition and commercialization challenges are equally clear.
The AI video generation track has entered a stage of fierce competition among major manufacturers. Head players such as Kuaishou Ling, ByteDance Seedance, and Tencent Hunyuan have all launched mature products, and industry competition continues to intensify. The embodied intelligence track has also attracted dozens of participants, and there are uncertainties in the landing progress and commercialization rhythm.
Compared with large text models, the commercialization threshold for video generation is higher: training data copyright, content security compliance, generation duration and picture consistency, reasoning cost control, etc., are all core factors restricting large-scale implementation. At the same time, the C-end payment habit for large video models is still in the cultivation period, and user retention and payment conversion have not been verified for a long time. The implementation cycle of embodied intelligence business is longer, and it is difficult to contribute large-scale revenue in the short term. The company's performance growth still highly depends on the single video generation track.
If this IPO fundraising of more than 500 million US dollars is successfully completed, Shengshu Technology will become the first enterprise in China's video generation track to land on the capital market. Whether the company's deep technical barriers can be transformed into sustained and stable profitability will be the core proposition of market attention after listing.