HomeArticle

What exactly do capital players see in the 1.5 billion-yuan bet on native full multimodality?

晓曦2026-07-29 11:31
Half a step ahead.

Zixiang Future, a native full-modal large model company, officially announced the completion of a new round of RMB 1.5 billion financing. This is its third round of financing in nearly three months, with cumulative financing exceeding RMB 2.1 billion and a post-money valuation of more than USD 1 billion, making it a new unicorn.

This round of financing was led by Sichuan Revitalization Science and Technology Innovation Fund under the National Social Security Fund, ICBC Capital, Hongyi Asset Management, and Dunhong Capital, with Shanghai Film New Vision Fund, Xiamen International Trade Capital, Huace Film and TV, and other institutions participating as follow-on investors, while old shareholders including Hefei Industrial Investment, Oriental Fortune Capital increased their stakes. State-owned capital, financial institutions, market-oriented funds and film & TV industry capital are all involved in this financing round.

What is particularly noteworthy is the bet from the film and TV industry capital: Industrial players like Shanghai Film Group and Huace Film and TV used to invest in content, IP and cinema chains in the past, but this time they have placed their bets on the upstream foundational model. Though it is a follow-on investment, it raises a real question — as AI begins to restructure content production, what exactly are the two leading film and TV industry funds betting on when they invest in a large model company?

Part01

Models and applications are a set of interlocking gears

In 2026, the AI industry returns to the cycle of "model is productivity". After Zhipu AI and MiniMax went public one after another, the balance of capital has shifted back from commercialization to model capability itself. Mei Tao, founder of Zixiang Future, judges that the industry is returning to 2023 — "The model is productivity, and the model is the product".

Different from the judgment of some "model-only theories", Zixiang has never defined itself as a pure model company since the first day of its establishment.

Mei Tao worked at Microsoft Research Asia for 12 years, leading the team to develop TGANs-C, one of the world's leading video generation models, which made him believe that there are still huge opportunities for model innovation. His five years of experience in charge of AI business and industrial implementation at JD made him see another truth: Technological leadership does not equal commercial success, and model capability will not automatically translate into market advantages.

This determines that Zixiang follows a two-wheel-driven path of "model + application": the underlying layer is the self-developed native full-modal architecture, and the upper layer is applications that can truly run into industrial processes. In Mei Tao's view, the two wheels are not two independent business lines, but a set of interlocking gears — applications provide data and feedback from real scenarios for the model, and the model provides application with capability momentum that other players cannot catch up with in the short term.

This two-wheel system has verified the commercial closed loop in scenarios such as marketing and film & TV. But what really determines the ceiling of Zixiang is the underlying architecture that iterates "half a step faster".

Part02

Must Get a Half-step Head Start

Mei Tao has a widely spread metaphor: "Large tech giants are like a machine gun with endless bullets; startups only have one magazine, and every bullet must be aimed at the future."

In his view, falling behind once in model evaluation is not fatal for a startup. The real trouble is being dragged by large tech giants into a war of attrition that tests who can last longer. The reserves of computing power, capital and talents determine that the two sides are not in the same weight class. The pragmatic way for startups to survive is to be half a step faster than large tech giants in underlying architecture innovation — every bullet must hit the direction that others have not yet reacted to.

This half-step advantage of Zixiang is based on a unified architecture called Universal Interactive Transformer (UiT), the native full-modal architecture.

The mainstream practice in the industry is that the image generation model first uses VAE to compress the image, and then sends it to the respective independent text encoder and generation model for separate processing. This method saves computing power, but the cost is that high-frequency details such as text strokes and material textures are easily lost during compression and reconstruction, and text and images are encoded separately, which may easily lead to semantic misalignment.

The native full-modal architecture (UiT) does not use external VAE or separately pre-trained text encoders. Instead, it puts original pixels, text Tokens and task conditions into the same space, and processes them with the same set of Transformers. After all modalities are connected, it can truly realize "Any to Any", supporting arbitrary input and arbitrary output, which is exactly the capability required by the world model — to understand, generate and predict different states of the real world in a unified architecture. This is a preemptive layout in terms of architectural paradigm.

This half-step preemptive layout is gradually delivering results. In May 2026, Zixiang's open-source model HiDream-O1-Image ranked first globally on the Artificial Analysis open-source list; in June, the commercial version HiDream-O1-Image-1.5 reached the third place in the world.

(Note: The screenshot of the list is taken on June 22, 2026)

Looking at the whole track, the weight of this half-step advantage will be clearer. Most manufacturers in the industry start from a single point of "video generation"; since its establishment, Zixiang has insisted on advancing both image generation and video generation in a "dual-mode" way, and through the innovation of native full-modal (UiT) architecture, it has realized the unification of image, video, voice, 3D interaction, motion and other modalities. This is the difference in development path, and it is also the real source of "half-step faster" — it is not betting on a single ranking on the list, but betting that the time window of the native full-modal path is longer than that of general large language models, with a much higher ceiling.

Part03

"Planet Test"

The value of the model path will eventually be perceived by users through agents. At the just-concluded 2026 WAIC, Zixiang Future released vivago R1, calling it an "infinite-duration content creation agent".

Its best proof comes from a tricky test question: a product evaluator asked vivago R1 to shoot a "Turkish bootleg Star Wars" — a super low-grade cult film with cardboard monsters, colanders as helmets and foam plastic stones.

The difficulty is subtle: AI's instinct is to render the picture as exquisite as possible, but this test requires it to maintain the cheap feeling of "ridiculously fake". In the one-minute-plus final film, Captain Ali with a moustache and red clothes, and the cardboard monster with a colander as a helmet keep the same image from beginning to end, and are not secretly rendered into exquisite CGI — the upper-layer Agent orchestration suppresses the aesthetic inertia of the underlying model.

This consistency does not come from luck. R1 runs on a structured pipeline: write story → revise script → split scenes → confirm art style → lock reference images of characters and scenes → generate shot by shot → stitch into the final film.

The so-called "infinite duration" also needs to be demystified: it does not spit out a two-hour movie at one time, but solidifies characters, props and scenes as reusable assets in the long-term production process, so that the story can continue to be extended. This exactly hits what industrial capital attaches great importance to — enabling IP to evolve from a static copyright inventory to a digital asset that can be called repeatedly.

vivago follows a hybrid path of "self-developed base model + aggregated scheduling + upper-layer orchestration". Its moat does not lie in "all the pictures are self-developed", but in orchestration capability and long-term consistency. Aggregation is a means, orchestration is the product, and the final user experience is the core battlefield that all players must compete for.

Part04

Running Into Industrial Processes

With this investment from industrial capital, what tests Zixiang Future is no longer just its ranking on the model list, but whether it can truly integrate into industrial processes.

For Gu Yuhao, General Manager of Shanghai Film New Vision Fund, investing in Zixiang Future is not only focusing on its product capability, but also "building a technical base for the future film and TV industrial system in advance". He wants to see whether this company can connect the underlying model, middle platform and vertical agents, integrate scripts, characters, scenes, shots and sounds into the same digital system, so that IP can be transformed from static copyright into reusable digital assets.

In Gu Yuhao's words, "Film and TV customers will not pay according to the ranking on the list, and they will ultimately judge by the delivery results." This is recognition, but also a more difficult exam: there is a long "industrialization gap" between an amazing demo and stable, large-scale delivery.

If Gu Yuhao's judgment gives a strategic answer to "why invest", then Zhang Situo, Vice President of Huace Film and TV, gives a judgment that has been verified in actual production lines.

In Zhang Situo's view, the first value created by AI does not lie in the creative end, but in the links that "take time and effort but are not technically demanding" — initial script screening, material sorting, multi-version editing, subtitle translation and localization. In the hit drama *The Peaceful Era* produced by Huace, AI is used to generate background materials for some scenes, which greatly shortens the production cycle and strongly supports Huace's global distribution for 68 million overseas subscribers.

More importantly, his standard for judging whether a model is worthy of cooperation, from the perspective of film and TV companies, is not "whether it can generate beautiful pictures", but "whether it has director thinking — understands shot language and narrative rhythm".

This sentence exactly echoes the structured planning of vivago R1 that "write the script first, then lock the reference images, and generate shot by shot". The "director thinking" at the model level is exactly the reason why industrial players are willing to place their bets.

Following this standard, Zhang Situo breaks down the competitiveness upgrading of film and TV companies into three steps: Efficiency Competition → Asset Competition → Ecology Competition. He values the second step the most — "In the past, the most valuable asset of film and TV companies was the copyright library, but now static copyright needs to be upgraded to 'computable content assets': scripts can be retrieved by AI, finished films can be disassembled and recombined, and production experience can be precipitated and reused."

This is almost two different expressions of the same idea as Gu Yuhao's judgment on "digital assets".

The two capitals also have surprisingly consistent views on the implementation pace. Huace's cooperation path starts from the direction of "fast to verify" such as AI high-quality short dramas, and moves towards the secondary development of IP digital materials and full-process AI collaboration for long head dramas in the medium and long term; Gu Yuhao believes that at this stage, AI is more suitable for promotion materials, pre-development, animation, manhua-adapted dramas and stylized short films, while live-action long films are much more difficult. The two film and TV capitals have given the same schedule: focus on auxiliary tools and lightweight content in the short term, and talk about restructuring the industry in the long run.

In terms of industry, the preference of state-owned capital and industrial capital has always been clear: investing in a company is expected to drive several industrial chains. Visual AI is exactly the layer of large models that is extremely close to the industry, with ready-made scenarios such as film & TV, cultural tourism, marketing, and e-commerce, and the results delivered by Zixiang can withstand item-by-item acceptance. The capital flows through the company into the industry, and can drive large-scale development with one fulcrum.

Part05

The Longer-term Goal is the World Model

Zixiang Future sets its longer-term goal on the world model, but Mei Tao always holds a prudent attitude towards this concept: "We do not think the industry has fully realized the world model today." In his view, an image represents the spatial state at a certain moment, and a video adds the time dimension; following this path, the model still needs to understand spatial structure, motion consequences, causal relationships and physical laws.

From this perspective, what Zixiang Future really wants to promote is not a single product, but a more underlying native full-modal path, which enables different modalities including text, image, video, audio and motion to be aligned at the bottom layer, and finally realizes arbitrary modality input and arbitrary modality output. So that the world model can truly realize the understanding, prediction and reconstruction of the real world.

From visual generation to native full-modal world model, the long-term proposition of this path has always been only one: to enable AI not only to generate content, but also to gradually learn how the world works, and have the ability to deduce and reconstruct the world.