HomeArticle

What exactly do capital players see in the 1.5 billion yuan bet on native full-modality?

凡泛2026-07-29 11:31
Half a step ahead

Zhi Xiang Future, a native full-modal large model company, has officially announced the completion of a new round of RMB 1.5 billion financing. This is its third round of financing in nearly three months, with total accumulated financing exceeding RMB 2.1 billion and a post-money valuation of over USD 1 billion, making it a new unicorn.

This round of financing is led by Sichuan Revitalization Science and Technology Innovation Fund under the National Social Security Fund, ICBC Capital, Hongyi Asset Management, and Dunhong Capital, with follow-on investments from Shanghai Film New Vision Fund, Xiamen Guomao Capital, Huace Film and Television, etc., and additional capital injection from existing shareholders including Hefei Industrial Investment, Oriental Fortune Capital, etc. State-owned capital, financial institutions, market-oriented funds and film and television industry capital are all involved in this deal.

What is particularly noteworthy is the bet from the film and television industry capital: industrial players such as Shanghai Film Group and Huace Film and Television used to invest in content, IP and cinema chains in the past, but this time they have placed their bets on the upstream foundational model. Although it is only a follow-on investment, it raises a real question — when AI begins to reconstruct content production, what exactly are the two leading film and television industry funds betting on by backing a large model company?

Part01

Models and applications are a set of interlocking gears

In 2026, the AI industry returns to the cycle of "model is productivity". After Zhipu AI and MiniMax went public one after another, the balance of capital has shifted back from commercialization to model capability itself. Mei Tao, founder of Zhi Xiang Future, judges that the industry is returning to 2023 — "The model is productivity, and the model is the product".

Different from the judgment of some "model-only theory" advocates, Zhi Xiang Future has never defined itself as a pure model company since its first day of establishment.

Mei Tao spent 12 years at Microsoft Research Asia, leading his team to develop TGANs-C, one of the world's leading video generation models, which made him believe that there are still huge opportunities for model innovation; his five-year experience in charge of AI business and industrial implementation at JD.com made him see another truth clearly: Technological leadership does not equal commercial success, and model capabilities will not automatically translate into market advantages.

This determines that Zhi Xiang Future takes a two-wheel-driven path of "model + application": the underlying layer is its self-developed native full-modal architecture, and the upper layer is applications that can truly be integrated into industrial processes. In Mei Tao's view, the two wheels are not two independent business lines, but a set of interlocking gears — applications provide data and feedback from real scenarios for the model, and the model provides application with capability momentum that other players cannot catch up with in the short term.

This two-wheel system has verified the commercial closed loop in scenarios such as marketing and film and television. But what truly determines the ceiling of Zhi Xiang Future is the underlying architecture that iterates "half a step faster".

Part02

Must get a head start by half a step

Mei Tao has a widely spread metaphor: "Large tech giants are like a machine gun with endless bullets; start-ups only have one magazine, and every bullet must be fired towards the future."

In his opinion, falling behind in one model evaluation is not fatal for a start-up. The real trouble is being dragged by large tech giants into a war of attrition that tests who can hold on longer. The reserves of computing power, capital and talents determine that the two sides are not at the same level. The pragmatic way for start-ups to survive is to be half a step faster than large tech giants in underlying architecture innovation — every bullet must be fired in a direction that others have not yet reacted to.

This half-step lead of Zhi Xiang Future is rooted in a unified architecture called Universal Interactive Transformer (UiT), a native full-modal system.

The mainstream practice in the industry is that the image generation model first uses VAE to compress the image, and then hands it over to separate independent text encoders and generation models for processing respectively. This method saves computing power, but at the cost of easily losing high-frequency details such as text strokes and material textures during compression and reconstruction; with text and images encoded separately, semantic alignment is also prone to deviation.

The native full-modal architecture (UiT) directly abandons external VAE and separately pre-trained text encoders. Instead, it puts original pixels, text Tokens and task conditions into the same space, and processes them with the same set of Transformers. After all modalities are connected, it can truly realize "Any to Any", which means any input supports any output. This is exactly the capability required by a world model — to understand, generate and predict different states of the real world in a unified architecture. This is a pre-emptive move in terms of architectural paradigm.

This half-step pre-emptive move is gradually delivering results. In May 2026, Zhi Xiang Future's open-source model HiDream-O1-Image ranked first globally on the Artificial Analysis open-source list; in June, the commercial version HiDream-O1-Image-1.5 rose to third place worldwide.

(Note: The screenshot of this list was taken on June 22, 2026)

Looking at the entire track, the value of this half-step lead will become clearer. Most players in the industry cut in from a single point of "video generation"; while Zhi Xiang Future has adhered to the simultaneous development of "dual modes" of image generation and video generation since its inception, and is working to unify modalities including image, video, audio, 3D interaction and motion through the innovation of the native full-modal (UiT) architecture. This is a difference in development path, and also the real source of the "half-step faster" advantage — it is not betting on a single ranking on the list, but on the fact that the time window for the native full-modal path is longer than that of general large language models, with a much higher ceiling.

Part03

"Planet Test"

The value of the model path will eventually be perceived by users through agents. At the just-concluded 2026 WAIC, Zhi Xiang Future released vivago R1, calling it an "infinite-duration content creation agent".

Its best proof came from a tricky test question: a product evaluator asked vivago R1 to make a "Turkish bootleg Star Wars" — an ultra-low budget cult film with cardboard monsters, colanders as helmets, and foam plastic stones as props.

The difficulty is subtle: AI's instinct is to render the picture to be as exquisite as possible, but this task requires it to maintain the cheap, "funnily fake" aesthetic. In the resulting one-minute-plus finished film, Captain Ali with a moustache and red clothes, and the cardboard monster with a colander as helmet, keep the exact same image from start to finish, and are not secretly rendered into exquisite CGI — the upper-layer Agent arrangement has suppressed the aesthetic inertia of the underlying model.

This consistency is not by luck. R1 runs a structured pipeline: write story → revise script → split scenes → define art style → lock reference images of characters and scenes → generate shot by shot → stitch into the finished film.

The so-called "infinite duration" also needs to be demystified: it does not spit out a two-hour movie at one time, but solidifies characters, props and scenes into reusable assets in the long-term production process, so that the story can continue to extend. This exactly hits the point that industrial capital attaches great importance to — enable IP to evolve from a static copyright inventory to digital assets that can be called repeatedly.

vivago adopts a hybrid path of "self-developed base model + aggregated scheduling + upper-layer arrangement". Its moat does not lie in "all the pictures are self-developed", but in arrangement capability and long-term consistency. Aggregation is a means, arrangement is the product, and the final user experience is the core battlefield that all players compete for.

Part04

Integrate into industrial processes

With this investment from industrial capital, what tests Zhi Xiang Future is no longer just its ranking on the model list, but whether it can truly be integrated into industrial processes.

For Gu Yuhao, General Manager of Shanghai Film New Vision Fund, investing in Zhi Xiang Future is not only about focusing on its product capabilities, but also about "building the technical base for the future film and television industrial system in advance". He wants to see whether this company can connect the underlying model, middle platform and vertical agents, integrate scripts, characters, scenes, shots and sounds into the same digital system, and turn IP from static copyright into reusable digital assets.

In Gu Yuhao's words, "Film and television customers will not pay according to the rankings on the list, they will only pay for the final delivery results." This is both a recognition and a more difficult exam: there is a long "industrialization gap" between an amazing demo and stable, large-scale delivery.

If Gu Yuhao gives a strategic qualitative explanation of "why invest", then Zhang Situo, Vice President of Huace Film and Television, gives a judgment that has already been verified on the production front line.

In Zhang Situo's view, the first value generated by AI is not in the creative end, but in those "time-consuming, labor-intensive but low-technical-content" links — initial script screening, material sorting, multi-version editing, subtitle translation and localization. In the hit TV drama *Tai Ping Nian* produced by Huace, AI is used to generate background materials for some scenes, which greatly shortens the production cycle and strongly supports Huace's global distribution for its 68 million overseas subscribers.

More importantly, his criteria for judging whether a model is worth cooperating with, from the perspective of film and television companies, is not "whether it can generate beautiful pictures", but "whether it has director thinking — understands shot language and narrative rhythm".

This view exactly echoes the structured planning of vivago R1, which is "write the script first, then lock the reference images, and generate shot by shot". The "director thinking" at the model level is exactly the reason why industrial players are willing to place their bets.

Following this standard, Zhang Situo breaks down the competitiveness upgrading of film and television companies into three steps: Efficiency Competition → Asset Competition → Ecosystem Competition. He values the second step the most — "In the past, the most valuable asset of film and television companies was the copyright library, but now static copyright needs to be upgraded to 'computable content assets': scripts can be retrieved by AI, finished films can be disassembled and recombined, and production experience can be precipitated and reused."

This is almost another expression of Gu Yuhao's judgment on "digital assets".

The two capitals also have surprisingly consistent views on the implementation pace. Huace's cooperation path starts from the direction of "quick to verify" such as AI high-quality short dramas, and moves towards secondary development of IP digital materials and full-process AI collaboration for long dramas in the medium and long term; Gu Yuhao believes that at the current stage, AI is more suitable for promotion materials, pre-development, animation, manhua-adapted dramas and stylized short films, while live-action long films are still much more difficult. The two film and television capitals have given the same timetable: focus on auxiliary and lightweight content in the short term, and talk about reconstruction in the long term.

In the industrial field, the preferences of state-owned capital and industrial capital have always been clear: investing in a company is expected to drive several industrial chains. Vision is exactly a layer of large model that is extremely close to the industry, with ready-made scenarios such as film and television, cultural tourism, marketing and e-commerce, and the results delivered by Zhi Xiang Future can withstand item-by-item verification. Capital flows through the company and falls into the industry, which can drive a large scale of development with one fulcrum.

Part05

The longer-term goal is the world model

Zhi Xiang Future sets its longer-term goal on the world model, but Mei Tao has always been restrained about this concept: "We do not think the industry has fully realized the world model today." In his view, an image is the spatial state of a certain moment, while a video adds the time dimension; following this path, the model also needs to understand spatial structure, motion consequences, causal relationships and physical laws.

From this perspective, what Zhi Xiang Future really wants to promote is not a single product, but a more underlying native full-modal path, which enables different modalities including text, image, video, audio and motion to be aligned at the underlying layer, and finally realize arbitrary-modal input and arbitrary-modal output, so that the world model can truly understand, predict and reconstruct the real world.

From visual generation to the native full-modal world model, the long-term proposition of this path has always been only one: to enable AI to not only generate content, but gradually learn to understand how the world operates, and have the ability to deduce and reconstruct the world.