The tech community has long been plagued by the drawbacks of closed-source models, and MiniMax aims to become the Deepseek in the multimodal field.
On August 3, 2026, Hong Kong stock MINIMAX-W (00100.HK) closed at HK$249.4, with an increase of over 10%. During the morning trading session that day, MiniMax officially announced before the market opening the release of H3, its new-generation multimodal generative model. The market has cast its vote of confidence for this product launch with real money.
A
The industry has long been plagued by closed-source models. The video generation track was previously dominated by leading models such as Seedance, Keling, etc. Pricing, computing power, and generation time have always been topics of high concern across the industry.
H3 released by MiniMax is exactly a brand-new option in the video generation market.
According to the official blog, H3 is positioned as a general-purpose full-modal generative model, which supports unified understanding of multimodal context composed of text, images, videos and audio, can output audio and video with native dual-channel sound, and supports a maximum duration of 15 seconds at 2K resolution.
In terms of performance, MiniMax H3 is in the same tier as Seedance 2.5.
Judging from the demos released in the early invitation testing phase, the performance of H3 is very impressive.
To give a specific example: input a character reference image + a reference video with Hitchcock-style camera movement + a reference audio clip of singing to H3, it can understand the composite instruction of "let the character in the image move according to the camera movement of the reference video and sing with the voice in the reference audio", and generate a 2K finished video with native dual-channel sound in one go.
This cross-modal composite reference capability used to be the signature advantage of Seedance, and now MiniMax has also mastered this capability.
In the authoritative Artificial Analysis video model ranking, as of July 31, 2026, the release date of H3, MiniMax H3 ranks first in the world in video editing capability, second and third in text-to-video and image-to-video capability respectively, and its comprehensive model capability is among the first tier in the world.
A prominent highlight of H3 is its complex text retention capability.
For example, for scenarios such as Chinese and English titles in film and television opening sequences, brand logos on dynamic posters, and subtitles in UI motion effects, H3 can make text truly participate in the visual composition: the text fits the perspective of objects, conforms to the irradiation direction of scene light, and moves with the camera and the main subject.
When a brand logo is placed on a glass, the model needs to understand the refraction of glass; when it is placed on a wall, it needs to understand the curvature and texture of the wall; when the camera zooms in or out, it needs to keep the font shape undistorted. This has gone beyond the technical boundary of "generating fonts", and is the product of the model's overall understanding of the scene.
Complex text retention is only an explicit observation point of the underlying capability of "better understanding of vision and the world". According to MiniMax's technical blog, the technical highlight of H3 is Contextual Omni Representation (organizing all modal information in the context and their interrelationships), which takes the entire material as a unified context, describes the relationship between multiple elements such as video and audio for the model, and builds a dedicated model and full-modal understanding pipeline for this purpose.
B
In addition to performance, H3 also inherits the most representative feature of MiniMax — sufficient and cost-effective output.
In terms of pricing, MiniMax H3 uses 2K resolution by default, with a pricing of only 0.8 RMB per second, which has strong price competitiveness among products of the same resolution tier.
In the past six months, when the industry discussed "cinema-level" video generation, most mainstream commercial models used 1080p as the default resolution. H3 directly raises the default output resolution to 2K, taking a solid step in the direction of high-definition output.
What supports this price advantage is the technical innovation of MiniMax H3.
During the training process of video generation models, each frame of image needs to be cut into small patches and fed to the model. The more patches there are, the higher the cost of training and inference. H3's self-developed VAE (H3-VAE) achieves an industry-leading image compression rate. According to the official blog, the same video can be represented with far fewer patches, and the sequence length is directly reduced by 4 times. Both training and inference costs are lowered as a result.
In terms of high resolution, the idea of most models in the industry to output 2K/4K content is to first generate a low-resolution version, and then use a dedicated super-resolution module to upscale the image. This module essentially "guesses" the missing details, so the small text, texture, light and shadow after upscaling are often blurry or distorted.
H3 does not take the low-definition video as the only basis, but treats it as a reference, and inputs it into the model together with the original multimodal context. It can reuse the existing generation capabilities of H3, and restore the information that is no longer visible in the low-definition video but still exists in the context.
This is also the underlying reason why H3 rarely fails in complex text scenarios (opening sequences, logos, subtitles).
In other words, H3 follows the path of "improving performance through architecture efficiency" and reduces computing power consumption through technological innovation. The significance of this technical path for commercialization lies in: when the computing power cost enters the compression channel, the gross profit margin ceiling of video generation will be raised again.
C
H3 is covering the entire post-video production chain represented by software such as AE and Premiere.
In the past, AI video models only delivered a piece of "material", which was usually a frame of image without dubbing, subtitles or packaging, and users had to complete the remaining work step by step in AE, Premiere, or CapCut. What makes H3 different is that it integrates the entire "post-production" process into the model: understanding creative intentions, generating main frames, matching sound effects and music, rendering subtitles, and completing motion effect packaging, all delivered directly in one single generation.
For the hook shot of an e-commerce product advertisement, the traditional workflow goes like this: generate a dynamic frame of the product on the shelf via text-to-video → synthesize the "50% off for a limited time" voiceover via TTS → apply a subtitle template in CapCut → add product lighting highlights and price number pop-up motion effects in AE.
With H3, you only need to input a product image and an advertisement script, and the finished video output in one go will come with shelf camera movement, voiceover dubbing, subtitles, price pop-up effects, and light & shadow atmosphere — every visual element is aligned with the space, light and rhythm of the frame, instead of being simply overlaid.
According to feedback from the early invitation testing phase, MiniMax H3 has commercial-grade multi-scenario content generation capabilities, performs well in instruction following, text and brand information presentation, V2V Motion Transfer and other aspects, and can realize accurate and controllable multimodal content editing and generation, which is widely applicable to commercial scenarios such as advertising, branding, e-commerce, product design, UI/UX, and games.
What does this capability mean for the market size?
Referring to Adobe's 2025 fiscal year financial report, the global revenue of its creative cloud business is about 14.5 billion US dollars, of which video post-production tools (AE, Premiere, etc.) account for a considerable proportion; the domestic video post-production outsourcing market in China is estimated to be in the range of 300-500 billion RMB.
If H3's AI-native post-production capability continues to be implemented in the B-end market, it may replace 10-30% of the traditional post-production processes.
MiniMax's official blog states that "video models will be able to understand more complete creative intentions, handle more complex content requirements, and gradually move from 'generating a video' to truly participating in the whole process of content production."
In the decade after 2011, Netflix displaced Blockbuster, Uber displaced taxi companies, and Airbnb displaced chain hotels: the quantitative accumulation of technical capabilities will reach a certain point and suddenly bring qualitative changes. Software has eaten those non-software industries. Now, AI is going to eat those non-AI software.
D
Video generation ushers in its first flagship open-source moment. On the day of H3's release, MiniMax announced that it will open the model weights of H3 within a few days.
In the past, the default option in the video generation field was closed-source models, including Seedance, Keling, Sora, Veo, all mainstream models without exception. The reason is straightforward: video models have high training costs and large computing power consumption, and the closed-source plus API billing model is the most straightforward commercial model.
Open source means that the industry will no longer be dominated by closed-source models alone.
For a period of time in the past, the entire video generation industry was dominated by a small number of closed-source models, and the commercialization path basically revolved around "API billing + membership subscription".
In this process, some noteworthy phenomena have emerged in the industry: some leading products have adjusted their pricing systems multiple times in a short period of time, resulting in large fluctuations in user usage costs; there is a pricing inversion between C-end products and B-end APIs; the enterprise-level access threshold is relatively high, and small and medium-sized teams often cannot directly obtain official API permissions, which has spawned gray intermediate services such as group buying and reselling.
Behind these phenomena is the commercial logic under the closed-source model — when a small number of models occupy technical advantages, the supplier has strong pricing power. The emergence of open-source models may change this pattern, shifting industry competition from a pure technical competition to a more open ecological competition.
High-value B-end scenarios that cannot be supported by closed-source models are also opened up.
Film and television companies usually will not upload brand character images and unreleased episode materials to public cloud APIs, advertisers will not accept random drift of brand visuals in each generation, and game companies' core IPs cannot be uploaded to third-party models. This group of customers was previously blocked by the "data sovereignty" issue of closed-source models. They are the highest-paying, most sticky, and most underserved group in video generation commercialization.
H3's open-source model + local private deployment capability perfectly meets the needs of this group of customers.
The video generation track has been defined by arrogance in the past six months: price hikes, waiting queues, whitelists, bundled packages, scalpers, and account reselling. Users are complaining, the market is silent, small and medium-sized companies are locked out, and B-end customers are tangled in data and asset security issues. This is not a healthy market, but an unbalanced market. H3 breaks this deadlock. With its flagship performance, 2K resolution, open-source nature, and a unit price of 0.8 RMB per second, H3 reduces the commercial starting cost to one third of that of mainstream models. All these are delivered to the entire developer ecosystem and small and medium-sized companies. The open-source era of video generation officially started on July 31.
For Hong Kong-listed MiniMax, the significance of H3 is more than just the release of another model. Since its listing in January, the market's valuation anchor for this company has been changing, and market sentiment has shifted from "idolization" to "prudence". H3 is the company's first proactive move after a period of ups and downs. The game table will not be completely reshuffled by just one model, but it at least proves one thing: MiniMax still has the ability to overturn the entire game from scratch.
This article is from the WeChat official account "Alphabeta Ranking" (ID: wujicaijing), written by the Alphabeta Ranking team, and authorized for release by 36Kr.