What does MiniMax rely on to achieve AGI?
The enthusiasm in the field of large AI models has been continuously pushed to new highs.
Recently, overseas players have been updating their large AI models at an astonishing speed. On September 1, Anthropic released Claude Fable 5.1, which is claimed to be "the world's most advanced model for programming and knowledge work"; on September 2, Meta launched Muse Spark 1.3, calling it "one of the biggest performance leaps in the history of models", while Google released Gemini 3.8 Flash; on September 3, OpenAI rolled out GPT-6 Astra. At the launch event, Greg Brockman, co-founder and President of OpenAI, said, "Welcome to the AGI era".
The so-called AGI, namely Artificial General Intelligence, refers to the advanced form of AI that enables machines to understand the world, reason autonomously and flexibly handle complex cross-domain tasks just like humans. Compared with previous Agents that rely on API interfaces from software vendors, GPT-6 Astra can simulate humans to directly read screen pixels, operate any software with a graphical user interface via mouse and keyboard, and connect reasoning, code execution, visual recognition, computer operation and tool calling into a single execution chain, enabling autonomous planning and completion of end-to-end complete workflows. This has begun to fit the definition of highly autonomous AGI.
Domestic and overseas AI players, in a broad sense, are all riding the wave of pursuing AGI, especially in the large model sector. Although different parties have varying definitions of AGI and their progress differs, most of them have now set the realization of AGI as one of their ultimate goals.
In contrast, domestic players have made relatively weak technical voices in early September, and the rhythm of model updates is not intensive. Alibaba updated Qwen3.8-Max on September 2, while iFlytek officially released Spark X2.5 on September 7. Another news related to domestic models is that HUMAIN, an AI company under Saudi Arabia's Public Investment Fund, released the world's first Arabic large model on September 3, which is post-trained based on MiniMax's open-source model M3.
It is worth noting that MiniMax is one of the pioneers in China's large model sector, founded at the end of 2021. It released its first text model abab1 in April 2022, earlier than OpenAI's launch of ChatGPT in November 2022. Moreover, MiniMax does not only focus on a single modality, but independently develops text, speech, video and music models at the same time, hoping to get closer to AGI faster in the form of full-modality integration. After being listed on the Hong Kong Stock Exchange on January 9, 2026, MiniMax's share price once soared, with a market cap peaking at over HK$410 billion (about RMB 350.9 billion).
However, after the peak, MiniMax's share price retreated sharply, and doubts from the outside world kept emerging. Its first-mover advantage failed to translate into absolute technological leadership. When the distance to AGI has become a common question across the whole industry, MiniMax has not yet shown a more leading position than other players.
01
More prominent cost-effectiveness
At present, the core highlight of MiniMax lies in its price rather than model quality.
In terms of technical routes, MiniMax has chosen full-modality investment which is rare among startups, covering four fields: text, speech, video and music. MiniMax's large language model belongs to the M series, which has been iterated to M3 on June 1; its video model is Hailuo, referred to as the H series, which has been iterated to H3 on July 31; it also has a speech model updated to Speech 2.8 and a music model updated to Music 3.0.
In a single field, including the most critical and fundamental dimension of large language models, MiniMax has not formed an overwhelming advantage.
According to the latest Intelligence Index ranking from the model evaluation platform Artificial Analysis, the top six large models with the highest global intelligence scores are all overseas models, including Claude Fable 5.1, GPT-6 Astra, Muse Spark 1.3, etc.; the top-ranked domestic models in sequence are Zhipu GLM-5.3 at the 7th place, Kimi K3 at the 9th place, Alibaba Qwen 3.8 2.4T A95B at the 13th place, and DeepSeek V4 Pro 0813 at the 15th place, with MiniMax-M3 ranking 17th.
On the SuperCLUE general evaluation list for Chinese large models, in July 2026, the top four models in total score are Qwen3.8, Kimi K3, DeepSeek V4 Pro, and Doubao-Seed-2.1-pro, and MiniMax-M3 ranks 9th.
In comparison, the H-series video models perform better. On the current Artificial Analysis video model list, MiniMax H3 ranks first in image-to-video generation, second in video editing, and third in text-to-video generation; other top-ranked video models include Google's Gemini Omni Flash and Veo 3.1, ByteDance's Seedance, Alibaba's Wan 3.0 and HappyHorse-1.1, and Kuaishou's Kling 3.0, etc.
MiniMax H3 integrates all tasks including text-to-video, image-to-video, video editing, and motion transfer into the same pre-training framework, which can describe task relationships uniformly with natural language, unlike the previous two generations of models that set independent pipelines for text-to-image, image editing, motion, timbre and other modules. That is to say, users only need to input one sentence, for example, referring to the camera movement of a certain video, making the character in a certain picture sing and match specific audio, and the model can automatically complete understanding, arrangement and generation. It supports arbitrary combination input of text, image, audio and video, and can mix up to 12 reference materials at a time.
This innovative architecture is called "Contextual Omni Representation", which enables MiniMax H3 to have the performance capability for complex text materials. It is worth mentioning that previous video generation models often produce garbled text, because these models make probabilistic pixel predictions for text, can mimic strokes, but cannot understand the text. However, MiniMax H3 allows text to truly participate in the visual composition, and can even fit the perspective relationship on the surface of objects, changing with camera movement and subject motion.
Although some of its technologies have outstanding performance, MiniMax H3 has not established absolute quality advantages in the video generation field, and is famous for its cost-effectiveness instead.
According to the official pricing of each model, calculated at 1080p resolution, the most expensive models in the world at present are Seedance 2.5 and Veo 3.1, with a price of about 3 yuan per second, while Wan 3.0, Kling 3.0 and Gemini Omni Flash are priced at about 1.1 yuan per second. In contrast, MiniMax H3 at a higher 2K resolution only costs 0.8 yuan per second. Of course, the actual usage logic of video generation is completely different from that of text generation. Creators often need to generate repeatedly and filter multiple times to pick out one usable material from several generated clips, so the actual cost is often several times the marked price.
In the eyes of many practitioners in the film and television industry, MiniMax H3 is a cost-effective choice. On the overseas community Reddit, most reviews of MiniMax H3 describe it as fast, cheap and the most reliable at its price range. Xiao Yan, a short drama practitioner, told Hike Finance that their company believes Seedance has the best overall effect, but its price is on the high side, so they only use it for the most important shots; MiniMax H3 is used more in actual production, which can achieve 80% to 90% of the effect of Seedance at a much lower price.
02
Shifting focus to B-end customers
When model capabilities are not sufficient to support technical premium, the pressure of commercialization is entirely shifted to scale and efficiency.
According to the financial report, MiniMax divides its business into two major segments by commercialization channel. The first segment is C-end AI-native products, including Talkie, Xingye, Conch AI, MiniMax Agent, etc., with revenue coming from subscription fees and in-app consumption; the second segment is B-end open platform and other enterprise services, which provide enterprise customers and developers with model API calls, token packages, MiniMax Code and other tools, essentially selling model capabilities based on consumption volume.
MiniMax's revenue used to be dominated by the C-end. Financial reports show that MiniMax's revenue in 2024 and 2025 was USD 30.523 million (about RMB 204 million) and USD 79.038 million (about RMB 530 million) respectively, of which AI native product revenue accounted for 71.4% and 67.2% respectively.
However, focusing heavily on the C-end means facing customer acquisition and operation challenges. Talkie/Xingye is essentially more like a gamified social product, corresponding to the operation logic of the internet industry. According to the observation of Hike Finance, this product allows users to build their own characters, adjust the appearance, voice, personality and background settings of the characters, and then save the characters in the form of AI character cards. Any user can have text or voice interaction with the character. AI is only the underlying technical support of the product, and the product revenue mainly comes from gacha, in-app purchases and subscriptions. Its commercial performance is highly dependent on operation promotion, content supply and community activity.
Conch AI and MiniMax Agent are more dependent on the original capabilities of the model. The explosion of such products often requires leaps and bounds of technological iteration to trigger natural traffic, and it is difficult to replicate the path of ChatGPT or DeepSeek. For ordinary users, there are many similar platforms and products, with almost zero conversion cost, so it is not easy for the company to increase the payment rate.
Now MiniMax's revenue focus has shifted to the B-end. According to the financial report, in the first half of 2026, the company's revenue was USD 116 million (about RMB 778 million), of which AI native product revenue was USD 42.6 million (about RMB 285 million), accounting for 36.6%, and the revenue from open platform and enterprise services increased by 700% year-on-year to USD 73.9 million (about RMB 495 million), accounting for 63.4%.
B-end API calls are closer to the essential logic of large model products: the model itself is the product, and revenue is directly linked to model capabilities and pricing, and MiniMax's competitiveness in this link still mainly comes from low prices.
According to the Cost per Intelligence Index Task calculated by Artificial Analysis, at present, overseas OpenAI and Anthropic, with their world-leading technical level, firmly occupy the high-end market, with Claude Fable 5 at USD 8.75 (about RMB 58.6) and GPT-6 Astra at USD 3.26 (about RMB 21.8); in the domestic market, Qwen3.8 2.4T A95B, GLM-5.3 and Kimi K3 are all around USD 2 (about RMB 13.4), DeepSeek V4 Pro is USD 0.67 (about RMB 4.4), and MiniMax M3 is as low as USD 0.51 (about RMB 3.4).
The low-price strategy has directly pushed down the unit token revenue, but the computing power cost of large model reasoning will not decrease synchronously with the pricing cut, and the reasoning overhead of full-modality models is higher than that of single text models. While the revenue side is continuously thinning by the price war, the cost side still has to bear high R&D investment and computing power expenditure, so the profit space is continuously compressed.
At the earnings call in August 2026, Yan Junjie, founder and CEO of MiniMax, said that MiniMax's ARR (Annual Recurring Revenue) in August exceeded USD 800 million (about RMB 5.36 billion); price reduction and gross margin improvement are not contradictory, as long as the unit token gross margin remains positive, the expansion of scale will bring higher absolute gross margin.
But in fact, MiniMax has always been in a loss state. According to the prospectus and financial reports, excluding the non-operating factors brought by listing, MiniMax's adjusted net losses in 2023, 2024 and 2025 were USD 89.074 million (about RMB 597 million), USD 244 million (about RMB 1.637 billion) and USD 251 million (about RMB 1.684 billion) respectively, and in the first half of 2026, the loss expanded 111.2% year-on-year to USD 293 million (about RMB 1.966 billion).
Brokerages including BOCOM International, Guotai Haitong and Orient Securities all predict that MiniMax's revenue will continue to grow, but it will remain in a net loss state from 2026 to 2028. A research report released by Orient Securities on August 22 mentioned that MiniMax's multi-modal layout has strategic scarcity, and the model has significant advantages in reasoning cost performance. Its revenue in 2026, 2027 and 2028 may grow to USD 407 million (about RMB 2.731 billion), USD 1.221 billion (about RMB 8.194 billion) and USD 3.053 billion (about RMB 20.488 billion) respectively, with adjusted net losses of USD 601 million (about RMB 4.033 billion), USD 688 million (about RMB 4.617 billion) and USD 393 million (about RMB 2.637 billion) respectively.
03
Volatility has its root causes
The multi-modal technical route itself is not problematic, but full-modality independent R&D means extremely heavy investment.
Financial reports show that MiniMax's R&D costs in 2024 and 2025 were USD 189 million (about RMB 1.268 billion) and USD 253 million (about RMB 1.697 billion) respectively; in the first half of 2026, R&D cost increased by 138.8% year-on-year to USD 297 million (about RMB 1.993 billion).
Only players with strong accumulation such as