With a monthly revenue reaching 900 million yuan, Zhipu AI totally deserves to toast to itself.
Compared with revenue, what people care more about is whether SOTA can be sustained.
On the evening of August 31, Zhipu, which has been listed for more than seven months, released its performance report for the first half of this year: its revenue in H1 reached 954 million RMB, a year-on-year increase of 399.7%, and the scale has exceeded the full-year revenue of 2025.
During the earnings call, the company also disclosed a more impressive figure: as of the end of August, Zhipu's ARR (Annual Recurring Revenue) reached 1.6 billion US dollars based on the monthly annualization standard, and exceeded 2 billion US dollars based on the more aggressive weekly annualization standard in the industry, which is 2.5 times that of MiniMax in the same period.
A clear leader has emerged among the six rising stars of large models. Its revenue explosion is impressive, and it is tightly bound to the model on the same track.
With four model iterations in 8 months, Zhipu has secured its position in the first echelon of open-source models relying on models such as GLM-5.2 and GLM-5.3. It is also a representative of "model as product" — in the first half of the year, cloud model API revenue accounted for more than 85%, and localized deployment is no longer the core business.
Zhipu's performance in H1 is ideal. Calculated back from ARR, its revenue in August alone reached 133 million US dollars, which seems to set a sufficiently high revenue tone for the next six months.
But compared with revenue, what people care more about is whether SOTA can be sustained.
At the earnings call, the focus was on a series of issues related to subsequent model iterations: can the dividend of model price increases continue to be realized? And the most critical one — after GLM-5.3, can Zhipu still stay in the first echelon?
Revenue surged 400% year-on-year, gross margin dropped by 26.4%
Maintaining the first echelon of models, betting on Coding early, and the combination of choice and effort drove Zhipu's revenue to skyrocket.
According to the financial report, Zhipu's total revenue in H1 was 954 million RMB, a year-on-year increase of 399.7%, which was five times that of the same period last year, and the scale has exceeded the full-year revenue of 2025.
However, according to previous forecasts from investment institutions, this level of revenue growth is actually slightly below expectations. Taking Bloomberg's forecast of 1.35 billion RMB for the first half of the year as an example, Zhipu's actual revenue is about 30% lower than this figure.
Why does this gap arise?
There are three main reasons: First, as explained by Bank of America Securities, the breakout time of the new model is later than predicted. GLM-5.2 was not widely popularized until mid-June, so this year's revenue may be highly concentrated in the second half of the year. Second, the capital market underestimated the impact of the large model price war. Third, the market predicted that Zhipu's localized deployment (high-margin project-based) revenue would still maintain a large proportion, while API and cloud business would grow rapidly. But in fact, Zhipu's localized deployment revenue in H1 decreased year-on-year, while the proportion of API revenue increased significantly.
The biggest change for Zhipu in H1 this year is its revenue structure. Previously, Zhipu was labeled as "localized deployment", but now Zhipu has completely taken "models" as its core revenue source.
According to the financial report, Zhipu's open platform and API revenue in H1 was 825 million RMB, a year-on-year increase of 2735.7%, and the proportion rose all the way from 15.2% in the same period last year and 26.3% in the whole of last year to 86.5%. In contrast, the revenue from model localized deployment has shrunk significantly, reaching 129 million RMB, a year-on-year decrease of 20.5%, and the proportion has shrunk from 73.7% to 13.5%.
The change in business structure comes half from the transformation of industry scenarios, and half from Zhipu's intentional adjustment.
On the one hand, Zhipu mentioned during its listing period that it would later scale down localized deployment for To B and To G ends. Considering that the localized deployment business is more inclined to customized projects, with problems such as high labor cost and slow payment collection, Zhipu itself is more inclined to build its own MaaS platform, and make its revenue composition healthier through more standardized businesses.
In short, compared with the previous customized business of building houses one by one for each order, the current API is a highway charged by traffic, which is a more ideal choice.
In the currently retained localized deployment business, compared with general large models, enterprises are more inclined to access agent services that "can get things done". The financial report shows that Zhipu's enterprise-level general large model revenue dropped from 148 million RMB to 67.04 million RMB, a decline of 54.6%, while enterprise-level agent revenue increased from 13.74 million RMB to 55.56 million RMB, a growth of 304.4%.
On the other hand, as Xiao Lei, Secretary of the Board of Directors of Zhipu, mentioned: "This is a natural evolution brought about by the improvement of model capabilities itself." With the improvement of model capabilities, in the current era of "model as product", the scenario of subscription services is feasible. This is reflected in the shift of MiniMax from C-end to B-end dominance, and the revenue of Anthropic overtaking OpenAI.
In short, the change of Zhipu's revenue structure is actually an inevitable result after the growth of "model intelligence".
Apart from revenue, profit quality is also very important. Short-term losses for AI model companies are not a big problem, but if the improvement of costs and expenses is not obvious, it will affect the profit schedule in the long run.
Taking the adjusted net loss as the standard, Zhipu's loss in H1 was 1.964 billion RMB, a year-on-year expansion of 12.1%. The main reason for this part is still R&D. The financial report shows that Zhipu's R&D expenditure in H1 was 2.131 billion RMB, a year-on-year increase of 33.6%, which is 2.2 times the revenue of the same period, and exceeded the adjusted net loss for the first time.
"The gross profit contributed by the business in H1 this year has begun to gradually feed back our R&D investment after initially covering management expenses and sales expenses," Xiao Lei said at the earnings call.
But in the short term, high R&D investment is still necessary, and the industry has reached a very consistent consensus on this point: the model is still in the stage of pursuing the upper limit of intelligence, and high R&D is the admission ticket. Compared with MiniMax's R&D expenditure of about 297 million US dollars in H1, the two base model companies maintain almost the same level of investment.
Focusing on gross margin, this level determines whether the product is "valuable". Compared with the decline in gross margin caused by the change of revenue structure, the more critical point is that after frequent price increases this year, the gross margin growth of Zhipu's API business itself has not kept up with the price.
The change of revenue structure has led to the further decline of the overall gross margin after the shrinkage of the high-margin localized business, from 50.0% in the same period last year to 26.4%. If we only look at the gross margin of the cloud business itself, it has turned from -0.4% in the same period last year to 24.6%, indicating that the model profitability has improved to a certain extent.
Looking at the gross margin of the cloud API part, it has turned from -0.4% in the same period last year to 24.6%, indicating that the model profitability has improved to a certain extent. However, compared with the gross margin of 22.4% in 2025, since the beginning of this year, the average pricing of Zhipu's API has increased by about 101%, but after the substantial price increase, the change reflected in the gross margin is extremely limited. Coupled with the recent situation that the "Niulai" model is free during the test period and the pricing is extremely low, the subsequent improvement of gross margin may not be too high.
On the whole, Zhipu's financial report shows a contradictory growth result. With the company's revenue growing rapidly and the pricing power of advanced models established, model price increases have not significantly improved the gross margin.
In general, Zhipu has delivered a good report. To measure the revenue in the second half of the year, it depends on whether the model can continue to stay in the first echelon.
2 billion US dollars ARR, bet on the SOTA of the model
For the performance in the second half of the year, Zhipu has given a quite ideal expectation — ARR of 1.6 billion US dollars in August.
At the earnings call, Zhipu also implicitly commented on MiniMax's method of calculating annualized revenue by multiplying single-week revenue by 52, and said that calculated in the same way, Zhipu's ARR in August has exceeded 2 billion US dollars.
After calculation, Zhipu's revenue in August alone reached 133 million US dollars (equivalent to about 894 million RMB), and its monthly revenue is close to the company's total revenue in the first half of this year.
Among large model startups, Zhipu has secured the top position in terms of revenue alone: compared with the ARR of 300 million US dollars disclosed by Moonshot AI in June and the nearly 500 million US dollars ARR of DeepSeek in August as reported by the media, Zhipu has opened a certain gap.
The surge in revenue is related to the release of new models (GLM-5.2, GLM-5.3, etc.). The leading models drive user growth, and the pricing power brought by the first echelon allows it to sell at higher prices.
Zhipu shared user growth data at the earnings call: as of August, the number of registered users of Zhipu's MaaS platform exceeded 7.4 million, an increase of 144% over the beginning of the year; paying daily active users increased by 603% over the beginning of the year; the daily call volume of the top ten users by revenue increased by 98 times over the beginning of the year.
But success comes from models, and failure also comes from models. The market is more concerned about whether Zhipu can continue to maintain its existing model advantages.
The most common question in the doubts is whether Zhipu only does "post-training"? When Moonshot AI launched Kimi K3 with 2.8T parameters and MiniMax is planning the 3T parameter model M3 Pro, Zhipu has not launched a model of the same level. GLM-5.3 is also completed based on the GLM-5.2 base, with only post-training work done.
In response, Tang Jie responded directly at the earnings call: Zhipu is not doing no pre-training, but the return of post-training is higher at present.
"Simply increasing parameters is not the essential path of Scaling," Tang Jie said.
As for why not increase the parameters, Tang Jie said that currently the domestic training data is generally 30 to 50 trillion tokens. In the case of limited computing power, simply stacking parameters according to the Scaling Law at this level does not bring much benefit.
Judging from the perspective of Scaling, looking at the whole industry, the importance of strengthening post-training is becoming more and more non-negligible. MiniMax also made a similar judgment at the earnings call some time ago: while improving performance, post-training requires a large amount of inference computing power — synthetic data, reinforcement learning, and automatic evaluation are all essentially inferences. The cheaper the inference is, the more post-training trajectories can be run under the same budget, and the higher the upper limit of model intelligence will be.
However, the importance of post-training does not mean that the status of pre-training can be replaced. As the "foundation" of the entire large model capability, it is also the most intuitive embodiment of the Scaling Law. After all, the expansion of model parameter scale and training data volume will bring predictable performance improvement, while post-training is more about "polishing" on the basis of the base model capability.
Compared with post-training, the threshold of pre-training is getting higher and higher at present: it not only requires the accumulation of the number of chips, but also involves the construction of a whole set of computing infrastructure such as heterogeneous chip adaptation, network storage, and computer room construction. Zhipu continues to increase investment in this line: build self-owned data centers, cooperate with manufacturers such as Sugon to expand computing power supply, and build a large-scale inference capability of 100,000-level domestic chips under the trend of domestic substitution. The official said that the unit cost of token inference has decreased by 80% compared with the beginning of the year.
Although these investments are heavy, they have accumulated the necessary foundation for Zhipu to return to the base model track.
"After achieving the ultimate post-training, the next stage of growth must return to the base model," said Liu Debing, Chairman of Zhipu. He revealed that the training of the next-generation base model has been advanced, and the directions include larger effective scale, longer native context and native multi-modal unified modeling, but the specific parameters have not been disclosed.
Zhipu makes trade-offs at this stage, and it is correct to focus on post-training. However, in the case that the base models of GLM-5.1, 5.2 and 5.3 are all derived from GLM-5, when the base model scale can be increased is also equally critical for Zhipu.
However, under the fierce competition of domestic models, "winning" does not only mean running faster than competitors.
Current users care about both the price and capability of the model. This leads model vendors to pursue SOTA status on the one hand, and take into account cost performance on the other hand.
For example, GLM-5.3 Flash is an attempt