Zhipu (Meeting Minutes): Revenue Structure Reversed, ARR Reached 1.6 Billion USD
The following is the FY26 interim earnings call minutes for Zhipu (02513.HK) compiled by Dolphin Research:
I. Core Earnings Highlights Review
1. ARR is disclosed for the first time, with two sets of calibers proactively provided
a. As of the end of August 2026, ARR reached USD 1.6 billion, calculated on a monthly annualized basis, which refers to the revenue of August multiplied by 12.
b. The management also pointed out another more aggressive calculation method in the industry — data of the latest week multiplied by 52; under this caliber, the latest ARR after the volume release of GLM-5.3 has exceeded USD 2 billion.
2. Revenue structure completes transition, with both volume and price rising
a. Total volume: Total revenue in the first half of the year was RMB 957 million, equivalent to USD 142 million, a year-on-year increase of nearly 400%.
b. Segment: Revenue from the open platform and API business was RMB 825 million, equivalent to USD 123 million, a year-on-year increase of more than 27 times; its proportion in total revenue rose from 15.2% in the same period last year and 26.3% at the end of last year to 86.5%.
c. Volume and price split: The average API pricing increased by about 101%, while the calling volume of coding plan increased by more than 23 times compared with the beginning of the year. The management emphasized accordingly that the volume growth is driven by model capabilities rather than price elasticity.
3. Gross margin and computing power efficiency improve simultaneously
a. The gross margin of the open platform and API business rose from -0.4% in the same period last year to 24.6%, up 25 percentage points year-on-year; it also increased by about 6 percentage points compared with 18.9% for the full year of last year.
b. The company first proposed the "computing power multiplier", which refers to the revenue of the open platform and API corresponding to every RMB 1 invested in computing power (including the training side and the inference side). This indicator has increased by 14 times compared with the first half of last year.
4. Loss narrows, the management states that gross profit has begun to feed back R&D
a. R&D expenditure in the first half of the year was RMB 2.13 billion, equivalent to about USD 317 million.
b. The loss during the period was RMB 2.072 billion, down 12.5% year-on-year, and the loss ratio (loss during the period divided by revenue) narrowed by 4.7 times year-on-year; the adjusted net loss was RMB 1.964 billion, and the adjusted loss ratio narrowed by 3.5 times year-on-year.
c. The management's judgment is based on the fact that the adjusted net loss of RMB 1.964 billion is already less than the R&D investment of RMB 2.13 billion, indicating that after covering administrative and sales expenses, gross profit has begun to feed back R&D.
II. Detailed Content of the Earnings Call
2.1 Core Information from Executive Statements
1. Capability Ladder and Commercial Formula
a. The management defines the evolution of large models as a five-level ladder that can only be completed in sequence: Chat, Coding, Agent, Co-worker, Autonomous AI. There are technical thresholds between each level, and the business model of the upper level will not be mature if the threshold cannot be crossed.
b. During the reporting period, the company's focus was between the second level (Coding) and the fourth level (Co-worker). Attempts have been made in fields such as cybersecurity, law, finance, and education, among which cybersecurity has the fastest progress, while the rest are still in the early stage and have not formed large-scale revenue.
c. The formula proposed by the management is: AGI commercial value = upper limit of intelligence × total token consumption. Each time the task boundary rises by one level, the market space expands by an order of magnitude.
2. GLM-5.3 and GLM-5.3 Flash
a. GLM-5.3 and 5.2 adopt the same architecture, same total parameters and activation parameters. The only variable is the post-training scale, and the end-to-end task completion rate has increased by more than 50%; training tasks have shifted from exercise-style to complete professional tasks that are close to the real work of experts.
b. GLM-5.3 Flash is oriented to high-frequency, large-scale, cost-sensitive scenarios, with a total parameter of 320B and an activation parameter of 18B. It adopts linear and mixed attention, the price is one-tenth of 5.2, and it outperforms 5.2 in both benchmark tests and practical applications.
c. On the first day of its launch in the form of an anonymous model, Flash topped the OpenRouter list and refreshed the token usage record at that time, driving the overall calling volume of the platform to increase by more than 20%.
3. Three-stage Evolution of Business Form
a. Before 2025 was the stage of localized deployment. What customers purchased were tools that needed to be integrated, customized, and delivered to their own environments, superimposed with requirements for data security compliance and independent controllability.
b. In early 2025, it entered the Coding stage, where the model shifted from knowledge-driven to task-driven; the company strategically shrank the software licensing business for localized deployment, and turned to callable, scalable, and metered intelligent services.
c. This year, it entered the parallel stage of Agent and Co-worker, and the business shifted from selling calls to selling subscriptions and end-to-end task results. The management's summary is: every time the model completes a capability leap, what customers purchase gets one step closer to the final economic value, and the revenue structure is rewritten accordingly.
4. Platform Users and Customer Quality
a. As of August 2026, the number of registered users on the MaaS platform has exceeded 7.4 million, up 144% from the beginning of the year; the number of paying daily active users has increased by 603% from the beginning of the year; the average daily calling volume of the top ten users by revenue this month has increased by 98 times from the beginning of the year.
b. In the last two months alone, driven by GLM-5.2, 5.3 and 5.3 Flash, the number of users increased by 1.6 million from 5.8 million at the end of June to 7.4 million.
c. According to the ARR caliber, the number of user groups with annualized contributions exceeding 100,000, 500,000, 1,000,000, 10,000,000, 25,000,000, and 250,000,000 US dollars are 115, 25, 37, 8, 2, and 2 respectively.
5. Security and Governance
a. The approach is to achieve synchronous growth of open source, openness and security capabilities: let the models enter as many real environments as possible through open source weights, MaaS platforms, coding plans and global developer networks, and establish reliability through permission control, process supervision, risk assessment, vulnerability disclosure and external verification.
b. The most sensitive capabilities are under controlled management. Before open sourcing, vulnerabilities are independently assessed by a third-party team, and responsibility disclosure is completed through the national vulnerability database; the principle of division of labor is that the benchmark score is self-reported by the company, and risk judgment is handed over to external teams, which will not change with the improvement of capabilities.
2.2 Q&A Session
Q: How long can the company's leading position in coding be maintained in the future, and what key variables will it face?
A: Leading position on a single benchmark will eventually be caught up. What we need to build is the capability of continuous leadership. The management stated that coding benchmarks are updated on a weekly or monthly basis and will soon be caught up, and there is noise in the rankings themselves. Therefore, the description of SOTA in the half-year report is that it is more meaningful to ask "how strong the capability of continuous leadership is" than "how long the lead lasts".
The supporting evidence is that in the past 11 months, the model family has completed continuous iterations of multiple generations from 4.6, 4.7 to 5.1, 5.2, 5.3, with the capability curve continuing to rise, while the cost per unit task remains at a relatively healthy level.
The management also explained that the company does not regard coding as a product line with rapid revenue growth, but as a necessary breakthrough point in technical logic — only by moving forward along the coding track can the model evolve to Agent, long-horizon tasks and Co-worker.
Q: How does the company plan to migrate the accumulated model capabilities to scenarios beyond coding, and what is the implementation path?
A: The migration has been verified to be valid in the cybersecurity field, with the CyberGym score rising from 77.2 to 84.5. The management explained that choosing coding as the starting point is because it provides a scalable, automated, low-cost verifiable environment — whether the code can run, whether the test passes, and whether the bug is fixed can all be judged objectively. This is crucial for reinforcement learning, and the difficulty of most knowledge work lies precisely in the inability to judge whether the task has been completed correctly.
Another benchmark score increased by 30 percentage points from 24.4 for 5.2 to 54.4, and the more the task requires complete long-horizon planning, the more significant the improvement of 5.3 over 5.2.
In the real environment, the company cooperates with multiple domestic security teams, and a total of 2436 vulnerabilities have been found in the real code base after expert screening and deduplication, of which more than 1000 are high-risk, covering 269 projects.
There are three criteria for scenario selection: whether the intellectual threshold is high enough, whether it can be effectively verified, and whether there is sufficient economic value after completion. Attempts are also being made in law, finance, and data analysis, but the reliability threshold and verification mechanism are different, and the commercialization structure will vary accordingly.
Q: What is the current progress of computing power expansion?
A: The company does not disclose the caliber of GPU card quantity, and advocates using "effective computing power" and "computing power multiplier" for measurement instead. The management stated that computing power supply is a common problem in the industry. The company has continued to expand at a healthy pace this year, with increasing investment in both the training side and the inference side.
The reason for not using the number of cards for measurement is that the sources and uses have been diversified: the sources include self-owned clusters, leasing and service procurement, and the uses cover pre-training, mid-training, post-training and inference. The efficiency of chips of different generations and architectures varies greatly, and the total number cannot reflect the real capability.
What the company pays attention to internally is how much computing power has been installed, can be stably scheduled, run for a long time, and finally converted into training progress and billable tokens. The management believes that the question to answer is not how many cards there are, but how fast the model iteration and how large the scale of commercial token supply these resources can support.
Q: What scale of model training and inference demand can the existing computing power resources support?
A: The inference side has reached the level of 100,000-card domestic chips, while the training side still faces structural mismatch of heterogeneous computing power. The 100,000-card domestic chip cluster has realized large-scale low-cost inference, and the corresponding unit inference cost has dropped by 80% compared with the beginning of the year.
When GLM-5.3 Flash was launched anonymously, it fully used the domestic cluster to provide services. The calling volume of about 60 trillion tokens in six days is a historical record for both OpenRouter and OpenCode. The management believes that domestic chips have proved that they "can run", and the deep water zone is to verify their economy.
To this end, the inference engine and service stack have been redeveloped since half a year ago, including quantization, PD separation, layer split, caching and communication optimization. The end-to-end performance under the same domestic hardware has been increased by 3 times compared with the baseline.
The problem on the training side is a structural problem rather than a quantity problem. It has very high requirements for the stability, network and software stack of the isomorphic cluster, while China's computing power is mainly heterogeneous.
The management expects the situation to improve after the volume release of advanced domestic chips in three to six months; the training of the new generation of base model is already in progress, and the computing power configuration will prioritize ensuring the key training window, and will not move all resources away due to the expansion of inference volume.
Q: The market believes that the model layer will eventually be commoditized, and value will shift to the harness or the orchestration layer of token distribution. Do you agree with this view?
A: We agree that a single-point leading advantage cannot constitute pricing power, which is maintained by four factors. Two price curves will emerge in the future. The four factors are iteration speed, real task completion rate, unit intelligence cost, and customer workflow depth.
In terms of iteration speed, 6 consecutive generations were launched in 11 months, the intelligence index rose from 32 to 60, the cost of a single flagship task remained at the order of USD 0.2, and that of Flash was about USD 0.045. Single-point scores can be replicated, but this capability curve is more difficult to replicate.
In terms of task completion rate, the convergence of public configurations will make the single-question capability close, but the longer the task, the greater the difference — continuous engineering tasks of several hours require the model to read the code base by itself, adjust tools, handle failures and re-plan the delivery.
In terms of value density, the value of 1 million tokens used for casual chat and for fixing production vulnerabilities is completely different for customers.
In terms of workflow depth, customers accumulate prompts, permission systems, evaluation standards and engineering integration around the model, leading to higher switching costs.
Verification data: While the average API selling price increased by about 101%, the calling volume still increased significantly, and the average daily calling volume of the top ten users increased by 98 times compared with the beginning of the year.
It is therefore judged that the price of intelligence at the same level will continue to decline, and models that can open up new task boundaries will still have premium and even room for price increase.
Q: What progress has the company made in overseas competition and overseas business expansion?
A: The overseas expansion form has shifted to open platforms and APIs, and CSP cooperation progress is expected in 1-2 months. The management stated that the company has had overseas business of significant scale since 2025, which was dominated by sovereign large model projects in the form of localized deployment in the early stage, and the previously disclosed Malaysia project belongs to this stage.
As the main business entity shifts to open platforms and APIs, GLM-5.3 Flash is an attempt under the new form. The management admits that Chinese models as a whole still have a certain gap with overseas leading models in cutting-edge capabilities.
However, if we judge the optimal solution from the two dimensions of intelligence and cost, GLM-5.3 as the representative of high-end intelligence and 5.3 Flash as the representative of inclusive cost-effectiveness have occupied two ecological niches, which is illustrated in the double Pareto optimal diagram in the half-year report.
The company will actively promote cooperation with overseas CSP platforms, and consider methods such as local hosting of open source models and revenue sharing in the future.
Q: What is the most core improvement direction of the next generation model, and what changes will take place in parameter scale, architecture and training methods?
A: Continue to scale up the base model, realize deep inference with small activation, strengthen post-training, and the end point is full self-training. The technical route has converged: the company previously invested in multi-modal, video and image generation research, but found that it does not help much to improve the upper limit of intelligence, so it focuses on text.
The base model will continue to scale, expand to store and better represent knowledge, and raise the ceiling of intelligence. However, the activation parameters cannot be simply expanded — excessive activation will slow down inference and increase costs, making it difficult to achieve inclusiveness. Therefore, the direction is to use small activation with new methods to achieve deeper inference.
Post