Zhipu AI Claims "Niu Lai": After Remaining Anonymous for Six Days, Tens of Trillions of Tokens Have Been Consumed by Global Developers.
After nearly a week of speculation, the mysterious model Ox Alpha finally has its confirmed owner.
On August 26, Bloomberg reported that Zhipu AI confirmed that Ox Alpha, which was launched anonymously recently, is its new-generation model of the GLM series. "Ox" corresponds to the Chinese word for "niu", which coincided with the popularity of the movie *Niu Lai*, so the Chinese developer community calls Ox Alpha the "Niu Lai large model".
Zhipu AI also stated that it will release the model weights on the evening of August 26.
However, Zhipu AI has not yet announced the official model number, parameter scale of Ox Alpha and its specific relationship with GLM-5.3.
From its anonymous launch on August 20 to being claimed by Zhipu AI on August 26, Ox Alpha has been used by global developers to consume tens of trillions of Tokens within six days.
01
Free and Anonymous, an "Ox" Took the Top Spot
On August 20, Ox Alpha appeared on the model aggregation platform OpenRouter as a "stealth model", namely an anonymous model.
At that time, OpenRouter only stated that it was an inference model developed and operated by an anonymous third party, mainly oriented to programming, long-cycle Agent tasks and production environments.
According to public specifications, Ox Alpha has a 1.0486 million Token context window, with a maximum single output length of 131,000 Tokens, supporting text, image and video input, as well as tool invocation and structured output. This means it can read large code repositories, project documents and long Agent operation records at one time.
It was free of charge that really detonated the usage volume.
On the day of the model's launch, the open-source AI programming tool OpenCode announced that Ox Alpha would be open for free for one week, providing a "nearly unlimited" call quota, and said the model provider had prepared a service capacity of 100 trillion Tokens per day.
100 trillion Tokens is equivalent to about 11.6 billion Tokens per second on average. This figure made developers flood in to test while asking: Which company on earth has such a huge inference resource?
However, 100 trillion Tokens is the theoretical service capacity, not the actual call volume, nor can it be directly equated with the deployed physical computing power. The calculation costs of input and output Tokens are different, and the code Agent will repeatedly read the same code base and historical context. A large amount of cache reuse can make the Token scale recorded by the platform far higher than the actual new calculation volume.
As of August 25, Ox Alpha processed about 23.2 trillion Tokens in the recent 7-day statistical window of OpenRouter, ranking first; in the same period, DeepSeek V4 Flash 0731 processed about 11.6 trillion Tokens. Less than a week after Ox Alpha was launched, its Token processing volume has reached twice that of the latter.
This set of figures shows that the model has obtained large-scale attempts, but the statistics of OpenRouter are the Token usage, not the number of users, requests, task success rate or revenue; free models are also naturally easier to get traffic.
02
Doubts About Computing Power
Before Zhipu AI claimed the model, computing power was once the main reason to oppose the view that "the model comes from Zhipu AI".
The community assumed at that time that Ox Alpha might be a flagship model with trillion-level parameters. To make such a model run for free for a week and provide a daily service capacity of 100 trillion Tokens requires huge chip clusters and power investment. Therefore, some people speculated that the model might come from Google or Microsoft, and others put forward the combination of "Zhipu AI's model hosted by a large manufacturer".
In July this year, Bloomberg reported that Zhipu AI had built a 1GW-level domestic chip AI data center and put part of it into operation. Data centers usually measure their scale by power capacity, but 1GW does not equal actual computing power. The report also did not disclose the number of chips currently in use, IT load and utilization rate. In the same period, Zhipu AI completed the acquisition of Zhongke Jiahe, a heterogeneous computing power software company, to supplement its capabilities in compilers, Runtime and inference engines.
If this information is true, it is not completely unexplainable that Ox Alpha processed 23.2 trillion Tokens in less than a week. However, 1GW is the designed power supply capacity of the data center, not the computing performance, which also includes the power consumption of refrigeration, power supply and distribution, and network facilities. Previous reports only said that the data center was "partially in operation", and did not disclose the actual number of enabled chips and their utilization rates.
Bloomberg confirmed the ownership of the model this time, but did not explain where Ox Alpha is deployed, what chips it uses, and whether there is third-party computing power support. Therefore, the 1GW data center can explain why Zhipu AI has the foundation to carry out large-scale stress tests, but still cannot verify that the peak capacity of 100 trillion Tokens per day has been actually deployed.
03
Developers Chased and Guessed Its Identity for Six Days
In addition to being free and having a large quota, the attention of Ox Alpha also comes from its performance in programming and Agent tasks.
A large number of calls occur in programming frameworks such as Claude Code, Hermes Agent, DeepSeek Harness and OpenCode. After the trial, Patrick Collison, CEO of Stripe, commented on X: "Very impressive."
A widely spread community small-sample test showed that Ox Alpha achieved 80% results in 10 DeepSWE software engineering tasks, higher than Claude Fable 5 and GPT-5.6 Sol in the same group of tests. However, this is only a small-sample test of 10 questions, not an official list in a unified environment. Later, some developers said that its performance in complex back-end and some visual tasks was average.
Other developers began to trace the source of the model.
In multiple sets of mixed texts of Chinese and English, codes and Emojis, the Token counts of Ox Alpha and GLM-5.3 show a highly consistent relationship, and the difference can basically be explained by the hidden system prompt words. In the controlled video test, its frame extraction method and the relationship between video duration and Token consumption are also similar to Zhipu AI's GLM vision model. The error codes and wording returned by the API abnormal parameters are also considered to have the characteristics of Zhipu AI's services.
The release time further reinforced this speculation. On August 14, Zhipu AI announced GLM-5.3 and said that it would complete the security assessment and make the weights public in about two weeks. Six days later, Ox Alpha appeared on OpenRouter anonymously; on August 26, Bloomberg revealed that Zhipu AI would release the Ox Alpha weights that night.
Zhipu AI has used a similar method before. In February this year, the anonymous model Pony Alpha was launched on OpenRouter, and was later revealed to be an early test version of GLM-5. From "Little Pony" to "Niu Lai", anonymous release has become a way for Zhipu AI to test new models.
This method can temporarily remove brand expectations and let developers evaluate the capabilities directly; free traffic can test the performance of the model, cache system, task scheduling and domestic heterogeneous clusters in the real Agent environment. The identity suspense itself also makes the community take the initiative to undertake the work of communication and testing.
However, what the final name of Ox Alpha is, what its relationship with GLM-5.3 is, what the official price is, and how many developers will continue to use it after the free period ends, will determine whether this anonymous release only brings short-term traffic, or becomes a new entry for Zhipu AI in the global developer market.
This article is from the WeChat official account "Tencent Tech", written by Xiao Jing, edited by Xu Qingyang, and published with authorization from 36Kr.