HomeArticle

Citi AI Tracker: Models are getting increasingly powerful, but chips and power supply can hardly keep up.

36氪的朋友们2026-07-27 15:23
Citi's analysis points out that the return on investment (ROI) in the AI industry is accelerating its shift to the infrastructure layer, and the future competitive moat will transform from computing power acquisition to efficiency and proprietary data.

Large AI models are becoming increasingly intelligent, but the physical world that supports them is being pushed to its limits.

In its latest report released on July 24, Citi wrote that Moonshot AI's Kimi K3 ranked third globally with a score of 57, only three points behind the top closed-source model Claude Fable 5. Meanwhile, the pricing of China's cutting-edge models represented by Kimi K3 jumped 45% in a week, the rental price of Blackwell GPUs has risen 27% since the beginning of the year, and some laboratories even spent 1 billion US dollars to purchase generator sets directly.

Citi's analysis points out that the return on investment (ROI) in the AI industry is accelerating its shift to the infrastructure layer; and the competitive moat for the next stage of model competition will completely shift from simply "obtaining computing power" to "efficient output and proprietary data".

01 Only 3 Points Gap Remains

Citi noted in its research report that Kimi K3 is currently the largest publicly released open-source model. A few weeks ago, the highest score for open-source models was only 51 points (Z.ai GLM-5.2), and Kimi K3 jumped directly to 57 points, second only to Claude Fable 5 (60 points) and OpenAI GPT-5.6 Sol (59 points).

The entire industry is accelerating. The median intelligence score of the top 20 model providers has risen from 34 points six weeks ago to 43 points. In the open-source camp, DeepSeek V4 Pro (44 points) has a per-million-token pricing, and Xiaomi's model (sub-score) is as low as 0.03 — approaching the cutting-edge intelligence level, but with a price two orders of magnitude lower.

The closed-source camp has not slowed down either. Google just released Gemini 3.6 Flash (July 21), and revealed at the same time that Gemini 3.5 Pro is still in testing, and the pre-training of Gemini 4 has already started — three generations of models are advancing in parallel. However, Citi believes that the release rhythm of "Flash first, Pro later" sends a signal: the advancement of cutting-edge models is becoming more difficult — which is consistent with the delays other companies have encountered in infrastructure construction, and also reflects the industry's situation of wanting to shorten the delivery cycle but failing to achieve it.

The larger and more powerful the model is, the resources required to run it are also expanding sharply.

02 The Bottleneck Has Shifted

Trillion-parameter models are rewriting the meaning of "computing power".

Citi points out that when these super-large models are actually running, more and more time is spent moving weights and KV-cache data across HBM memory and GPU interconnection networks, rather than on the matrix operations themselves. The bottleneck has shifted from "how fast it can compute" (FLOPs) to "whether it can move the data" — memory bandwidth, GPU interconnection, and power supply.

The demand for GPUs remains strong, and the rental price of the Blackwell architecture has risen 27% since the beginning of the year. But simply stacking GPUs is no longer enough.

Some laboratories have begun to directly enter the upstream power generation sector: SpaceX spent 1 billion US dollars to purchase 1GW of mobile turbine units (July 15), and Georgia Power signed a service agreement in the same week (July 22). AI laboratories buying power generation equipment — something almost unthinkable a year ago — is now happening.

Model pricing directly reflects how tight the production capacity is. The mixed pricing of China's cutting-edge models has jumped to 0.87 per million tokens, with a 45% increase both week-on-week and month-on-month, marking the first major fluctuation in two months. The average price of global cutting-edge models rose 6.8% week-on-week and 11.3% month-on-month. The United States and Europe are relatively stable (-0.5% week-on-week), but the month-on-month increase has also reached 4.1%.

Citi judges that as incremental infrastructure goes online one after another, the pricing pressure will eventually ease. At that time, valuable data and task-specific performance will form a more lasting moat than the acquisition of computing power.

03 Agents Escape the Sandbox

Models are getting stronger. But the agents running on the models are also becoming more dangerous.

The Citi report uses a thought-provoking statement: the realistic risk of autonomous AI agents escaping the sandbox has escalated from "interrupting lunch" in April to "penetrating Hugging Face infrastructure" in July.

The stronger and more popular open-source models are, the wider the spread of security risks will be. The debate on AI regulation is still ongoing — both NVIDIA (July 24) and US Treasury Secretary Bessent (July 22) have made statements recently — but it is difficult to reach a conclusion in the short term. Even if the regulatory framework has not yet been implemented, enterprises maintaining their own model operation architectures are already facing higher compliance thresholds.

What is more tricky is that the providers training these models themselves still cannot fully inspect why the models act in such a way. The problem of interpretability has not been solved so far.

But security concerns have not slowed down the commercialization process of agents. Data from METR (July 21) shows that the economic gap between AI agents and humans is narrowing. When agents become more autonomous and closer to economic feasibility, token consumption will continue to accelerate — which in turn pushes up infrastructure demand.

The stronger the capability, the tighter the constraints. The tighter the constraints, the greater the demand for infrastructure. This cycle shows no sign of slowing down.

This article is from the WeChat public account "Hard AI", the author is a researcher focusing on technology production and research, and is republished by 36Kr with authorization.