The AI sector is most short of exit buyers.
Recently, ByteDance secured a USD 29.6 billion syndicated loan, Alibaba completed a HKD 80 billion share placement, and a few months prior, Google and Amazon raised over USD 800 billion in total financing, while NVIDIA announced a "grand plan" worth USD 500 billion. All major tech corporations around the world are scrambling to raise capital, and they are investing the funds in one single direction: AI infrastructure.
From a macro perspective, every underlying opportunity that reshapes the industrial landscape in modern industrial history is essentially a top-tier infrastructure upgrade.
The vigorous railway mania and the Gold Rush in the western United States in the 19th century, the real estate urbanization drive in China, and the current global AI infrastructure boom — although the process of massive capital injection is likely to be constantly marred by market disputes and accompanied by bubbles, the infrastructure finally precipitated will firmly underpin industrial development for decades to come.
However, compared with previous global infrastructure waves, this round of AI infrastructure has a wider coverage, an unprecedented scale of capital investment, far faster industrial iteration, and naturally more disputes.
In July, AI infrastructure experienced a sharp fluctuation, as the capital market began to question the gap between the huge investment in AI infrastructure and the returns it could generate. The extensive era of mindlessly stacking GPUs and frantically expanding clusters has not ended yet, but the market is starting to wonder: after trillions of dollars are invested in computing power, can we really make profits in the future?
This marks a signal of changes in both the capital side and the industrial side. The difference from last year is that this year, people can clearly feel that ordinary users have begun to use various desktop Agents on a regular basis. AI is no longer a simple chat tool, but has truly become an agent that can run automatically and perform tasks in cycles. This year is the real first year of Agent.
The widespread popularization of C-end agents has also made Token consumption continuously exceed expectations. In March this year, the average daily Token call volume in China exceeded 140 trillion, which was more than 1,000 times higher than that at the beginning of 2024.
However, there is a rather awkward problem at present. Everyone is shouting that there is a shortage of GPUs and competing for computing power, but in fact, a large number of GPUs are idling, lying unused and wasted every day. Carnegie Mellon University conducted an actual test, running 756 GPUs continuously for 31 days, and found that nearly 20% of the cluster's time was spent in "waiting".
For example, for OpenAI's inference business, more than half of the energy consumption is wasted on idling; the situation of domestic intelligent computing centers is even more extreme, with the average GPU utilization rate less than 30%. In the past few years, AI infrastructure grew wildly, and the expansion of scale could cover up the efficiency problem. But now that the time for returns is approaching, the competition lies in who can reduce the Token cost and increase the computing power utilization rate.
Efficiency has become the new key to success for AI Infra. Therefore, the urgent issues now are how to revitalize the idling computing power and how to reduce the inference cost. As the AI Infra market enters the deep water zone, industry noise is gradually increasing.
The AI track is expanding too broadly, with chips, power, hardware, software, models, and applications stacked layer upon layer. New financings, new stories, and new concepts emerge every day, and each party is telling a self-consistent version that serves its own interests. As the book *The Signal and the Noise* puts it: in the era of information explosion, noise is always more abundant and faster than signals.
The AI industry is in exactly this state. Everyone is telling fanatical long-term stories, but no one dares to be certain about short-term certainty. For the hundreds of billions of dollars invested in infrastructure, what is the ratio of long-term dividends to short-term bubbles? Does AI Infra have real barriers? Will the cooling of market sentiment affect the financing and exit in the primary market? What are the differences and gaps between China and the United States in AI Infra? What opportunities are there for overtaking on curves? These questions require us to filter out the "noise" and find the real "signals". With these trends and questions, ChinaVenture, together with Zhongguancun Science City Corporation and Zhongguancun Venture Street, co-hosted the in-depth closed-door salon themed "The First Year of Agent: Reconstructing New Forces of AI Infra".
"The Seven Sisters Stick Together, While China's Market Is Fragmented"
In this AI Infra salon, Jiuzhang Yunji, Terminus, and Convergent Intelligence are all core direct Token suppliers in China. With their self-developed scheduling, heterogeneous integration and full-link optimization capabilities, they have their own ideas to solve problems such as GPU idling, low computing power utilization, and high inference cost.
Coinciding with the large-scale outbreak of Agent and the explosive growth of Token demand across the network, enterprises that can effectively improve efficiency, reduce costs, and ensure stable computing power supply have already become the core underlying support for the refined upgrading of AI Infra.
Jiuzhang Yunji has deployed 12 intelligent computing center nodes across China, and has expanded to the Southeast Asian market, with its business covering three major sectors: new intelligent computing cloud, training factory and Token factory.
Shang Mingdong, Co-founder and COO of Jiuzhang Yunji observed an industry trend: the computing power supply in the US market is highly concentrated, and leading cloud vendors and AI infrastructure companies have formed close ecological collaboration; while the Chinese market is more driven by actual demand, showing more scattered and industry-close characteristics.
The core team members of Convergent Intelligence have scientific research and engineering backgrounds from the Institute of High Performance Computing, Department of Computer Science, Tsinghua University, inheriting more than ten years of technical accumulation from the relevant team, and most of the R&D core team members have Tsinghua backgrounds. When it was established at the end of 2023, when the "Hundred Model Battle" was at its peak, Convergent Intelligence chose to clearly focus on large model inference, which was rare among startup companies in the industry at that time.
Wu Wenjie, President and CFO of Convergent Intelligence was an investor and researcher in this industry three years ago, and now she is engaged in entrepreneurship. She feels that: "Real entrepreneurs and practitioners are deeply aware that there is still a big gap between domestic GPUs and overseas high-performance GPUs, and it is getting increasingly difficult to obtain these overseas GPUs."
Liu Yue, Head of AI Infra Business of Terminus focuses on the scenario side. Terminus was founded in Chongqing in 2015. Starting from AIoT scenarios, it has grown into an AI-IoT unicorn, with shareholders including state-owned central enterprises, well-known investment institutions such as IDG Capital and AL Capital, as well as leading industrial capitals such as JD, iFLYTEK, and SenseTime. Up to now, its products have been deployed in more than 170 cities around the world, serving more than 900 customers.
His core view is: Terminus's strategy is to use vertical scenarios to drive the product iteration of AI Infra and Agent, which is different from the mainstream path on the market that focuses on general Token services and computing power leasing. "US vendors are more concentrated in infrastructure and upper-layer general models; the domestic market is different, with high scenario density, complete industrial categories, and a flourishing application layer."
The gap at the chip level also has another form. Liu Chen, Managing Director of Haisong Capital gave an example: now there are some new companies that specialize in inference chips, and even directly write the architecture of a certain model company into the chip, which is equivalent to making a dedicated chip for a specific model, with operating efficiency 100 times higher than existing solutions. The obvious cost is that once you switch to another model or the model is iterated, this chip will no longer be usable.
Liu Chen further explained that there are many intermediate zones between absolute versatility and absolute efficiency, and each chip only solves a specific scenario in the intermediate zone. "It is difficult to solve all problems through one company or one product." Large factories cannot acquire the entire company, so they acquire talents and patents instead.
For the path of domestic general-purpose GPUs, the experience of Imagination Microelectronics itself is quite illustrative. Imagination Microelectronics was founded in Chongqing in September 2020. Its founder Tang Zhimin participated in the R&D of Loongson 1 and Loongson 2, and is also one of the founders of Haiguang Information. 80% of its R&D personnel come from Haiguang Tongxin, and its products on sale have reached the third generation. The company has raised a total of about 3 billion yuan in financing, and its valuation once reached 15 billion yuan, but the capital pressure has always been hanging over it. It got relief through several billion yuan of strategic investment from investors in 2025, and in April this year, it signed an agreement with CITIC Securities to start all preparations before listing.
Wang Yaoyu from Imagination Microelectronics judged that NVIDIA is not a pure hardware company. Its success lies in software, as it has built the CUDA ecosystem, and the docking of various hardware and software based on the ecosystem forms the industry barrier. "At present, China does not yet have the ability to achieve this level. It is indeed very difficult for the domestic GPU market to achieve full universality, because there is no unified industry interface." His proposed path is to take customization first: "Different from the difficulty of universality, this path is achievable for domestic chip companies, and it should be faster than pursuing universality. With the rapid expansion of demand for AI Agents, domestic GPUs can meet the rapidly growing market demand through customization."
Moreover, the gap in chip supply between China and overseas is further widening compared with the past. Chen Qiuwu, Co-founder & CTO of AIGCode presented a set of intuitive comparison data at the salon scene: the scale of computing power cards produced by NVIDIA last year was 100 times that of Ascend, and this gap further expanded to about 200 times this year.
As a technical team that has long been deeply adapted to domestic computing power, AIGCode has been working on the training and inference performance optimization of domestic chips. Its service targets not only include Ascend, but also recently it has successively taken over the model training and inference optimization requirements of large domestic factories. The team's domestic large model adaptation capabilities are also iterating rapidly. At the end of 2025, it has cooperated with HiSilicon to complete the pre-training of the 130B parameter model, and now it can complete the training task of the 700B parameter large model on the domestic hardware base.
However, the hard gap in hardware production capacity still exists objectively. Even if software optimization can increase the effective computing power of a single card or cluster, the insufficient total hardware supply is still an unavoidable practical constraint for domestic AI Infra.
Zhang Ruiqi, Managing Partner of Junde Danmu used to work in the Lenovo Capital Angel Fund. Dollar funds naturally have a dual-currency perspective. His judgment is that China and the United States will gradually form two different systems. "In the short term, packaging and heat dissipation will be relatively important in China, because the process technology is not mature yet and the power consumption is relatively high, while the United States has always been pursuing more advanced chip computing power."
In terms of entrepreneurial ecology, he made an analogy: the US venture capital ecology is a bit like pizza, with two layers and relatively simple; while China's is more like a thousand-layer cake, much richer.
His observation on the capital side is that the "Seven Sisters" in the US stock market launched more than 300 mergers and acquisitions last year, while the number of domestic mergers and acquisitions is limited, which is the gap in capital volume. "Chinese entrepreneurs are constrained by the overall insufficient capital of the country. Compared with the capital volume obtained by US venture capital institutions, the gap is at least at the level of the currency unit, that is, the exchange rate from US dollar to RMB."
Jiang Xu, Executive Director of CMC Capital focuses on storage and interconnection in data centers. She believes that as AI computing power evolves from single-card, single-machine to 10,000-card-level clusters, the bottleneck is gradually extending from single-point computing capability to how data flows efficiently among computing, network and storage. The larger the computing power scale, the higher the importance of data migration and resource collaboration, and storage and interconnection will increasingly become the key infrastructure that determines the actual computing power utilization rate of the cluster.
This trend is more obvious in super nodes and large-scale clusters. Both Scale Up and Scale Out require not only higher multi-card interconnection capabilities, but also higher requirements for network architecture, storage access and heterogeneous computing power collaboration. Even if the performance of a single card continues to improve, if the interconnection bandwidth, communication efficiency or storage access speed cannot match, the newly added computing power cannot be fully released. Therefore, storage and interconnection are not only supporting links of computing power infrastructure, but also the key to converting the theoretical performance of chips into the effective computing power of the cluster.
From an investment perspective, Jiang Xu believes that the opportunity for AI Infra lies in the incremental demand for storage and interconnection generated by the expansion of cluster scale. Especially in links such as high-speed interconnection, high-performance storage and hardware-software collaboration, their value comes from the direct impact on cluster efficiency: on the one hand, they support larger-scale computing power expansion, and on the other hand, they reduce the computing power loss caused by communication and data access. Under the background of the continuous construction of the domestic AI computing power ecosystem and the continuous improvement of cluster scale, such infrastructure capabilities that connect computing and data are worthy of key attention.
Joule Heat Transfer is deeply engaged in the thermal management field of high-power equipment, with its business fully covering four core directions: commercial aerospace, laser communication, 5G equipment and AI computing power. Liu Xu, General Manager of Joule Heat Transfer emphasized that heat dissipation is the key defense line to ensure the stable operation of computing power clusters — every 10°C rise in the temperature of electronic equipment will double the failure rate.
In terms of technical strength, Liu Xu pointed out that the domestic heat dissipation technology level has been among the global first echelon. Taking a certain type of high-end heat dissipation system as an example, there are currently only 3 systems in the world that can achieve stable operation, and China accounts for one of them.
However, in terms of industrial implementation scale, the demand traction in the domestic market is still insufficient. Compared with the deployment volume of millions of units in the United States, the current domestic scale is about 200,000 units. This shows that although China has taken the lead in core technologies, further breakthroughs are still needed in large-scale application and commercial closed loop.
True Investment mainly focuses on two directions: AI applications and cutting-edge technology, and aerospace and advanced manufacturing are also within its investment scope. However, Li Jianwei, Managing Partner of True Investment frankly admitted that he missed the early stages of large models and chips. He did not see a clear inflection point until last year, and the most solid evidence appeared in programming. "Since October last year, at least Coding has achieved AGI", and the quality, hallucination rate and research depth of research reports are all improving.
He intuitively feels that global computing power is in an extremely tight state. In the United States and Europe, suitable computing power sites are quickly snapped up, and Chinese enterprises are competing with overseas enterprises for limited resources. Being able to find a stable computing power site with a scale of 50MW and relatively friendly electricity price conditions is already a very sought-after scarce resource in the current market.
Li Jianwei also raised a question about domestic computing power and NVIDIA: GLM5.3 on OpenRouter has a computing power call volume of about 100T per day, which is said to be supported by domestic computing power. If this data can be verified, will it subvert the mainstream narrative that the market has long believed that AI must rely on chips with advanced process technology?
Where Are the Opportunities
Capital is already shifting to inference. In the first quarter of 2026, about two-thirds of AI capital expenditure has turned to inference, focusing on running models rather than training models. RadixArk, the developer of SGLang, received a USD 100 million seed round with a valuation of USD 400 million. The cycle from open source to financing has been compressed to several months, which means that the competition in the Infra layer is moving forward from technology competition to ecological competition.
It is true that the demand for inference is growing, but the growth is not necessarily a linear surge. It will be continuously compressed by algorithms and software. The real opponent of Infra practitioners is likely to be the speed at which model companies themselves reduce Token consumption.
"The United States maintains a leading position in cutting-edge technology exploration, while China is better at solidifying the technology and calculating the cost clearly, so that AI capabilities can enter thousands of industries in a low-cost and large-scale way."
Shang Mingdong believes that training is advance preparation, and inference is the actual usage process. Following the main line of inference, he sees that storage is becoming a new opportunity point — by optimizing the data access architecture to reduce repeated calculations, HBM, SSD and enterprise-level storage are expected to derive new technical forms.
In Convergent Intelligence's evaluation system, Tokens are divided into 5 levels from L1 to L5, among which L3, L4 and L5 are defined as high-quality Tokens for actual enterprise application needs. The company adopts the technical route of "fewer models, deeper optimization", and continuously optimizes a small number of mainstream large models with clear production demands. "We will choose the top intelligent large models