首页文章详情

Have skyrocketing memory prices forced major tech companies to become hardware scavengers? Even Google has to dismantle old servers to deal with the urgent shortage?

江瀚视野2026-09-20 10:58
Memory shortages are leading to price hikes, and Google is coping with the situation by dismantling old servers to recycle memory.

Recently, the continuous surge in memory prices has affected everyone. Not only are ordinary consumers complaining that they can't even afford the three essential items for the new school semester, but even the deep-pocketed major internet giants are deeply troubled by the memory price hike. Some media even broke the news that Google, a well-known tech giant, has been forced to become a "scavenger" by the sky-high memory prices, and has started dismantling old server memory modules to meet urgent needs. How exactly should we view this situation?

1. Sky-high Memory Prices Have Forced Tech Giants to Become Scavengers

According to reports from Kuai Technology, the continuous tight supply of DRAM has pushed AI giants to the point of dismantling old servers to recycle memory modules.

Nikhil Cherian, Senior Director of Supply Chain Infrastructure at Google, recently revealed that the AI industry has rapidly shifted from being computing power-constrained to memory-constrained, and high-performance memory currently accounts for about 75% of the material cost of a single AI server.

According to Cherian, to break through the memory bottleneck, Google is making efforts on both hardware and software. At the hardware level, it has even built an internal recycling supply chain to dismantle decommissioned servers and recover usable DDR4 memory modules.

Cherian admitted that Google has specially designed hardware adapters to connect the previous generation of DDR4 memory to the new generation of AI servers, and is specially reintroducing decommissioned servers to remove the DDR4 modules inside.

At the software level, Google continues to optimize the underlying function libraries, model architecture and KV Cache compression technology, trying to reduce the memory consumption per unit of computing power.

This year, Google launched two TPU ASICs, namely TPU8t for training and TPU8i for inference. Each TPU8i chip is equipped with 288GB HBM3e, provides 8.6 TB/s memory bandwidth, and integrates 384MB on-chip SRAM for storing active KV cache to reduce the demand for querying external system memory.

The TPU8i server uses Google's self-developed Axion processor based on Arm architecture to replace the traditional x86 host CPU, and relies on the high-speed DDR5 memory architecture to handle host-level tasks such as data preprocessing.

2. Why Even Google Has to Dismantle Old Server Memory for Reuse?

At first glance, Google acting as a scavenger sounds quite surreal. Even a tech giant with a trillion-dollar market cap like Google has started to "pick up scraps" and dismantle memory sticks from decommissioned servers. Many people's first reaction is: Is Google short of money? Or is this just a public relations gimmick?

First of all, the competition focus of the AI industry has shifted from computing power to memory. For a long time in the past, when the industry discussed the bottleneck of the AI industry, everyone was always talking about GPUs, and all eyes were on graphics chips, assuming that as long as enough high-performance GPUs were obtained, the computing power problem of large models could be easily solved. But as the industry developed in reality, people gradually realized that no matter how powerful the GPU is, if the memory cannot keep up, a large number of model weights and context data cannot be stored, and the powerful computing power cannot exert its due value. As large models become larger and larger, and the context window becomes longer and longer, each inference and each round of training will consume a huge amount of memory space. Memory is no longer an insignificant supporting role in servers, but has directly become the core shortboard restricting the expansion of AI businesses. That's why the shortboard of the barrel effect has shifted from the previous graphics chips to the more realistic memory chips.

For an enterprise of Google's scale, it needs to run a huge number of models and build computing clusters all over the world. Memory is not an optional extra, but a real rigid demand. The cost of buying brand new server memory has risen to an unacceptable level, and memory already accounts for an extremely high proportion of the material cost of AI servers. Continuing to purchase large quantities of brand new memory will directly erode a large amount of profit margin of the cloud computing business. The DDR4 memory in decommissioned servers does not fail completely in hardware. It is only that the whole machine of the old generation of servers has reached its service life and is forced to be taken offline and eliminated. Friends who are familiar with computer hardware know that unless the product is completely upgraded and cannot be used, the service life of a memory stick is often very long. The Shenzhou laptop I bought when I was pursuing my graduate degree in 2013 still uses old DDR3L memory, but it can still work normally today. If the screen of the laptop had not aged badly, this laptop would still have quite good performance.

For Google, this kind of reuse is even simpler. Google can specially design hardware adapter boards to modify old memory to fit new generation AI servers, which essentially reactivates the hardware assets that have been settled in its own system. Many people will wonder that DDR4 has a performance gap compared with the new generation of DDR5, so why force reuse? Looking at the reality of the industry, not all AI tasks require extreme memory bandwidth, and some inference scenarios can fully accept old-specification memory. Sufficient performance is the optimal solution.

Secondly, the structural imbalance on the supply side is the root cause of the problem. Where is the problem now? The supply side represented by the three core giants Samsung, SK Hynix and Micron, driven by huge profits and caught up in industrial trends, has allocated the vast majority of advanced process production capacity and packaging resources to HBM. This is completely reasonable in terms of business logic, because under the boom of AI, the profit margin and order certainty of HBM far exceed other products. But the consequence of this strategic choice is disastrous: a large amount of production capacity that originally served the general market has been squeezed out, resulting in a serious supply shortage in the consumer market.

This shortage is not a simple imbalance between supply and demand, but a structural shortage. In order to pursue high-profit HBM orders, the giants have voluntarily given up the supply guarantee for the general market, breaking the ecological balance of the entire industrial chain. For demand-side players like Google, this mismatch is extremely painful. Although they need top-level HBM to support core computing power, they also cannot do without the support of general-purpose memory in their huge infrastructure systems. When the entire market presents a situation of "overcrowded high-end production capacity and cut-off supply for mid-to-low end products", even with abundant capital, Google can hardly buy cost-effective general-purpose memory in the open market, especially the familiar consumer-level and server-level memory.

Worse still, this structural shortage makes the supply chain extremely fragile. Once HBM production capacity occupies an absolute dominant position, any slight production capacity fluctuation will trigger a chain reaction, and the shortage of general-purpose memory will limit the expansion capability of the entire system. Therefore, Google dismantling old servers is to some extent a forced move caused by this structural production capacity mismatch.

Thirdly, Google acting as a "scavenger" at this time is also for cost control. Although the practice of Google dismantling server memory is used by many "scavengers" among domestic internet practitioners, we must not think that this is just a makeshift folk method for engineers. Behind it lies the philosophy of extreme cost control. On the one hand, Google is building its own internal circular economy system. Many servers have reached the end of their service life, but their core components, especially memory particles, are far from the end of their physical life. The traditional scrapping process is a huge waste of resources. Now, by dismantling these parts for recycling and reuse, the cost is digested within its own ecological closed loop. This is not just about saving money, but about improving asset turnover and residual value utilization.

On the other hand, which is also the most easily overlooked point, this is to continuously improve the resilience level of its supply chain. The production capacity of external suppliers may be cut off at any time due to various force majeure or games between giants, and Google cannot put the lifeline of its large models in the hands of others. By establishing this internal recycling mechanism, Google is equivalent to building a "reservoir" in its huge data centers. I am not afraid of water shortage outside, because there are still stocks in my own pool. The value of this supply chain resilience in extreme market environments is far greater than the saved hardware procurement costs. To put it bluntly, this is to exchange space for time, and use existing stock to support incremental business.

Fourth, the systemic risk of memory shortage is growing. Even Google, a top-tier giant with abundant capital and strong supply chain discourse power, has been forced to be so cost-conscious that it even needs to rely on dismantling old parts to maintain operation. This actually sends out a very dangerous industrial signal. It means that the current problems in the memory market have reached a precarious state. We often say that the duck knows when the river water gets warm. In the industrial chain, the reaction of large manufacturers is that duck. If even Google starts to worry about memory, how harsh the living environment faced by small and medium-sized AI enterprises will be?

If memory prices continue to rise unrestrainedly, it will have a huge negative impact on the development of the entire industry, and may even become the last straw that crushes some AI innovation companies. Industrial upgrading cannot become a tool for a few upstream suppliers to harvest the entire industrial chain. When the infrastructure cost is so high that even the largest application side finds it unbearable, the bubble is not far from bursting.

The current AI industry is like building a highway leading to the future, but the asphalt necessary for road construction has been hyped up to sky-high prices by several suppliers. Even the largest construction party can only dig up the asphalt from old roads to fill the gaps. How far can this road be built? If the structural contradiction in the memory market cannot be balanced as soon as possible, the large-scale implementation and commercial closed loop of the AI industry will likely come to an abrupt halt at the checkpoint of "sky-high memory". This problem deserves the vigilance of the entire industry.

When we are forced by rising memory prices to see the three essential items for the new school semester become more and more expensive, a single mobile phone can cost as much as all home appliances combined, and even Google is forced to become a "scavenger", isn't this memory shortage problem worth our deep thinking?

This article is from the WeChat official account "Jiang Han's Vision Observation", the author is Jiang Han's Vision Observation, and it is published by 36Kr with authorization.