HomeArticle

NVIDIA's AI servers will see a 15% price hike, leading to a $5 billion sharp surge in the cost of a 1GW data center.

量子位2026-08-24 09:49
Soaring memory costs are driving up hardware prices.

Nvidia is raising prices again!

According to Bloomberg, some major Nvidia customers have been notified that the price of AI servers to be delivered early next year may be increased by more than 15%.

The involved products include the existing Grace Blackwell and the next-generation flagship Vera Rubin.

The specific price increase will vary depending on GPU generation and memory configuration.

The Information reports that some GB300 and Vera Rubin 200 systems are expected to see a price increase of about 17%.

Calculated based on the current price model used in the report, for a 1GW-scale AI data center, this round of price hikes alone may add at least 5 billion US dollars in costs.

Sigh...

Note that this is not the first time this year that we have seen the price of Nvidia-related hardware go up.

Earlier there were consumer-grade RTX 50 series graphics cards, as well as the professional card RTX PRO 6000 Blackwell.

Now, AI servers have also joined the price hike queue.

Prices are rising across the board, Nvidia, what on earth does this mean????

Nvidia AI server prices will rise by more than 15% next year

Some of Nvidia's largest customers have been informed that the price of servers for its AI chips will rise by more than 15% due to soaring memory chip costs.

According to The Information, two people who received notifications from server manufacturers revealed that some of Nvidia's flagship AI server chip systems are expected to see a price increase of about 17%.

This price increase will take effect for products delivered at the beginning of 2027, mainly involving the Grace Blackwell 300 and Vera Rubin 200.

What does a 17% price increase mean?

According to estimates, the current price of a 72-GPU Vera Rubin rack is approximately 7 million US dollars; after the price increase, a single rack may reach around 8 million US dollars.

In this case, with just this round of AI server price hikes, building a 1GW-level AI data center may add at least 5 billion US dollars in additional construction costs.

Moreover, server prices have recently been in frequent fluctuations.

Previous interviews with server vendors and GPU cloud manufacturers found that the price of some components that make up Nvidia AI servers can even fluctuate by 40% within a week.

These include TSMC wafers, advanced packaging, network equipment, heat dissipation systems and memory.

An executive at a GPU cloud manufacturer said that the server racks they purchased recently once saw a 2% to 3% price increase every week.

Another executive mentioned that NVMe storage in GB300 servers was previously one of the parts with the most obvious price fluctuations, and currently some racks are already 10% to 15% higher than their so-called normal benchmark price.

In fact, since 2026, Nvidia-related hardware has seen three rounds of sharp price increases.

The first round took place in the consumer market.

Since this summer, the market prices of multiple RTX 50 series graphics cards have risen significantly.

The median price of RTX 5060 Ti 16GB has increased by up to about 39%, RTX 5070 by about 36%, RTX 5060 by about 27%, and RTX 5090 has also increased by about 9%.

Attention Please, these statistics refer to retail market prices, which include factors such as supply and demand, channel inventory, and cannot all be counted as Nvidia's official price adjustments.

The second round of price hikes targeted professional GPUs.

The professional card RTX PRO 6000 Blackwell with 96GB of video memory has also experienced significant price changes.

Previously, the quotation on Nvidia's official mall has been raised all the way from the early price, with an absolute increase of 2750 US dollars.

Since its launch, the price trajectory of RTX PRO 6000 has been shocking:

2025 early ~ 8565 USD → June 2026 13250 USD → August 2026 16000 USD.

The third round of price hikes is the current AI server price increase.

From gaming graphics cards to professional GPUs, and then to multi-million-dollar per rack data center systems, price pressure is moving upward along Nvidia's product line.

More notably, the next-generation Rubin has even begun to reassess its memory configuration.

TrendForce disclosed that starting from the third quarter of this year, Nvidia has expanded the original evaluation scope of Rubin Ultra, which was dominated by 12-Hi HBM4E, to multiple solutions including 8-Hi HBM4E, 12-Hi HBM4 and 8-Hi HBM4, and the final specification has not yet been determined.

The reason is that DRAM supply will remain tight in 2027, and there are still uncertainties in the verification time and mass production yield of 12-Hi HBM4E.

There are also reports that Nvidia has begun to reassess the HBM configuration of Rubin Ultra.

Who can we blame?

We can only blame the tight memory supply.

To what extent has the memory market price risen

Looking at this year's memory market, you will probably understand what is happening.

TrendForce predicted at the end of March this year that the traditional DRAM contract price in the second quarter of 2026 will increase by 58% to 63% quarter-on-quarter.

NAND Flash has risen even more sharply, with a projected quarter-on-quarter increase of 70% to 75%.

In the third quarter, the growth rate has slowed down, but prices are still continuing to rise.

Among them, server DRAM contract prices are expected to rise 13% to 18% quarter-on-quarter. TrendForce also judges that from the second half of 2026 to the second half of 2027, server DRAM prices will likely continue to rise quarter by quarter.

The problems on the supply side are more intractable.

DRAM used in servers, PCs, and mobile phones shares wafer production capacity with HBM in AI chips.

The more AI servers there are, the greater the demand for HBM, and the more willing memory manufacturers are to prioritize advanced processes and wafer production capacity to HBM.

HBM itself consumes more wafers than ordinary DRAM.

It requires larger dies, more stacking layers and more complex packaging processes. The same wafer input cannot bring a proportional bit output.

As a result, while HBM expands production rapidly, the production capacity available for ordinary server memory is further squeezed.

TrendForce expects that the bit supply of server RDIMM in 2027 will only increase by about 15% to 20% year-on-year, significantly lagging behind the growth of server CPU shipments.

At the same time, HBM and SOCAMM are continuing to increase their consumption of DRAM wafer production capacity.

This has forced some customers to take the initiative to adjust their configurations.

Since the first half of this year, some cloud service providers and server OEMs have adjusted 96GB and 128GB RDIMMs in some systems to 32GB and 64GB modules, hoping to reduce procurement costs.

Nvidia itself is also making similar trade-offs.

Due to tight LPDDR5X supply, Nvidia has decided to halve the SOCAMM memory capacity in the next-generation Vera Rubin Superchip.

Calculated based on the initial capacity allocation of Samsung, SK Hynix and Micron, the amount of LPDRAM that suppliers can currently meet is only about 60% of Nvidia's estimated demand.

AI has been frantically consuming GPUs over the past few years, GPUs have driven a surge in HBM demand, and HBM continues to squeeze DRAM production capacity... By 2026, memory prices have begun to push up the costs of GPUs and AI servers in turn.

Now, the boomerang that was thrown out has finally hit Nvidia itself hard.

Nvidia has taken a step forward in addressing power issues

However, the rise in memory costs is only the apparent reason. Behind it are the improvement of high-end system integration, tight supply chains, and the competition between cloud vendors and large model companies for delivery cycles.

In particular, electricity, land, grid connection, heat dissipation, networks, and whether the entire park can be completed on time have all begun to affect when GPUs can generate revenue.

Nvidia has recently been extending its reach into this layer.

On August 21, Nvidia announced that it has made a minority equity investment in Cloverleaf Infrastructure, a U.S. data center infrastructure developer.

The specific amount has not been announced, but the Wall Street Journal previously reported that this investment could reach hundreds of millions of US dollars.

What Cloverleaf does is help data center developers find suitable land, communicate with utility companies and energy suppliers, secure electricity and other infrastructure resources, and turn these conditions into projects that can build data centers.

The company was only established in 2024 and has already advanced multiple GW-level projects.

After the cooperation, Cloverleaf will also deploy the Nvidia DSX platform to use it to optimize data center site selection, power, cooling and computing infrastructure planning.

This is also one of the reasons why Nvidia has frequently entered the upper reaches of infrastructure recently.

A few days ago, Nvidia just announced a $1.5 billion investment in SB Energy, a SoftBank-owned data center developer.

Further back, it has also cooperated with institutions such as Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR, hoping to leverage more than 500 billion US dollars in third-party funds in the long term for the construction of AI infrastructure.

What Nvidia is worrying about now has expanded all the way from "whether the chips can be manufactured" to "whether customers have money to buy", "whether there is land to build data centers", "where the electricity comes from", and "when the GPUs can be powered on after delivery".

At the same time —

It is reported that Amazon, Microsoft, Google and Meta are all ramping up their own AI chips, hoping to reduce their dependence on Nvidia.

One More Thing

With hardware sales reaching their limit, Nvidia is also looking for higher efficiency for GPUs.

In its newly announced AVO research, the Agent first optimizes the GPU kernel by itself, and then migrates to the ARC-AGI-3 test.

The official disclosed that AVO explored more than 500 optimization directions, submitted 40 kernel versions, and achieved a maximum speedup of 10.5% over FlashAttention-4 on DGX B200.

Reference Links:

[1]https://x.com/KobeissiLetter/status/2091243942595445165

[2]https://www.straitstimes.com/business/nvidia-customers-notified-about-ai-related-price-hikes-above-15-bloomberg

[3]https://sg.finance.yahoo.com/news/nvidia-customers-face-over-15-211744754.html?guccounter=1&guce_referrer=aHR0cHM6Ly93d3cuZ29vZ2xlLmNvbS5oay8&guce_referrer_sig=AQAAACrczxkZWDLqa82E33Tul0CxSc3AS3XD5QwVQ0rWJHGrgCthVxLwz6JR4Iv78gsNNfCcn4DEcCMC-a3SxQWCLezzf43xr9NXpvmXQg-IAe5ggDay3Ss0mCNW1Iz4-kuNGHPvWMff9imnnlr2VVymvHsZNfGj48TsCljCyHEcHs4W

[4]https://www.reuters.com/business/nvidia-customers-notified-about-ai-related-price-hikes-above-15-bloomberg-news-2026-08-