HomeArticle

Google conspires to "weld shut" Gemini, mysterious chip exposed with energy efficiency crushing TPU by ten times

新智元2026-07-21 11:47
a high-stakes gamble

Woke up to a bombshell from Google!

According to an exclusive scoop from The Information, Google is developing a mysterious server chip codenamed "Frozen v2" — which permanently etches part of the Gemini model's underlying architecture into silicon wafers.

It's not running on the chip; it's grown into the chip.

Google employees involved in the project estimate that, measured by the number of Tokens processed per watt, this chip will be 6 to 10 times more efficient than Google's latest-generation TPU. The earliest deployment is expected in 2028.

As soon as the news broke, Alphabet's stock price rose 3.7% on Monday, sweeping away the gloom of the previous few weeks.

"Freezing" a Model

The name "Frozen" is literal: "freeze" the model into the silicon wafer.

How radical is this? Whether it's NVIDIA's GPU or Google's TPU, they are essentially general-purpose chips — capable of running any model.

This sounds flexible, but it comes at a huge cost: The chip must make massive real-time decisions during operation, constantly moving data, and repeatedly querying where to go and how to compute.

It's like a kitchen that can cook any dish, but every time you make a dish, you have to look up the recipe, find the tools, and adjust the heat from scratch. The efficiency is easy to imagine.

Frozen v2 is built exclusively for the Gemini model.

It directly builds the critical decision paths for Gemini into the circuit — the chip no longer needs to "think" about what to do, because the paths are already cast into the silicon wafer.

Fewer steps to execute, less data to move, fewer decisions to make. It's equivalent to building a custom stove for a single dish, with the burner size, pot position, and heat settings all tailored specifically.

Every redundant step eliminated translates directly to tangible energy efficiency gains.

In fact, this idea has a predecessor. The original Frozen chip was led by Jeff Dean, Chief Scientist at Google DeepMind, with an even more extreme approach: directly burning the model weights into the chip.

What are weights? They are the core set of parameters that determine how the model answers questions.

But that plan was shelved — once the weights are burned in, the chip can only serve a specific version of Gemini. As soon as the model is updated, the chip becomes obsolete, with an absurdly short lifecycle.

Frozen v2 is a compromise version: no weight burning, only architecture solidification. The underlying computing blueprint is etched in, but the weights can still be updated.

This means that as long as Gemini's architecture doesn't undergo major overhauls, the chip can continue to work, and new version parameters can still be loaded into it.

However, Google is still weighing internally how much to lock into the silicon — the more locked in, the higher the efficiency, but the lower the flexibility. Striking this balance is the most critical engineering judgment for the entire project.

Forced by the Compute Shortage

Why is Google taking this risky move?

The answer is simple yet harsh: Computing power is truly running out.

It's not a "slightly tight" kind of shortage — it's a shortage so severe that Google is turning down revenue that's practically knocking on its door.

The compute shortage has sparked fierce internal conflicts at Google, forcing Google Cloud to start rejecting orders from external customers — which is almost the most extreme signal in a commercial company.

In June, Google signed a contract: paying SpaceX $920 million per month to lease the computing power of 110,000 NVIDIA GPUs, with the contract extending to 2029.

One of the world's largest cloud computing companies actually has to lease computing power from a rocket company — this scenario alone shows how crazy the AI arms race has become.

Google is not alone.

The entire industry is betting heavily on inference chips: startups like SambaNova and d-Matrix are working on them, so are OpenAI and Microsoft. Amazon has Trainium and Inferentia, and Microsoft has Maia.

Even NVIDIA can't sit still — last December, it spent $20 billion to license technology from inference chip company Groq.

Everyone has only one goal: Squeeze more Tokens out of every watt of electricity.

The path of "hardwiring the model into the chip" was previously only pursued by Canadian startup Taalas — which has raised over $200 million from investors including Quiet Capital and Fidelity.

Now, Google is stepping into the arena itself.

A High-Stakes Gamble

But 10x energy efficiency doesn't come for free. The price can be summed up in two words: rigidity.

Frozen v2 can only continue to work if future generations of Gemini still use the same underlying architecture as when the chip was designed.

In other words, Google is betting — betting that it won't need to drastically overhaul Gemini's development approach in the coming years.

From pure Transformer to MoE Mixture of Experts, from pure text to native multimodality, from dense inference to sparse activation — the iteration speed of model architectures rivals the most frenzied era of the mobile phone industry.

Google is well aware of this risk. That's why Frozen v2 won't be produced at the same scale as TPUs; it's more like a limited experiment: if the bet pays off, it will be a money printer, squeezing several times more Tokens per watt, and the saved electricity and cabinet space in data centers will be pure profit. If the bet fails, it will be considered tuition, teaching engineers how to build more specialized chips and accumulating experience for the next generation.

The timing is thought-provoking. Just last week, Bloomberg reported that Gemini 3.5 Pro has been delayed three times due to unmet targets in multiple metrics including programming capabilities.

With its model side under attack from all directions, Google is suddenly pushing more chips onto its deepest moat — chips.

From developing custom TPUs in the past to reduce reliance on NVIDIA, to today's move of welding the model directly into silicon, Google is moving deeper down the same path: The end of hardware-software co-design is hardware-software integration.

AI competition is beginning to compete at the physical level.

References:

https://x.com/kimmonismus/status/2079193746944790666?s=20

https://castlecrypto.gg/news/alphabet-shares-rise-following-reports-of-ultra-efficient-ai-chip/

This article is from the WeChat official account "AI Era", author: ASI Revelation, editor: Solomon, published with authorization from 36Kr.