HomeArticle

Google is developing new chips, with the Gemini architecture directly etched into the silicon wafer, delivering up to a 10x improvement in efficiency.

36氪的朋友们2026-07-21 09:25
Google is developing a new server chip tailored specifically for its Gemini AI model, which directly solidifies the model architecture into the silicon wafer to significantly enhance AI inference efficiency.

Google is developing a new custom server chip purpose-built exclusively for its Gemini AI model, embedding the model architecture directly into silicon to drastically boost AI inference efficiency.

According to a Monday report from tech media outlet The Information, this chip, internally called "Frozen v2," could deliver 6 to 10 times the efficiency of Google's existing in-house chips, with deployment possible as early as 2028.

Two people with direct knowledge of the matter told the outlet that the project is driven by Google's severe internal AI compute shortage — a shortage that has sparked internal frictions and forced Google Cloud to turn down partnership requests from some external clients. Frozen v2 pre-embeds portions of the Gemini model's decision-making logic into the chip, cutting runtime computation steps and data movement to boost response speed and reduce power consumption per unit of compute.

A Google spokesperson said in a statement that the company's teams "continuously research and explore new innovations to deliver the highest performance and efficiency to users and customers," adding that "not every project will make it to mass production, but this rigorous exploration is central to our full-stack approach."

Buoyed by the news, Google's stock price rose as much as 3.7% during Monday's trading session, before closing up 1.52%.

01

Dedicated Inference Chip: An Architecture Bet Prioritizing Efficiency

The core design philosophy of Frozen v2 lies in permanently etching the underlying architecture of the Gemini model into the chip's silicon substrate, eliminating the need for extensive real-time decision-making that general-purpose chips must perform when handling inference requests.

The report notes that measured by the number of tokens processed per unit of power, the chip is projected to be 6 to 10 times more efficient than Google's latest generation of existing in-house AI chips.

This logic stands in stark contrast to general-purpose AI chips. Google's existing Tensor Processing Units (TPUs), similar to NVIDIA's GPUs, are designed to be compatible with a wide range of AI models, which means they must handle a large number of dynamic decision-making processes when running a specific model. By pre-hardcoding some decisions, Frozen v2 achieves higher operational efficiency at the cost of reduced general-purpose compatibility.

Google is currently evaluating how much information should be locked into the chip to strike a balance between flexibility and efficiency. Sources indicate that the Frozen v2 chip will only be compatible with subsequent Gemini versions that share the same underlying architecture as the design target, meaning Google is effectively betting that it will stick with its current Gemini model architecture for the long term. Google states that the chip will support updates to model weights — the parameter settings that determine how the model responds to queries.

02

Project Evolution: From "Full Hardcoding" to "Flexible Hardcoding"

The Frozen project is not a brand-new concept.

According to the report, it originated as the original "Frozen" initiative, initially led by Jeff Dean, Chief Scientist at Google DeepMind, which planned to burn the model weights themselves directly onto the chip. However, this design would have limited the chip to serving only one specific version of the Gemini model, resulting in an unacceptably short lifecycle, and the plan has since been shelved.

Frozen v2 adjusts this approach, shifting to a "flexible hardcoding" strategy — hardcoding the model architecture rather than the specific weights, which retains the efficiency advantage while extending the chip's usable lifespan.

In terms of industry benchmarks, Canadian semiconductor startup Taalas has adopted a similar design philosophy of hardcoding specific AI models into silicon, and has raised over $200 million from investors including Quiet Capital and Fidelity. In comparison, companies such as SambaNova, d-Matrix, OpenAI, and Microsoft are also developing inference chips, though their technical paths differ from Google's. NVIDIA, meanwhile, announced a $20 billion technology licensing deal with inference chip startup Groq last December to strengthen its position in the inference market.

03

Positioning: A New Branch Beyond TPUs, Not a Replacement

Google explicitly positions Frozen v2 as a new product line separate from its existing TPU family, rather than a replacement for its current in-house chip ecosystem. TPUs, like NVIDIA GPUs, are general-purpose AI chips compatible with multiple models; Frozen v2 is built exclusively for Gemini, and the two product lines will develop in parallel.

Google currently has no plans to scale Frozen v2 production to match the volume of its TPU lineup.

Sources say the relatively limited deployment scope allows Google to bring in external design and manufacturing partners later in the project timeline, instead of requiring large-scale advance coordination as with major product launches. Furthermore, Google is treating this generation of products in part as an engineering experiment — to give its engineers hands-on experience building more specialized chips, especially as AI model architectures gradually stabilize.

Google's existing in-house chip strategy is already delivering results.

This year, Google launched its 8th-generation TPU and began selling it directly to cloud customers, challenging NVIDIA's market dominance. Google has signed a multi-billion-dollar TPU lease agreement with Meta, and is actively pursuing other cloud service provider clients. Sources note that developing chips in-house has already allowed Google to run Gemini models at lower cost — and if Frozen v2 succeeds, it could further amplify this cost advantage.

This article is from the WeChat Official Account "Wall Street CN Max", written by Bu Shuqing, published with authorization from 36Kr.