The killer weapon forced out by 180 billion dollars: Google's Frozen v2 burns Gemini into hardware, boosting tokens per watt by 10 times, with its stock price rising by 3.3% first
Citing people familiar with the matter, The Information reports that Google is designing a server chip with the internal codename "Frozen v2", which plans to integrate part of the Gemini model directly into the hardware to improve model inference efficiency.
This chip could be deployed as early as 2028. According to current internal expectations, measured by the number of Tokens generated per unit of energy consumption, its efficiency may reach 6 to 10 times that of Google's latest self-developed AI chip.
However, this project is still in the design stage, and the engineering team has not yet finalized how much model information will be solidified into the chip.
Google has not officially confirmed that Frozen v2 will definitely go into mass production. It only told TechCrunch that the company has been researching and experimenting with new technologies, emphasizing that "not every project will eventually be put into production".
Therefore, the 6 to 10 times efficiency improvement is currently closer to an internal design goal, rather than a publicly tested product indicator. Even if the existing plan is not finally implemented, this project still reveals an important change in Google's chip strategy: in addition to continuously iterating TPUs, Google has begun to try to design more dedicated hardware for specific Gemini models.
Frozen v2 is not the next-generation TPU
The key to understanding Frozen v2 is that it is not the codename for Google's next-generation TPU.
Reuters, citing reports from The Information, states that the goal of the Frozen project is to build a new chip system independent of TPUs, rather than to replace TPUs.
Google is still advancing the TPU roadmap at the same time. One next-generation processor, codenamed "Icefish", may also enter mass production in 2028. According to relevant reports, the main computing part of Icefish is planned to be manufactured by TSMC. Google is also discussing the use of Samsung's 2-nanometer process to produce components connecting memory, and co-designing the chip with MediaTek.
Google may have two categories of AI chips with different positioning at the same time.
One category is TPUs. It remains an AI accelerator oriented to workloads such as matrix computing, training, inference, and reinforcement learning. It can run models of different scales and architectures, and can also be provided to external customers through Google Cloud.
The other category is Frozen. According to the information disclosed at this stage, it may map part of the Gemini model structure, parameter information, or specific computing patterns deeper into the hardware, sacrificing a certain degree of versatility in exchange for higher inference efficiency.
From an industrial perspective, this is equivalent to pushing the specialization of AI chips one step further.
GPUs mainly rely on general-purpose parallel computing capabilities to cover different AI models; TPUs are specially designed for tensor computing in machine learning; Frozen v2 may further optimize for a specific model family. What it faces is no longer just "what kind of AI computing" but "how Gemini specifically performs its computing".
The closer a chip is to a fixed workload, the more opportunities it theoretically has to reduce unnecessary data movement, control logic, and computing overhead. But at the same time, the scope of application of the hardware will also narrow. Whether Frozen v2 is viable depends on whether Google can find parts of the Gemini model that are stable enough, large enough in scale, and worthy of hardware implementation during the chip design phase.
TPUs are becoming more open, while Frozen may become more closed
The timing of Frozen v2's emergence is also noteworthy.
In April 2026, Google released the 8th-generation TPU, and for the first time clearly divided the same generation of products into two architectures: TPU 8t and TPU 8i. TPU 8t is mainly used for large-scale pre-training, while TPU 8i is optimized for sampling, inference, and model serving.
Google disclosed that the TPU 8t can deliver up to 2.7 times the performance per dollar of Ironwood in large-scale training tasks; the TPU 8i can improve performance per dollar by up to 80% in low-latency large-scale Mixture of Experts model inference; both chips can achieve up to twice the performance per watt of Ironwood. The TPU 8i is also equipped with 288GB of high-bandwidth memory and 384MB of on-chip SRAM to reduce data access costs in latency-sensitive inference.
After the 8th-generation TPU, the attention of the developer community quickly shifted from the chip itself to a more practical question: To what extent can Google's complete technology stack covering models, chips, and data centers be transformed into Gemini's product capabilities?
In a related discussion on Hacker News, a representative view is that Gemini has proven that model capabilities do not necessarily have to be achieved by continuously expanding scale.
A user with the ID himata4113 speculated that the model scale of Gemini 3 Pro and Flash may only be one-fifth to one-tenth that of Claude Opus and GPT-5-level models. In his view, Gemini generates significantly fewer Tokens, yet can still perform at a similar level to Opus and GPT in original problem solving that does not rely on tools and search.
However, this efficiency advantage has not been fully translated into agent capabilities.
himata4113 believes that Gemini has not invested enough in the inference and execution stages. The model often generates incorrect tool calls and performs poorly on agent tasks. In other words, Gemini may already be an efficient problem-solving model, but it is not yet a sufficiently reliable task executor.
He still believes that sooner or later Google will launch a product that is one generation ahead of the state-of-the-art at that time, "making everyone shine". But the premise is that Google can truly transition from the prototype stage to a formal product, instead of continuously releasing preview versions. In his view, many of the AI products Google has launched so far still look like hastily released prototypes, mainly used to demonstrate progress to investors and validate concepts.
A user with the ID deaux expressed strong skepticism about this, especially disagreeing with the claim that Gemini Pro is five to ten times smaller in scale. He put forward the opposite speculation: Google's hardware advantage may not help it run smaller models, but allow it to run larger models at lower costs and higher speeds.
deaux believes that the scale of Gemini 3 Pro may be smaller than GPT-5.4 and Claude Opus 4.6, but a gap of more than five times "seems too large". From the perspective of actual usage experience, Gemini 3 Pro is instead one of the most "intelligent" models in all aspects, especially in the humanities field.
In his view, Gemini 3 Pro is rich in knowledge and leads in generating a large amount of natural text that is close to human expression. The advantage is even more obvious in some niche languages. In his ranking of multilingual capabilities, the top four are all from Google: Gemini 3 Pro, Gemini 3 Flash, Gemini 2.5 Pro, and Gemini 2.5 Flash.
Even the largest models of OpenAI and Anthropic can hardly compete with Gemini on these language tasks.
However, deaux also pointed out that Gemini's shortcomings are equally obvious. Its mathematical capabilities are relatively weak, and its agent capabilities lag even further. The Gemini chat app itself is also "more than one generation behind", with a product form that is not much different from ChatGPT when it was first released three years ago.
Although many users believe that Gemini's agent products are still immature, the community's judgment on Google's long-term advantages is relatively consistent: Google's biggest bargaining chip is not to launch another chatbot, but to embed AI into search and its existing product ecosystem.
User onlyrealcuzzo believes that Google is obviously prioritizing the development of AI technologies that can enhance or even partially replace traditional search, because search is Google's lifeblood and also the scenario where profitability is easiest to achieve.
Google has a user base that is a billion people larger than its other competitors. Even if all other large model products and local mobile applications are added together, their user request volume may still not be comparable to Google Search.
asah added that Google can also integrate Gemini into Google Workspace, YouTube, Maps, Google Cloud, and a large number of already successfully running applications and behind-the-scenes infrastructures.
This forward-looking layout in infrastructure has also accelerated the commercialization process of Google's chips.
Alphabet said at its Q1 2026 earnings call that Google will begin delivering TPU hardware to some customers, allowing them to deploy it in their own data centers, rather than just renting computing power through Google Cloud.
Related revenue will start to be recognized in the second half of 2026, with most of it expected to be reflected in 2027. Part of the internal infrastructure is gradually transforming into a computing platform for external output. To attract more model companies and enterprise customers, TPUs need to support software stacks such as JAX, PyTorch, vLLM, and XLA, and minimize the cost of model migration and adaptation.
Google's emphasis on the native PyTorch experience and cross-generation code compatibility in the 8th-generation TPU is essentially to improve the versatility of TPUs. It is exactly the opposite direction.
It does not necessarily need to serve a large number of external models, nor does it need to build a software ecosystem similar to CUDA or Cloud TPU. As long as it can run Gemini stably and be called on a large scale in Google Search, the Gemini app, Workspace, advertising, and cloud services, Google can spread out its chip R&D costs.
Therefore, Google's chip strategy is forming a "dual-track system": TPUs are responsible for expanding the versatility and commercial reach of the infrastructure, while Frozen is responsible for driving down the unit operating cost of the core internal model to a lower level.
After $180 billion in spending, Google must answer the return question
The most notable thing about Frozen v2 is not the yet-to-be-validated figure of "6 to 10 times", but the metric Google chose to measure it — the number of Tokens generated per unit of energy consumption.
In the previous phase of AI chip competition, manufacturers were more accustomed to emphasizing floating-point computing power, parameter scale, chip interconnection bandwidth, and training cluster scale. But as large models enter the large-scale inference phase, the peak computing power of a single chip cannot directly determine commercial efficiency.
For model developers, a more realistic question is: Under the same power and server budget, how many user requests can be processed, how many valid Tokens can be generated, and what kind of response speed can be maintained.
Especially in Agent scenarios, a single user request may trigger planning, search, tool calls, code execution, result verification, and multiple rounds of reflection. On the surface, it only completes one task, but behind it, it may generate far more Token consumption than a chatbot. The efficiency of chips in data movement, memory access, and low-batch inference begins to directly affect product gross margins.
The design of Google's 8th-generation TPU has already reflected this change. TPU 8t mainly addresses training throughput and cluster scaling issues, while TPU 8i increases memory bandwidth and on-chip SRAM for low-latency inference and multi-Agent collaboration.
If Frozen v2 continues to solidify part of Gemini's computing patterns into hardware, it is equivalent to pushing this division of labor further to the model layer. It has special value. Gemini is not just a standalone app, it is entering Search, Ads, Workspace, and Google Cloud. As long as the model is deployed in enough products, even if each call saves a small amount of energy consumption and computing time, after accumulating to billions of requests, it may translate into a significant difference in infrastructure costs.
Frozen v2 has attracted attention from the capital market, which is also related to Alphabet's rapidly expanding capital expenditures.
In 2025, Alphabet's capital expenditure was $91.4 billion, of which about 60% was invested in servers and 40% in data centers and network equipment. By the first quarter of 2026, the company raised its full-year capital expenditure guidance to $180 billion to $190 billion, nearly double that of 2025. Alphabet also said that capital expenditures in 2027 will be significantly higher than in 2026.
Alphabet management stated that internal and external demand for AI computing power has reached unprecedented levels, and Google Cloud is still in an environment of tight computing power supply.
Related reports on Frozen v2 further mention that the shortage of computing power has triggered pressure on Google's internal resource allocation, causing Google Cloud to reject some external customer orders.
According to Reuters, after the news of this new Google chip was released, Alphabet's stock price rose by about 3.3% in early trading on July 20, and by press time, Alphabet's stock price had risen by 2%. The market is not just concerned that Google has one more chip, but whether it has found a way to reduce the capital intensity and operating costs of AI infrastructure.