Sudden halt, Gemini 3.5 Pro faces a difficult birth, Google falls into the trap of disappointment
Just yesterday, the entire AI community was immersed in a state of euphoria.
Rumors poured in from all sides: Google's ultimate killer app — Gemini 3.5 Pro, codenamed "Cappuccino," would officially launch within 48 hours!
Boasting an ultra-large 2 million context window and a brand-new "Deep Thinking" reasoning mode, it is said that internal evaluations have already outperformed GPT-5.6 Sol and Claude Fable 5.
Clearly, this is a blockbuster product that is about to disrupt the AI landscape.
Everyone was excitedly counting down, gearing up to witness history.
However, after waking up from a night's sleep, the situation took a sudden turn.
An exclusive report from Bloomberg poured a bucket of cold water over everyone's enthusiasm: the launch of Gemini 3.5 Pro has been delayed — not just for a few days, but for months on end!
A launch that was supposed to go down in history has been paused by Google itself.
What on earth is the reason behind this?
The 48-Hour Frenzy and the Sudden Emergency Brake
Just yesterday, social platforms were flooded with spoilers about Gemini 3.5 Pro.
Codename: Cappuccino.
Ultra-long context: 2 million tokens.
Deep Thinking: The newly added "Deep Think" mode pushes its performance in mathematics, programming, and logical reasoning to unprecedented heights.
Comprehensive evolution: Capabilities in code writing, agent workflows, front-end UI design, and SVG graphic generation have been significantly enhanced.
Insiders predict that this will be Google's "ultimate weapon" for a full-scale counterattack against OpenAI and Anthropic.
The entire community was abuzz. Everyone was looking forward to the legendary launch date of July 17th.
However, this morning, a report from a Bloomberg journalist instantly dashed all hopes.
Insiders say that the development of Gemini 3.5 Pro has fallen behind schedule by months. The core issue is that the model failed to meet Google's strict internal standards in key capabilities, especially AI coding performance.
Just at the end of last month, Google urgently updated its training data in a final sprint to improve coding capabilities, but the results were "disappointing".
These two words mark the end of the 48-hour frenzy.
Google's stock price fell immediately after the news broke, with a maximum drop of 4.43% at one point.
While OpenAI and Meta's new models are advancing rapidly in coding capabilities, the difficult development of Gemini 3.5 Pro has directly triggered severe anxiety within Google.
Engineers, AI researchers, and executives are deeply frustrated, growing increasingly worried that Google is losing the not-so-thick moat it once had.
Google's "Tacitus Trap": Why Can't the Entire Company Build the Most Powerful AI?
Why did this highly anticipated "ace in the hole" fail to deliver?
This report reveals the deep-seated difficulties within Google, which is a microcosm of a sprawling tech empire in the era of transformation.
Innovation Speed Dragged Down by Bureaucracy
The report highlights a key detail: Google has a complex internal hierarchy with numerous stakeholders.
Launching a single model requires balancing the needs of its massive product lines, including Search, Maps, and YouTube.
This decision-making pattern of "trying to have it all" leads to scattered resources and slow decision-making.
A former employee used a vivid metaphor: "Trying to get the leadership of every department to pull in the same direction is like attempting to boil the entire ocean."
The result is frequent shifts in directives, redundant work across multiple departments, and a failure to form a unified force.
While OpenAI and Anthropic are advancing at the speed of startups, Google's massive "giant ship" is stagnating due to internal coordination hurdles.
One netizen commented sharply: "Google needs to slash its bloated bureaucracy to make progress in this field."
The Waterloo of AI Coding: Engineers' Pure-Blood Complexity and Insatiable Demand for Computing Power
Moreover, why did coding capabilities become the weak link? Behind this lies deeper contradictions within Google.
On one hand, Google boasts one of the world's most elite engineering cultures, which has fostered a "pure-blood" mindset.
Many veteran engineers adhere to the belief that "all critical code should be written manually." This distrust of AI-generated code has limited engineers from using Gemini to assist in development, over fears that proprietary code might leak into training datasets.
When Google finally recognized the importance of AI coding and decided to mandate AI-assisted code writing, a new problem emerged — insufficient computing power.
The report notes that when engineers tried to use internal AI tools, they frequently encountered computing power capacity limits.
The most ironic detail in the entire report is that in a company projected to spend $180 billion to $190 billion in capital expenditures this year, its own engineers cannot even access GPUs!
Wall Street data shows that Google's capital expenditures in the first quarter of this year reached $35.7 billion, more than doubling from the same period last year. After pouring so much money into purchasing chips and building data centers, what is the outcome?
Facing this chaos, Google is now trying to remedy the situation.
The Chief AI Architect is unifying AI programming tools across all departments under the underlying architecture of Google Antigravity, and a dedicated AI programming team has been established within DeepMind — but it may be too little, too late.
Internal Competition and the Vicious Cycle of Talent Drain
Google is not unaware of these problems. It has top-tier research labs including Google DeepMind, its cloud division Google Cloud, and the Android team, and has even set up multiple internal groups to tackle AI coding challenges.
However, this "internal competition" mechanism also leads to unnecessary internal friction.
Different teams work in isolation, with overlapping products and fluctuating strategies. Worse still, this chaos and frustration have directly caused top talent to leave the company.
The report states that a large number of researchers, disappointed by Google's lagging performance, have jumped ship to Anthropic and OpenAI.
This forms a terrifying closed loop: bureaucracy leads to inefficiency → inefficiency leads to lagging products → lagging products lead to talent drain → talent drain further exacerbates technological backwardness.
The delay of Gemini 3.5 Pro is the inevitable outcome of this loop.
The Entire Industry Sounds the Alarm: Tech Giants Fall Into the "Next-Gen Large Model Disappointment Trap"
Ethan Mollick of the Wharton School, while sharing the report, put forward a thought-provoking observation —
This is not just Google's tragedy, but a "cyclical tech winter" that the entire Silicon Valley is experiencing.
Mollick sharply points out that Google's current setback perfectly replicates the pain that Meta Llama 4 and xAI Grok 4 went through earlier.
He named this phenomenon the "Next-Gen Large Model Disappointment Trap."
Despite investing massive amounts of capital and computing power to train next-generation models, their actual performance improvements fall far short of expectations, leading to a noticeable decline in market leadership.
In the past, the industry believed in the Scaling Law. However, when model size expands to a certain threshold, the "brute force aesthetics" of simply piling on computing power and data begins to fail.
Data bottleneck: High-quality human text data has almost been "exhausted," and the effectiveness of synthetic data remains to be verified.
Algorithm bottleneck: The existing Transformer architecture and its variants may be approaching their performance limits. Diminishing returns: To achieve tiny performance improvements, exponentially increasing computing power costs are required.
In this game of tech giants, only OpenAI, with its Orion/GPT-4.5, has temporarily escaped this trap without any major setbacks.
It is certain that as model sizes approach physical and engineering limits, the difficulty of iterating cutting-edge models is rising sharply.
The delay of Gemini 3.5 Pro has made everyone wake up to reality —
We are now in a plateau phase. The era of breakneck progress, where "one day in AI equals one year in the human world," is coming to an end.
For the entire industry, this may actually be a good thing. When the noise fades away, people will truly begin to reflect on the real value of AI.
As for Google, the time and patience the market has left for it may be running out.
References:
https://x.com/Mr_Salio/status/207736089707741624811
https://x.com/emollick/status/2077849021150888408
https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals
This article is from the WeChat Official Account "AI Era", authored by ASI Revelation, and published with authorization from 36Kr.