首页文章详情

Why has Gemini, which was absolutely crushing it last year, gone completely silent?

差评2026-07-24 08:09
I admit that Gemini was extremely strong last year, but by now, its public reputation has been widely referred to as the "North American soybean bun"...

I admit that Gemini was impressive last year, but by now, its public reputation has completely tanked, becoming a total laughingstock in North America.

When you ask it to solve a math problem, it somehow spits out nonsensical, garbage results right after you finish the prompt.

When you ask it to generate a photo of a beach, it somehow frames the user as a pedophile out of nowhere.

Thanks to these absurd, nonsensical outputs, Gemini has been widely mocked as a total joke across the entire North American internet.

Beyond the sharp decline in the model's own public reputation, Google has also seen extremely frequent personnel changes among its top AI researchers recently.

The co-author of Transformer, whom Google poached for a whopping $2.7 billion two years ago, has now left the company and joined OpenAI.

Even the Nobel laureate who has worked diligently on AlphaFold for 9 years is now reportedly in active talks with Anthropic for a potential move.

Everyone was eagerly waiting for the launch of Gemini 3.5 Pro to help Google turn the tide, but the release was delayed from May all the way to July with no sign of launch.

A few days ago, there was finally some movement, but what Google released was not the flagship Pro model, but a budget-focused Flash model marketed as "affordable".

So what's the actual capability level of this new model?

To put it simply, while other competing models are positioning themselves against Fable and GPT-5.6 Sol, Gemini 3.5 Flash is only benchmarking against GPT-5.6 Luna and Claude Sonnet 5.

Just from this benchmarking choice alone, we can already tell Google's low expectations for this new model...

In the fields that everyone values the most right now — AI coding and intelligent agents — Gemini's performance has plummeted dramatically.

You might not even find its name listed on the first page of most mainstream leaderboards.

In real daily use, no one would actually use Gemini for serious coding work, right? I once assigned it a project migration task, and all it did was offer empty emotional support, spout tons of pleasant empty words, and never actually finished any of the assigned work.

Even Google's most proud strength — its world knowledge capability — swings wildly between impressive and completely useless.

Its strengths lie in its broad knowledge coverage: DeepSeek even explicitly admitted in its research paper that it lags behind Gemini in the breadth of world knowledge.

DeepSeek-V4-Pro-Max significantly outperforms all open-source models on SimpleQA-Verified, but still falls behind the leading closed-source model Gemini-3.1-Pro.

When you ask about obscure, niche historical facts, Gemini is often the only model that can recall those details correctly.

However, as soon as you talk about new things that emerged in the past two years, Gemini's responses become extremely weird and nonsensical.

For example, when we were writing this article, we asked Gemini to help outline the structure.

To our surprise, with zero related context provided, Gemini randomly spit out a reference to Claude 3.5 out of nowhere.

That's right, in Gemini's default knowledge base, its information about Claude only goes up to the 3.5 version.

This is because all current Gemini models, whether it's the 3.5 Flash released in May this year, or Gemini 2.5 Pro released in June last year,

The built-in knowledge base of all these models only contains information up to January 2025.

That's why it has no idea what Claude 3.7, released in February 2025, is — it only knows Claude 3.5 released in 2024, so it confidently makes up completely incorrect facts.

This means that when you chat with Gemini, if the topic covers events that happened in the past year and a half, Gemini will completely lose its mind and start spitting out garbage outputs.

If you want it to access the latest information, you have to turn on web search. But very often, after enabling web search, the model's performance drops because it ingests too much conflicting data, leading to contradictory, nonsensical results that go nowhere.

Meanwhile, the other two members of the "Big Three" — GPT and Claude — updated their knowledge bases to the end of last year long ago. Only Google was still clinging to its outdated old knowledge base, until the recently released Gemini 3.6 Flash finally updated the company's knowledge base to March 2026.

It's almost unbelievable that Google, a company that started out as a search engine, is now getting outperformed in the information retrieval field it used to dominate.

So what on earth went wrong with Gemini?

After digging through reports from multiple authoritative media outlets, we found that there is no new secret under the sun: Gemini's recent underperformance is actually a direct reflection of the chaotic organizational structure of Google's entire AI division.

Unlike pure large language model companies such as OpenAI and Anthropic, for Google, Gemini is far more than just a simple chatbot.

It is both a cutting-edge model developed by DeepMind, a weapon for Google Search to defend its search market share, an API product sold to enterprise customers by Google Cloud, and a feature that needs to be integrated into Android, Chrome, Workspace, Pixel, Google Home...

Every single business line wants AI to become its own next-generation user entry point, and every team believes their own use cases are the most important.

So who exactly is the top priority here?

The Information previously reported that colleagues from Google AI Labs once developed an AI note-taking app called NotebookLM. As soon as the product launched, it immediately made internal teams working on Gmail and Google Docs extremely unhappy.

Who gave you permission to do this? Do you have no respect for internal workplace protocols? You're directly encroaching on our core business!

There were even unconfirmed rumors that the Workspace team once considered shutting down NotebookLM entirely.

On top of that, the Google Cloud team and the DeepMind team have also had multiple internal conflicts over resource allocation and product positioning.