HomeArticle

Is the entire internet getting hyped up again after Google releases its new model?

王智远2026-09-03 13:36
Testing a large model on the very day of its release is nothing but fooling people.

Google has released a new model, Gemini 3.8 Flash.

When I woke up in the morning, my subscription feed was already flooded with real-world tests of it. There are all kinds of conclusions: some say it's absurdly powerful, some say it's nothing special, and others are just being sarcastic.

I would have clicked through to read them a few months ago, but these past two months I just scroll past; content creators are actually working really hard, but this thing can't be measured structurally at all.

I've had an unresolved question for a long time: for a trillion-parameter model released for only 48 hours, on what grounds can a single person judge whether it's good or bad? To figure that out, you have to cross three hurdles.

The first hurdle: it doesn't leave you any time.

This time, Google only waited 20 days after the previous generation. Counting backwards, 3.6 was released on July 21, 3.7 on August 13, and 3.8 on September 2. Three generations in 43 days, the third model in six weeks.

Sundar Pichai promised shareholders on the earnings call that the pace would be fast enough to roll out a new model almost every month. The term "product roadmap" is no longer sufficient to describe this; it's basically a shift schedule.

You haven't even gotten familiar with the last version, and the arguments over its evaluation results haven't ended, before the next one is already here. What's more, Google itself admitted that this 3.8 was further trained on the basis of 3.7, with no changes to its specifications at all.

It's equivalent to the same student taking extra classes for another semester, without replacing the student.

The progress is real. It scored 73.7% on the DeepSWE programming leaderboard, while Claude Opus 5 scored 74.0%, a gap of only 0.3 percentage points, almost a tie.

With this report card on the table, anyone who says Google is standing still is being unreasonable.

But that's exactly where the problem lies. How long does it take to validate a model? At the very least, real users need to use it for real tasks for weeks or months, but Google only gives you 20 days.

The validation cycle can't keep up with the iteration cycle. You just start to figure out some patterns, and it's already obsolete. Those real-test articles for 3.7 from three weeks ago now read like archaeological records.

The second hurdle: you can't fully grasp all its capabilities.

With trillion-level parameters, it's so large that you can't measure it at all with your naked eyes or manual operations. Google hasn't specified exactly how large 3.8 is. Anyway, cutting-edge models are always at this order of magnitude, and no one can clearly state where the boundary of its capability space lies.

I once saw a foreign engineer who builds models on GitHub say a hard truth: no one knows how good a new model is when it's released, even the lab that built it is just guessing.

But when you actually get your hands on it, how much of it can you really touch? A few conversations, asking it to write a low-code website, make a totally unrelated casual game, and then ask it a few brain teasers.

These are just a few random points poked in that huge capability space, with a sampling volume close to zero.

What's more troublesome is that the test results are still unstable. Content creator AI Pulse Daily used the same prompt to make three models generate a New York city scene: Opus 5 placed 1840 trees, the version of Google's model from three weeks ago placed 10 trees, and the newly released 3.8 only placed 6 trees.

For the crosswalks and water ripples that were explicitly requested in the prompt, 3.8 completely skipped them, but instead added five camera presets on its own initiative.

Another group of people tested it with similar tasks, and 3.8 was fast and cheap: it finished four scenarios for only $0.12, while Opus 5 cost $1.86.

It's the exact same model, but the conclusions are completely opposite, and the difference all lies in the test questions chosen by the testers.

The third hurdle: what you get to touch is not the full picture of it.

The demo at Google's launch event used one single prompt to generate a playable 3D wizard game, with puzzles and environmental narratives fully built, and then made a runnable DOS version of Google Maps with even street view navigation. It does look really impressive.

But these are the upper limits achieved by Google after repeated fine-tuning using its own framework and its own environment.

This morning I also saw an article that tested it in the opposite way, ran it myself, and the conclusion was exactly the opposite of the official one. With three generations updated in six weeks, there are four words in the article: it remains as mediocre as ever.

What do ordinary people get?

A web page, an API, a bare chat window, without that set of frameworks, no engineers waiting to support you, and what you send is an out-of-context natural language query.

What Google shows you is the effect fed and polished by the whole team. The window you have in your hand has no such support. What you can test on launch day is only the part they deliberately show off.

Look at these three hurdles: not enough time, not enough capability to measure, and what you touch is not even the full model. What's the point of all this?

......

It's useless, but the output of content never stops for a single day, because releasing the model and releasing this set of content is never meant to be a validation activity in the first place.

I specifically dug into this detail, and found that the opening of this carnival started several days earlier than the media articles.

Wall Street leaked the news the day before the official release.

There is an internal coding platform at Google called Jetski, where engineers did head-to-head tests between 3.8 and Anthropic's Opus, and most people preferred Google's new model.

There were no prompts, no task sets, no test methods in the news, just a single conclusion. But just because of that one sentence, Alphabet's after-hours stock price rose by 0.6% on the day of the report.

Flip back a few days earlier, Business Insider broke the story first.

Google employees tested the preview version of 3.8 on Jetski, some said it was significantly better than 3.7, and added a line at the end that a comprehensive assessment is still too early. No repost later kept that sentence.

One day earlier, the 3.8 model briefly appeared on the official website before being taken down, and the page turned into 404. Fortunately, someone took a screenshot to keep evidence.

Flip further back, in a Tencent paper on August 10, "Gemini 3.8 Flash" was already listed in the evaluation referee list, when 3.7 had not even been officially released to the public.

Look, before the official release even starts, the tests are already running on the dissemination pipeline.

Manufacturers know very well that holding a launch event alone is not enough. They need the entire industry to run real tests for them, the more articles the better, praise or criticism is fine, as long as the timing is right. The "real test" you are looking forward to is just a dish they served on the table.

The articles are also perfectly aligned in timing.

Right after midnight, the first real-test article went online. By 8:30 in the morning, all five top tech media had published their articles. The titles all followed the same pattern: "just released", "lightning fast", "another new model", "unconventional iteration", "being ridiculed".

They are all fighting for that "just" timing: the "just" at midnight and the "just" at noon are not the same thing. The same batch of tests by netizens were reposted over and over in a few hours, with titles changed from "cost-performance killer" to "being widely ridiculed", while the main content is almost identical.

Why are they so perfectly aligned?

The traffic window on launch day only lasts 48 hours. Once the window closes, no one will read those articles anymore. People who make a living from this are earning their income from this narrow window.

I also checked if there are any sober people out there. Yes, an AI practitioner in the English-speaking circle publicly said that she refuses to evaluate a model that has been online for only 15 minutes, and will wait a week for the model to stabilize first.

But these people don't make a living from launch day traffic, so their voices can't squeeze into the launch day headlines.

Okay, we've gone through the media part. Dig deeper, even professional institutions are also present.

On the very day of release, Artificial Analysis published its score: 59. Datacurve's programming leaderboard also released results the same day: 74%, tied for first place with Opus 5.

These two institutions are serious about evaluation, their task sets are public, results are reproducible, and few people question them. But the more serious they are, the more it illustrates the problem: even they rush to submit results on launch day.

In fact, these institutions release new scores every day. The day before, Anthropic's Fable 5.1 just got the highest score of 66 in the history of Artificial Analysis, but no one saw it go viral.

Only the scores on launch day get reposted. The release schedule dominates everyone, even the most serious evaluations can only get an audience if they are published within that window.

There's something even funnier. Arena's own note on its leaderboard says that the ranking reflects users' subjective preferences, not an absolute capability test. The platform itself has already laid its cards on the table.

Dig to the last layer: the readers.

Mutually exclusive conclusions can make headlines at the same time, because there are audiences for both sides. Overwhelming praise and overwhelming criticism are both ready-made materials for people to pick sides. As for whether the model itself is good or not, no one cares. What people care about is that the side they bet on wins.

I checked the comment sections under those real-test articles on Xiaohongshu, and the people arguing the most fiercely are all those who haven't decided which model to use yet. After winning the argument, they feel as if they are already using the winning model.

Reading these real tests is a bit like reading sports commentaries. There's nothing wrong with just having fun. The only difference is pretense: sports commentaries clearly state that they are just watching for fun, while these "real tests" are labeled as such, making you think you are making a rational judgment. After the fun is over and you've picked your side, no one remembers to ask whether the model is actually good or not.

It's a whole show: manufacturers write the script, media play supporting roles, institutions rush to meet deadlines, readers pick sides. Everyone has their own role, and after the whole performance ends, no one actually gets a real validation. Is that abstract?

......

It's not abstract. Real validations have never been absent, they just never happen on launch day. They are either held behind closed doors, or come late, and refuse to show up on the carnival day.

If you follow the real validations, they fall into two categories: the first category happens before the official release.

OpenAI gave GPT-5.6 Sol to the US government for preview in advance. Agencies under the Ministry of Commerce have been able to get unreleased models from Google, Microsoft and xAI for pre-deployment assessments since May.

OpenAI released a statement in June itself, saying that showing the model to the government in advance is to make the public launch smoother, it's a short-term arrangement: pass the government's check first, then pass the market's check.

For this time, Google's Cyber version is only open to audited security teams. OpenAI's Astra even had a restricted release because its cybersecurity capabilities crossed the line.

Elon Musk also proposed that peers should audit each other, with the government providing final backing.

Behind every launch day, there is a round of unseen exams. These exams are real and worth doing, but they are not public, and their conclusions won't make judgments for you. They are only responsible for checking whether the model is safe enough.

The second category happens after the official release.

On the day GPT-5 went online, the whole network was full of negative reviews, and the consensus was that this model was bad. Three months later, the public opinion changed completely, and it became a very capable tool for agent programming. Many people later realized that they were too hasty in criticizing it back then.

Those who criticized GPT-5 back then were not being unreasonable.

People judge models based on their feelings, but once the model is smarter than you, your feelings will fail. You can't even see where its strengths lie, so naturally you can't get to the point. That's the root cause of why the launch day consensus always turns out wrong.

I myself have fallen into this pit:

When DeepSeek's V4 Pro was released, the whole screen was full of praise. After I used it for a while, the backlash came, and people in the circle started to say it was bad.

But the model I use the most in daily life is actually V4 Flash. Its agent scheduling is more efficient than Pro, it costs fewer tokens, and its hit rate for running loop tasks can basically reach 98%, which even Pro sometimes can't achieve.

The Pro that was praised to the sky on launch day and the Flash that is underestimated in daily use have completely reversed roles.

It takes three months for people to reverse their previous wrong judgments, but a new model release only takes 20 days. By the time you draw a real conclusion about one model, the next carnival has already started, and the whole public conversation is reset.

People in the circle have summarized this rhythm: it's a machine specially built to create a sense of breakthrough, which arrives on time every few weeks. Whether a real breakthrough has actually happened is another question.

Can community testing find out real problems? Yes. But that requires running real tasks continuously for many days, which is not at the same level as the 48-hour concentrated bombardment of tests on launch day.

Back to the original question: what exactly are those real tests all over the launch day measuring?

What is truly produced on launch day is a report of human attention. Which article gets clicks, how long do people stay after clicking, what the comment section looks like, and which group the article is reposted to: these are the things that are accurately measured that day.

This measurement system is extremely precise. After an article is published, the open rate, the finish reading rate, and the position where people scroll past are all data. What kind of title can make people click in? What kind of conclusion can make people argue? Manufacturers know all of this very well.

After the measurement, everyone takes what they need: the heat curve, traffic bills, materials for picking sides, all have their own destinations. The model is just an excuse. While you are choosing the model, they are also studying and choosing you.

Instead of humans testing the model, it's more like the model is testing humans.

This set of playbooks was already performed in the smartphone industry long ago. Back then, smartphone evaluation also went down this path: review units for media, exclusive first articles, material support from manufacturers, the whole process ran perfectly.

By June this year, CCTV exposed this practice: manufacturers gave content creators special "cheating phones" that were faked at three layers from hardware to the cloud. The phones the creators tested were not the same as the ones ordinary consumers bought.

But at least smartphones have frame rates that can be measured, which is why faking is possible. AI models don't even have a clear definition of what to measure, so they don't even need to bother faking.

What form will the liquidation of this carnival take? When will it come? No one knows. What's certain is that time will give the answer to whether a model is good or not, but that answer will never make it to the front page on launch day.

Next time someone recommends a model to you using launch day test results, just reply to them:

Can your carbon-based brain really tell whether a trillion-parameter large model is good to use or not?

Data Sources:

[1]. Official Google Gemini 3.8 Flash model card and DeepMind published blog, Wall Street Journal and Business Insider reports on Google