HomeArticle

Google and Meta have been confirmed to be involved in ranking fraud.

新智元2026-09-12 16:17
Gemini plunged 70 points, tumbling directly from the 2nd place to the third from the bottom.

Benchmarking has been a little off lately.

On September 8, analysis firm SemiAnalysis posted five consecutive tweets, directly naming two trillion-dollar giants, Google and Meta —

Gemini 3.8 Flash and Muse Spark 1.3 are "the two models with the most obvious traces of benchmark gaming".

The evidence is very clear.

On the Terminal-Bench 2.1 leaderboard that tests the actual working capability of Agents, Gemini 3.8 Flash scored 89.4, ranking 2nd out of 182 models, even surpassing GPT-6 Astra.

However, on Terminal-Bench 4.0 which was just launched on August 29, the same model only got 19.1 points, falling directly to 12th place among 14 test subjects, with its performance plummeting to 20% of its previous score.

At the same time, GPT-6 Astra dropped from 88.4 to 57.7, which still retains 65% of its original score.

Just 27 minutes later, Alexandr Wang, Chief AI Officer of Meta, personally appeared in the comment section, stating straightforwardly: This is a stupid argument.

He believes that GPT-5.6 Sol scored 88.8 on version 2.1 and dropped to 37.3 on version 4.0, which is a larger gap, but no one accuses Sol of benchmark gaming.

At the same time, he also emphasized that Meta has never claimed that Muse Spark 1.3 is as powerful as Astra or Fable 5.1, but its cost-effectiveness is obviously much higher.

Swap the test paper, and the score drops sharply to a fraction of the original

Briefly speaking, Terminal-Bench measures the actual working capability of Agents.

The test method is to assign a terminal and a vague goal to the model, and let it handle all the rest independently. The system needs to plan the execution path, call tools, write scripts, and debug by itself when errors occur.

On version 2.1 that has been gamed by the industry for more than half a year, the scores of these two models are indeed very impressive.

Gemini 3.8 Flash scored 89.4 and ranked 2nd, Muse Spark 1.3 scored 88.8 and ranked 4th, followed by GPT-6 Astra with 88.4 points and Claude Fable 5.1 with 85.02 points.

Surprisingly, a low-cost Flash-tier model and Meta's new product focused on cost-effectiveness both outperformed the flagship models from two top AI labs by a large margin.

Then version 4.0 was released.

The new test set contains 66 questions in total. It not only removes 8 old questions that have been "completely gamed", but also overhauls 19 questions, and adds an 8-hour timeout limit to all questions.

The anti-cheating mechanism is also set to a very high intensity.

The organizer not only formulated 35 scoring rules, but also added a special "adversarial cheating test", deliberately letting Agents try to exploit loopholes in the reward mechanism. Any question whose loophole can be exploited will be directly removed from the valid test set.

As soon as the new leaderboard was released, the situation changed drastically.

Claude Mythos 5.1 (max) scored 60.9 to stay ahead, GPT-6 Astra and Claude Fable 5.1 followed closely with 57.7 and 55.8 points respectively, followed by Claude Opus 5 with 52.3 points. The previous generation GPT-5.6 Sol and Terra scored 37.3 and 23.6 points respectively.

In contrast, the score of Gemini 3.8 Flash plummeted directly to 19.1. According to the chart released by SemiAnalysis, the score of Muse Spark 1.3 is 33.3.

To sum it up:

On the old leaderboard, Gemini 3.8 Flash could even outperform GPT-6 Astra; on the new leaderboard, its score is less than one third of the latter, and it is even surpassed by the previous generation GPT-5.6 Terra.

There is no need to copy answers, just buy a set of "secret test papers"

Despite the dispute, the second tweet from SemiAnalysis pointed out a real insider secret of the industry.

In their view, the tasks of Terminal Bench 2.1 are completely public. Meta and Google are certainly not foolish enough to directly use the original questions for training, but they will definitely purchase data, specifically training data that is designed to be infinitely close to the question type of TB 2.1. The final effect is no different from cheating.

And this is the advanced way of benchmark gaming in 2026.

The old "data contamination" method was very crude, which directly mixed the original questions into the pre-training corpus to let the model memorize them by rote.

This kind of behavior is very easy to detect. Once the data is cleaned, the inflated scores will be exposed immediately.

The industry has suffered a lot from this. For example, nearly 40% of the samples in HumanEval were contaminated, and GSM8K's score dropped by 13 points after decontamination. The most extreme case is SWE-bench: after re-testing with a private code base, the score was directly cut in half.

Therefore, major companies no longer use original questions now. Instead, they buy "mock test papers that look like real exam questions", which are closed task systems that provide models with opportunities for trial and error and corresponding rewards.

At present, this has become a mature business with clear price tags.

A single training task is priced at 200 to 2000 US dollars, and complex software engineering tasks can even be priced as high as 20,000 US dollars.

If you want to clone a website to build a UI training ground, it costs about 20,000 US dollars.

If you require a product of the same scale as Slack to be reproduced, the price starts at 300,000 US dollars.

For customers who require exclusive buyout, the price can be 4 to 5 times higher.

According to the survey by Epoch AI, Anthropic has internally discussed investing 1 billion US dollars in this business every year. Single-quarter contracts often reach six or seven figures, and some researchers revealed that the average price is between 300,000 and 500,000 US dollars per quarter.

There are at least 35 companies on the supply side engaged in this business, most of which are start-up teams with less than 20 employees. Established data giants are also all transforming: Scale AI's revenue exceeded 1.4 billion US dollars before Meta took a stake in it, and Surge's ARR is also approaching 1 billion US dollars.

The current situation is very surreal.

On the one hand, there is a completely public leaderboard question bank, and on the other hand, there is a mature industrial chain of selling test questions. If you want the model to get a high score on a certain leaderboard, you really don't need to go through the trouble of cheating, just pay and place an order.

Acting as both the test question seller and the invigilator

Subsequently, SemiAnalysis directly named a company called Datacurve.

On the DeepSWE 1.1 leaderboard that specifically tests long-cycle programming capabilities, Muse Spark 1.3 (max) took the first place with a score of 75.4, GPT-6 Astra and Claude Opus 5 followed closely, and Gemini 3.8 Flash ranked 4th with a score of 73.8.

The two models accused of benchmark gaming took the first and fourth places respectively.

Presumably, you have already spotted the trick behind this.

This leaderboard is run by Datacurve, and Datacurve's profitable business happens to be selling "expert-level coding data and reinforcement learning environments" to cutting-edge labs.

DeepSWE is not only Datacurve's own benchmark, but the benchmarking process also uses the evaluation environment operated by Datacurve itself.

Acting as both the tutoring agency that sells test questions and the invigilator that designs the test papers, this situation is completely confirmed.

The public leaderboards are about to expire

Finally, SemiAnalysis naturally delivered a verdict on the entire industry's public benchmark evaluation.

They believe that this is ultimately the fate of all high-quality public benchmarks. TB 4.0 is no exception. The reason why it still seems accurate now is only because it has just been released for two weeks.

As long as its questions are still public, it will not take long for all the labs that claim to be "cutting-edge" to game it thoroughly.

The ultimate solution proposed by SemiAnalysis is very straightforward: to develop more high-quality private benchmarks.

The direction is correct, but the cost is also very high.

Once all evaluations are privatized, public leaderboards will completely become marketing press releases for each company.

The real, discriminative performance data will only circulate privately between labs and private evaluation institutions.

Major tech companies are certainly happy with this, after all, no one can memorize the test questions in advance for a closed-book exam.

But for the vast number of developers, the public ruler that can fairly measure the performance of all models will probably never be found again.

References:

https://x.com/SemiAnalysis_/status/2097112791471522292

Edited by: Moses

This article is from the WeChat official account "AI Era", Author: ASI Revelation, published with authorization from 36Kr.