19 AIs fought a fierce battle in StarCraft, GPT-6 won all the matches, yet still failed to defeat novice human players.
Thinking too much will drag down even the smartest AI.
43 minutes, 11138 reasoning tokens, 6 batches of instructions — 0 combat units.
This is a desperate report card handed over by Grok 4.6 in StarCraft: Brood War.
In the game, Grok, which cranked up the reasoning to the highest xhigh gear, pondered hard from the start to the end of the game, and failed to produce even a single soldier that could go into battle.
On the other end of the same leaderboard, GPT-6 Astra also went all out, winning all 18 battles with a 100% win rate.
In the past two days, this StarCraft large model leaderboard called Brood War Bench has gone viral on Hacker News, where Silicon Valley programmers gather.
Top 10 on the Brood War Bench leaderboard. Astra ran at the highest xhigh reasoning mode and won all 18 matches, while Grok ranked at the bottom across all three gears.
Peter Steinberger, the father of OpenClaw who is now an OpenAI employee, did not forget to repost to promote his own model: the new leaderboard is here.
AI has been breaking century-old problems in the mathematics circle one after another, and some people exclaimed that the singularity is coming, with such a strong momentum that even Terence Tao publicly spoke out to slow down the pace.
But Ben Swerdlow, the author of this leaderboard, threw out a completely different judgment at the beginning:
Even the undefeated champion GPT-6 Astra cannot beat a human novice.
Someone in the Hacker News comment section pointed out the key point in one sentence: "They are not unable to play, they just need too much time. If you pause the game while the model is thinking, it can play."
Others directly mocked: "It's as if these models are not general-purpose intelligences."
Why would the smartest AI on the entire network be dragged to death by its own thinking in StarCraft?
Are we still far from AGI?
The Game Will Never Wait For You
The birth of this leaderboard is somewhat dramatic.
At first, the developer Ben only made a StarCraft version fully operated by agents, and invited several friends to play together.
These people hardly ever touched StarCraft before, but they played no worse than veteran players who have been playing for many years.
Their secret is ridiculously simple: just type a word "attack", and leave all the work of producing soldiers, assembling troops and launching battles to AI.
This made Ben can't help but wonder: if humans shut up completely and let these models play by themselves, how long can they survive?
Thus, 19 AI players from the three families of GPT, Claude and Grok all appeared on the stage, and a 171-round AI melee kicked off.
The competition runs on Freestyle's cloud virtual machine, with a separate machine for each round, starting at the same time. All game data and every step of AI's thinking records are saved.
When doing math problems or writing code, large models can be in a daze for 5 minutes before submitting the answers at one time.
But it doesn't work for StarCraft.
There is no turn-based system that AI is used to, only a non-stop real-time battlefield.
Every second you spend thinking, the workers (the worker units responsible for mining, called so by English-speaking players) on the opposite side are mining frantically, the barracks are producing troops rapidly, and the main force is already rushing towards the high ground of your base.
Older generation models still play this real-time strategy game with the turn-based routine, and as a result, their minds are still running, but their base has already been flattened.
Under Steinberger's post, a netizen pointed out the truth in a word:
It's interesting. No wonder AI always prefers to make turn-based games when developing games.
Grok Ponders for 43 Minutes
Not A Single Soldier Goes Into Battle
It is Grok that brings this turn-based thinking to the extreme.
In the match numbered G043, Grok can be called a giant in thought but a dwarf in action.
It frantically outputs large sections of strategic reasoning, but in the 43-minute game, it only issued 6 batches of instructions in total, taking an average of more than 7 minutes to take an action.
This is no accident.
In G003, Grok produced 3 marines, but they never walked to the enemy base; in G002, it produced 2 Zealots, who also failed to cross the map.
The soldiers were produced, but they just couldn't reach the opposite side.
At the highest reasoning gear, it had 2 wins and 15 losses, which was a terrible result.
The data curve is very intuitive: when the game reached 11 and a half minutes, Astra already had a strong army, while Grok had an average of less than 7 workers and a few soldiers on the field.
Interestingly, there is actually a huge amount of deposits in Grok's treasury that far exceeds that of Astra.
The money is in the pocket, the brain is running, but it just refuses to take action no matter what.
Astra's workers and Zealots demolished Grok's main base. Grok had no combat units in hand, and there were still 928 crystal minerals in its account.
On the cruel battlefield of StarCraft, thinking for a long time is not called careful consideration, but hanging up and waiting for death.
Astra's Winning Move
Make The Opponent Think To Death
If Grok is dragged to death by thinking too much, then Astra is specifically forcing others to think too much and dragging the opponent to death alive.
The most brilliant move of the Codex family that Astra belongs to is harassment.
It often only sends one of the most basic mining workers to cross the entire map and wander around the enemy's base. In the eyes of human players, this kind of worker with zero combat effectiveness can be driven away by casually pulling two engineers.
But in the AI vs AI match, this move is absolutely dimensionality reduction strike.
As soon as the opposite AI sees this worker, it immediately falls into deep thought, spending dozens of seconds frantically calculating "how should I deal with this worker".
During these precious dozens of seconds, it does nothing at all, and the base is completely shut down.
Astra's harassment force demolished Grok's main base. Grok had no combat units in hand, and there were still 928 crystal minerals in its account.
In Real-Time Strategy (RTS) games, harassing the opponent is to destroy the economy; when it comes to AI vs AI, the biggest lethality of harassment is to directly drag the opposite AI to crash.
Of course, Astra also has weaknesses.
Its interior is like a dysfunctional makeshift team, with the economy manager, the soldier producer, and the battle manager doing their own things.
As soon as the new recruits are produced, they are pulled to the front line to die in vain by the battle-managing AI, and they don't know what it means to assemble troops at all.
However, the author also found that as long as he personally acts as the commander-in-chief and connects several sub-agents, Astra will immediately know how to accumulate troops and choose the right time to fight.
The Codex family also has a tenacious spirit.
In one round, Astra's brother GPT-5.6 Terra's troops were completely wiped out, and the main base was also demolished.
It lifted off the only remaining Command Center (Terran building can take off), drifted all the way to the corner at the other end of the map, and stubbornly survived for another 6 minutes before being eliminated.
Fable Is The One That Plays StarCraft Most Like A Human Player
The one that made the author want to cheer for it in several rounds is actually the top student Fable in the game.
Its playing style is the most standard, it engages in economic development steadily, develops the tech tree down-to-earth, and rarely takes opportunistic tricks.
In one round, it produced Mutalisks; in another round, it built a whole row of advanced Protoss buildings one by one, even the Templar Archives required to produce High Templars was completed.
When the game reached 11 and a half minutes, Fable surpassed Astra in several data indicators including workers, troops, buildings and technologies.
But a good student may not necessarily get the first place.
Although Fable's win rate is as high as 83.3% ranking third, it still can't catch up with Astra.
In one round, it filled the table with all kinds of technologies, but before it could convert them into combat effectiveness, it was beaten to death by Claude Opus 5 with random punches.
Fable was full of barracks, factories and academies, and there were more than 2800 crystal minerals in its account, but it was flattened all the way by the marines of its own Opus 5.
The gap in rhythm is also very large: the median duration of Astra's matches is only 8 minutes and 10 seconds, while for Fable, who loves to develop slowly, this number is dragged to 10 minutes and 37 seconds.
Three large models, three different souls.
Codex is like a street rogue full of bad tricks, Grok is like a stubborn top student who sticks to the first question in the examination room, and Fable is that honest person who follows the strategy guide word by word.
The Longer You Think
Does Not Necessarily Mean The Better You Fight
Since thinking too much will lead to failure, is it better to keep your mind as empty as possible?
Not exactly.
For GPT-5.6 Sol, when the reasoning is turned to the highest gear, the win rate is at the bottom, which is even worse than that of the medium gear.
But for Astra, the deeper it thinks, the higher the chance of winning, and the highest gear directly achieves a 100% win rate.
The more interesting part is the cost bill.
When Astra is running at the highest reasoning gear, it only operates 12.6 times per minute, and spends an average of only 10.54 US dollars per round; when switched to the lowest gear, the operation frequency doubles to 25.7 times per minute, but it costs 21.07 US dollars per round instead.
At the highest gear, it makes half as many moves, spends half as much money, and wins the most. This shows that the new generation of models has obviously evolved.
They understand that thinking has a cost, and they also know when to stop thinking and get to work.
AI Beat Professional Players 7 Years Ago
Why Can't It Beat Novices Today?
Some people will definitely ask: 7 years ago, didn't DeepMind's AlphaStar already beat professional players?
That's right. In early 2019, DeepMind announced the record of AlphaStar: two 5-0 sweeps, crushing professional players TLO and MaNa.
In 2019, the second match between AlphaStar and professional player MaNa. The lower part shows AI's decision-making process: reading the situation, judging where to attack, what troops to produce, and predicting the win or loss at the same time.
In 10 matches, professional players didn't win a single round.
Later, DeepMind added the same restrictions as humans to it: it can only view the local map through the screen lens, and the number of operations it can perform per second is also reduced.
Even so, it anonymously mixed into the European ladder, all the way to the Grandmaster rank, surpassing 99.8% of human players.
On one side it ranks top 0.2% on the ladder, on the other side it can't even beat a novice who knows the Photon Cannon Rush tactic.
Has AI's technology regressed?
In