Is Jev, who has been wildly hyped up across the entire internet, really that amazing?
Nearly 40 million people tuned in, with over 70,000 likes. How could an AI that can't even speak stir up the entire AI community?
Fellow enthusiasts who follow the AI circle must know that a model named Jev has gone viral recently, and its official announcement tweet has gained nearly 40 million views.
Its founder is no ordinary figure — he is a former OpenAI researcher, who can be regarded as a key member of the team that developed ChatGPT.
What makes Jev special is that it is a "non-speaking model". The official named it the System One model, an AI model that is "only responsible for making judgments, not for generating text or speaking".
If other large models are designed to mimic human thinking, Jev can be said to mimic human intuition, which will make the choice it deems most correct in the shortest possible time.
The official also stated that a single call to Jev only takes 70 to 500 milliseconds, and its price is incredibly low: it only costs $0.042 per million input tokens, and outputs are completely free of charge, so cheap that it's almost immeasurable!
But as soon as this model was released, netizens were completely divided in their debates.
Some people believe Jev represents the future. It is so fast that many people have started to use it to build some very creative applications, for example, connecting it to WeChat to analyze the intentions of their supervisors.
Some users used it to complete a speedrun of Minecraft in 8 minutes and 43 seconds, and it can also play games like Subway Surfers and Super Mario.
Some netizens even said they built a real-time trading bot with Jev, only to lose 31,680 dollars.
But at this point, another group of netizens raised doubts: From what you described, isn't this something the old-generation classifiers can also do? This is nothing more than pure marketing hype.
So, is an AI that can only answer multiple-choice questions really that powerful?
I also set up three games to test Jev's performance. To cut to the chase: It can play the games, but it really doesn't perform that well...
Didn't the user above say he could do a speedrun? So I used his open-source project, set up a Minecraft server, let Jev join as a player, and watched the whole process.
However, the project note said we can't use Jev alone, we need to equip it with a smarter model as a strategist, so I paired it with GLM 5.3.
After several attempts, it actually cleared the game: it went to the Nether, entered the End, and finally blew up the Ender Dragon with beds, finishing in 12 minutes and 39 seconds. That looks pretty impressive, right?
But after checking the source code, the actual speedrun is far from what we expected. The author stated that the code has already marked the mandatory paths on the map in advance, and a fixed module is used for the dragon fight...
Most importantly, among Jev's 274 decisions, 250 of them are questions with only one single option. This is a completely pre-scripted game. After seeing all this, I just feel like I was made a fool of.
No wonder I saw it swimming in lava for a while, and spinning in place for another. Its performance is just total garbage.
Then I decided to write my own test case, starting with something simple. For example, the card game Dou Dizhu is essentially a game of selecting cards, right? So I made a 1v1 version of Dou Dizhu to play against it. Don't ask me why it's 1v1, I was really worried its brain couldn't handle more complex scenarios.
Then I started a round with Jev, and to my embarrassment I lost. At the final stage of the game, I played a pair of 9s, and it immediately threw out a Royal Bomb to beat me instantly. That's my bad for being too unskilled.
But what's interesting is that I let it replay the same final stage several times. Out of six attempts, it chose the Royal Bomb four times, and switched to a pair of Ks twice.
This is probably because in this specific final scenario, the winning probability of the Royal Bomb and the pair of Ks are very close, so tiny differences in each calculation may push the choice to the other side.
From this perspective, it makes different choices for the same question every time, which is quite similar to humans who often go back on their decisions.
Since they are all card games, why not let it play Slay the Spire? This game is also full of multiple-choice questions from start to finish. So I let Jev fight on its own, without any auxiliary tools at all.
It played ten rounds in a row and failed to clear the game. Its best run reached the 19th floor, and on average it died around the 11th floor, which is actually not that bad.
One of the rounds was very interesting: it had 16 HP left, 3 energy points, three defense cards in hand, and the monster was about to deal 18 damage to it.
It actually chose to end the turn, and there was no follow-up — it died immediately.
I checked the logs and found out that the three defense cards were listed as three separate options. Jev scored them 17 points, 20 points and 18 points respectively, while the "End Turn" option got 21 points, so it chose to play no cards at all.
This shows that Jev's choices can easily be influenced by the environment designed by humans.
This is determined by how Jev works: when it answers multiple-choice questions, it does not directly output a single answer, but outputs a probability distribution for all options.
When our program listed the three defense cards as three separate options, its simple "brain" couldn't adjust flexibly. As a result, the probability that should have been concentrated on "taking defensive actions" was split into three parts, which was lower than the probability of "ending the turn", leading directly to its death.
Which actions count as one single option, and whether options should be merged, are all determined by the program we designed, which is called the harness.
So Jev's performance depends half on the model itself, and the other half on the rigorous design of developers. It seems that people who use Jev have to take responsibility for their own mistakes.
But among the more than 1,000 operations just now, half of the decisions were made within 400 milliseconds. Isn't that incredibly fast?
After completing these three experiments, my impression of Jev is very clear now.
Among all the large models that are competing to improve thinking ability and expand parameter scale, Jev is indeed very unique: it does not generate text, runs extremely fast, and only makes choices.
But it is really not as magical as many netizens praised. It cannot do