HomeArticle

The "mute model" Jev has gone viral, what changes will it bring to the AI industry?

深流研究所2026-09-30 13:52
When each model performs its own dedicated function, an Agent entry that can invoke multiple models according to different tasks will become far more important.

The Jev model has been a viral hit for over a week. Discussions around it, ranging from its various creative use cases to its underlying paradigm, are still ongoing.

Jev does not chat with people, nor can it help people write articles or code. It only does one thing: when developers give it a set of options and the current state, it is responsible for giving a judgment.

Diogo Almeida, the creator of Jev, used to be a researcher at OpenAI and participated in research related to InstructGPT, ChatGPT and GPT-4. He used to teach large language models to learn to chat, but now he has built a model that refuses to chat.

Why did Almeida change his direction? What anxieties in the current large model industry does Jev address? Will narrowing capabilities to the extreme become a future trend in the model industry?

Some tasks do not need to involve the most powerful model

The rise of Jev first taps into one of the most practical problems in the Agent era: model invocation needs to be faster and cheaper.

The founder of the open source project Browser Use even built a browser Agent with Jev to check a flight ticket from Zurich to London on the real Google Flights page. The whole process took 7.1 seconds, the single cost was $0.0039, and the median latency of Jev in the loop was about 178 milliseconds.

The price of Jev is also low enough. At present, the input price of Jev on Vercel AI Gateway is about $0.04 per million tokens with free output, so developers can call it without any burden.

Slowness and high cost have become common anxieties in the industry, which is related to the spread of Agent from code scenarios to a wider range of office scenarios.

OpenAI's data shows that from February to June this year, the number of weekly active enterprise Codex users in the legal, sales and recruitment, and marketing fields increased to 108 times, 41 times and 26 times of the original respectively.

When models become infrastructure that continuously executes tasks, the invocation frequency and Token consumption will also increase accordingly. While model manufacturers are pushing up the upper limit of flagship model capabilities, they also need to figure out how to reduce the cost and latency of model invocation in the face of new demands for large-scale Agent invocations.

Therefore, a differentiation trend has emerged at the model layer. Nowadays, the same manufacturer will launch models of different tiers and different prices, leading to a phenomenon where flash models are emerging in large numbers.

Google launched Gemini 3.5 Flash in May, targeting daily development and work scenarios. On July 31, DeepSeek released V4 Flash. In late August, Qwen3.8-Flash-Next and GLM-5.3-Flash were unveiled one after another.

These models do not necessarily pursue the highest performance on all benchmarks, but they are not "castrated versions" that cut off capabilities. They are new tiers specially optimized for high-frequency and simple tasks, with more cost-effective prices.

Take DeepSeek as an example. It recently updated V4 Flash to V4.1 Flash, and announced that before the launch of the next-generation Pro model, API requests for V4 Pro will be routed to V4.1 Flash and billed at the Flash price.

In the past, when users encountered tasks, their first reaction was to look for the most powerful general-purpose model, the more capable the better. The main line of manufacturers was also to make a flagship model as versatile as possible.

But now, model manufacturers are starting to re-segment capabilities according to task complexity, invocation frequency, response speed, cost and other factors, launching different tiers to let different models take charge of different work.

This differentiation is still advancing. From the perspective of usage scenarios, there have long been specially optimized models for code, image, speech and other directions. Scenarios such as office work and games are also gradually seeing models that better fit their own characteristics.

After all, the task structure, context, speed requirements and cost requirements of different scenarios are different, so models have room to make trade-offs for specific links.

Jev is an extreme sample in the differentiation trend

While other models are still adjusting capabilities and prices within the general-purpose framework, Jev simply separates a type of capability from the general-purpose model and makes it an independent model.

Developers give Jev a current state and a set of preset questions, and it returns Choice, Score or Boolean with corresponding probabilities. Jev does not explain why it makes such a choice, and the program can directly proceed to the next step after getting the result.

This is very important for Agents. In an Agent, the output of the model often directly determines the next action. At this time, what the system needs may not be a complete natural language answer, but just a clear signal.

Behind this design is Almeida's reflection on RLHF, the training method of large models. He believes that human feedback will push the model to generate answers that are more in line with human preferences, but it also easily makes the model develop another habit: that is, when facing problems that it is not very sure about, it still tends to give fluent, complete and definite answers.

Jev tries to make the model write uncertainty into its output results. If the model judges that there is an 80% chance that a thing is correct, then in long-term statistics, its judgment results should be close to 80% correct. After the system gets this probability, it can set its own threshold: automatically pass tasks with high confidence, transfer tasks with insufficient confidence to humans, or to a more powerful model.

However, enabling the model to achieve reliable automated decision-making is the direction Almeida wants to reach, and Jev has not achieved that yet.

Some developer tests have found that for the same judgment, if you ask it once in the positive way and then once in the reverse way, the sum of the two sets of probabilities given by Jev sometimes exceeds 1. A properly calibrated probability system should not behave like this. Jev does not generate natural language, which avoids the risk of fabricating a decent answer, but the probabilities it gives can still be wrong.

The Agent entry that invokes different models becomes more important

Even though Jev still has many problems, it has shown the industry another possibility: The intelligence in an Agent can also be split, and a complex task does not have to be handled by one general large model.

For example, when complex reasoning is required, a powerful model can be invoked, while a large number of simple judgments and screenings can be handed over to models like Jev that are faster and cheaper.

There is also a hidden meaning in the name Jev. There is a concept called "Jevons Paradox" in economics, which means that after a technology becomes more efficient and cheaper, its usage will increase instead. Then, when Jev makes the action of "judgment" faster and cheaper, Agents are more likely to call it frequently when executing tasks, thus optimizing the efficiency and cost of the overall task.

In other words, tasks themselves will gradually become the core of Agents. Invoking different models according to tasks may make Agents more efficient and cost-effective in execution.

The relationship between models thus has another possibility. In the past, models competed to be more powerful, and when one model's capabilities improved, it would replace the previous one. Now, with tasks as the core, different models can take charge of the parts they are better at and collaborate in the same Agent.

If this division of labor continues to develop, an Agent entry that can invoke multiple models according to tasks will become even more important.

There are already Agent products moving in this direction. Domestically, Tencent WorkBuddy has supported switching between different mainstream models. WorkBuddy also has an Auto mode, where the system will automatically select the most suitable model for response according to the type, complexity and context characteristics of the user's task. When users do not know which model to choose, or face mixed tasks, they can let the system decide.

Cursor overseas is also doing similar things. The Cursor Router it launched in August this year will first judge according to task type, complexity, context and other information before each request is sent to the model, and then select the most suitable one from multiple models. The basic principle is that simple tasks are handed over to fast and cheap models, while difficult and long-cycle tasks are handed over to cutting-edge models.

The ideas of these products are consistent, that is, to let the system automatically complete the scheduling of models according to tasks. Users only need to put forward tasks, and the entry will arrange which model to use, so as to improve the efficiency of Agent usage and reduce costs.

In the long run, as models gradually move towards their respective roles, an Agent entry that can flexibly allocate them will become increasingly important.

This article is from the WeChat Official Account "Deep Flow Research Institute", the author is Xiao Ying, and it is released with authorization from 36Kr.