HomeArticle

After Jev, the Chinese team began to dig deep into the "intuition layer" of AI.

极客公园2026-10-04 18:58
Large models are responsible for thinking, small models are responsible for judgment, and Chinese teams will not be absent from this emerging business.

On September 15, TypeSafe AI, founded by former OpenAI researcher Diogo Almeida, released Jev. This model does not generate a single word of text, and is only designed for judgment tasks.

Two weeks later, the entire industry is doing the same thing. OpenAI launched the Decisions API at its Developer Day, Cloudflare open-sourced Clef on October 1, and Amazon also released Strands Decider.

China's tech circle is also not lagging behind. Shanghai AI Laboratory open-sourced the multimodal decision model Intern-Decision.

On September 30, StartLux, a Shanghai-based company founded less than five months ago, released the open-source StartLux-Decision, claiming that it outperforms Jev in public evaluations.

A new category has evolved from "first launch" to "full competition" within two weeks, a speed rarely seen even in the large model industry.

The leader behind StartLux is Chen Danian, co-founder of Shanda Interactive Entertainment.

The "intuition" ignited by Jev

To understand this boom, we first need to clarify what decision models are actually for.

TypeSafe calls Jev a "System One model", named after Kahneman's book *Thinking, Fast and Slow*. System One refers to the part of human thinking that operates fast and intuitively, while System Two is responsible for deliberate, slow thinking. Over the past two years, almost all large language models have been iterating in the direction of System Two, with longer reasoning chains and deeper thinking capabilities.

However, things change when Agents actually run in practice. In a typical Agent task, what happens most frequently is not deep reasoning, but a large number of trivial judgments: which team this work order should be assigned to, whether this command should be allowed to pass, which button to click next on the page, and whether the task is actually completed.

In the past, these judgments were either handled incidentally by large models, which generate a paragraph of text before parsing the result, or hard-coded into rules that cannot cover long-tail scenarios. The former approach is slow and expensive, while the latter is not flexible or intelligent enough.

Jev's solution is to only answer questions with preset options, including single-choice questions, scoring tasks, or yes/no questions, and directly return structured results with probability values. Its pricing is $0.042 per million input tokens, with free output.

It can be said that Jev has separated "judgment" from "generation", turning it into a low-cost, high-frequency infrastructure.

The market reaction is very direct. Vercel disclosed that within 24 hours of Jev's launch, 13% of its paying users on AI Gateway started using it, twice the adoption rate of any previous model release. TypeSafe named the model after economist William Stanley Jevons, alluding to the famous Jevons paradox: the cheaper a product is, the greater its demand will be.

However, Jev has an obvious gap. It is a closed-source hosted model with no public weights and cannot be deployed locally, so users can only call TypeSafe's own API.

This gap has exactly become the entry point for latecomers.

Chinese teams follow up rapidly

Domestic teams are also rapidly developing Jev-like models.

Shanghai AI Laboratory's Intern-Decision has open-sourced three versions with 0.8B, 2B and 4B parameters, focusing on multimodal capabilities that can understand images and interfaces before making decisions. The team's published data shows that its reasoning speed is 2 to 3 times faster than Jev.

StartLux, a startup company, has launched five specifications ranging from 0.8B to 27B parameters, equipped with complete quantization files, directly targeting local deployment scenarios.

The score of StartLux model exceeds Jev | Image source: Github

According to the data released by StartLux, the 27B version outperforms Jev 1.13 in 31 out of 38 tests on Decision Index 0.2.1, with a comprehensive score of 63.88 versus 57.91. The team also demonstrated a set of chess games, winning 35 out of 36 matches.

StartLux model plays chess | Source: Github

These figures need to be viewed objectively. They are self-test results completed by the team based on public tools and the snapshot of the ranking list on September 28. More importantly, the "first place" in this track has an extremely short shelf life. The day after StartLux's release, Cloudflare claimed that its own Clef took the lead on the same ranking list. As for chess playing, it is never the main battlefield of decision models, but more like an eye-catching demonstration.

Compared with who temporarily ranks first, it is more noteworthy that Chinese teams almost unanimously chose open source and localized deployment.

This is no coincidence. Jev is closed-source and can only run on the cloud, which itself is a threshold for domestic developers who are sensitive to data and have concerns about cross-border calls. Open source weights that can run on local machines perfectly fill the vacancy left by Jev.

There is another detail: Cloudflare's Clef is post-trained based on Qwen, Amazon's Strands Decider also uses Qwen as the base, and APUS's solution also runs on Qwen. In this round of decision model boom, Chinese open source base models have become the common "backbone" for global players.

Why bet on "local deployment"

Let's go back to StartLux itself.

This company was established this year, with Chen Danian as CEO and Dr. Guo Quanwei as CTO. Chen Danian's resume is familiar to all Chinese internet practitioners: he developed the internet access billing software ENCounter in 1998, co-founded Shanda with Chen Tianqiao the next year, and later developed WiFi Master Key.

StartLux positions itself as "focusing on local models", so that model capabilities and data stay on users' own devices. Its last notable move was StartLux-27B, post-trained based on Qwen3.6-27B, which scored 39.25 in the MCP special test of China Academy of Information and Communications Technology, exceeding DeepSeek-V4-Flash with 284B parameters. It should be noted that this is a special test focusing on tool calling, and does not represent a comprehensive lead in general capabilities.

In StartLux's plan, "local AI" is not just a small model that can be downloaded to a computer, but a complete personal AI system. The 27B model undertakes general capabilities, the decision model is responsible for high-frequency judgments, the upper layer consists of Agents and memory modules, and the bottom layer is the quantization and reasoning system.

In this framework, decision models are more meaningful for local scenarios than for cloud scenarios. Cloud computing power can be scaled up arbitrarily, but personal computers have hard upper limits on video memory, memory and power consumption. If every small judgment requires calling a 27B model, local Agents simply cannot run. According to the data released by StartLux, the 4B version after Q4 quantization is about 2.71GB, and its decision consistency with the original full-weight model still reaches 98.3%.

For a company that bets on local AI, the decision model is not a new category to chase hot spots, but a key component that determines whether the local system can run smoothly.

This is also the real significance of Jev to StartLux. To be precise, Jev did not choose the direction for StartLux. Judging from the timeline, the project establishment of StartLux's decision model is obviously later than that of Jev. What Jev proved is another thing: the path of "splitting intelligence into different levels with each performing its own duties" has been voted by the market with real money, which exactly matches the local architecture logic that StartLux has been emphasizing all along.

The figure of "3 days" is also worth mentioning. StartLux attributes this to its Auto Research system, which allows AI to participate in data construction, training, evaluation and failure analysis. Whether this method can be reused repeatedly across different model categories can better demonstrate the real strength of this company than any ranking result.

Will the judgment layer become extremely low-cost?

However, it is too early to draw a conclusion on how deep the moat of decision models is.

Jev claims that it has been researching in stealth mode for two years, OpenAI developed the Decisions API based on its own small model in just two weeks, StartLux completed the development in 3 days, and Cloudflare and Amazon also followed up within half a month. When Jev was released, some analysts predicted that large manufacturers would soon launch their own decision models, and Jev would eventually have to fight a price war with its own "more advanced version".

A category that can be replicated in just two weeks is difficult to form a moat on its own.

For StartLux, the real test is not at the decision model layer. Guo Quanwei himself admitted that long-task stability, exception recovery, context management and overall experience are far more difficult than simply getting the model to run.

According to the plan, StartLux's first public-facing local intelligent experience version is expected to be launched within this year.

From the time-based billing internet access tool to WiFi Master Key, many products Chen Danian has made in the past are high-frequency, underlying businesses that win by scale. Decision models are also high-frequency, underlying businesses.

The question is, when "judgment" is so cheap that everyone can develop it, will the value finally fall on the model itself, or on the people who assemble these models into usable products?

This article is from the WeChat official account "GeekPark" (ID: geekpark), written by Hualin Wuwang, published with authorization from 36Kr.