Jev was dethroned, and the StartLux China open-source decision-making model surged to the top spot.
On September 30, a decision model was officially released and quickly drew huge attention: in 38 benchmark tests of Decision Index 0.2.1, it outperformed the currently most popular Jev in 31 items, with a comprehensive score of 63.88 versus 57.91, leading the top of the public ranking by nearly 6 points; what's more, it is completely open-source.
Practical comparison can better illustrate its strength — in the chess matches against Jev, its response speed is comparable to its rival, but it won 35 out of 36 games.
The company that launched this model is exactly the red-hot Shanghai AI star enterprise StartLux. Founded less than 5 months ago, this company has delivered its second industry-shocking achievement: last time, the local model StartLux-27B surpassed the 284B DeepSeek-V4-Flash in tests by the China Academy of Information and Communications Technology under the Ministry of Industry and Information Technology, and basically matched the performance of DeepSeek-V4-Pro with only about 1/60 of its parameter count. This time, it is the fully open-source StartLux-Decision — based on the comprehensive scores of public evaluations, it is already the world's strongest decision model to date.
Outperforming Jev
This time, StartLux's Decision model has launched 5 versions with 0.8B, 2B, 4B, 9B and 27B parameters in one go, and provides quantized models that support local deployment.
GitHub link: https://github.com/StartLuxLabs/Startlux-Decision
Hugging Face model collection download: https://huggingface.co/collections/startlux-models/startlux-decision-6abba92b301b573fa154d493
How strong is its capability? Let's first look at a chess game.
The opponent is Jev 1.13, which has drawn much attention recently. Both sides see the same chessboard state, do not use additional search, and each move is directly selected by the model.
Eventually, StartLux-Decision-27B completed checkmate. In this game, it also performed better in indicators such as average loss, best move hit rate and number of mistakes. In such chess games, StartLux won 35 out of 36 matches.
Now let's look at the specific test situation. First, the overall performance of StartLux-Decision speaks for itself, you can check the numbers directly.
Then look at the scores of the 27B version.
In the independent Decision Index 0.2.1, the 27B version of StartLux-Decision scored 63.88, outperforming the latest Jev version Jev 1.13 in 31 out of 38 benchmarks; in another set of 7 tests, the average accuracy of StartLux-Decision-27B reached 91.82%. It should be noted that the above results are self-test results completed by the StartLux team based on public evaluation tools and the snapshot of the public Decision Index ranking as of September 28, 2026; the 7 public tests come from the evaluation set of the Shanghai AI Lab Intern-Decision team, which uses a completely independent standard from the Decision Index.
Why Do Agents Need
A Dedicated Model for "Judgment"?
Decision models have been rising in popularity recently. So why do Agents need Decision models to take charge of judgment exclusively?
First let's see what StartLux-Decision actually does.
In a customer service ticket Demo demonstrated by StartLux, the system receives a message like this: "My order has been deducted twice repeatedly, but the system still shows that it is unpaid."
Traditional generative AI usually first understands the problem and generates a piece of analysis, while StartLux-Decision can complete three clear judgments in one single request: team diversion, processing timeliness and severity level.
Customer service ticket Demo: Complete team diversion, urgency judgment and severity level judgment in one request.
This reflects the most intuitive difference between the Decision Model and ordinary generative models: it does not need to generate natural language first and then be parsed by programs, but directly makes selections, yes/no judgments or grade scoring for the options provided by developers.
When this type of capability is integrated into Agents, its effect will be more obvious.
Take the following shopping Demo as an example. The user's request to the Agent is: "Buy the cheapest 8-pack AA batteries with free shipping, and deliver them to my home address."
The system comprehensively considers the model, quantity, price and delivery conditions. In each round, StartLux-Decision judges the next operation and whether the task is completed based on the current page state; after executing the operation, it updates the page state and continues to judge, until the order is placed and the task conditions are checked.
Shopping Demo: Filter products that meet the conditions of AA specification, 8-pack, free shipping and lowest price, and complete the order placement.
A similar pattern also appears in the office collaboration Demo. The task is to invite specified members to the Design team and set them as Editor. The model needs to select the correct team, member and role in sequence, and then judge whether the invitation process is completed.
Team management Demo: Select the Design team, invite specified members, and grant Editor permissions.
The common point of these Demos is that the answer space is relatively clear. Browser controls, customer service departments, and tool lists are all predefined. Different from the open content generation that LLMs are good at (such as writing emails and making plans), the requirements for a large number of Agent steps are very short (such as Yes/No, one out of three options, five-level scoring or tool selection).
If we look back, we will find that a very interesting cycle has taken place here.
In the traditional software era, engineers would split businesses into a large number of rules. Amounts exceeding a certain threshold enter a certain process, a certain state triggers a certain operation, and different users enter different permission branches. The system is very fine-grained, but a large amount of logic needs to be written manually in advance.
After the emergence of LLMs, a large number of such explicit logics are absorbed by the models. Developers began to directly submit natural language and environment states to large models, letting the same model complete understanding, classification, judgment, planning and generation. The system becomes more general and more flexible.
In the Agent stage, AI begins to divide labor again — deterministic processes are handed over to code, search to Embedding, complex planning to Reasoning Model, and high-frequency selection begins to use dedicated models such as Decision Model.
Computing power is also being allocated in a more fine-grained manner accordingly.
StartLux announced a set of speed tests. Under the conditions of single H200 GPU, BF16, local HTTP and short requests, StartLux-Decision-4B takes an average of 26 milliseconds to answer three questions at once; the 0.8B and 2B versions take 12.2 milliseconds and 15.5 milliseconds respectively.
The team also uses CUDA Graph, fast linear attention operator, and combines multiple questions into one forward calculation to reduce repeated calculation and call overhead.
When an Agent task contains dozens of judgment nodes, the amount of calculation required for each step will soon affect the response speed and operating cost of the entire Agent.
This also explains why Decision Model has heated up rapidly in a short period of time. Many of the things it handles seem trivial: selecting a tool, judging whether a task is finished, checking which process a piece of information should enter, deciding where to click next on the page. Once Agents start running for a long time and at scale, these "small judgments" may become the most frequently called layer in the entire system.
Local Intelligence of StartLux Begins to Form Hierarchy
This September, Machine Heart once had a discussion with Guo Quanwei, Co-founder and CTO of StartLux, about "Local × RSI" (original article: How to Implement RSI Locally? We Talked to StartLux CTO). At that time, a very important question was: what exactly is the "local AI" mentioned by StartLux?
Guo Quanwei gave a broad definition: it includes models, quantization, inference systems, hardware adaptation, as well as upper-layer Agent Harness, tools, search and continuous learning. The team's goal is to deploy these capabilities on users' local devices as much as possible.
This is essentially different from simply running an offline model locally. A long-running local Agent must face many system-level challenges: video memory and memory constraints, model precision selection, scheduling of different tasks among 4B/9B/27B models, long context management, recovery mechanism after step interruption, Memory access strategy, and overhead balance of high-level judgment.
StartLux-Decision is exactly the component launched to fill a specific tier in this system.
In StartLux's system division, StartLux-27B undertakes the general capability layer, Decision Model focuses on high-frequency decision-making, Agents and Memory are deployed at the upper layer, and the bottom layer is supported by quantization, inference systems and hardware adaptation.
This architectural design originates from the realistic computing power constraints: the hardware resources of personal computers have clear boundaries.
Different from the cloud that can expand model scale through GPU clusters, local systems must operate for a long time under limited video memory, memory, bandwidth and power consumption. Therefore, "how to allocate unit computing power efficiently" has become a core issue.
Previous post-training experiments of StartLux on the 27B model have reflected this logic. Guo Quanwei described model parameters as a high-dimensional space: the scale determines the capacity, while training reconstructs the capability distribution within the capacity. Experiments show that after post-training for Agent capabilities, GPQA rose from 83.33 to 90.40, DeepSearchQA rose from 47.04 to 59.39, and Tau3-Banking also improved; but at the same time, GAIA dropped from 57.57 to 45.70, and IFEval also declined.
This phenomenon indicates that under a fixed parameter count, there is a redistribution of resources between different capabilities. The training process essentially determines what tasks the parameters are more inclined to handle.
Decision Model goes a step further: handing over frequently called specific tasks with clear targets to dedicated lightweight models for processing.
This is also the reason why StartLux launched five specifications of 0.8B, 2B, 4B, 9B and 27B, to cover different devices and scenarios. At the same time, it provides three GGUF quantized files of BF16, Q8_0 and Q4_K_M.
Take the 4B specification as an example, the Q8_0 file is about 4.48GB in size. In the 231 public JevBench quantization retest results announced by the team, its decision consistency with the original weight reaches 100%, both answering 204 questions correctly; after being compressed to the Q4_K_M version of about 2.71GB, the decision consistency is 98.3%, answering 201 questions correctly.
In the actual local Agent system, this specification differentiation has clear engineering significance.
A long-running personal AI system may schedule multiple models at the same time: low-latency discrimination is completed by a small Decision Model, complex reasoning is delivered to large models, retrieval is carried by professional modules, long-term memory depends on Memory, and cross-application operations are coordinated by Agent Harness. Models are gradually transformed into different processing units in the computing system.
In addition, probability output is also its key feature. StartLux-Decision can output the probability distribution of candidate options, enabling the system to establish routing and diversion strategies. If the model's confidence in a known task reaches the standard, the system directly executes the subsequent logic; if the probabilities of multiple candidates are close, it can supplement information, transfer to a more powerful model or trigger manual confirmation.
However, the probability itself still needs to be calibrated in combination with specific business scenarios. For high-risk operations such as payment, deletion and permission modification, the original authorization confirmation mechanism is indispensable.
This marks that local intelligence is evolving from a simple "model iteration" to a comprehensive "system engineering".
StartLux once set a clear goal: to make users willing to entrust files, research and office tasks to local AI for a long time, ensuring that the system works continuously for several hours and maintains a high task success rate.
Guo Quanwei pointed out that at present, the progress of models, quantization and local inference systems is relatively fast, and the real difficulties lie in long-task stability, exception recovery, context management and overall experience. The Decision Model solves a specific link in this chain, providing a more cost-effective computing form for high-frequency judgments with clear candidate sets.
Complete One Auto Research Pipeline in 3 Days
Let's talk about this very noticeable number: 3 days. According to the information announced by the team, it took about 3 days from determining the direction of the Decision Model to completing the first round of model R&D and verification. For a model project, this cycle is very compact.
This benefits from the Auto Research R&D system that StartLux has built for a long time under the RSI concept: AI has entered R&D links such as data construction, training, evaluation and failure analysis, which has been continuously proven in StartLux's recent two model R&D processes.
In the previous exclusive interview with Machine Heart, Guo Quanwei explained the data that "70% of the experiment execution involves AI participation": if only counting mechanical engineering such as configuration generation, training start, Eval operation and result sorting, the automation rate has exceeded 95%; and 70% refers to the proportion of AI model participation in the core research and decision-making links.
There is an essential difference in technical difficulty between the two. Automated training startup is relatively mature, while the difficulty lies in the analysis and decision-making after training: why did the model fail? What kind of data should be supplemented? Is the problem in the algorithm or the reasoning sequence? What variables need to be adjusted in the next experiment? Is this route worth continuing to invest computing power in?
StartLux defines this high-level decision-making process as Auto Research.
Gu