HomeArticle

Frontline | Local Large Model StartLux-27B Passes MCP-Universe Evaluation: Ranked Second in Overall Performance, Tops Multiple Specialized Indicators

王毓婵2026-08-31 19:23
Local models are likely to account for 80% of the large model market within 3 years.

Recently, an inspection report released by the Artificial Intelligence Research Institute of the China Academy of Information and Communications Technology (CAICT) under the Ministry of Industry and Information Technology shows that the local large model StartLux-V1.0-27B-Preview developed by Shanghai Origin Starshine Technology Co., Ltd. (hereinafter referred to as StartLux) ranked second in overall performance with 27B parameters in the MCP special test, a benchmark test for trusted AI large models, surpassing the third-ranked DeepSeek-V4-Flash, and its capability range has entered the capability interval of trillion-parameter models represented by DeepSeek-V4-Pro.

The MCP special test, the benchmark test for trusted AI large models, covers 6 special tasks including location navigation, web search, browser automation, financial analysis, code repository management, 3D design, as well as a comprehensive evaluation, totaling 7 inspection items. It focuses on examining the comprehensive performance of large models in multi-tool collaboration, complex task execution and real environment interaction, and some indicators and data of this special test refer to the open-source project MCP-Universe.

The tested models include DeepSeek-V4-Pro (1.6T), DeepSeek-V4-Flash-0731 (284B), Step-3.7-Flash (198B), StartLux-27B-260715 (27B), Qwen-3.6-27B (27B) and AgentCPM-Explore (4B).

The results show that StartLux-V1.0-27B-Preview gets a total score of 39.25 and ranks second, exceeding the 284B DeepSeek-V4-Flash-0731 and the 198B Step-3.7-Flash. Under the same parameter scale, StartLux-V1.0-27B-Preview is 5.34 percentage points higher than Qwen-3.6-27B. Its performance in individual tasks is even more outstanding: it ranks first in location navigation, and also ranks first either tied with the 1.6T-parameter DeepSeek-V4-Pro or exclusively in tasks such as browser automation and financial analysis.

Test Results

Test Results

StartLux uses Qwen3.6-27B as the base for targeted post-training enhancement, pushing a 27B local model to the same performance level as the 1.6T flagship model, relying on a brand-new, multi-dimensional, verifiable and scalable model optimization and upgrading technology independently designed by the StartLux team.

During the model training process, the team adopted the AI trains AI (Auto Research) method to carry out training experiments independently, and continuously optimized the training strategy through feedback. This method represents the current cutting-edge direction of global model R&D. StartLux-V1.0-27B-Preview is the first domestic local Agent model that adopts this method for post-training.

Around the world, local models are gradually entering the industry's field of vision.

Google launched the Gemma 4 series in April, among which the 31B Dense version ranks third in the open-source model leaderboard and can run locally on consumer-grade graphics cards after quantization. In early August, Meta released and open-sourced the 30B local model Muse Glimmer. NVIDIA also launched the open-source model Nemotron 3.5 Lightning in August, which is a 30B-parameter MoE model that can run directly on local devices such as RTX PCs.

Chen Danian, founder of Shanda Network and Chairman of Yusheng Science, recently predicted that local large models will occupy 80% of the large model market within 3 years. Roman Orús, co-founder and Chief Scientific Officer of Multiverse Computing, a quantum and AI software company, also stated that the industry is entering a period where local large language models will become real competitors to cloud services.

At present, StartLux-V1.0-27B-Preview can run on consumer-grade personal computers. It is understood that StartLux plans to launch its first-generation local intelligent solutions within this year. Meanwhile, the StartLux team is steadily advancing the model training process, and is also conducting in-depth research and optimization on new architectures such as diffusion language models.