When large AI models are repriced: Unisound releases U2, ushering in the "DeepSeek Moment"
The large language model industry has long been immersed in a consensus that is almost taken for granted as correct.
Models need massive parameters to be powerful; they need sufficiently long context to have comprehensive capabilities; they need sufficiently complex reasoning chains to demonstrate their intelligence level.
Therefore, in the past few years, from hundreds of billions of parameters to trillions of parameters, from hundreds of thousands of tokens of context to millions of tokens, from single-turn responses to increasingly lengthy reasoning, large model companies have continuously pushed the upper limit of technology, and the capital market is also willing to pay for this stronger imagination. Model leaderboards are updated frequently, training costs are driven higher and higher, and GPUs have become one of the most expensive means of production.
However, as the passion for blindly stacking parameters fades, the industry begins to face an unavoidable awkward reality: dense models with hundreds of billions or even trillions of parameters have simultaneously pushed up astronomical training and reasoning costs, creating extremely high deployment barriers.
Whether for large-scale enterprises or individual geeks, there is a huge gap between the ideal "emergence of intelligence" and the reality of "unaffordable and unusable".
The more frenzied the first half is, the more stark the second half will be.
At this moment, some more far-sighted players have realized that a paradigm shift is taking place, and generative AI is fully evolving into productive AI.
It is necessary to integrate sufficiently hardcore intelligent capabilities into the real industrial capillary network smoothly at a lower overall cost and with a more stable delivery method.
At this node that determines the development trend of the industry, an AI veteran has provided a solution to reconstruct efficiency through a hardcore base model upgrade.
Today, Unisound officially releases its new generation of general large language model base — U2.
This is not only the most important base model technology iteration since Unisound went public, but also a key milestone for its full transformation into an "Agentic Native Large Model Company".
While the industry is still chasing the chat stunts of generative AI, Unisound has proactively put forward the new concept of "productive AI" long before, and the meaning behind it is very clear: the ultimate value of AI is not to generate content, but to solve complex tasks in the real world.
The proposal of this concept makes Unisound a well-deserved pioneer: when peers are still talking about "intelligence emergence", Unisound has already begun to think about "getting work done"; when other companies are competing on parameter scale, it has already seen through the commercial essence of intelligence density and token value. It is this forward-looking cognitive stance that has allowed Unisound to obtain the discourse power to define the rules before the official competition of large models kicks off.
This enterprise that has been deeply rooted in the AI industry for more than ten years has not chosen to join the blind parameter consumption war. Instead, at this node, it uses the new underlying logic represented by U2 to announce to the industry the recalculation of the commercial value of large models. With more than ten years of know-how accumulated in vertical scenarios, Unisound has built an almost unreplicable moat, which is the fundamental confidence for it to firmly occupy the core position in the first echelon of domestic large models.
Beyond Voice
To understand the second half of the large model industry, you must first see clearly the players at the table.
In the domestic large model camp where many players compete fiercely, Unisound, founded in 2012, is a rather unique existence. It started with speech recognition in the early stage, and was once active in scenarios such as smart healthcare, smart home, and in-vehicle cockpit. Over the past ten years, it has gone through the complete technical cycle from statistical learning, deep learning to the era of large models, so it is often regarded as a somewhat "old-school" AI player.
For a long time in the past, because the word "sound" is included in its name, the outside world habitually labeled it as a speech recognition vendor. At the peak of the large model boom, when the attention of the outside world was attracted by new internet players and the "six rising stars" who received billions of dollars in financing and were active in the spotlight, Unisound, which was in the quiet period of its Hong Kong stock listing, seemed rather low-key.
Fortunately, the era of generative AI that can only chat ended in 2025, and everyone realized that productive AI that can get things done is the real priority. At this moment, the industry suddenly finds that Unisound's old business that was underestimated in the past has become its broadest moat in the era of Agent.
"Behind the voice is language, and behind the language is intention. What we hear is not the voice itself, but the consciousness behind the voice." Huang Wei, founder of Unisound, explained the meaning of "sound" in the name of Unisound.
In his understanding, human-computer interaction always has three levels: the first level is "understanding what is heard", that is, speech recognition, which converts voice into text; the second level is "understanding the intention": when a user says "I'm cold", he does not want a simple response, but hopes that the air conditioner will automatically adjust the temperature and the curtains will automatically close; the third level is to understand the deeper consciousness and scenarios: when an elderly person living alone says casually "I have nothing to do today", can AI recognize the loneliness from the tone and pauses, and actively trigger companionship or reminder services.
From speech recognition and natural language understanding to today's large models and Agents, what Unisound has been doing is always one thing: to make machines truly understand humans and help humans get things done.
In the real physical world of human-computer interaction, to make machines truly serve humans, a series of engineering problems such as multi-turn interaction, long-link tasks, complex environmental noise, and human-machine collaboration must be solved.
The scenario know-how and multi-turn interaction engineering experience accumulated in these serious and complex vertical scenarios are exactly the natural soil for the survival of native Agents.
Based on this deep insight, the internally defined positioning of Unisound's newly released general large base model U2 is Agent-Native Model, whose size, training objectives and optimization directions are all designed around "task execution".
In terms of technical path, Unisound did not follow the industry's common route of "completing model training first and then attaching an Agent framework externally", but proposed a more radical idea:
First, the native Agent model + Harness collaborative evolution mechanism. In the past, most Agent systems were more like adding an outer shell to a general chat model — the model was only responsible for generating responses, and tasks such as planning, tool calling, and task execution were all handed over to the external framework, and the model itself did not really "understand" these. However, U2 has directly internalized the full capabilities of how to plan, execute, and verify results into the model layer during the training phase. During the training process, the model and Harness (task execution scaffold) continue to evolve collaboratively: as the main structure of the model becomes more and more complex, the support nodes and verification accuracy of the scaffold also extend and become more precise; and the more sophisticated and stricter scaffold in turn ensures the robustness of each layer of logic of the model, forming a continuously self-reinforcing cycle.
Second, the systematic application of process supervision and curriculum learning. In order to make Agent as efficient as people who get things done neatly, U2 introduces the "curriculum learning" method in the training phase, allowing the model to progress gradually from easy to difficult tasks, from short to long context, and from simple to complex tool calling. In the trajectory of long-horizon tasks, U2 introduces an advanced process supervision method, using a more optimized model to disassemble, evaluate and correct errors at each key node of task execution. U2 can not only see the final result, but also optimize every intermediate execution path, realizing rapid convergence of learning.
Third, the industry-level data ratio that is more inclined to serve the real economy and hardcore industries. While many large models still rely heavily on general internet corpora for generalization training, Unisound chooses to actively reduce the proportion of corpora from low-value scenarios such as entertainment, tilt more data resources to high-value industry scenarios such as healthcare, medical insurance, insurance, government affairs, and industry, and conduct training combined with desensitized data of real scenarios accumulated over years of business implementation. It is worth mentioning that Unisound uses desensitized data of real scenarios that has been precipitated over years of business operations and is difficult to replicate for synthesis and training, directly serving the real economy and hardcore industries.
After the underlying capability is reconstructed, Unisound U2 shows strong performance competitiveness without blindly stacking parameters. In instruction-following evaluations such as IFBench, U2's performance ranks among the top in the industry; in Claw-related evaluations, its Agent and tool calling capabilities show strong advantages; in hardcore knowledge reasoning and long-context tasks such as GPQA, U2 also demonstrates the ability to challenge the world's top large models; in GDPval, which measures the delivery capability for real office and knowledge work, U2 scores 72.5 points, demonstrating solid professional office capabilities.
Most importantly, U2 completely breaks the curse that "top-tier performance must be bound to super-large parameters". It rejects parameter bloat, and through the ultimate Mixture of Experts (MoE) architecture and algorithm optimization, it is committed to compressing capabilities comparable to the world's top level into a smaller parameter scale, pursuing to be powerful yet lightweight, powerful yet cost-effective.
The low-key and restrained AI veteran has joined the first echelon of domestic large models as a leader.
How does the business logic form a closed loop?
As a technology company with more than ten years of industry experience, Unisound understands better than the "newcomers" who have just entered the industry for only a few years that while the generational leap of technology is worth striving for, the business logic cannot be ignored either.
In the past, the large model industry was accustomed to discussing tokens from a single perspective of hardware and computing power. When the entire industry was competing on who generates more tokens and who has higher computing efficiency, Huang Wei, founder of Unisound, calculated a more insightful business account: "If the same 1 million tokens are all carrying small talk and nonsense, no matter how high the computing efficiency is, it has no commercial value."
Based on this cognition, Unisound has for the first time proposed a highly disruptive commercial formula in the industry:
AI Commercial Value = Intelligence Density × Token Value.
To break it down, Intelligence Density means using smaller parameters and lower overall resource input to reach a sufficiently high level of intelligence. Token Value emphasizes that every call of the model must be directly converted into a measurable business result — either reducing risks or improving productivity.
The U2 model released today is the ultimate implementation carrier of this cognition and thinking. To ensure that every penny of customers is spent where it matters, U2 has achieved almost strict optimization at the underlying technology level.
The Agent + Harness collaborative evolution mechanism mentioned earlier is exactly the solution to this problem. Through the co-evolution of the model and the tool harness, U2 can complete task planning, tool calling, execution and verification with fewer interaction turns, reducing a large amount of token waste caused by repeated trial and error, and further improving the task completion rate.
At the same time, U2 adopts a Sparse Mixture of Experts (MoE) architecture at the underlying layer. Compared with traditional dense models that need to activate all parameters, MoE only activates the most relevant part of expert models for different tasks. According to the information disclosed by Unisound, U2 only activates about one-tenth of the parameters to participate in calculation each time when processing tasks, and the rest of the parameters "sleep on demand". This means that the actual amount of calculation during the operation of the model is far less than its full scale, which significantly reduces the computing power cost required for reasoning while maintaining high performance.
What is more special is U2's redesign of the thinking process. Some large models tend to expand lengthy reasoning as the thinking process during complex reasoning — writing out the complete intermediate process step by step. Although this method improves interpretability, it also brings another problem: users are paying for a large number of token costs that do not generate final value. U2 prioritizes efficient exploration in the latent space, avoiding decoding every intermediate step of thinking into visible tokens; when the task enters a critical stage, the model switches to explicit reasoning, completing logic calibration, process verification and final decision-making through a readable and verifiable reasoning process. Unisound calls it "implicit thinking reasoning + explicit thinking verification".
"If these 1 million tokens are all small talk and nonsense, no matter how high the efficiency is, there is little commercial value." Huang Wei once said.
This strategy of "pursuing high-value tokens with high intelligence density" soon received real-money feedback in the commercialization battlefield, and the results proved the unlimited ceiling brought by the new formula.
The latest data shows that benefiting from the explosive growth of high-quality scenario-based token demand, Unisound's ARR from token call revenue in May increased by 600% month-on-month, and according to the current order momentum, it will continue to maintain a strong high growth trend in June, with the expected ARR reaching 15 million US dollars.
In the second half of the large model industry, the ceiling of Unisound's business scale has been fully and completely opened.
Behind this is an essential leap in Unisound's business model.
For a long time, traditional ToB companies have been trapped in the quagmire of project-based operations — long delivery cycles, high degree of customization, and they only earn hard money from one-off deals. However, with the release of the large model U2, through the continuous output of high-value tokens, Unisound's revenue model has been deeply bound to the AI usage intensity of customers. As long as customers continuously call AI in real business processes, revenue will generate high-frequency, high-gross-margin recurring purchases just like tap water from an open faucet.
Today, this efficient closed business loop is being accelerated to land through Unisound's unique Two-Wheel Drive Layout:
On the To B side (Zoya Agent Platform), Unisound takes U2 as the core base to expand its territory in vertical industries. With an extremely high task completion rate, the company has recently successively won bids for a series of industry scenarios with strict accuracy requirements such as healthcare, medical insurance, transportation, customer service, and employee ID cards. These high-value industries not only continuously contribute high customer unit prices, but also continuously feed back the model base with real and high-quality business data precipitated in the scenarios, forming a virtuous cycle where the more it is used, the smarter it gets and the higher the intelligence density becomes.
On the To C and developer side (Public Cloud MaaS), Unisound is fully expanding relying on the OPC Ecosystem. Through lower-threshold and more cost-effective model API calling capabilities, it continuously and stably harvests high-frequency token traffic and revenue for a wide range of independent developers and C-end application ecosystems.
Instead of blindly following the limit of Scaling Law, Unisound finds an ecological niche that perfectly matches its own resources and endowments. With its hardcore technology hematopoietic capability of ARR surging 6 times in a single month, Unisound proves that in the second half of the large model industry, the players who can calculate the efficiency account clearly and grasp the closed business loop will have a real unlimited future.
The official competition has just begun
Looking back at the past year since its listing on the Hong Kong Stock Exchange, Unisound has delivered a hardcore answer sheet that does not follow the crowd, does not engage in empty talk, reconstructs efficiency with technology, and proves its value with business results. When the whole industry is trapped in parameter inflation and paying for high-cost computing power experiments, this AI veteran has realized technological foresight and found commercial value with a clear strategic orientation and more than ten years of deep cultivation in the industry.
The release of the U2 large model base is even a pattern-breaking declaration. With the dual-high solution of "high intelligence density" and "high value token", it has formally established Unisound's stable position as a core vendor and first-tier player of domestic general large models. In the fully launched era of native Agents, large models are no longer unaffordable luxuries, but practical productivity tools that get things done efficiently.
"2023 to 2025 is the warm-up match of large