From conversational intelligence to decision-making prowess: Baichuan M3 redefines the benchmark for medical large language models.
On January 13, Baichuan Intelligence released and open-sourced Baichuan-M3, a new-generation medical-enhanced large language model. On HealthBench, the authoritative medical evaluation dataset led by OpenAI, and its difficult subset, this model achieved the highest overall score across the globe, significantly outperforming GPT-5.2. It also reached the current lowest level in the pure model evaluation of medical hallucination rate. In the SCAN-bench evaluation focusing on full-process clinical capabilities, M3 ranked first in multiple core indicators such as medical history collection, auxiliary examination and diagnosis, demonstrating its comprehensively leading medical reasoning and consultation capabilities.
In addition, M3 is the first model equipped with native "end-to-end" formal medical consultation capability. It can actively ask follow-up questions like a doctor, dig deeper layer by layer to identify key medical history and risk signals, and then conduct in-depth medical reasoning based on complete information. Evaluations show that its consultation capability is significantly higher than the average level of real human doctors.
However, the significance of this release is far more than another breakthrough in technical rankings. More importantly, Baichuan-M3 has pushed medical large language models to a new stage: it no longer stays at the level of dialogue and expression, but begins to truly have the ability to support the complete diagnosis and treatment process, and can participate in medical decision-making itself. For this reason, its significance far exceeds that of other models, and the technological progress of large language models has finally been fully transformed into practical value that can be implemented on a large scale in the medical and health field.
"Helping patients generate the value of auxiliary decision-making is meaningful," shared Wang Xiaochuan, Founder & CEO of Baichuan Intelligence, at the launch event.
In the medical scenario with the highest requirements for safety and responsibility, such changes do not happen by chance. It means that someone has chosen a slower, more difficult and less opportunistic path, gradually pushing the model capability from demonstrating intelligence to undertaking decision-making.
Why did Baichuan get to this point? Why did this breakthrough occur in the medical field, rather than in more popular tracks such as coding, search or intelligent agents? And why at this very moment have these long-accumulated technical choices and engineering routes begun to converge to a clear result?
The Evaluation Criteria for Medical Large Language Models Are Being Rewritten
Almost from the very beginning of the birth of artificial intelligence, people have regarded medical care as one of the industries that is most likely and most worthy of being transformed by AI.
Before the advent of HealthBench, AI capabilities related to the medical industry were barely comparable. All model vendors could claim that their models understand medicine and can answer medical questions, but there was no unified evaluation coordinate system, let alone horizontal comparison.
In May this year, OpenAI launched HealthBench. This set of standards brings together a large number of multi-round dialogue samples designed based on real clinical scenarios, enabling medical capabilities to be quantitatively evaluated with public criteria. Therefore, for a long period of time, it was almost equivalent to the highest standard for medical large language models, and became the common coordinate system for all model vendors to demonstrate their medical capabilities.
For this reason, for quite a long time, the default consensus has been that whoever scores higher on HealthBench has a better understanding of medical care. This is not because HealthBench covers all the complexity of medical care, but before its emergence, the industry did not even have a standard of its own.
From a certain point in time, the industry trend has changed. From the middle of last year to now, as domestic medical assistants such as Afu and Xiaohe Doctor have been launched one after another, OpenAI released ChatGPT Health, and Anthropic launched Claude for Healthcare. Medical care is no longer just a benchmark used to test how smart the model is, but has become a product direction that large language model vendors must invest in directly; the model also has to face the issue of whether its answers can be used as the basis for decision-making.
This is no longer just a ranking issue.
It is also at this stage that the boundaries of HealthBench begin to emerge. It is still important, but no longer sufficient. It can still prove whether the model has medical knowledge and professional expression ability, but cannot answer a more core question: whether the model is qualified to enter the real medical decision-making process.
Clinical decision-making never starts from a standardized question, but from highly incomplete and even chaotic information. Patients often cannot clarify the key points, symptoms overlap with each other, and different risks are mixed together. The real difficulty lies not in "how to give the answer", but in "how to ask the question". A large part of a doctor's professional ability is reflected in the judgment of information priority: which high-risk signals must be ruled out immediately, which ones can be postponed; which information cannot lead to a conclusion if missing, and which ones are only supplementary references.
It is precisely at this point that Baichuan has made a significantly different choice from the mainstream route. On the one hand, it has not given up the competition in the HealthBench system, and continues to pursue the best performance under the existing authoritative standards; on the other hand, it has launched SCAN-bench at the same time, trying to make up for the long-neglected dimension of modeling and evaluating the complete clinical process itself.
Centering on the SCAN principle, Baichuan drew on the OSCE method that has been widely used in medical education for a long time, and cooperated with more than 150 front-line doctors to build the SCAN-bench evaluation system. This system takes real clinical experience as the "standard answer", splits the diagnosis and treatment process into three stages: medical history collection, auxiliary examination and accurate diagnosis, conducts assessments in a dynamic and multi-round manner, and fully simulates the whole process of a doctor from receiving a patient to confirming the diagnosis. Compared with HealthBench, SCAN-bench is a new paradigm of more comprehensive end-to-end dynamic evaluation for the whole process.
In other words, when the industry is still competing for who is better at "answering", Baichuan has turned its attention to another more fundamental question: can the model "ask" like a doctor?
This is exactly what makes this M3 release truly special: it forms a closed loop in the capability structure, which can not only reason, but also avoid making up facts randomly, and know how to ask for all necessary information. The reasoning capability solves the problem of "whether it can make judgments", the low hallucination feature solves the problem of "whether it is trustworthy", and the consultation capability solves the problem of "whether it is qualified to enter the decision-making process".
When all three conditions are met, the medical large language model can be regarded as evolving from an intelligent system that can talk to a system that can be entrusted with part of the responsibility for medical decision-making.
From the results, M3 is still a model with multiple firsts. Its top ranking on HealthBench means that it has achieved a comprehensive breakthrough under the medical capability standard system defined by OpenAI itself; and in the more challenging HealthBench Hard subset that focuses on complex clinical decision-making capabilities, M3 won the championship with a score of 44.4%, systematically surpassing GPT-5.2 for the first time. This result is more convincing, because what it verifies is not only whether the answer is professional, but also the stability and reliability of the model in scenarios with high uncertainty and high reasoning difficulty.
At the same time, M3 has achieved the world's lowest hallucination rate without using any tools, which means that safety is internalized as the model's own capability, rather than relying on external retrieval, rule constraints or engineering patches for compensation. More importantly, in the SCAN-bench evaluation targeting the complete clinical process, M3 also ranked first, especially in the most core consultation session, it significantly outperformed the GPT series models and the baseline level of human doctors, which shows that the model has truly supplemented the core capability of clinical information acquisition, a capability that has long been neglected but determines the upper limit of medical decision-making.
The Real Watershed of AI Medical Care
If we say that in the past two years, the industry has mostly been working on making models "talk" like doctors, the judgment given by M3 this time is: expression alone is not enough, the model must have the thinking structure of a doctor.
A large number of "AI doctors" are still at the level of role-playing, with smooth dialogue and professional tone, but their questions are mostly designed to make the dialogue look complete, rather than to collect key information for clinical decision-making. The model often follows the patient's description to continue the conversation, but rarely does risk stratification, checks for red flag signs, and designs questions reversely around the diagnosis and treatment path like a real doctor. As a result, the dialogue looks professional, but is not sufficient to support serious judgments, and can only end up with a safe conclusion like "please seek medical treatment as soon as possible".
This is the essential difference between "being able to talk" and "being able to make clinical decisions", and it is also the background for Baichuan to propose "formal medical consultation" and the "SCAN principle". Wang Xiaochuan shared at the launch event, "In the medical industry, patients often cannot fully express themselves, they only know superficial symptoms, so they need to consult doctors to clarify the development of past medical conditions through consultation. With sufficient data, we can do a good job in subsequent testing, diagnosis and conclusion. Current large language models do not have such capabilities."
What Baichuan wants to do is to decompose the working methods that clinicians have long relied on experience to complete into engineering goals that can be learned by the model, evaluated, and directly optimized through reinforcement learning.
Specifically in terms of engineering, Baichuan did not choose to pile up functions, but focused on solving three most fundamental problems.
The first is the fully dynamic reinforcement learning system.
In the M2 stage, reinforcement learning relied more on relatively static verification rules. After the model capability improved to a certain level, the verification system itself became the upper limit. In M3, the Verifier is designed as a system that can co-evolve with the model capability: when the model exposes new error patterns, the verifier generates new constraints; old, low-value rules are eliminated, and high-value rules are continuously strengthened. The rules and the model jointly raise the upper limit, solving the problem that capability is easy to hit a ceiling in the later stage of iteration.
The second is the SPAR algorithm.
Medical consultation is naturally an extremely long decision-making chain. If you only check whether the final diagnosis is correct, the model cannot figure out where the problem is: whether the medical history is not asked clearly, or the examination suggestion is wrong, or the reasoning path is deviated. Through step-by-step penalty and relative baseline mechanism, SPAR decomposes the long-chain decision-making into local processes that can be held accountable, allowing the model to learn to ask key questions accurately and sufficiently within a limited number of dialogue rounds, rather than by extending the number of conversation rounds.
The third is Fact-aware RL. In medical scenarios, the stronger the reasoning capability is, the easier it is for the model to be "overconfident"; the more certain the model expresses, the more dangerous it will be once the factual basis is not solid. The traditional method often relies on external retrieval or rule systems to correct deviations, while M3 directly takes low hallucination as the optimization goal of reinforcement learning, making fact consistency part of the model's own capability. At the same time, through dynamic weight adjustment, the model is prevented from degenerating into a conservative state of saying less to make fewer mistakes, so that strong reasoning and high reliability can be achieved at the same time.
Behind these three sets of designs, they actually point to the same goal: no compromise between capability and safety, strong reasoning and high reliability. Baichuan wants both, and make them become coordinated indicators in the same engineering system.
In this way, AI medical care has truly crossed that watershed.
From Health Assistant to Decision Support
When the model capability completes the closed loop of reasoning, low hallucination and qualified consultation, the focus of Baichuan will inevitably begin to shift: from the display of the model itself, to the support of capabilities in real medical scenarios.
This is also why, from external observations, we can find that the product rhythm of Baichuan Xiaobaiying has accelerated significantly recently, with various functions being supplemented one after another, gradually building a system skeleton that can undertake the medical workflow. What the model needs is no longer a display window, but a carrier that can precipitate information, support long-term use, and connect to the real decision-making chain.
In this way, the difference between the "formal medical care" that Baichuan adheres to and a large number of "general health" products in the industry has become particularly clear.
Products represented by Afu and Xiaohe Doctor are more similar to health consultation, medical science popularization, triage suggestions and emotional companionship, which solve the problems of information asymmetry and patients' anxiety before seeking medical treatment.
What Baichuan tries to enter is a completely different chain: doctors can use it to deduce their consultation and diagnosis ideas, and patients and their families can also use the application to systematically understand the medical logic behind diagnosis, treatment, examination and prognosis.
This is a decision support path with high risk, high responsibility and high value density: in this scenario, the model no longer only provides reference information or emotional comfort, every judgment it makes may affect the patient's next choice; every question it raises determines whether key information is fully collected; every conclusion it forms must be verifiable and can be truly incorporated into the medical decision-making process.
The fundamental difference is that while most products in the industry still stay at the level of helping users collect health information, Baichuan has chosen a more difficult, slower but higher-ceiling path.
Looking back at the timeline of Baichuan's bet on medical care, its choice is a judgment of advance layout.
At the communication meeting, Wang Xiaochuan summarized his judgment on several core pain points of the medical industry: high-quality doctor resources have long been in short supply, and medical services are highly uneven among different regions and groups; the United States has a family doctor system to undertake primary diagnosis and treatment, while in China, patients are more concentrated in tertiary hospitals, further squeezing high-quality medical resources. Based on long-term observation of these real contradictions, Baichuan has aimed at solving the problems of medical care itself from the very beginning.
In 2023, at the hottest stage of the large language model industry, Baichuan did not choose to prioritize entering tracks such as coding, search and content creation that are easier to verify commercial value, but clearly took medical care as its core direction. This was not an opportunistic choice at that time: medical data is sensitive, scenarios are complex, responsibility boundaries are vague, and product implementation cycles are long, making it difficult to get quick feedback. "At that time, we were also questioned by many people in the industry," Wang Xiaochuan told us.
At the beginning of 2026, OpenAI released ChatGPT Health, and Anthropic also officially launched Claude for Healthcare. Leading international model vendors began to enter the medical field collectively, and all companies around the world realized that medical care is the must-win battlefield for large language models.
In this race, as the only large language model enterprise in China that focuses on medical care, Baichuan has continuously made breakthroughs in core capabilities such as low hallucination rate, end-to-end consultation and complex clinical reasoning, and has achieved generational leadership on the medical large language model base. It has evolved from a "follower" to a "leader" of the industry and a "definer" of the new paradigm, and is shouldering the banner of China's AI medical development with its hardcore strength.