Agent acts as the "chief assistant" of doctors: it has an extremely fast speed in reading medical images, and human-computer collaboration remains the optimal solution.
Recently, the 5th Academic Conference on Medical Imaging AI was held in Shanghai. At the event, an open "Human-Machine Film Reading Competition" became the focus of discussions at the conference.
In China, AI medical imaging enterprises have long been fond of holding "Human vs. Artificial Intelligence" film reading competitions to demonstrate the "judgment capability" of their own AI products. Different from the common internal evaluations in the industry this time, the conference organizer collected 10 real imaging cases that "doctors find completely unprecedented" from tertiary hospitals across the country, set up 7 categories of participating groups in the form of "doctor + AI collaboration" or independent writing by AI agents. The participants completed the writing of 10 diagnostic reports in advance, and then doctor experts scored on site, to intuitively reflect the real differences in report quality and output efficiency.
Putting the model directly into real-time combat under the eyes of dozens of top experts is extremely stressful. After all, every output report will be scrutinized with full attention. Shukun Technology, an AI participant, told 36Kr that this was a "closed-book exam". "The organizing committee did not tell us the specific details of the cases, so the AI had no way to learn in advance. We were also very nervous before the competition, not knowing how it would perform."
After the competition, the CT Full-Chest Agent and DR Full-Body Agent that participated independently by Shukun Technology were awarded the "Unprecedented · Courage Certificate" issued by the conference
According to the results publicly announced on site, the combination of senior doctors and multi-task agents had the highest average score of 85.2 points; the combination of junior doctors and agents scored 83.8 points, and even got the highest score of the whole competition in some cases. The independently working agent scored 66.2 points, which is lower than the scores of reports written independently by doctors and human-machine collaboration groups, but it only took 20 minutes and 32 seconds, the fastest among all groups (the time taken by other groups ranged from 66 to 88 minutes).
The time consumption and scores of 7 groups of participants for writing 10 diagnostic reports (where "junior" and "senior" refer to junior doctors and senior doctors respectively)
A junior participating doctor assigned to "write manually" using the traditional case template admitted frankly that when working without AI tools and relying on traditional templates to write reports, junior doctors often expose their shortcomings in experience, are prone to omissions in the links of lesion identification and sign characterization, and the overall quality of reports fluctuates greatly. At present, this is also the common work situation faced by a large number of primary-level and young radiologists in China.
In the groups that cooperated with multi-task agents, doctors can get the sign list and preliminary draft of the report initially sorted out by AI, which has higher coverage of lesion detection and can reduce the risk of missed detection caused by human eye film reading.
However, new practical problems have emerged subsequently. The content output by the agent will be mixed with some false positive results, and doctors need to go back to the original images one by one for recheck and verification. When the judgment logic given by AI is inconsistent with their own clinical cognition, they need to invest extra time in screening and judgment, which instead prolongs the overall working time for some cases. "At this stage, there is no clear AI operation specification in clinical practice. How should we adopt the AI output content and how to grasp the boundary? This is also a very practical problem in front-line daily work."
At the same time, there is an interesting phenomenon in the competition: many review experts misjudged the source of the report when interpreting the scoring results. For example, one expert concluded that a certain report was completed by a senior doctor using a traditional template because "the writing style and wording are too familiar", but in fact it was written by a senior doctor using a vertical AI model. This means that at this stage, AI-generated reports are getting closer and closer to the expression habits of human doctors in terms of terminology and structure, and the boundary between human and machine is gradually blurring.
The reason why this actual measurement has attracted much attention from enterprises and doctors at the conference site is that it takes place at a key node of technical iteration in the medical imaging industry.
In the past few years, domestic AI medical imaging enterprises have been mainly focusing on vertical models, with one model focusing on a single task. Pulmonary nodules, coronary arteries, cerebral hemorrhage and other corresponding independent products solve the problems of specific lesion detection and quantitative measurement. Many products have obtained medical device registration certificates and have been put into use in hospitals. However, the limitations of vertical models have also been gradually exposed. For example, a CT report in the radiology department needs to cover multiple organs of the whole body and dozens of types of abnormal signs. Doctors need to jump back and forth between multiple AI modules to call tools, which makes the tools fragmented and difficult to fit the complete report workflow of the radiology department.
With the penetration of large language model technology, in the past one or two years, many domestic medical imaging manufacturers have collectively turned to the direction of multimodal and medical agents, trying to integrate scattered single-task models, simulate the complete film reading logic of radiologists, realize one-time image input, complete multi-organ and multi-disease analysis, and directly generate a complete preliminary draft of the report.
During this conference, Shukun Technology also released two clinical agent products, CT Full-Chest and DR Full-Body, which are the products under this round of technology transformation. According to Mao Xinsheng, founder and chairman of Shukun Technology, the company will also "release a medical world model" in the near future.
The general expectation of the industry is to free doctors from repetitive work of sign listing, measurement and preliminary draft writing, so that they can shift their work focus to review and comprehensive judgment. However, when technical concepts are implemented into the real clinical environment, the effect is often different from that in the demonstration environment. This "Human-Machine Competition" is intended to fully present this gap.
Mao Xinsheng believes that the previous generation of vertical models are good at "recognition" tasks. Trained with labeled data, they can complete lesion detection, segmentation and quantification, with clear capability boundaries. Their advantage is strong controllability, while their disadvantage is solidified task capability. For each new type of disease, it is necessary to re-do data labeling and train new models, which makes it difficult to generalize and deal with the endless atypical and rare cases in clinical practice. The change brought by large models is that the models have the capabilities of reasoning, induction and language organization. They no longer only output lesion coordinates and measurement values, but can imitate radiodiagnosis thinking, complete the dismantling of film reading process, sign induction, and organize complete medical texts, with the ability to open up the whole link from image signals to the final report.
To make an analogy between the two generations of technologies, the AI in the past is like a "specialized intern", who can only treat one kind of disease and needs doctors to guide them step by step. Today's AI agent is more like a "chief assistant". After uploading a film, it can triage by itself, label by itself, and write the preliminary draft of the report by itself, and doctors only need to review and sign.
The direct result brought by this change is the reconstruction of the workflow: doctors have changed from "people who write reports" to "people who review reports".
Extending to the product side, in the past, most vertical AI products could only be purchased as single-point auxiliary tools to solve a specific screening demand in the department, and their purchase value was concentrated on reducing missed diagnosis and improving the efficiency of specific diseases. Agent products aim at the complete workflow of the radiology department, with the goal of intervening in the whole process of report generation. Theoretically, they can serve high-load departments in tertiary hospitals at the same time, and also help primary-level hospitals make up for their diagnostic capabilities.
However, this also brings new practical problems: on the one hand, the output content of the agent is longer and the logical chain is more complex, and the error forms are no longer simple missed detection, but there will be reasoning errors and description deviations, which put forward higher requirements for product verification, registration and review. On the other hand, how hospitals should calculate the clinical value of such products and establish supporting quality control and doctor usage specifications, the whole system has not kept up with the changes of technology.
Mao Xinsheng also mentioned that even if technological breakthroughs are achieved, the implementation of medical AI products still relies on the polishing of massive real-world multi-center cases. Models trained only with single-center data are difficult to adapt to the complex and diverse clinical scenarios in hospitals at different levels in China.
In terms of commercialization, the competition logic has also changed. "In the past, everyone competed for who got the certificate first, but now it is who can really make hospitals pay for AI services, continue to use and renew the contract. These are two completely different evaluation systems. Therefore, the real challenge lies in the policy level, because medical AI is essentially a "digital employee", but the current charging system still regards it as ordinary software." Mao Xinsheng believes.
In other words, large AI models and agents have opened up new imagination space for AI medical imaging and solved the fragmentation pain point of traditional vertical models, but technology implementation is not directly equivalent to the realization of clinical value. After going through the popularization stage of single-point tools, today's AI medical imaging has come to the juncture of workflow reconstruction. With technical capabilities taking a big step forward, the supporting clinical specifications, value evaluation and supervision paths are all in the process of simultaneous exploration.