HomeArticle

The biggest lie of AI4S is the misconception that with large models, people can solve all problems in new drug R&D.

晓曦2026-09-29 16:00
Understanding scientific laws is always the overarching premise.

In late September, at the 11th Huawei Connect conference, the Staris AI Computing team, which was established just over a year ago, was exceptionally busy.

Huawei Connect has always served as the annual window for the Ascend ecosystem to make a centralized appearance, as well as a key vantage point for the industry to explore AI infrastructure and the implementation of various solutions. Over the three days, the core team of Staris AI Computing was fully deployed:

At the "Deeply Rooted in Pharmaceuticals, Empowering Intelligence to Take Root" forum, COO Ding Pengcheng introduced how AsterFire Enterprise, the company's enterprise-level AI4S agent platform, helps R&D departments of pharmaceutical companies improve efficiency and accelerate wet-dry iteration.

COO Ding Pengcheng sharing on site at the AsterFire Enterprise launch (Source: Staris AI Computing)

CEO Chen Zhenyu attended the "Intelligent Transformation of Manufacturing, Innovation to Break Through" Global Manufacturing Summit, participated in the release of industry solutions, and received the "Rising Star Partnership Award" on behalf of the company at the "Manufacturing and Large Enterprise Partner Night".

Chief Scientist Xia Yijie focused on the developer session, launching SPONGE Mokda, a molecular science intelligent infrastructure platform. As envisioned by the team, after combining self-developed scientific algorithms with domestic computing power base, the efficiency of experimental decision-making and the utilization rate of R&D resources are expected to reach a new level...

CEO Chen Zhenyu receiving the "Rising Star Partnership Award" on behalf of the enterprise (Source: Staris AI Computing)

Stepping on the AI4S track that is a popular trend in the primary market, Staris AI Computing has deliberately kept a low profile before this conference. In 2025, after being incubated and established from the research group of Professor Gao Yiqin of Peking University, the team devoted almost all its energy to product polishing and project implementation. Ding Pengcheng judges that the capability of AI agents has reached "the critical point of systematically solving various scientific problems", but the real industrial value ultimately needs to be verified in real R&D scenarios. Therefore, rapid practical application is the best strategy.

More than a year later, the team has launched a complete product line covering the whole process from "assisting individual research, to team collaboration and enterprise-level R&D asset precipitation". Up to now, the personal version of AsterFire Go has accumulated tens of thousands of active users; the enterprise version of AsterFire Enterprise has been implemented in dozens of benchmark enterprises, covering more than 100 R&D scenarios in the biomedicine, petrochemical, new materials and other industries, with partners including Huawei Pharmaceutical Corps, Sinochem, Sugon and many other enterprises.

Ding Pengcheng believes that Staris AI Computing has gained initial recognition in the industry quickly because from the very beginning of its strategic orientation, the company has not confined itself to the concept of AI4S tool development, but positioned itself as a scientific computing provider that builds the underlying system to support scientific research innovation (from problem definition to verification completion). "At this dimension, we will actively embrace all technologies and tools that can promote scientific progress, not limited to AI4S."

This may sound like a difference in expression order, but behind it lies a completely different understanding of the industry.

1. The booming AI4S, where is the bottleneck

In the past few years, the excitement brought by AI4S to the industry has been extremely real. Large models have continuously lowered the threshold for work such as molecular generation, protein design, and material prediction, but after the excitement fades, a more realistic problem emerges: so many solutions are generated, what comes next?

Take the new drug R&D scenario as an example. When candidate drug molecules enter the experimental verification stage, it often means higher cost and longer cycle. Therefore, the key to determining whether a project can move forward is not how many ideas there are, but whether it can quickly judge "which ideas are worth verifying". It is like building a house: AI is a highly inspired "designer" that can draw many sketches in a short time, but cannot tell which sketch is usable.

This is one of the core bottlenecks of current AI4S tools: generative AI has greatly reduced the cost of proposing hypotheses, but has not simultaneously reduced the cost of proving hypotheses.

Scientific computing is exactly the key to solving this problem. Simply put, scientific computing uses computers to simulate the physical and chemical laws of the real world, and "runs" the process first in the virtual space. For example, using molecular dynamics to track the binding process of proteins and drug molecules, using free energy calculation to accurately measure the binding strength, etc. These are methods that have been used in scientific research for many years, and many mature commercial companies have grown up based on them. For example, Schrödinger, an overseas scientific computing software company, has gained a firm foothold in the pharmaceutical industry relying on this set of professional tools.

However, in the current AI4S boom, this matter is easily overlooked. After all, "generating 100 molecules" sounds much more attractive than "verifying 10 molecules".

But from the perspective of industrial implementation, the value of the latter is no less than that of the former. First of all, the two have their own strengths: AI is "divergent", while scientific computing is more like "convergent". Large models are good at finding correlations from massive data and quickly proposing candidates in the huge molecular space. But essentially, the answers they give are the most probable solutions in statistics, and do not guarantee that they are valid in physical sense. Scientific computing uses the first principle of physics to simulate and verify, and is responsible for constraining and screening the answers.

Especially when experimental data is scarce, scientific computing will become an important supplementary source of information. The development logic of scientific research tools is different from that of conventional AI office products, and it is impossible to accumulate data through rapid iteration and unlimited scaling. High-quality experimental data is expensive, scarce and scattered, and under the influence of the traditional scientific research assessment system, a large number of failed results are not even completely preserved. When the amount of data is not enough to support AI models, methods such as molecular dynamics that do not fully rely on data often become additional sources of information.

In addition, scientific computing can better explain "why". In many cases, the thinking process of AI is like a black box: it may draw a conclusion that "a certain molecule may be effective", but cannot explain why it is effective. The results of scientific computing have clear physical meanings. For example, in a certain scientific research problem, what is the key force, how does the conformational change occur? This information is critical for scientists to understand mechanisms, optimize molecules and design experiments.

"Based on this, we believe that the key to competition in the next stage of scientific research scenarios is to connect problem definition, candidate generation, scientific computing, experimental verification and feedback iteration into a closed loop, on the premise of fitting the actual process of enterprise R&D decision-making and execution better." Ding Pengcheng explained.

This is the original intention why the team attaches so much importance to scientific computing capability, and it is also the starting point of Staris AI Computing.

2. R&D pipeline that understands users

According to the introduction, Professor Gao Yiqin's team has long been engaged in the research of scientific problems in the fields of computational chemistry and computational biology, and has accumulated a lot of experience in algorithm and model development at the scientific tool level. "What the research group has been doing all along is to combine scientific computing and deep learning to solve problems in the scientific field."

Before the outbreak of large models, the threshold for implementing this work was very high. Ding Pengcheng introduced that in the past, when carrying out horizontal project cooperation with enterprises, the most time-consuming part was actually sorting out the technical requirements clearly. If this part was not done well, the overall efficiency would be very low.

But the capability of large models is expected to change this. "In many cases, once a scientific research problem is correctly defined, it is actually half solved. Therefore, if we can understand user needs through natural language interaction, convert user demands into scientific problems, and then break down the scientific problems into workflows that can be executed step by step by tools and algorithms, we may have the opportunity to solve this pain point systematically."

The entrepreneurship started from this point. However, the road was not smooth at the very beginning.

In May 2025, the Staris AI Computing team released the first generation of GaliLeo scientific computing platform, which integrated many independent and professional computing tools. But the problem was that ordinary users logging into the platform, facing a row of unfamiliar software icons, often did not know which one to use at all. "It was very much like a group of scientific researchers who made things they thought were great, but users did not know how to use them."

This experience made the team deeply realize the importance of "understanding users".

For example, to solve the problem that "users do not know which tool to choose", the team designed Meta-Agent, which is equivalent to a front-end triage desk. After users input their own needs, Meta-Agent can automatically guide them to the corresponding tools and push adapted parameters; for another example, many users are not familiar with subdivided scientific theories and tools, and are not good at sorting out demands into standardized computing problems. Therefore, the team added a "problem definition sorting and guidance module", which converts users' vague demands into standard scientific computing problems through simple conversations.

Ding Pengcheng mentioned that these modifications have now been integrated into the AsterFire series of products. "This series follows the development path of scientific research agents, and its core goal is to embed our accumulated scientific technologies and tools into the real R&D processes in academia and industry. However, scientific research users and enterprise users have different focuses: the former focuses on cutting-edge scientific exploration and paper writing; the latter is highly pragmatic, pursuing efficiency improvement and obtaining implementable solutions. Considering this, we made distinctions at the user group level and launched two product lines, AsterFire Go and AsterFire Enterprise."

Among them, AsterFire Go mainly serves individual researchers and lightweight development teams. Users describe the research objectives in natural language, and the system will complete the whole process of information research, task splitting, tool matching, code execution, computing power calling and result analysis. "It lowers the threshold for individuals to use professional computing tools, and also undertakes the role of accumulating workflows and cultivating user habits."

AsterFire Enterprise introduces the same set of capabilities into enterprises. It connects internal enterprise data, literature, experimental records, models and other resources into the same project space. After R&D personnel initiate a task, the system not only returns the results, but also completely saves the data sources, calculation parameters and operation records, so as to clearly trace the basis of each step.

The last product, SPONGE Mokda, is a set of molecular computing and design workstation for life science R&D, covering molecular modeling, docking, dynamics simulation, etc., with multiple sets of self-developed algorithms integrated.

Chief Scientist Xia Yijie sharing on site at the SPONGE Mokda launch (Source: Staris AI Computing)

"It is actually targeting the professional scientific computing market that has long been occupied by mature international commercial software. Therefore, our self-developed algorithms need to be comparable to or even surpass similar overseas software to prove our value. Take the self-developed SPONGE-FEP as an example. FEP (Free Energy Perturbation) is one of the most rigorous binding strength calculation methods in new drug R&D. It calculates the relative change of the binding strength between drug molecules and target proteins through molecular dynamics simulation. This indicator directly determines whether candidate molecules are worth being sent to the laboratory, and it is a rigid demand of the industry. The maturity of the SPONGE-FEP algorithm means that the company has the foundation to compete with mature international software in underlying computing capabilities." Ding Pengcheng said.

The enzyme evolution project cooperated by Staris AI Computing and a CDMO company can reflect the value of this combination to a certain extent. In the past, after this CDMO company received a protein optimization demand, experts would select mutation sites by referring to literature and personal experience, carry out experiments and iterate. But the effect of manual iteration is usually obvious in the first three rounds, and after the fifth and sixth rounds, it will reach the performance bottleneck.

"In the project they handed over to us, the activity was increased by about 5-6 times after manual iteration, which did not meet the customer's requirements. After we took over, we encoded the results of each round of experiments into information that the model could use, and then predicted the next batch of candidates worth trying. Finally, AI iterated three more rounds after human experts reached the bottleneck, and the activity was increased by 4-5 times on the basis of manual results; compared with the initial version of the project, the activity was increased by more than 20 times in total."

Interestingly, the team also conducted a set of controlled experiments. Assuming that there is no human accumulated data in the early stage and AI is used directly, what will the effect be? The answer is: the activity was increased by 16.5 times in the first round; after three months of continuous iteration, the overall improvement reached 21 times.

"This project is still ongoing, and we hope to continue testing the performance limit of the AI method." Ding Pengcheng introduced.

3. In the global AI4S race, what is the moat of Staris AI Computing?

At present, AI4S has become a global technology competition track. Overseas, Google DeepMind continues to make breakthroughs in biomolecular structure prediction with the AlphaFold series; NVIDIA seizes the AI4S computing infrastructure market through BioNeMo. In China, various players have gradually realized that relying only on the tool calling capability of agents may achieve short-term commercial implementation, but it is difficult to form long-term barriers. Therefore, they all begin to emphasize evolving towards a complete R&D closed loop, expecting to connect AI Agent, professional models and automated experimental robots.

In such a competitive landscape, what exactly is the uniqueness of Staris AI Computing?

At present, teams focusing on the field of scientific research agents are roughly divided into two categories: one type has an AI background, and tends to pursue model intelligence and general capabilities; the other type has a science background, and pays more attention to whether a specific scientific problem can be accurately described and truly solved.

Staris AI Computing obviously belongs to the latter. Ding Pengcheng believes that the team's previous focus on complex scientific problem research brings value that is not just a specific algorithm or model, but a kind of judgment. This judgment is reflected in many details. For example, when doing data augmentation, people with an AI background will think of general methods such as translation and cropping; while people with a scientific research background will design solutions starting from the translation invariance and symmetry of the system. The two may point to the same operation, but the logical depth behind them is completely different.

"After receiving an enterprise demand, we not only care about what scientific problem it corresponds to, to what extent the existing methods can solve it, and how to verify the results, but also need to combine the actual business model of the enterprise to implement and solve the problem. Only in this way can our capabilities be truly reflected in products and services."

The second advantage lies in the engineering capability of turning scientific capabilities into enterprise products. According to the information provided by Staris AI Computing, the team includes both scientists from universities and senior engineers from leading companies such as Microsoft, ByteDance and Alibaba. At present, the proportion of R&D and product engineering personnel in the team is roughly 50/50. The scientific team ensures that the calculation methods are professional enough and the results are reliable enough, and the engineering team is responsible for packaging scientific research achievements into deployable, traceable and reusable enterprise-level products. Both are indispensable.

The fact that Huawei Pharmaceutical Corps chose to cooperate with Staris AI Computing is, to some extent, a recognition of this comprehensive capability. Huawei Pharmaceutical Corps mainly serves the digital and intelligent upgrading of large pharmaceutical enterprises, and hopes to find partners with "excellent technology and the ability to solve real problems in practice". "At present, the solution jointly launched by us and Huawei mainly focuses on preclinical R&D scenarios, and will be implemented in pharmaceutical enterprises soon. We hope to build benchmark customer cases through this to prove our strength." Ding Pengcheng revealed.

From the very beginning of entrepreneurship, the Staris AI Computing team has taken "accelerating the progress of human