HomeArticle

Anthropic and ByteDance step into AI drug discovery: the "next Coding" is stuck in data bottlenecks

明亮公司2026-09-21 15:32
Large model companies and AI pharmaceutical companies are competing for the same type of asset — experimental data that can be used to train the model repeatedly.

According to Reuters' information on September 18, Anthropic has been operating a wet lab in the San Francisco Bay Area. This lab is not dedicated exclusively to drug discovery; its function is to expose Claude to cells, reagents and experimental equipment, and test the model's ability to direct robots to complete biological experiments. Generally speaking, "dry lab" refers to using models to design and screen candidate molecules, while "wet lab" synthesizes, tests and verifies these results in a real laboratory.

Two days ago, Anthropic and Novo Nordisk announced a partnership. According to the announcements from both parties, Claude Science will be applied in drug discovery and R&D software. Previously, Anthropic also cooperated with Bristol Myers Squibb to deploy Claude in research, clinical development, production and commercial departments. In April this year, Anthropic acquired AI biotech company Coefficient Bio for approximately 400 million US dollars.

In the very same week, Anew Labs, a new lab spun off from ByteDance, completed a $290 million first round of financing with a post-money valuation of about $1.5 billion. ByteDance still retains a 56% stake, and the team, algorithm platform and R&D pipelines under development have been injected into the new company.

These moves continue the path of AI drug discovery over the past few years, and have also brought new participants. In the past, AI pharmaceutical companies developed vertical models, and then cooperated with pharmaceutical companies or CROs to complete experiments. Nowadays, large model companies have begun to build their own laboratories, trying to directly grasp the feedback loop between models and experiments.

Efficiency improvement of new drug R&D is still concentrated in the early stage

Traditional drug R&D is generally summarized as the "double ten law": it often takes more than ten years for a new drug from discovery to marketing, with an investment of about 1 billion US dollars (some statistical calibers exceed 2 billion US dollars), and the final success rate is about 10%.

Citi further broke down the early discovery link in its AI pharmaceutical research report in September 2026. To find a preclinical candidate through traditional methods, it is usually necessary to synthesize and test about 5,000 molecules, which takes 4 to 6 years. The AI path may reduce the test scale to about 330 molecules, and the time to around 17 months.

However, this set of numbers only covers the process from design to preclinical candidates. According to Citi's calculation, the complete R&D cycle of a first-in-class drug is about 13 years, and AI is expected to shorten it to 8 years, but Phase I-III clinical trials still take 5 to 7 years.

In other words, AI first compresses the time of search and experimental iteration, and clinical trials will not be shortened by the same magnitude.

Insilico Medicine also provided a case. According to company data cited by Orient Securities, Insilico Medicine (3696.HK) took 18 months for a candidate drug for idiopathic pulmonary fibrosis from target confirmation to entering the clinical stage, and the discovery phase cost about 2.6 million US dollars. In contrast, the traditional path takes an average of about 4.5 years, and the cost usually ranges from tens of millions of dollars to hundreds of millions of dollars.

The hit rate also needs to be distinguished by calibers. The protein design experiment announced by Anthropic in August this year shows that Claude generated 1320 designs for 15 targets, 354 of which passed the binding experiment, with a total hit rate of about 26.8%, and experimental hits were obtained for 14 targets. It should be pointed out that this number measures whether the generated protein can bind to the target, rather than the probability that the candidate drug enters the clinic or is finally marketed.

Industry statistics cited in an expert exchange minutes (hereinafter referred to as "Expert A") show that the success rates of stages from target to hit (hit compound), hit to hit-to-lead (lead compound), and lead compound optimization are about 80%, 75% and 85% respectively. After entering Phase I, Phase II and Phase III, the stage success rates are about 54%, 34% and 70%, and the overall success rate of the whole process is about 11% to 14%. Among them, Phase II clinical trials are still the main loss link. This means that the increase in early hit rate does not equal the synchronous increase in clinical success rate.

Therefore, Citi summarizes the current judgment basis as "faster is proven, better is not" — operational efficiency has been verified, while the quality of single candidate drug and clinical success rate are still uncertain. From this perspective, AI drug discovery can realize the high-frequency search and feedback loop in Coding, but the "test procedure" of drugs are cells, animals and human bodies, and the feedback cycle is extended from seconds to weeks, months and years.

Route differences between large model companies, AI pharmaceutical companies and large pharmaceutical companies

The starting point of AI pharmaceutical companies is a single problem in drug R&D. They first do computer-aided drug design, molecular docking or structure prediction, and then add deep learning and generative models. According to the sorting of UBS and Expert A, the core capabilities of such companies are usually domain models, medicinal chemistry experience, dry and wet experiment closed loop and self-owned pipelines. This means that their models serve a certain target or a certain type of molecule, and the tool chain is also built around specific tasks.

Large model companies, on the other hand, pursue the general capabilities of models more.

In another brokerage interview minutes, "Expert B" summarized this route as "a single basic model plus a vertical tool chain". The large model does not retrain a set of small models for each link, but is responsible for disassembling tasks, selecting tools and sorting out results, and then calling structure prediction, molecular design and database tools. Claude Science integrates more than 60 scientific research skills and connectors, which is exactly the product form of this kind of "scientific workbench". It is more like using the capabilities of large model + Agent to complete the "dry-wet closed loop".

The efficiency goals of the two types of companies are also different.

Previously, the goal of AI pharmaceutical companies was to compress the cycle of projects from project initiation to candidate substances from 4.5 years to 12 to 18 months. The above expert believes that the goal set by large model companies is closer to a single cycle: to shorten the design of a single target from weeks or months to seconds or minutes, and then compress one round of wet experiment iteration from 6 months to less than 14 days.

The difference between the two routes is also reflected in the data sources. It can even be said that data is the most core asset in the AI pharmaceutical field.

In order to get through the whole process, AI pharmaceutical companies usually build part of their own wet experiment capabilities, or purchase verification services from pharmaceutical companies and CROs around specific pipelines, and data accumulates along with the projects. In contrast, large model companies pay more "attention" to the general workbench, often building a general platform first, and then connecting their own laboratories with external service providers such as Twist Bioscience (TWST.US), Adaptyv Bio, GenScript (01548.HK) and so on.

Specifically in terms of data, what large model companies want to obtain is not only the data of one pipeline, but also the experimental records of success and failure (especially the data of failure). The model needs to judge which tools are effective and which designs fail from these records, and then adjust the next round of experiments.

In contrast, the data of traditional AI pharmaceutical companies is closer to specific drug assets, what large model companies pursue is a data infrastructure that can be used across targets and tasks.

However, general models cannot replace pharmaceutical R&D experience.

Research reports from Western Securities and other institutions mention that large model companies lack proprietary clinical data, failure data and interdisciplinary talents, and the relationship between model output and clinical endpoints is still weak. The previous generation of AI pharmaceutical companies are closer to drug pipelines, while large model companies are closer to R&D infrastructure. The former needs to supplement models and capital, while the latter needs to supplement experiments, medicinal chemistry and clinical transformation capabilities.

Traditional pharmaceutical companies are in the middle of the two. Pharmaceutical companies have preclinical, clinical and production data, and also know how to advance candidate drugs to IND and clinical trials.

However, Expert A pointed out that large pharmaceutical companies usually organize R&D by pipelines, and different teams may use different software, experimental processes and recording methods, historical data needs to be cleaned and re-annotated before it can be used for model training. Although pharmaceutical companies have the most data, they may not have formed a unified data infrastructure.

Capital investment direction, data and data infrastructure

As models enter laboratories, R&D budgets have also begun to be redistributed.

Expert B said that Anthropic plans to invest 3%-4% of its annual recurring revenue in AI drug discovery, of which 60%-65% is used for model training and virtual simulation, and 35%-40% is used for biotech talents, wet lab assets and CRO cooperation. The expert also said that about 60%-70% of the R&D budget in the traditional pharmaceutical industry flows to wet experiments.

If so, this allocation ratio reflects the different cost structures between large model companies and pharmaceutical companies. Pharmaceutical companies invest more funds in experiments, clinical trials and production, while large model companies first invest in computing power and basic models, and then use experimental budgets to purchase high-quality feedback.

After the model design speed is improved, the number of candidate solutions entering the laboratory will increase accordingly. Therefore, the proportion of wet experiments in the R&D budget may decrease, but the absolute expenditure may not necessarily decrease. This also explains why large model companies are building their own laboratories while expanding cooperation with CROs and life science service companies.

Laboratories also determine what data model companies can get. When the model only provides software to pharmaceutical companies, the experimental results usually remain inside the clients. Public papers generally only record successful results, while failure data is rarely made public. However, molecules that fail to be synthesized, cell experiments that do not show the expected phenotype, and changes in experimental conditions will all affect the judgment of the next round of models.

In this context, Anthropic is trying to connect tools, equipment and experimental data. According to company information, the Model Hardware Standard, which opened research preview at the end of August, can connect microscopes, liquid workstations, robotic arms and plate readers. Genentech has used it to complete protein concentration detection with the collaboration of multiple devices, and the team from the University of Washington has accelerated a dose-response experiment by about three times.

At the same time, Anthropic does not keep all experiments in-house. According to its protein design research, Adaptyv Bio and Twist Bioscience undertake external verification. Large-scale synthesis, animal experiments and clinical research can still be handed over to CROs, CDMOs and pharmaceutical companies.

In China, ByteDance invested in Huixiang Technology (the main entity is "Shanghai Huixiang Information Technology Co., Ltd.") as early as 2022. According to the website of Huixiang Technology, it has successfully helped Insilico Medicine complete the construction of the large-scale 6th generation intelligent robot laboratory in China in Suzhou.

The function of self-built laboratories is to master key experiments, equipment scheduling and data standards, rather than replicating a complete pharmaceutical factory. Some insiders even directly put forward a "radical view": "Large AI companies will start to acquire large pharmaceutical companies. Not for drugs, but for data."

Tweet from Simon Maechling, Innovation Manager of Bayer Crop Science (Source: X)

Simon Maechling, Innovation Manager of Bayer Crop Science, believes that the current bottleneck (of AI drug discovery) no longer lies in the model itself, but in the ability to obtain unique biological data, experimental conditions, verification systems and real-world clinical results. In the past, it was biotech companies that purchased AI technologies; but now, AI companies may begin to acquire biotech assets.

"This is about the transfer of industry discourse power and power pattern." He posted on X, "In the field of life sciences, the most valuable assets may no longer be specific molecular structures, but that proprietary biological verification system that can teach AI 'what works and what doesn't'."

Pipelines are still the final basis for valuation

From the perspective of commercialization paths, the business models of AI pharmaceutical companies are roughly divided into AI-CRO, software tools, and platform plus self-developed pipelines. Expert A believes that simply providing AI-CRO services is difficult to cover the investment of models and experiments, and To B software faces sales and delivery costs. In contrast, self-developed pipelines can realize value through upfront license fees, R&D milestones and sales sharing.

Whether the platform can continuously generate candidate drugs determines how much premium the market is willing to give. This is the expert's judgment on the business model, not a verified industry conclusion.

After large model companies enter the laboratory, they may also extend along the same direction. If only models are provided, the revenue mainly comes from software licenses and project services. After participating in experiments and pipelines, the company may also obtain drug licensing revenue, but the return cycle and failure risk will both increase.

Reuters' information shows that Anthropic is not currently conducting clinical trials and says it will not directly compete with pharmaceutical companies and biotechs that are responsible for drug marketing. This means that its recent focus is still on preclinical research and R&D infrastructure. It can design experiments, generate data and screen candidates, but clinical development and commercialization still need to rely on pharmaceutical companies.

In China, ByteDance's Anew Labs does not only retain the structure prediction and molecular design platform, the company is also advancing its pipelines.

R&D pipelines of Anew Labs (Source: Anew Labs official website)

XtalPi (2228.HK) combines models with robotic experiments, uses automated workstations to complete chemical synthesis and detection, and then returns the results to the models.

Insilico Medicine provides a case of platform extending to drug assets. The company does not stay at software licensing, but uses its own platform to advance pipelines. According to company information, its idiopathic pulmonary fibrosis candidate drug Rentosertib has entered Phase III clinical trials.

Citi's research report believes that the Phase III result of this candidate drug from Insilico Medicine is regarded as an important verification node for AI drug discovery. If the clinical result is positive, the market's pricing basis for AI drug discovery may shift from R&D efficiency to drug assets; if it fails, the valuation of the platform will still be mainly based on service revenue, cooperation transactions and preclinical outputs.

For investors, the final observed indicators are similar, including the number of candidate compounds, IND applications, assets entering clinical trials, upfront cooperation payments and milestones. Models and experimental platforms can increase the speed of pipeline output, and pipelines are still the most direct verification of platform capabilities.

As mentioned earlier by "Suchbright Company", from the perspective of valuation logic, Insilico Medicine is more like an option portfolio composed of "core pipelines + early assets", the revenue mainly comes from upfront license fees and milestones, but value realization still depends on clinical data. XtalPi's revenue grows with the number of customers and experimental throughput, with higher operational certainty, but linear service revenue also limits the valuation elasticity.

This also creates another difference between AI drug discovery and AI Coding. Programming products can be priced by