Shortly after the AlphaFold team was disbanded, Claude ventured into the protein design field and successfully hit 14 out of 15 targets.
Claude has forayed into protein design and delivered an impressive performance in its very first attempt.
Just now, Anthropic released a new blog post that detailed the whole process of its Mythos and Opus models advancing into the AI4S track side by side.
Scrolling quickly to the results page, you will find the outcomes are indeed outstanding:
This time Claude designed 1320 protein candidates from scratch, which were then sent to external laboratories for production and verification.
Surprisingly, 354 of them successfully bound to the targets, 14 out of 15 targets were hit, with the highest hit rate reaching 35.1%.
It is worth noting that the typical hit rate of current protein design projects is usually only 10%~15%.
More notably, it did not stop its work after finishing the design of new molecules.
Facing the raw NMR and LC-MS files provided by the laboratory, Claude only read a two-sentence task instruction, and completed the processing in 23 minutes and 19 minutes respectively.
The final measured sample purity was 96.4%, while the result given by the laboratory was 96.33%.
Only a difference of 0.07 percentage points.
That means Claude's results are highly consistent with the laboratory's results.
Not Prediction, But De Novo Protein Design
Speaking of proteins, many people may immediately think of the AlphaFold project from Google DeepMind.
To clarify first, AlphaFold is used to predict protein structures, and its core function is to solve the problem of "given an amino acid sequence, what structure will it fold into".
But the task Claude received this time is completely different:
The target protein is presented here, please design a new non-existent protein from scratch that can bind to it precisely.
The target Claude needs to design is called minibinder, that is, miniature binding protein.
According to public available information, minibinder is an artificially designed small protein, usually only 50~70 amino acids in length, with a compact structure and single function.
Its task is very clear, that is, to bind to a specific site on the target protein with extremely high accuracy and intensity.
Do not underestimate this "binding" action. Many drugs take effect relying on this precise binding at the molecular level —
Bind to the bad protein to deactivate it, bind to the good protein to activate it, or directly pull the drug molecule to the lesion site for release.
The problem is that this work is extremely labor-intensive.
Protein designers not only need to schedule a series of professional models, but also repeatedly generate, optimize and screen candidates, and the whole process relies heavily on experience.
Each target requires experts to spend weeks or even months to screen out a few truly effective candidates from a large number of candidate molecules.
This time Anthropic simply handed over the entire toolbox to Claude.
The models participating in the experiment are Mythos Preview and Opus 4.8, which can access papers, network resources and multiple professional protein models, and are also allocated with sufficient GPU computing power.
After humans input the initial task, Claude ran the process autonomously. The final overall result is:
Claude hit 14 out of 15 targets, of which at least 6 produced high-affinity binding proteins, and at least 4 matched or exceeded the previous best results.
There are two specific operation modes:
Multi-target mode: Let Claude process multiple targets simultaneously in a single 48-hour task, and can call up to 12500 H100 GPU hours.
Single-target mode: Focus on only one target at a time, all tasks run in parallel for 24 hours, and each target can use up to 2500 H100 GPU hours.
Experiments prove that focusing on one task is indeed better than multitasking.
In the multi-target mode, the overall hit rates of Mythos Preview and Opus 4.8 are 26.7% and 22.6% respectively.
After switching to the single-target mode, the hit rate of Mythos Preview directly rose to 35.1%.
Its performance on some targets even surpassed the previous best results achieved by humans.
Take RBX1 as an example, it is a target involved in regulating protein degradation. In the previous design competition held by Adaptyv Bio, the overall hit rate of all participants was only 3.7%, but when Mythos Preview challenged it alone, the hit rate reached 40%.
Not only is it easier to hit the target, but the top-ranked design of Claude is also a high-affinity binding protein, whose binding intensity exceeds the championship design selected from 245 participating designs in that competition.
There is also a more difficult target TNFα, which is an inflammatory signaling protein released by the immune system.
Since the appropriate binding site is hidden in the groove formed by two proteins, many expert teams have encountered setbacks on this target before.
As a result, Mythos Preview failed to solve this problem, but Opus 4.8 successfully conquered it:
Opus 4.8 not only designed multiple effective binding proteins, some of them can also bind to TNFα of humans, cynomolgus monkeys and mice at the same time.
This cross-species binding ability is very critical, because when entering animal experiments in the future, there is no need to redesign for different species.
However, a problem arises:
Why did the more powerful Mythos Preview fail while Opus 4.8 succeeded?
Anthropic admitted that it has not figured out the reason yet.
It can only be said that the scientific research capabilities of large models are still uneven.
When Claude continued to challenge the more difficult-to-design β-sheet structure, it successfully produced 15 effective binding proteins on 6 targets.
But when switching to the maltose-binding protein MBP, Claude failed completely —
None of the 90 designs was confirmed to be successful, and only one showed a weak signal.
Therefore, it is still too early to say that "inputting a target, Claude can automatically output a new drug", but Claude has proved that:
It can autonomously schedule professional models and complete the design process that used to require experts to spend a lot of time organizing and screening.
This is already a huge step forward.
After designing the new molecules, Claude Started processing experimental data
However, Anthropic's experiments did not stop here.
In addition to testing whether Claude can design new things, they also wanted to verify:
Facing the already synthesized compounds, can Claude independently read the experimental data, and judge what the compound is and its purity.
The model participating in this part of the experiment is the widely available Claude Opus 5.
The tasks come from two very conventional but labor-intensive jobs in analytical chemistry:
Nuclear Magnetic Resonance Spectroscopy (NMR): Observe the signals of hydrogen atoms in the molecule to confirm "whether the synthesized product is the target molecule".
Liquid Chromatography-Mass Spectrometry (LC-MS): Separate different components in the sample, then measure their molecular weight and content to judge "whether the sample is pure and what else is in it".
This kind of work requires chemists to manually process raw data in proprietary formats, perform calibration, peak picking, integration, and verification step by step, and finally write a report.
There are many steps and the work relies heavily on experience, and a slight deviation may lead to wrong conclusions.
This time, Anthropic only gave Claude the raw files returned by the laboratory, as well as a two-sentence task instruction.
There was no vendor software, and no operator to guide it (p.s. the relevant data and prompts have been open sourced).
As a result, Claude completed the NMR and LC-MS analysis in 23 minutes and 19 minutes respectively.
For the NMR part, Claude sorted out 18 signal peaks, and calculated the number of hydrogen atoms corresponding to each peak, and the error of the results did not exceed 0.08 hydrogen atoms compared with the laboratory results.
It also found that 4 of the broad peaks may come from hydrogen atoms connected to nitrogen or oxygen, so it suggested adding deuterated water for further verification.
This is a common elimination method. After adding deuterated water, the signal of some specific hydrogen atoms will weaken or disappear, so chemists can judge the molecular structure accordingly.
Coincidentally, the laboratory also carried out the same verification later.
After verification, Claude initially thought that all 4 signals disappeared, but after self-checking, it found the error and actively corrected it to only 2 signals disappeared, and the final conclusion was consistent with that of the laboratory.
The LC-MS part is more straightforward.
Facing the vendor file without public format description, Claude first figured out how the data was encoded, and then reproduced all 2664 scan records to confirm that the file was read correctly.
After that, it output the chromatogram, mass spectrum, purity table and molecular weight, and even wrote a set of reusable parsing code.
Finally, Claude measured the sample purity as 96.4%, and the laboratory result was 96.33%.
Only a difference of 0.07 percentage points.
Anthropic wrote in the conclusion:
The two experiments together point to one thing: AI is reducing the professional threshold, cost and time required for life science research. In chemical analysis, Claude has begun to take over the processing work that used to rely on manual labor; in protein design, Claude can complete the design of binders end-to-end with very little input, and some results match or even exceed the previous best designs.
Although it is still far from real new drug research and development, at least AI has shown us the possibility of accelerating this process.
There is also an easter egg at the end
At the end of the blog, Anthropic also emphasized a key point for scientists.
This protein experiment used Mythos Preview and Opus 4.8. As for the more capable Fable 5, it is currently not open to ordinary users for such life science tasks.
The reason is straightforward: the model that can design drug proteins may also be used to carry out dangerous biological research.
However, Anthropic has announced that it is preparing an exclusive access program for scientists, and more information will be released later.
Full technical report: https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf
Open source address of relevant prompts and data: https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/tree/main
Reference links:
[1]https://www.anthropic.com/research/Claude-accelerates-protein-design
[2]https://x.com/AnthropicAI/status/2089842387845804246
This article is from the WeChat official account "QbitAI", author: Yishui, published with authorization from 36Kr.