HomeArticle

No more waiting for natural evolution: A Nobel laureate bets 94.6 million on the path that AlphaFold did not take

新智元2026-09-28 09:54
After deciphering 200 million kinds of proteins, humans have begun to create them on their own.

AlphaFold has predicted the structures of more than 200 million proteins in nature.

David Baker wanted to take a different path: instead of predicting what nature has created, he would create things that nature has never made.

One morning in August this year, research assistant Jack Boylan pulled up the experimental results from the previous night.

Next to the DNA sequencer in the lab, a large number of artificially designed DNA sequences that have never evolved in nature are soaked in test tubes. Which one is usable and which one is waste is entirely read and reported at one time by this machine.

In a recent experiment, they put 6 million of such designed sequences into the same test tube for simultaneous testing. Using traditional methods, testing this batch alone would take several years.

The institution where Boylan works is called AI BioDesign, which was officially unveiled in Seattle on September 3, just a few minutes' walk from the headquarters of the Allen Institute.

It is backed by $94.6 million invested by the Paul G. Allen Legacy System, led by 2024 Nobel Prize in Chemistry laureate Baker and genomicist Jay Shendure.

Left: David Baker, 2024 Nobel Prize in Chemistry laureate. Right: Jay Shendure, genomicist. The two co-lead AI BioDesign.

In this era where computing power determines everything, they do not plan to use this nearly 100 million dollars to train a larger model, but to focus on making breakthroughs in that test tube.

Hijack the cellular machinery to achieve massive concurrency in test tubes

Their approach is to build a fast-running closed loop:

The AI puts forward designs, the lab manufactures and measures them, the results are fed back to the model, and the model determines what to manufacture in the next round.

This sounds like the familiar "Design-Build-Measure-Learn" framework, but the real difference lies in how this cycle operates.

There is no automated workshop full of robotic arms here. "The scale does not come from robots, but from parallel processing inside the test tube," said Jesse Gray, Executive Director.

The so-called parallelism inside the test tube, in Shendure's words, is to hijack the cellular assembly line left by evolution to humans, and use it to create and measure millions of molecules.

The specific method is: put millions of designed sequences into the same test tube, let the cells produce each of them, then run them under the same conditions, and finally send them to the sequencer to read at once which ones are usable and which ones are waste.

In this way, when the model learns the "rules of design", it is faced with a huge number of examples, rather than the limited small set produced naturally.

What do these designed molecules look like?

The figure below shows a binding protein previously designed by Baker's lab: the purple and pink parts are designed by AI, and the white part is amylin that it completely wraps around.

Surface rendering of the AI-designed amylin binding protein.

Experiments no longer verify conclusions, but only feed the model

What is more interesting is the KPI of this production line.

The lab is divided into teams of five or six people, each focusing on a specific biological design problem, with a machine learning team on standby for support at any time.

How to calculate the performance of each round of experiments? It is not only about how many usable molecules are produced, but also about how much the model can learn from this round.

The logic of traditional laboratories is to first put forward a hypothesis, and then design experiments to verify it.

AI BioDesign reverses this order: experiments are first "practice exercises" for the model, which are carried out in cycles of "Design-Build-Measure-Learn".

At the end of each round, the machine learning team will select the batch of sequences that the model is most uncertain about and most likely to learn new things from based on the results, and determine what to test in the next round. The tested data is returned to the model, and the updated model outputs the next batch of designs.

The goal is to maximize the amount of information brought back by each experiment, rather than rushing to produce a finished product as soon as possible.

In the past, humans directed the lab bench, but now the lab bench starts to follow the arrangements of AI.

Jack Boylan (left), research assistant at the Allen Institute, and Jesse Gray, Executive Director of AI BioDesign, in front of the DNA sequencer in the lab.

He created proteins that do not exist in nature as early as 2003

This is not a story about "AI creating proteins that do not exist in nature".

Baker achieved this as early as 23 years ago.

In 2003, his team used the Rosetta software to design a protein named Top7: 93 amino acids, with a folding pattern that had no precedent in all known proteins at that time.

Half of the 2024 Nobel Prize in Chemistry was awarded to Baker for computational protein design, and the other half to Demis Hassabis and John Jumper from the AlphaFold team for predicting protein structures.

On October 9, 2024, David Baker received a call from the Nobel Committee.

If AlphaFold solves the problem of "reading": predicting what shape this string of amino acids will fold into, then Baker solves the problem of "writing": what shape I should create to achieve a certain function.

Later, the "reading" half has been highly praised in recent years. AlphaFold2 and other large models of the same period adopt the same strategy: they all feed on the data that humans have already accumulated.

The "writing" half cannot take this shortcut. Things that do not exist in nature cannot be found in the database at all, and can only be created and measured by ourselves.

As a result, it has long been stuck in a realistic quagmire:

AI can generate hundreds of thousands of candidate designs overnight, while a traditional lab is considered highly productive if it completes hundreds of them in a year.

The huge gap between design and verification is the ceiling that this field has long been difficult to break through.

In Baker's words, the real breakthrough of this project is that for the first time the speed of AI begins to match the experimental capabilities of synthetic biology.

One side generates millions of candidate designs a day, and the other side tests millions of them in one test tube. The two gears are finally engaged.

Not chasing general large models, focusing heavily on experimental infrastructure

While everyone is chasing a general large model that can answer all questions and even simulate the entire virtual cell, Baker and AI BioDesign have chosen another path:

Build a number of small models that only focus on specific difficult problems, and then generate a large amount of new data for each problem to train them.

The logic is not hard to understand.

The samples left by natural evolution are only a small subset of solutions it has tried over billions of years. Using this small subset to learn the universal rules of design inherently has huge gaps.

To fill this gap, we have to generate whatever data we lack. And generating data does not rely on adding more graphics cards, but on test tubes, sequencers and lab benches.

In the previous stage, everyone competed for computing power infrastructure, and in the next stage, the competition will be for experimental infrastructure.

What is more counter-intuitive is that even if this bet pays off, all the achievements will be given away for free. Models, datasets, experimental methods, reagents and benchmarks are all fully open to the public.

What if others use these data to develop profitable drugs?

CEO Rui Costa said casually: "If many companies use this data to make the world a better place, that will be our luck."

There are also hidden concerns amid this rapid progress. Mass manufacturing of molecules that do not exist in nature must guard against unexpected ecological impacts and malicious abuse.

There is an unavoidable link in this chain: synthesizing DNA. No matter how many designs you draw, someone has to turn them into physical products in the end.

The industry's screening system is in charge of this part. When you place an order to synthesize a piece of DNA, the system will compare it with the known toxin and pathogen genes, and intercept it if there is a match.

A study published in *Science* in October 2025 shows that the Microsoft team used publicly available protein design tools to rewrite 72 proteins of concern into about 72,000 synthetic homologous sequences, with largely preserved structures and functions but completely different sequences, leading to a large number of missed detections by four mainstream screening software.

The team later cooperated with four synthesis companies to patch the system, and the detection rate has risen significantly.

The red team process of AI protein design by Microsoft team: generate synthetic homologous sequences from proteins of concern, and then send them to screening software for verification.

The problem is that patches can only fix the loopholes that have been discovered. Design capabilities are advancing exponentially, but the gatekeepers are always lagging behind.

Baker's proposition is: all artificially synthesized DNA should be registered and archived, with sequences recorded and the information of the person placing the order documented.

He believes this is a practical anti-abuse threshold.

Showdown in 18 months, opening a new drawer of biology

The bet has been placed.

Costa set a timeline of 18 to 24 months to deliver real progress, and make it open to researchers all over the world within five years.

Baker's vision list includes: therapies for new diseases can be developed in weeks instead of decades, plastics in the ocean can be broken down by enzymes, and biocomputers consume far less power than silicon chips.

These are all just directions for now. Baker himself also admits that in this process, a batch of usable molecules whose working mechanism cannot be clearly explained will emerge first.

We can create them, but we cannot clearly explain why they work.

Charles Darwin once marveled at the diversity of nature with the phrase "endless forms most beautiful".

But in Baker's view, those most beautiful forms are only a small subset of solutions tried over billions of years, and most of the drawers have never been opened.

Now, he is pulling open the drawer with the help of AI, flipping 6 million compartments at a time.

AlphaFold has predicted the structures of more than 200 million proteins, and the path of "understanding nature" has come to an end.

Baker bets that the next bottleneck is not the model, but the data: a production line that can continuously generate new data.

References:

https://www.wired.com/story/nobel-prize-protein-design-now-using-ai-to-create-molecules-beyond-nature/?utm_source=chatgpt.com

https://alleninstitute.org/news/ai-biodesign-accelerator-combines-experimental-biology-and-artificial-intelligence-to-learn-natures-design-rules?utm_source=chatgpt.com

https://www.nobelprize.org/prizes/chemistry/2024/press-release/b/?utm_source=chatgpt.com

This article is from the WeChat official account "AI_era" (ID: AI_era), written by Yuan Yu, and authorized for release by 36Kr.