HomeArticle

Just now, Google DeepMind has cracked the human "Book of Life", completing full calculation of all 9 billion gene mutations.

新智元2026-09-09 15:00
Google DeepMind releases AlphaGenome Atlas, constructing a genome-wide variant prediction map.

This is absolutely stunning!

Just now, Google DeepMind officially released its major "Alpha" achievement — the AlphaGenome Atlas.

It is a searchable predictive database empowered by cutting-edge AI that covers the entire genome.

It has completed an unprecedented project in the history of life sciences —

It has calculated all 9 billion potential "single-base mutations" in the human genome, without missing a single one.

This marks another full-scale mapping of the human life code by AI after AlphaFold mapped the "protein universe".

What does 9 billion mean?

The human genome has approximately 3 billion base sites, with 3 types of variations at each site, and the Atlas has calculated every single one of them without omission.

In the past, it was almost impossible to bring all these 9 billion mutations into the laboratory for research.

More impressively, for each mutation, it also provides approximately 27,000 predictive indicators to determine how it affects gene expression, RNA splicing, chromatin accessibility, and the 3D spatial conformation of DNA, covering hundreds of cell types across humans and mice.

9 billion multiplied by 27,000 results in over 240 trillion predictive values, which add up to a total size of 1PB when fully packaged.

The volume of the AlphaGenome Atlas is 30 times that of the AlphaFold database.

Demis Hassabis could not help but exclaim, "AlphaFold mapped the protein universe, and now the AlphaGenome Atlas begins to map the human genome."

The most critical point is that this atlas has almost zero threshold for researchers:

No need to configure a local environment, no need to write a single line of code — just open the browser, enter the site, and the results will come out in seconds.

From protein structures to the effects of DNA variations, DeepMind has laid out another life map.

Cracking the "Book of Life": Humanity Waited 23 Years

DNA is known as the "language of life", written by four bases: A, T, C, and G.

In 1953, humanity first observed the double helix structure of DNA. In 2003, the Human Genome Project was completed, which took a full 13 years and 2.7 billion US dollars to read out all 3 billion base letters.

But reading all the letters is far from equal to understanding them.

Of these 3 billion letters, only a mere 2% are responsible for encoding proteins; the remaining 98% seemingly do not encode anything, but in fact strictly control when, in which cells, and at what intensity each gene is expressed.

The vast majority of variations related to diseases and traits are precisely hidden in this 98% non-coding region.

If one of these letters is altered, what consequences will it cause exactly?

A change in one base is a type of "single-base mutation"

Among the billions of base pairs in the human genome, there are approximately 9 billion potential single-nucleotide substitutions.

If verified one by one in the laboratory, it is a completely impossible task to complete.

AlphaGenome completely breaks through the old limits: it reads 1 million bases at a time, uses convolutional layers to capture local patterns, and uses Transformer to transmit the long-range dependencies of the entire sequence, directly predicting thousands of epigenetic and transcriptional tracks with single-base precision.

To evaluate a variation, you only need to input the sequences before and after the mutation into the model to run them separately and compare the differences, and you can get the result in one second.

But even if each calculation only takes one second, multiplying by 9 billion, it would still take 285 years to calculate all results.

What the Atlas does is directly compress this long 285 years into a ready-made static lookup table.

80x Speedup: Compressing 285 Years of Calculation Into a Single Table

With 9 billion types of mutations, how do we find the most worthy ones for research?

This time, the AlphaGenome Atlas has built a multi-level, high-resolution linked decryption system:

  • The bottom layer is raw predictions

For each base variation, the Atlas generates 27,000 molecular effect predictions that span the key dimensions of gene regulation, covering hundreds of human and mouse cell types and tissues.

It reveals how mutations affect chromatin accessibility, transcription, and splicing across different tissues and different developmental stages.

  • The middle layer is a unified measurement standard

To help scientists filter out "high-risk variations" in one second, DeepMind combines AlphaGenome (which excels at regulatory mechanisms) and AlphaMissense (which excels at predicting the pathogenicity of missense mutations), and extracts an intuitive unified metabolic/pathogenic score — the AVI indicator (AlphaGenome Variant Impact).

The higher the score, the greater the harmful risk, and the same measurement standard is applied to both coding and non-coding regions.

As a result, it has conquered the 98% "dark matter" of the non-coding region, which is exactly where the vast majority of mutation associations with human complex traits and diseases are hidden.

  • Fully transparent attribution decomposition

AI cannot just be a "black box that outputs scores".

The Atlas also deconstructs the AVI score of each variation into interpretable biological contribution items, accurately telling scientists —

Whether the high score of this mutation is caused by destroying RNA splicing, interfering with chromatin accessibility, or altering the abundance of gene expression.

  • The top layer is a full-scale vocabulary table

The team extracted 2601 high-frequency DNA patterns from massive predictions, and labeled 253 billion instances across the whole genome. You can immediately check which transcription factor recognizes which sequence in which type of cell.

These layers of data are fully connected vertically. Starting from a variation, researchers can click the mouse a few times to fully trace which regulatory element it destroys, the core motif in the element, and the corresponding transcription factor.

The workload that used to require multiple research groups to collaborate for several years is now condensed into a few mouse clicks.

Žiga Avsec, head of DeepMind's Genome team, revealed that in order to process this 1PB of data, the team increased the speed of the entire computing pipeline by 80 times.

Solving Unsolved Cases: Identifying the Pathogenic Cause for Children with Epilepsy

How powerful is this system in real-world practice?

Before the official release, DeepMind's early cooperation with the world's top scientific research institutions has already brought stunning results.

One of the most troublesome scenarios in rare disease research is that researchers have obtained genetic data, but cannot find the key change that explains the condition for a long time.

The team of Anne O'Donnell-Luria and Laura Covill from the Broad Institute once treated a child patient who had epileptic seizures since infancy, accompanied by spasms, developmental delay and hypotonia.

However, previous genetic analysis failed to make a definite diagnosis, and this case became a long-standing unsolved medical mystery.

This time, they fed all the candidate variations screened from this child patient to the AVI system for re-ranking.

As a result, the top-ranked variation is a G-to-A mutation in the intron of the DNM1 gene on chromosome 9.

The Atlas even restored the "crime scene" of DNM1: it created an incorrect splice site (the RNA splicing instruction went wrong), which led to abnormal extension of the translated protein.

Subsequent wet experiments fully confirmed the mechanism predicted by AI, and successfully solved this molecular unsolved case that caused epileptic encephalopathy.

526 Candidate Variations Narrowed Down to 15 Base Target Regions

Another use of the Atlas is to help researchers analyze rare variations in the population.

Single rare variations occur infrequently, and their effects are easily submerged in the background noise formed by a large number of harmless changes.

Gareth Hawkes from the University of Exeter applied the Atlas to the whole genome data of more than 54,000 participants in the UK Biobank.

The study grouped and analyzed rare variations according to their predicted molecular effects.

According to DeepMind, this method discovered 22% more non-coding genetic associations, and helped locate regulatory variations that affect circulating protein abundance, involving PLA2G7 and EGLN1, a gene related to cellular oxygen sensing.

In another analysis focused on Body Mass Index (BMI), Hawkes focused on the 1% non-coding variations with the largest impact predicted by the Atlas, and identified 19 genetic regions, narrowing the scope for subsequent targeted research.

The Atlas also sorted out more than 2500 recurring short DNA sequences and their positions.

They are like "words" in the genome, helping researchers find the binding positions of regulatory proteins, and trace which regulatory links are disrupted by non-coding variations.

Life Sciences Now Has a Directly Accessible Predictive Map

All the above new discoveries share a striking common feature —

No research team has ever run AlphaGenome from scratch in their local server room, and all breakthroughs are achieved entirely by "directly looking up the table".

This is the first time in human history that researchers anywhere in the world can see the complete map of the human genome and its variations just by opening a browser.

In the past, to understand a non-coding mutation, researchers often needed GPU resources, senior bioinformatics engineers, and spent a lot of time debugging complex pipelines.

Now, all these technical thresholds have been eliminated by this static atlas.

In addition to the web portal and API, the Atlas is also integrated into the Google Antigravity agent platform as a native tool skill.

Prior to this, DeepMind has packaged more than 30 scientific databases including AlphaGenome and UniProt into Agent skill packs.

In the official demonstration, when researchers say a sentence to the AI, the agent will automatically check the sites, rank the risk levels, draw the structural diagram, and provide a verifiable hypothesis chain.

The Era of Guesswork Is Over

After 23 years of development, humanity has finally obtained the interpretation guide for this "Book of Life".

AI has pre-annotated every potential typo in the 3 billion characters and the possible molecular consequences, and handed over a complete errata sheet to the world.

Looking back at DeepMind's development path in the field of biological computing:

In 2020, AlphaFold conquered the 3D folding of proteins; in 2022, the AlphaFold database opened 200 million protein structures to the world; in 2023, AlphaMissense completed the pathogenicity scoring of 71 million protein variations; in 2025, AlphaGenome understood the non-coding regulatory regions.

Until today, the Atlas has completed the exhaustive mapping of all possible single-base mutations in the human genome.

In just five years, DeepMind has gone from analyzing the physical form of a single protein to exhausting all single-base changes in the entire human genome.

From "understanding life" to "predicting life".

The next era of life sciences has already begun.

References: https://x.com/googledeepmind/status/2097325048109384166

This article is from the WeChat official account "AI Era", author: ASI Revelation; editors: Moses, Taozi, published with authorization from 36Kr.