Meta-backed Biohub spearheads the construction of AI pharmaceutical data infrastructure, with its total scale increased to 1.8 billion US dollars.
Following the popularity of its personal AI assistant Muse, Meta (META.US)'s latest foray into the biotech sector has also drawn market attention. On October 7, Biohub, a biotech institution backed by Meta, is expanding a "Virtual Biology Initiative" with a total scale of approximately 1.8 billion U.S. dollars. Biohub also announced that the U.S. Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs and Meta will join the initiative. NVIDIA will provide computing infrastructure. Multiple cell and genome research institutions will participate in data production and standard formulation.
The 1.8 billion U.S. dollars is not all newly added cash: among the total, Biohub will invest 500 million U.S. dollars; the U.S. Department of Energy (DOE) will invest more than 500 million U.S. dollars in the next five years; the National Institutes of Health (NIH) will coordinate the use of data and scientific research facilities formed by previous federal investments of over 500 million U.S. dollars; Google DeepMind, Isomorphic Labs and Meta will invest a total of 300 million U.S. dollars.
Founded and supported by Meta founder Mark Zuckerberg and his wife Priscilla Chan, Biohub is an independent non-profit research institution that does not belong to Meta. However, Meta is not only an investor, but also a source of part of Biohub's AI technologies and talents.
01
Biohub: Aims to Generate Training Data for Virtual Cells
One of Biohub's core goals is to build a "virtual cell", which is a set of AI systems that learn from genome, protein, imaging and cell perturbation data, used to predict the responses of cells under the influence of drugs, gene editing or environmental changes. Researchers can test drugs, gene editing or other solutions in the model first, and then select projects that are worthy of being advanced to laboratory research.
What this scaled-up "Virtual Biology Initiative" intends to do is to solve the data problem of AI biotechnology. Existing biological data is scattered in papers, experimental records and public databases, with different formats and experimental conditions, and few failed results are publicly available.
Biohub plans to establish unified standards, identifiers and access portals through this initiative, and produce more intervention data (intervention data refers to records of cell responses after exposure to drugs, gene editing or environmental changes).
As previously mentioned by *Suchbright*, such models are currently constrained by data limitations. Biological data is scattered in papers, experimental records and public databases. Different institutions use different samples, instruments and recording methods. Public data is dominated by successful results, and results such as unbound and synthetic failures are rarely included in papers.
For AI-driven new drug discovery, models also require failure data. Only when the model has seen sufficient numbers of both successful and failed results can it possibly predict the effects of new drugs or new targets. Moreover, failed results can explain why a certain design cannot be expressed, why non-specific binding occurs, or why it only works in specific cellular environments.
Under this initiative, Biohub will allocate 400 million U.S. dollars for cryo-electron tomography, in vivo microscopy and bioengineering tools, and another 100 million U.S. dollars to support external research. The cooperating U.S. Department of Energy will provide supercomputing, imaging facilities and automated laboratories, while NIH will provide disease data and public databases.
This process is similar to the logic of AI pharmaceutical companies building wet laboratories: the model proposes a solution, the laboratory conducts verification, and the results are used for the next round of training. The difference is that Biohub hopes to open data and tools to the public as scientific research infrastructure, rather than limiting their use to its own drug pipeline.
02
Meta's Layout in Biotech: Extending from Protein Models to Cellular Data
Meta entered the field of computational biology relatively early through its FAIR lab. FAIR developed the ESM protein language model. Alex Rives, the head of the ESM project, later founded EvolutionaryScale, and the team joined Biohub in 2025, with Rives serving as Biohub's Director of Science.
In May 2026, Biohub released ESMC, ESMFold2 and ESM Atlas. ESM Atlas covers 6.8 billion protein sequences and 1.1 billion predicted structures. Biohub also used ESMFold2 to design protein binders for multiple cancer and immune targets, and some of the designs have completed laboratory verification.
Source: Official website of Biohub
Meta itself retains research at the atomic and molecular levels.
In 2025, Meta FAIR released Open Molecules 2025 and Universal Model for Atoms, which are used to predict atomic interactions in molecules and materials. After participating in the Biohub initiative, Meta's public layout has covered protein models, atomic models and cellular data infrastructure.
However, Meta currently has no public drug clinical pipeline, and its position is closer to a provider of models, data and computing power. Biohub is responsible for experiments and data production, while downstream pharmaceutical and biotech companies are responsible for candidate drug and clinical development.
As previously mentioned by *Suchbright*, OpenAI, Anthropic, Google and ByteDance are also moving closer to the experimental segment.
OpenAI launched GPT-Rosalind and invested in Red Queen Bio, Chai Discovery and Valthos; Anthropic built a molecular biology laboratory in the Bay Area; Google expanded from AlphaFold to Isomorphic Labs; ByteDance spun off Anew Labs and retained controlling stakes after financing.
The paths of these companies are different. Google is approaching drug pipelines through Isomorphic Labs; OpenAI adopts the approach of model development, cooperation and investment; Anthropic directly builds wet lab capabilities; ByteDance retains its AI pharmaceutical assets; Meta and Biohub are more focused on data and foundational models.
The signal conveyed by these moves is that competition in AI-driven pharmaceuticals is extending from models to data production. Models can expand the search scope, but cannot replace experiments. Whoever can continuously obtain data in unified formats that includes failed results will have more opportunities to improve the model's ability to predict new targets and new drugs.
Biohub still needs to prove that virtual cells can predict unseen drugs, patients and experimental conditions. Judging from this initiative, the investment of AI giants in biotechnology has begun to shift from individual models to data, experimental and computing infrastructure.
This article is from the WeChat official account "Suchbright" (ID: suchbright), author: 24/7 Chief Editor, published with authorization from 36Kr.