
499 Illinois Street in San Francisco’s Mission Bay, home to Biohub, in May 2025. Image: 9yz / Wikimedia Commons, CC BY 4.0, cropped
Biohub, the Zuckerberg-backed research institute, has announced a $1.8 billion push with the US government, Google DeepMind, Isomorphic Labs and Meta to build the data needed for AI models that can predict how living cells behave.
The organisations call it “the largest coordinated commitment to generating AI-ready biological data to date”. The goal is a so-called virtual cell: an AI model good enough that scientists could run some experiments on a computer instead of in the lab, and the open datasets to train it.
Where the $1.8 billion comes from
The headline figure adds up four pots, and not all of it is new money:
- US Department of Energy: more than $500 million over five years, through its Genesis Mission, for lab measurement, modelling and computing, drawing on national lab supercomputers, cryo-electron microscopes and autonomous labs.
- National Institutes of Health: datasets and repositories built with more than $500 million of past federal spending, which Biohub will help standardise for training AI models.
- Google DeepMind, Isomorphic Labs and Meta: $300 million between them for the Virtual Biology Initiative.
- Biohub: its founding $500 million commitment from April, with $400 million for new measurement tools such as cryo-electron tomography and microscopes that can image millions to billions of living cells, and $100 million for research outside Biohub.
The Allen Institute, Broad Institute, Gladstone Institutes, the Human Cell Atlas, the Human Protein Atlas and the Wellcome Sanger Institute have also signed up to help organise the science, Nvidia will provide computing and software support, and Renaissance Philanthropy is raising more money for data generation. The announcement doesn’t say how the $300 million is split between the three companies.
“One of the most important challenges for the next era of science”
The problem the partners are trying to solve is data. Today’s biology models are trained on whatever experiments happen to exist, and very few record how many different cell types respond to drugs, genetic changes and other interventions. The initiative plans to measure exactly that, across far more cell types and conditions than have been studied, and publish the results as an open resource with shared standards and a single point of access.
Alex Rives, Biohub’s head of science, said:
An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally.
Alex Rives, head of science, Biohub
He called the virtual cell “one of the most important challenges for the next era of science” and invited researchers around the world to join.
Google and the government sign on
Pushmeet Kohli, who runs AI for science at Google DeepMind, said the challenge won’t be solved “without open, experimental biological data at an unprecedented scale, showing how living cells behave and respond to changes”. Max Jaderberg, president of DeepMind’s drug discovery sister company Isomorphic Labs, said the data “will push the industry closer to the next significant breakthrough for biology”.
For the government, the DOE’s under secretary for science, Darío Gil, called it “a critical step forward in leveraging artificial intelligence for public benefit”, and the NIH’s Nicole Kleinstreuer said the models could bring “substantially faster timelines for medical breakthroughs” than lab experiments alone. All of those benefits are the partners’ own expectations: the announcement sets no dates for the first datasets or models.
Google DeepMind already works at the edge of AI and biology: last week it showed a way to hide a watermark inside AI-designed proteins, and AI models have started making their own finds, such as when Claude turned up a new CRISPR-like system.
Why it matters
AI has cracked protein structures because decades of shared data already existed; cells have no equivalent yet. Pooling government labs, big tech money and the world’s main genome institutes into one open dataset is the most serious attempt so far to build it, and because the data is meant to be open, any lab, not just the funders, should be able to train on it.
Sources: Biohub.


