On October 7, 2026, Biohub, the nonprofit research organization backed by Mark Zuckerberg and Priscilla Chan, announced that its Virtual Biology Initiative has grown into an international coalition with a headline figure of about $1.8 billion. The goal is a virtual cell: an AI model trained on large volumes of open, standardized experimental data that predicts how a cell responds to a drug, a gene edit or a disease state. The partners include the US Department of Energy (DOE), the National Institutes of Health (NIH), Google DeepMind, Isomorphic Labs and Meta, with NVIDIA supplying computing. The most important question about the announcement is less about the model than about the data, and about how much of the $1.8 billion is new money.
Meta
NVIDIAWhat a virtual cell model actually is
A virtual cell model is a machine learning system that represents a cell's state and predicts how that state changes when something is done to it. Biohub's April 2026 announcement describes accurate predictive models of the cell as tools that could reveal fundamental mechanisms and the causes of disease, letting researchers ask and answer questions digitally at a scale laboratories cannot match. Arc Institute co-founder Patrick Hsu framed the payoff as showing "which levers to pull" to move a cell from a disease state to a healthy one. NIH Deputy Director Nicole Kleinstreuer, quoted in the October release, described the target as universal cell models with enough biological complexity to predict how any cell responds.
That word "universal" is the ambitious part. Today's single-cell foundation models are typically trained on snapshots of cells in particular tissues and conditions. A universal model would need to generalize across cell types, patients and interventions it has never seen, which is a much harder standard than fitting the data it was trained on.
Why data scale is the bottleneck
Biohub's own framing is blunt. Head of Science Alex Rives said in April that building AI that represents biology's complexity requires "orders of magnitude more data than exists today." In January 2026, Tahoe Therapeutics co-founder Johnny Yu put it more compactly: "Virtual cells are reaching a turning point, and data is the bottleneck."
Virtual cells are reaching a turning point, and data is the bottleneck.
Johnny Yu, Tahoe Therapeutics, on the Tahoe, Arc and Biohub perturbation dataset (January 2026)
The scale gap can be seen in the datasets that exist. The Tahoe, Arc and Biohub perturbation dataset announced in January 2026 contains more than 120 million single-cell data points covering about 225,000 perturbation interactions, which the release says is over four times as perturbation-rich as the earlier Tahoe-100M dataset. Separately, the Billion Cells Project, launched in 2025, coordinates 17 single-cell sequencing projects across institutions including MIT, Stanford, UC San Francisco and ETH Zurich. Its name states the ambition: the field is currently counting data in the hundreds of millions of cells and aiming for billions. Biohub's October release does not publish a target cell count, but it describes new microscopy capability as imaging "millions to billions of cells" in living tissue.
The point is not raw cell count alone. Perturbation data, where a cell is deliberately changed and the response measured, is the scarce kind, because it is what a predictive model needs to learn cause and effect rather than correlation. Arc's Silvana Konermann said in January that diverse, high-volume, high-quality perturbational data is still scarce.

Who is paying, and how much is new money
The coalition's funding is a mix of different things, and the chart below separates them.
Who contributes what to the roughly $1.8B headline
Dollar figures in millions, as stated in Biohub's October 7, 2026 announcement. The four bars sum to about $1.8B. NVIDIA contributes compute and expertise with no stated dollar value, so it is not charted.
Founding commitment (April 2026)
Over five years, Genesis Mission (more than $500M)
Combined, no per-company split
Existing datasets from more than $500M of prior investment
According to Biohub's release, the contributions are:
- Biohub: a $500 million founding commitment announced in April, split into $400 million for data generation and new measurement technologies and $100 million to nucleate a coordinated worldwide data effort with outside researchers.
- DOE: more than $500 million over five years through the Genesis Mission, covering lab measurement, modeling and computation, drawing on national laboratory facilities.
- Google DeepMind, Isomorphic Labs and Meta: $300 million combined, with no per-company split published.
- NIH: coordination of datasets, repositories and knowledge bases built with more than $500 million of prior federal investment, which Biohub will help standardize for AI training.
- NVIDIA: accelerated computing, domain software and expertise, with no dollar value given. Renaissance Philanthropy is also named as helping expand funding for data generation.
Adding the four dollar figures gives roughly $1.8 billion, which is our arithmetic, and AllSci reached the same tally. The consequence is that the NIH portion is existing data rather than a fresh check, so the amount of new money is smaller than the headline. Coverage that read the release closely says the same: the figure mixes new pledges, existing datasets and prior federal spending, and Biohub has not published a line-by-line valuation. There is also a small discrepancy to note: the release headline says "nearly $2 billion," while its body text says $1.8 billion.
The public-private structure and the one-year head start
The structure pairs a nonprofit anchor and two government agencies with three companies that have commercial interests in biology. Biohub says the output will be an open resource with shared standards, common identifiers and a single access point. The release itself states no exclusive-access terms.
Press coverage adds one. Axios, as relayed by Streamline Feed and Future Party, reported that commercial partners get one year of exclusive access to the data they help generate before it is released publicly, and that government-funded work is expected to carry no such restriction. We could not open the Axios article directly, so this detail rests on secondary reports, and detailed data agreements and release schedules have not been published. If the embargo is as described, it is a familiar trade: a limited commercial advantage in exchange for private participation and funding, with public release as the end state. Whether one year is short or long depends on how fast the partners can turn data into models and drugs, which is exactly what is unproven.
How it compares with other AI-biology bets
The reference point for most people is AlphaFold. Demis Hassabis and John Jumper of Google DeepMind shared the 2024 Nobel Prize in Chemistry for AlphaFold, which predicts protein structure from amino acid sequence, and DeepMind says more than 2 million researchers in 190 countries have used it. AlphaFold addressed a problem with a clean input and a clean output: sequence in, structure out, with a long-established benchmark. A virtual cell has neither. The state of a cell is high-dimensional, the response to a perturbation depends on context, and there is no single agreed measurement of success. That is a reason to expect the path to be slower and less tidy than the protein case, though that is our reading rather than a claim from the partners.
Isomorphic Labs, DeepMind's drug discovery spin-out, is the commercial branch of that lineage. Startup Fortune reports it raised a $2.1 billion Series B in May led by Thrive Capital, which makes the $300 million combined pledge small next to the capital Isomorphic alone has raised. Meta's role is less described in the release. Biohub's own earlier work, including protein models such as ESMC and datasets such as CELLxGENE, shows the same data-first approach: Biohub is building the training material in the open rather than a single closed model.
Realistic timelines
Reuters, as summarized by secondary outlets, reports that a first dataset is expected in about a year, a participant projection rather than a delivery guarantee. The DOE and Biohub commitments each run five years, and the release gives no overall completion date or performance milestones. Researchers plan to assess model capabilities after large datasets emerge, with no date attached.
For readers tracking the field, the useful markers are practical ones:
- Whether the first dataset ships on schedule, and under what license after any embargo.
- Whether independent benchmarks show virtual cell models beating simple baselines on held-out perturbations.
- Whether the standards and common identifiers actually make partner data interoperable.
- How Genesis Mission compute and NVIDIA infrastructure are allocated across partners.
Nothing in the announcement changes what a model can do today. It changes the odds that the data needed to find out will exist, in public, within the decade.
Sources:
- International, cross-sector collaboration commits nearly $2 billion to build foundational data for AI models to predict and treat disease (Biohub, Oct 7, 2026)
- Virtual Biology Initiative: $500M for AI-powered biology (Biohub, April 29, 2026)
- Biohub, Arc Institute, Tahoe partner on largest perturbation dataset (Biohub, Jan 12, 2026)
- Biohub, DOE, NIH and AI giants launch USD 1.8 billion push to build virtual cell models (AllSci, Oct 7, 2026)
- What Biohub's $1.8bn AI Biology Push Will Build (Streamline Feed)
- Google, Meta and the US government put $1.8 billion behind a virtual human cell (Startup Fortune)
- Can AI Make A Cell? (Future Party)
- Demis Hassabis and John Jumper awarded Nobel Prize in Chemistry (Google DeepMind)
- Frontier AI models and datasets from Biohub
Image credits
Header and in-body image: entrance to 499 Illinois Street, San Francisco, which houses the Chan Zuckerberg Biohub, by 9yz via Wikimedia Commons, licensed under CC BY 4.0.