AI Frontier Post
AI News

$1.8B to build a virtual cell: Biohub, DOE and NIH commit to the largest AI-ready biology data push ever

Biohub, the U.S. Department of Energy and the National Institutes of Health announced a combined $1.8 billion commitment to generate the AI-ready biological data needed for predictive models of life — the largest coordinated investment of its kind to date, with Google DeepMind, Isomorphic Labs and Meta adding $300 million more.

Biohub, the U.S. Department of Energy and the National Institutes of Health announced today a combined $1.8 billion commitment — in funding, data, computation and new measurement technology — to generate the AI-ready biological data needed for predictive models of life. Biohub describes it as the largest coordinated commitment to generating AI-ready biological data to date. Google DeepMind, Isomorphic Labs and Meta are collectively adding $300 million more.

The money funds a major expansion of the Virtual Biology Initiative, the international effort Biohub launched in April to coordinate data generation across institutions and disciplines. The stated end goal: an open data commons that lets researchers worldwide build and use AI models to ask, predict and answer biological questions digitally — a “virtual cell” that could model how any cell responds to a drug, a mutation or a disease.

The premise is that biology’s AI bottleneck is data, not algorithms. At the April launch, Biohub said a high-accuracy predictive model of the cell would require orders of magnitude more data than exists today. Biohub’s head of science Alex Rives called the virtual cell “one of the most important challenges for the next era of science,” arguing it will take coordinated data generation at national and international scale — no single institution can produce it alone.

Where the money goes

The Department of Energy will invest more than $500 million over five years through the Genesis Mission, its cross-agency push to harness AI for scientific discovery: data collection, AI analytics, measurement and imaging, modeling and computation — drawing on exascale supercomputing, X-ray and neutron scattering, cryo-electron microscopy and tomography, and autonomous laboratories across the National Laboratory system. DOE Under Secretary for Science Darío Gil framed the partnership as a new standard for open science that accelerates medicine and biotechnology.

NIH, through its Bio Genesis Mission, will coordinate the contribution of datasets, repositories and knowledge bases built on more than $500 million in prior federal investment, and will work with Biohub to standardize those datasets for AI model training. The agency has set a goal of doubling the pace of biomedical innovation within five to ten years. NIH deputy director Nicole Kleinstreuer said combining resources could accelerate universal cell models with enough biological complexity to predict how any cell responds to an intervention — on timelines substantially faster than laboratory experiments alone.

Industry is in as well. Google DeepMind, Isomorphic Labs and Meta are collectively putting $300 million into the technologies and multi-modal datasets the initiative needs. Isomorphic Labs president Max Jaderberg said generating the data for predictive systems biology requires scaling past what any single organization can produce today, and that Isomorphic is joining as a founding member. Google DeepMind’s Pushmeet Kohli said the investment would help create the open, standardized data commons researchers need to model biology. NVIDIA is contributing accelerated computing infrastructure, domain-specific software and technical expertise, and Renaissance Philanthropy is helping expand funding for data generation.

Fluorescence microscopy of a single cell, with its actin cytoskeleton in red, mitochondria in green and nucleus in blue.
Fluorescence microscopy of cells — the kind of high-resolution imaging the initiative aims to scale across millions of cell types and conditions. Image: NIGMS / NIH.

A coalition built like the Human Genome Project

The initiative is deliberately modeled on the big-science playbook. Institutions with experience organizing international collaborations dating back to the Human Genome Project — the Allen Institute, Broad Institute, Gladstone Institutes, the Human Cell Atlas, the Human Protein Atlas and the Wellcome Sanger Institute — are joining to help nucleate the scientific community around shared strategies. Biohub says it is building the unglamorous but decisive layer that makes partner datasets interoperable: shared standards, common identifiers and a single point of access.

That builds on Biohub’s open-data track record — Tabula Sapiens, OpenCell and Zebrahub for data generation, CELLxGENE and the CryoET Data Portal for community infrastructure. Biohub, a nonprofit combining frontier AI and biology, anchored the initiative with $500 million: $400 million for new measurement technologies — cryo-electron tomography resolving near-atomic detail inside cells, microscopy imaging millions to billions of cells in living tissue, engineering tools to build and perturb biology from molecules to whole organisms — plus $100 million for research outside its walls.

Close-up fluorescence microscopy of a cell showing detailed internal structure in red, green and blue.
New measurement technologies — from cryo-electron tomography to high-throughput microscopy — are central to the initiative’s $400 million technology arm. Image: NIGMS / NIH.

Why it matters

Funders are treating biological data as shared infrastructure — closer to telescopes or particle accelerators than to any single lab’s project. If virtual-cell models work, researchers could run experiments digitally before touching a pipette, compressing drug discovery and disease research timelines. That is the bet behind NIH’s goal of doubling the pace of biomedical innovation.

The open part matters as much as the scale. The output is promised as an open resource for the global research community, not a proprietary dataset — a commons anyone can train on. The risk is execution: committed money is not generated data, and coordinating agencies, institutes and companies across countries over years is the hard part. Standards and identifiers will decide whether this becomes one dataset or many.

Sources: Biohub (official announcement, Oct. 7, 2026); Unite.AI; News-Medical.