• 9 mins read
  • Published

A $1.8 Billion Bet on Predictive Models of Living Cells

Noel Sharkey Technology, AI and robotics editor Science.Report

Post by Noel Sharkey

A $1.8 Billion Bet on Predictive Models of Living Cells Science.Report © science.report
A $1.8 Billion Bet on Predictive Models of Living Cells © science.report

Federal agencies and major technology companies are assembling nearly $1.8 billion in funding and resources for open biological datasets designed to train models that predict cellular responses

Nearly $1.8 billion in declared resources is being assembled to give predictive cell models the one resource they currently lack at sufficient scale: standardized experimental data. On October 7, 2026, the Chan Zuckerberg Biohub announced an expansion of its five-year Virtual Biology Initiative, adding the U.S. Department of Energy (DOE), the National Institutes of Health (NIH), Meta, Google DeepMind and Isomorphic Labs to the original Biohub commitment of $500 million. The coalition plans to combine biological measurements, repositories and high-performance computing in an effort to develop models that could forecast how living cells respond to interventions.

The project is an infrastructure program rather than a finished artificial intelligence system. It has not produced a universal virtual cell that can reliably predict arbitrary biological behavior. Its immediate objective is to measure cellular responses under many more conditions, including drug exposures and genetic changes, then standardize those observations well enough for machine-learning researchers to train and test predictive models. Biohub describes the package as a combination of funding, computing, data and measurement technologies, not as an already operating universal cell simulator. The Biohub announcement makes that distinction central.

  • The data problem

    Biohub's Virtual Biology Initiative is built around a straightforward constraint. A model cannot learn cellular responses from a catalog that merely names cell types or records static images; it needs measurements showing what changes when conditions, signals or interventions change. Those measurements must also be standardized across instruments, laboratories, cell types and experimental methods.

    In practice, this means combining several forms of evidence rather than relying on a single assay. Molecular measurements can describe gene activity, protein abundance or structural changes, while imaging can preserve information about morphology and the relationships among neighboring cells. Spatial transcriptomics is particularly relevant because it links molecular activity to the position of cells within intact tissue. That spatial context can matter when the behavior of a cell depends on its local environment, rather than on its molecular profile in isolation.

    Public descriptions of the initiative place its ambition on a steep scale curve: existing datasets containing hundreds of millions of cells could be expanded toward billions and, eventually, trillions of cellular observations. Such numbers describe measurements, not necessarily unique biological states or independent experiments. A very large collection can still be biased toward particular tissues, donors, instruments or perturbations, so sample count alone is not a measure of predictive validity.

    Biohub says the resulting datasets will be openly available to researchers. Commercial participants that help finance data collection may receive a one-year period of preferential access to the datasets they helped create. Researchers could use the resources to build systems that estimate a cell's response before conducting some laboratory experiments, potentially narrowing the number of physical tests required for a particular question. That is a proposed use of the datasets, not a demonstrated replacement for laboratory biology.

    Alex Rives, Biohub's Head of Science, has described accurate predictive biology models as a way to move some experiments into digital environments and accelerate work on disease and treatments. The claim is technically plausible only if the underlying measurements capture the relevant biological variables and the models generalize beyond the experiments used to train them. The announcement provides no evidence that this generalization has already been achieved.

  • Who is contributing

    Google DeepMind, Isomorphic Labs and Meta are collectively investing $300 million in technologies for biological measurement and multimodal datasets intended for predictive models. The U.S. Department of Energy will contribute more than $500 million over five years for biological measurements, modeling and computing, while the National Institutes of Health will coordinate datasets, repositories and related resources developed through more than $500 million in previous federal investment.

    Biohub had already committed $500 million when it launched the Virtual Biology Initiative in April 2026. Of that amount, $400 million is designated for measurement technologies and $100 million for research outside Biohub. The expanded effort connects those commitments with federal laboratory infrastructure and new private investment rather than presenting a completed modeling platform. It therefore marks a shift from a primarily philanthropic program toward a coalition involving federal agencies, nonprofit research organizations and private AI companies.

    The measurement program includes cryo-electron tomography, which can reveal near-atomic structures inside cells, and microscopy systems intended to image millions to billions of cells in living tissue. DOE facilities add exascale supercomputers, X-ray and neutron scattering, cryo-electron microscopy and tomography, and autonomous laboratories. NVIDIA will provide accelerated computing infrastructure, specialized software and technical expertise. These instruments do not measure the same biological features: structural imaging, molecular profiling and functional perturbation experiments each provide different dimensions of cellular state.

    The NIH contribution is linked to its Bio Genesis Mission. It includes national repositories and programs developing biological atlases, common data standards and datasets intended for computational modeling. The wider group includes the Allen Institute, Broad Institute, Gladstone Institutes, Human Cell Atlas, Human Protein Atlas and Wellcome Sanger Institute. This federated structure resembles the broader data challenge recognized by organizations such as MIT and Stanford: useful biological AI requires not only algorithms, but also carefully documented measurements that can be compared across laboratories.

  • Standardization is the test

    Biohub plans to work with NIH on common identifiers, shared standards and a single access point for datasets produced by different institutions. That technical work may determine whether the spending produces a usable training resource or a collection of incompatible archives. A dataset is more useful for predictive modeling when its records include consistent descriptions of cell identity, perturbation, timing, sample preparation, instrument settings and quality-control results.

    Biological measurements are not interchangeable by default. Differences in instruments, sample preparation, experimental conditions and cell populations can make two datasets appear to describe the same process while measuring different aspects of it. A model trained on such material could reproduce laboratory-specific patterns rather than learn a response that holds across settings. This is why the data architecture, metadata and independent validation plan are as important as the number of cells measured.

    Peer-reviewed biology in journals such as Nature and Cell routinely distinguishes discovery cohorts from independent validation cohorts; the same logic is necessary here even though the initiative has not yet reported a predictive benchmark. A credible evaluation would need predefined test conditions, clear error metrics and experiments performed independently of the data used for training. The current announcement reports neither a model version nor prediction error, confidence intervals, p-values or an external validation result.

    The initiative's strongest concrete output at this stage is therefore not predictive accuracy but an attempt to define the data architecture needed to measure it. The available material does not report a completed virtual-cell system or an independent evaluation. It also does not establish that digitally predicted responses will match results in living tissue. The distinction is especially important for medicine, where a model can assist hypothesis generation without being evidence that a treatment is safe or effective in humans.

  • Scale without proof

    The financial commitment is substantial: $300 million from Google DeepMind, Isomorphic Labs and Meta; more than $500 million from DOE over five years; more than $500 million in earlier NIH investment; and Biohub's existing $500 million commitment, including $400 million for measurement technologies and $100 million for outside research. These figures describe funding and resources committed to the effort, not a measured improvement in predictive performance.

    That distinction matters because biological prediction has a difficult verification problem. A model may fit observations from a particular experiment and still fail when a cell type, instrument or intervention changes. Open access can help researchers inspect and compare the data, but it does not by itself guarantee representative samples, consistent measurement quality or successful transfer to new biological conditions. Nor does a larger training corpus automatically resolve differences between healthy tissue, diseased tissue, engineered cell systems and living organisms.

    The initiative could become important if it produces datasets that support reproducible tests across institutions and if those tests show reliable predictions before laboratory validation. Until then, the central achievement is coordination: government laboratories, federal repositories, nonprofit researchers and technology companies are committing resources to build the evidence base required by the models they hope to create.

    For readers evaluating claims about virtual cells, the key concept is training data rather than model size. Training data supplies examples from which a model estimates relationships, while inference is the later process of applying those learned relationships to a new case. More measurements can improve coverage, but only when the measurements are comparable and include the conditions that matter to the prediction. This funding is best understood as a bet on that foundation, not proof that biology has become digitally predictable.

  • Related articles