The Chan Zuckerberg Initiative-backed research institute Biohub announced on 7 October that it, the US Department of Energy and the National Institutes of Health are committing $1.8 billion to generate the data needed to train predictive AI models of biology — what Biohub’s own announcement calls the largest coordinated commitment to AI-ready biological data to date. Google DeepMind, Isomorphic Labs and Meta are joining with a collective $300 million through the Virtual Biology Initiative, per the press release carried on PR Newswire.
The goal is a shared, open resource: enormous datasets measuring how cells respond to interventions, standardised so researchers can train models that answer biological questions digitally — predicting how a cell reacts to a drug or a mutation without running the physical experiment first. “An accurate predictive model of biology could dramatically accelerate scientific discovery by enabling scientists to perform experiments digitally,” Biohub head of science Alex Rives said in the announcement.
Who pays for what
The breakdown, as assembled by The Next Web from the announcement and follow-up interviews: the Department of Energy is putting in more than $500 million over five years for lab measurement, modelling and computation, drawing on exascale supercomputers and self-running laboratories through its Genesis Mission science push. The NIH contributes datasets and repositories built with more than $500 million in earlier federal funding, which Biohub will standardise for AI training. That joins Biohub’s own $500 million, pledged in April when the Virtual Biology Initiative launched. The Allen Institute, the Broad Institute, the UK’s Wellcome Sanger Institute, the Human Cell Atlas and Human Protein Atlas consortia are also participating, and NVIDIA is providing computing and software.
The scientific bet underneath the money is about scale. Current cell datasets hold hundreds of millions of cells, Rives told Reuters; an accurate predictive model will need billions, and eventually trillions. Reuters’ report on the announcement confirms the government and lab-industry framing of the deal. The partners aim for a first dataset in roughly a year and predictive models they consider accurate within five years. This is, in effect, the physics-style playbook — massive curated datasets plus scaling laws — applied to living systems, and it extends the wave of AI-for-biology bets we have tracked, from DeepMind weather models that out-predict traditional forecasting to AI narrowing rare-disease diagnosis and AlphaFold-era researchers moving between labs.
The one-year head start, stated plainly
The awkward part of an “open resource” funded partly by commercial AI labs: The Next Web reports that commercial funders will get one year of exclusive access to the data before it opens to everyone — an embargo Rives defended to Axios as “some incentive for commercial players to be a part of this.” Government-funded work carries no such restriction. Isomorphic Labs is Alphabet’s drug-design company, Meta has its own protein-structure research line, and a year of exclusivity on the world’s richest new cell data is a genuine commercial advantage — the pattern of public-private data deals where the public gets the resource eventually, and the private side gets the head start. Researchers on the outside will want the embargo clock to be publicly visible, not quietly negotiated.
There is also a competitive reading worth stating: OpenAI’s foundation recently launched a $125 million grant programme for biology datasets, and Anthropic has built its own wet lab. Biology is becoming a data moat race between frontier labs, and the US government just became the largest single customer of that race. For New Zealand’s research community, the open slice of this data — after the embargo — is exactly the kind of resource that lets a country without exascale computers still contribute to computational biology; the value will go to groups that can move fast on the data the moment it opens.
❓ FAQ
What is the Virtual Biology Initiative? It is Biohub’s programme to build the data, measurement tools and models needed for AI systems that can predict cell behaviour. The 7 October announcement expands it with US government and lab-industry partners, per Biohub.
How much money is involved and from whom? $1.8 billion total in funding, data, computation and measurement technology. DOE is investing more than $500 million over five years, the NIH brings more than $500 million in prior federal investment in datasets, and Google DeepMind, Isomorphic Labs and Meta are collectively investing $300 million. Biohub itself pledged $500 million in April.
When will the data be public? Commercial funders get one year of exclusive access before datasets open to the research community, according to reporting by The Next Web and interviews Rives gave to Axios and Reuters. Government-funded work carries no such restriction.
Why does AI biology need this much data? Current cell datasets hold hundreds of millions of cells; Rives says accurate predictive models will require billions to trillions of cell measurements across far more cell types and conditions than have been studied.
🔍 THE BOTTOM LINE
The headline number will get the attention, but the structure is the story: governments funding measurement, AI labs funding for a head start, and a shared data layer that turns biology into something a model can be trained on. If the five-year timeline holds, the way new drugs and diagnostics get discovered changes from bench experiments to compute — and the race to own the data layer of biology is now an open contest between the US government’s research machine and the frontier labs.
📰 Sources
- Biohub — Virtual Biology Initiative expansion (official announcement)
- The Next Web — Biohub, Meta, Google DeepMind and the US pool $1.8bn for AI biology data
- PR Newswire — International cross-sector collaboration commits nearly $2 billion (official release)
- Reuters (via Yahoo Finance) — US government, Google join Zuckerberg-backed Biohub in $1.8 billion push for AI biology data
- Axios — Rives interview on the commercial embargo (cited as reported by The Next Web and Reuters)