The nonprofit arm of OpenAI is putting real money behind a problem that has quietly limited how useful artificial intelligence can be in medicine: there simply is not enough high-quality, openly available biological data for models to learn from. On September 15, 2026, the OpenAI Foundation launched Public Data for Health, its second major science program, committing more than $125 million in initial grants to nonprofits and universities that build, curate and preserve scientific datasets researchers and AI systems can actually use.
A Data Problem Bigger Than a Model Problem
Much of the recent conversation about AI in medicine has focused on model capability, how well a chatbot can pass a board exam or summarize a chart. Public Data for Health is aimed at a more upstream bottleneck: the raw experimental measurements that let any model reason accurately about biology in the first place are scattered, incomplete, or locked away in proprietary silos, and some risk disappearing entirely when a lab closes or a company goes bankrupt. The Foundation’s bet is that funding the creation and preservation of open datasets will do more for medical AI over the next decade than another round of model fine-tuning.
Where the First Money Is Going
The largest single grant, $40 million, goes to the University of North Carolina’s Lineberger Comprehensive Cancer Center, which will use the funding to launch an Initiative for Generative Immunotherapy. The goal is to generate the kind of public, multimodal data needed to make personalized cancer vaccines more effective. Today’s neoantigen vaccines are built by sequencing a patient’s tumor to predict which tumor-specific targets sit on its surface, then guessing how strongly that patient’s immune system will react to each one; UNC’s project aims to replace that proxy with direct, shareable measurement data. Two other grantees round out the initial cohort: OpenADMET, which runs competitions aimed at improving how well AI models predict a drug’s effects in the body, and CTD Commons, a project focused on preserving regulatory and drug-development knowledge that would otherwise stay fragmented across companies and agencies. A smaller, $500,000 grant to the nonprofit 1Day Sooner funds an unusual strategy: acquiring regulatory and manufacturing data from bankrupt biotech companies through bankruptcy proceedings, before that institutional knowledge is lost for good.
Building on an Earlier Bet on Alzheimer’s
Public Data for Health is not the Foundation’s first move into life sciences. In April 2026, it launched AI for Alzheimer’s, a narrower program targeting a single disease. The new initiative is explicitly broader, designed to seed open data infrastructure across many diseases rather than one, on the theory that the biggest returns come from fixing the data pipeline that every future disease-specific effort will depend on. OpenAI has framed the two programs as complementary steps in a longer-term life-sciences strategy that sits alongside its more commercial health moves, including plugins that let ChatGPT pull from public healthcare data sources for clinicians.
Why Open Data Matters for Smaller Labs, Too
Proponents argue the program addresses a structural inequity in AI-driven medical research: large pharmaceutical companies and well-funded academic centers can generate or license proprietary datasets that give their models an edge, while smaller labs, public-health researchers and startups in lower-resource settings cannot. By funding datasets that are broadly available rather than exclusively licensed, the Foundation says it hopes to level that playing field and accelerate research that would otherwise stall for lack of comparable, standardized data. UNC’s cancer vaccine work is being pitched as a case study, sequencing and immune-response data that could, in principle, benefit any lab working on personalized immunotherapy, not just UNC’s own pipeline.
Skeptics Question Scale, Incentives and Follow-Through
Not everyone is convinced $125 million, spread across a handful of grantees, will meaningfully dent a data gap that has persisted for decades and involves millions of unpublished or poorly annotated datasets across thousands of institutions. Some researchers who track corporate philanthropy in AI have also raised a familiar concern: a program funded and shaped by a company that stands to benefit commercially from better biological training data is not a purely disinterested actor, even when it is structured as a nonprofit grant. Others note that dataset projects, unlike a single clinical trial, can take years to show whether the data they produce is actually usable, standardized and adopted by outside researchers, meaning the real test of Public Data for Health will not arrive with the announcement but with whether independent labs cite and build on these datasets years from now.
What to Watch Next
The Foundation has said it expects to name additional grantees as the program matures, and the UNC cancer vaccine initiative in particular will be watched closely as an early proof point: if generative immunotherapy data produced under the grant measurably improves how neoantigen vaccines are designed, it could become a template other disease areas try to replicate. For an industry that has spent the past two years racing to build ever larger models, the more understated bet embedded in Public Data for Health is that better data, not just better algorithms, is what finally turns AI’s promise in medicine into something patients can rely on.
Photo: shameersrk / PIXABAY via Pixabay