Boston Children’s Hospital and OpenAI announced on June 18, 2026 that a specialized AI model, o3 Deep Research, helped identify 18 new diagnoses among 376 pediatric patients whose rare genetic diseases had gone unexplained despite prior genomic testing and expert review. The findings, published in NEJM AI, mark one of the largest demonstrations yet of large language models being used to systematically reanalyze the backlog of “unsolved” genomic cases that pile up in children’s hospitals around the country.
A Genome Reanalysis Pipeline Built Around o3 Deep Research
The study, led by researchers at Boston Children’s Hospital’s Manton Center for Orphan Disease Research, took 376 de-identified pediatric cases that had already gone through standard genetic testing and expert clinical review without producing an answer. Clinicians fed the o3 Deep Research model case notes, patient symptom descriptions, and filtered gene lists from the original workups. The model synthesized that information against current genetic data and medical literature to nominate candidate diagnoses, which human geneticists then had to independently confirm. A finding only counted as an official diagnosis after expert review, additional targeted testing, classification of the variant as pathogenic or likely pathogenic, confirmation by a CLIA-certified laboratory, and formal return of the result to the family by the clinical team.
From 376 Unsolved Cases to 18 New Diagnoses
The reanalysis produced 18 confirmed new diagnoses, an additional diagnostic yield of roughly 4.8% on cases that had already stumped specialists. The breakdown spanned several categories: 10 new diagnoses came out of 100 neurodevelopmental disorder cases, four came from 61 neuromuscular disease cases, and two each came from cohorts of sudden unexplained pediatric death and early-onset psychosis cases. A further seven cases turned out to be “rediscoveries,” meaning the correct answer had technically been identified somewhere in the scientific literature already but had never made its way back to the family’s medical record, a gap researchers say reflects how fragmented data-sharing still is across the rare disease research community. Boston Children’s said in a separate announcement in late May 2026 that its broader AI-assisted genomics effort has now produced more than 40 rare disease diagnoses that were previously considered unsolvable.
Why Rare Disease Diagnosis Takes So Long
Roughly 7,000 rare diseases affect an estimated 30 million Americans, and patients often endure what geneticists call a “diagnostic odyssey” lasting years, bouncing between specialists and repeated genetic panels while symptoms worsen without a name for what is wrong. Even when whole-genome or whole-exome sequencing is performed, a single patient’s raw data can contain millions of variants, and a case can go unsolved simply because no human analyst had time to connect a patient’s symptoms to a paper published after the original workup, or to a gene newly linked to disease. The Manton Center, which works with more than 3,500 patients across all 50 states, built this study to test whether an AI system could revisit that backlog faster than overstretched genetic counselors could manually recheck each case by hand.
“A Total Game Changer”: The Case for AI-Assisted Reanalysis
Catherine Brownstein, scientific director of genetic investigations at the Manton Center, called the result “a total game changer,” describing the roughly 5% diagnostic yield as “a huge number” given how many prior analyses had already failed to crack these same cases, and noting that each diagnosis “means an answer for a family” that may have been searching for years. Adam Rodman, an outside physician at Beth Israel Deaconess Medical Center who was not part of the study, called the yield “truly meaningful” and described the approach as a significant screening tool for genetics teams drowning in unresolved cases. OpenAI technical researcher Suyash Shringarpure pointed to how often the missing piece of a diagnosis is simply timing: a case can go unsolved when it first reaches a lab, only for a relevant paper to be published a year later that an AI system can instantly surface. Kyra Benton, a 20-year-old patient who received a diagnosis of myofibrillar myopathy after 15 years without an answer, said she was initially skeptical of AI being involved in her care but acknowledged the technology “can really change people’s lives.”
Why Experts Urge Caution Before Scaling Up
Not every outside voice treated the results as a breakthrough to scale up immediately. Chunhua Weng, a bioinformatics professor at Columbia University who was not involved in the study, called the paper “wonderful” but stressed that large language model results “still require rigorous human review” and warned the field needs “careful attention to trustworthiness” before leaning on AI-generated candidate diagnoses in routine clinical practice. The study’s own authors were careful to frame the tool as a screening aid rather than a diagnostic replacement, writing that the results are “not a panacea” and noting that a diagnosis is frequently only the first step toward treatment, not the end of a family’s journey. The seven “rediscovery” cases underscored a structural problem too: even when the right answer about a gene-disease link exists somewhere in the published literature, there is no reliable mechanism to automatically route that discovery back to the patients and labs who need it.
What Happens Next
Boston Children’s and OpenAI have not released a timeline for a prospective, larger-scale trial, but researchers say the next step is testing whether periodic AI reanalysis can be built into standard genetics workflows at other hospitals rather than run as a one-time research project. The team’s stated ambition is to democratize access to this kind of reanalysis for institutions without large genomics research centers of their own, so unsolved cases don’t simply sit untouched for years. Open questions include how often reanalysis should be repeated as medical literature grows, who pays for the additional confirmatory testing each AI-flagged case requires, and how the data-sharing gaps behind the seven “rediscovery” cases get fixed. For now, the 18 new diagnoses stand as a concrete, human-reviewed data point in a larger debate over how much clinical weight hospitals should put on AI-generated hypotheses before treating them as medical fact.
Photo: deepakrit / PIXABAY via Pixabay