A study published April 13, 2026 in JAMA Network Open by researchers at Mass General Brigham delivered a sobering counterpoint to a year of celebratory AI-in-medicine headlines: when large language models are asked to reason through a case the way a physician does — gathering information and narrowing a differential diagnosis before all the facts are in — they fail to produce an appropriate differential more than 80% of the time.
How the study was built
The research team tested 21 different AI models against 29 standardized medical case scenarios drawn from the MSD Manual, a peer-reviewed clinical reference widely used to train medical students and residents, generating 16,254 total responses, according to coverage in Becker’s Hospital Review and Medical Xpress. Rather than simply handing each model a fully worked-up case and asking for the final answer, the researchers structured the test to mimic real clinical workflow, in which a clinician must decide what to ask and what to rule out before every lab result and imaging study is already on the chart.
The central finding: good at the answer, bad at getting there
The results split sharply along that line. When every piece of pertinent clinical information was already provided, the tested models reached the correct final diagnosis more than 90% of the time. But at the earlier, reasoning-driven stage — building an appropriate differential diagnosis from partial information, the step that determines what tests get ordered and what gets ruled out first — the same models failed more than 80% of the time. Marc Succi, the study’s co-author, was blunt in the paper’s framing: "off-the-shelf large language models are not ready for unsupervised clinical-grade deployment," he said, according to Mass General Brigham’s own release on the findings.
Why the distinction matters clinically
Differential diagnosis is not a footnote to clinical reasoning — physicians and medical educators widely regard it as the central skill of the diagnostic process, the mechanism by which a clinician avoids anchoring on the first plausible explanation and instead systematically considers and eliminates alternatives. A model that reliably lands on the right final answer once handed a complete case file is solving a different, easier problem than the one physicians actually face at the bedside, where information arrives incrementally and imperfectly.
A contrasting result from Harvard
The Mass General Brigham findings arrived roughly three weeks after a separate, more optimistic study out of Harvard Medical School and Beth Israel Deaconess Medical Center, reported by Fortune in May 2026, in which OpenAI’s o1-preview model was compared against internal medicine attending physicians on emergency-room diagnostic benchmarks and reportedly outperformed the human baseline. Researchers on that study noted the model "eclipsed both prior models and our physician baselines," while also flagging that it tended to recommend unnecessary additional testing. Taken together, the two studies illustrate how sensitive AI diagnostic performance is to exactly what task is being measured and how it’s evaluated — a nuance that can get lost when a single splashy benchmark result is reported as evidence that "AI beats doctors."
What clinicians and skeptics are saying
Patient-safety advocates and practicing physicians have pointed to the Mass General Brigham results as validation of a cautious deployment posture: using AI as a decision-support layer that a clinician reviews, rather than a standalone diagnostic tool a patient might query directly through a symptom-checker chatbot. That caution runs directly counter to product narratives from some AI health-chatbot vendors marketing direct-to-consumer symptom-checking tools, and it complicates the broader case for models like Aidoc’s newly FDA-cleared CT triage system and other narrow imaging AI, which perform a fundamentally different, more constrained task than open-ended differential diagnosis and shouldn’t be judged by the same yardstick.
What’s next
The Mass General Brigham team has signaled interest in follow-on work testing whether structured prompting, retrieval-augmented tools, or hybrid human-AI workflows can close the differential-diagnosis gap, rather than treating current-generation chatbot performance as a ceiling. For hospital systems evaluating any generative-AI diagnostic tool, the study is a concrete argument for keeping a licensed clinician in the loop at the reasoning stage of a workup, not just at the point where a final diagnosis gets signed off. The study also raises a practical question for hospital IT and compliance teams evaluating any AI diagnostic tool for procurement: benchmark claims about model accuracy need to specify which stage of the diagnostic process was actually tested, since a headline accuracy figure measured against a fully worked-up case tells buyers very little about how a tool will perform earlier in a real patient encounter, when the information available is incomplete and the stakes of an early misstep are highest.
Photo: Tessy Agbonome / PEXELS via Pexels