Uncategorized

OpenEvidence Rolls Out a Family of Medical AI Models, Keeps Its Perfect-MedQA-Score Model Under Wraps

OpenEvidence, used daily by more than 40% of U.S. physicians, split its clinical AI chatbot into three named models for speed and depth while holding back a fourth, Darwin, which it says scored perfectly on the MedQA benchmark, citing dual-use safety risks.

OpenEvidence, the physician-facing AI platform sometimes called "ChatGPT for doctors," announced on September 3, 2026 that it is splitting its single chatbot into a family of four named, purpose-built models. Three are shipping now, free, to verified clinicians on openevidence.com and its iOS and Android apps. The fourth, code-named Darwin, is being held back in a research preview, with the company citing dual-use risks in virology and genetics research even as it claims Darwin is the first AI model in history to post a perfect score on MedQA, the leading independent benchmark for medical AI.

Three models, three speeds, one bar for accuracy

The models are named for pioneers of clinical medicine and epidemiology. Osler, named for William Osler, who moved medical education from the lecture hall to the bedside, answers in about five seconds and is built for the pace of a live patient encounter. Sackett, named for David Sackett, often called the father of evidence-based medicine, takes roughly 30 seconds and is tuned for questions that hinge on weighing conflicting evidence. Snow, named for John Snow, who traced London’s 1854 Broad Street cholera outbreak to a single contaminated water pump, is the deepest production model, taking around five minutes to run a fuller sweep of the medical literature before answering. OpenEvidence says all three are held to the same clinical-accuracy bar; the difference is how long each is willing to think and how much literature it searches before responding.

Darwin: a perfect MedQA score, and a model too dangerous to fully release

The most striking claim in the announcement centers on Darwin, which OpenEvidence says achieved a perfect score on MedQA, a benchmark drawn from U.S. medical licensing exam-style questions that has become a standard yardstick for medical AI since it began tripping up early chatbots several years ago. Rather than releasing Darwin broadly, OpenEvidence is limiting it to a research preview, pointing to dual-use concerns in areas like virology and genetics, where the same reasoning that helps a model synthesize complex disease biology could, in principle, be misused. It is a rare instance of a health-AI company voluntarily throttling its own most capable model rather than racing it to market.

Why OpenEvidence can move this fast

The model family launch builds on a user base that has scaled unusually quickly for a clinical tool. OpenEvidence says more than 860,000 licensed U.S. clinicians now use the platform, with over 40% of U.S. physicians using it daily across more than 10,000 hospitals and medical centers. The company says over 100 million Americans were treated in 2025 by a doctor who used OpenEvidence at some point in their care. Monthly clinical consultations on the platform climbed from roughly 3 million a year earlier to about 18 million in December 2025, then topped 20 million by January 2026 — a trajectory that gave the company both the usage data and the revenue to fund a broader model lineup instead of a single general-purpose chatbot.

The money behind the models

That growth has been reflected in OpenEvidence’s finances and fundraising. The company closed a $250 million Series D in January 2026 at a $12 billion valuation, led by Thrive Capital and DST Global, roughly doubling its $6 billion valuation from just three months earlier. By the time of the September model-family launch, OpenEvidence was generating close to $300 million in annualized revenue, about $25 million a month, roughly double what it was earning seven months prior. Total funding raised now exceeds $735 million. The scale of that war chest is part of why the company can afford to hold Darwin back rather than monetize it immediately — a luxury few AI startups in a competitive field can claim.

Two views on splitting one model into four

Supporters argue the tiered approach mirrors how doctors actually think: a quick gut-check between patients calls for a different tool than a literature review before a tumor-board presentation, and giving clinicians an explicit choice of speed versus depth is more honest than hiding that trade-off inside one black-box chatbot. Skeptics counter that naming models after historical medical giants and touting a "perfect" benchmark score risks overselling systems that still make mistakes, and that withholding Darwin while publicizing its perfect score functions as marketing for a future release as much as a genuine safety measure. Both camps agree on one thing: with more than 40% of U.S. physicians touching the platform daily, whatever OpenEvidence ships next will be tested in exam rooms almost immediately, not just in a lab.

What to watch next

The near-term questions are whether Osler, Sackett and Snow measurably change how clinicians use AI at the bedside versus during chart review, whether independent researchers can replicate Darwin’s MedQA result once it moves beyond preview access, and how regulators and hospital systems respond to a commercial AI vendor publicly withholding its most capable model on safety grounds. With OpenEvidence reportedly weighing a further raise near a $20 billion valuation, the model family launch looks less like a one-off product update and more like the opening move in a longer contest over how much clinical reasoning doctors are willing to outsource to AI — and how fast.