Google’s Breast-Scan AI Can Match Radiologists, But Researchers Warn Not to Read Too Much Into That Yet
New Imperial College London and Google research shows an AI model matching or exceeding radiologists in detecting breast cancer, including in tie-breaking arbitration cases, but the researchers themselves, along with independent experts, are cautioning against reading the results as proof AI is ready to replace human screening.
Imperial College London researchers, working with Google, published new findings in March 2026 showing that a Google-developed AI model can match or exceed radiologists at detecting cancer in breast scans, a result that’s generating real enthusiasm in cancer-imaging circles, tempered by pointed reminders from the same research teams that matching radiologists in a study is not the same as being ready to replace them in a clinic.
How the Model Was Actually Tested
Rather than a single simple accuracy test, the Imperial-Google collaboration ran the AI through several distinct clinical roles to see how it performed against real screening workflows, not just theoretical accuracy benchmarks. In one arm, the AI acted purely as a second reader alongside a human radiologist. In another, it was tested in arbitration, the process used when two human radiologists disagree on a scan and a third opinion is needed to break the tie, an arm covering 50,000 women, the first time AI has been tested in that specific decision-making role. The AI performed comparably to human arbiters in that setting, a more demanding test than simple pattern-matching against a known-cancer dataset, since arbitration cases are disproportionately the ambiguous, hard-to-call scans where humans themselves disagree.
The Google Model’s Methodology
The underlying model, developed by Google’s health AI team, was trained and evaluated on mammography data spanning multiple NHS Trusts and screening cohorts, deliberately structured to test generalization across different equipment and patient populations rather than a single hospital’s imaging pipeline. Dr. Susan Thomas of Google described the milestone in terms of collaboration rather than substitution, saying it’s the first time doctors and AI have worked alongside each other in a clinical setting like this, positioning the model explicitly as a second reader rather than an autonomous diagnostician. That framing matters, because Imperial’s own researchers were careful to caveat where the model’s strength actually lies.
What Independent Experts Are Saying
Dr. Hutan Ashrafian, one of the Imperial researchers involved, called the results the closest AI has ever come to helping reduce breast cancer deaths within the NHS, a strong statement, but one explicitly framed around potential rather than a settled conclusion. Other researchers in the field have been more cautious in the wake of the announcement, noting that a single model performing well on retrospective UK screening data doesn’t automatically generalize to other populations, other scanner manufacturers, or other countries’ screening intervals. There’s also a well-documented history in medical AI of models that perform impressively in published papers but underperform once deployed against messier real-world data, variable image quality, different demographic mixes, and edge cases that don’t show up cleanly in a curated research dataset.
The Skeptic’s View
Radiology researchers not involved in the Google-Imperial work have pushed back gently on some of the more triumphant framing of AI-beats-radiologists headlines that followed the announcement. Screening AI has produced a string of impressive-sounding papers over the past several years, several of which have not translated cleanly into national rollout once regulators, clinicians and patient advocates scrutinized the details, including questions about whether recall rate thresholds hold up prospectively, whether the AI’s performance advantage shrinks once radiologists know they’re being second-guessed by software, and whether the reported gains persist across the full diversity of breast tissue density and patient age in a national population rather than a curated study cohort. The Royal College of Radiologists has repeatedly stressed that any shift in screening protocol needs expert oversight, precisely because early enthusiasm in AI screening research has, in prior cycles, outpaced what held up under full clinical scrutiny.
Why the Distinction Matters
The nuance here is important: this research is about model performance and methodology under controlled study conditions, distinct from the question of whether the NHS should actually change its screening protocol nationally, which is a separate, ongoing policy and workforce decision being tested through other channels including the NHS’s own large-scale AI screening trials. Conflating a strong study result with a decision that the NHS is switching to AI-only screening has been a recurring source of media overstatement that researchers on this project have specifically tried to avoid in their public statements.
What’s Next
Imperial and Google say further validation work is underway to test the model across more diverse imaging equipment and patient populations before any recommendation on clinical deployment. For now, the honest summary from the researchers themselves is narrower than the headlines: a promising second-reader and arbitration tool, tested rigorously in a specific research context, with real work still ahead before it becomes a deployed clinical standard rather than a well-performing study result.
Photo: MART PRODUCTION / PEXELS via Pexels