Artificial intelligence software is increasingly good at spotting the faint mammographic clues that radiologists miss, according to a review published in early September 2026 by researchers at UCLA Health’s Jonsson Comprehensive Cancer Center. But the same researchers are urging caution before declaring the technology a breakthrough for patients, arguing that catching a cancer after the fact in a research study is a very different thing from proving AI actually gets women diagnosed earlier in real-world screening.
What the review looked at
Led by Dr. Tiffany Yu, an assistant professor of radiology at the David Geffen School of Medicine at UCLA, along with co-authors Dr. Javier Galvan, Dr. Hannah Milch, and Dr. Niki Nourmohammadi, the team examined published studies on commercially available AI tools used to interpret screening mammograms, focusing specifically on their ability to detect interval cancers, tumors diagnosed in the window between a normal-looking screening mammogram and a patient’s next scheduled screening. Interval cancers are considered a particularly important target because they represent cases where a cancer was likely present, but not caught, at the prior screening visit.
Some striking, if inconsistent, numbers
Retrospectively reviewing prior mammograms with AI, researchers found the software could identify subtle, missed signs of cancer in anywhere from 5% to 78% of interval cancer cases, depending heavily on which AI system and which study methodology was used, a range wide enough to itself raise questions about how reliable any single figure really is. A separate 2026 prospective trial involving more than 105,000 women offered a more controlled comparison: AI-supported screening produced an interval cancer rate of 1.55 per 1,000 women screened, versus 1.76 per 1,000 for standard double reading by two radiologists, alongside a sensitivity of 80.5% for the AI-assisted approach compared with 73.8% for standard reading, while specificity held steady at 98.5% for both groups. The AI approach also cut radiologist workload by 44.3% without compromising cancer-detection performance, according to the trial.
Predicting risk before a lump ever forms
Beyond just catching visible signs of cancer, some AI tools are being used to assign women a risk score based on subtle patterns in mammogram images that predate any visible abnormality. The UCLA review found that AI assigned its highest risk scores to 23.1% of women who went on to develop interval cancer as far back as three screening rounds, roughly six years, before their eventual diagnosis, with that share rising to 39.4% in the screening round immediately before diagnosis. That capability feeds into a broader industry push, reflected in the National Comprehensive Cancer Network’s updated 2026 guidelines, which now recognize mammogram-derived AI risk scores using a 1.7% five-year risk threshold to help decide which women should get closer monitoring.
Why the researchers are hitting the brakes
Dr. Yu framed the ultimate goal plainly: \”The ultimate goal of screening mammography is to eliminate these interval cancers.\” But co-author Dr. Hannah Milch drew a sharper distinction between the retrospective studies driving headlines and what would actually help patients, saying \”finding a cancer retrospectively is very different from demonstrating that using AI during routine screening would have led to an earlier diagnosis.\” The review flagged that most of the supporting studies were retrospective rather than prospective, used wildly different AI systems and screening intervals, and left open questions about what should happen clinically when AI flags a high risk score but no radiologist can see anything abnormal on the image, a scenario that could drive unnecessary follow-up procedures and patient anxiety without a clear clinical protocol.
Enthusiasm from radiology, caution from methodologists
Radiologists and AI vendors point to the prospective 105,000-woman trial as evidence the technology is already delivering measurable gains, higher sensitivity, lower workload, and comparable specificity, without waiting for a perfect long-term outcomes study. Research methodologists and some breast-imaging specialists counter that the trial was explicitly designed to show AI performs no worse than standard reading, not to prove it reduces the number of women eventually diagnosed with an interval cancer, a much higher bar that has yet to be cleared in U.S. screening populations. That gap between \”not worse\” and \”definitively better\” is exactly the kind of distinction that can get lost as AI mammography tools are marketed to hospitals and patients.
What needs to happen next
The UCLA team is calling for future studies to follow patients prospectively over time, explicitly measure the effect on recall rates and false positives, track radiologist workload in real clinical settings rather than research conditions, and monitor long-term outcomes rather than short-term detection statistics. They are also pushing for formal post-market surveillance of commercial AI systems once they are deployed in hospitals, arguing that a tool’s performance in a curated research dataset does not guarantee the same results once it is running across the varied mammography equipment, patient populations, and screening intervals found in everyday clinical practice. Whether AI mammography ultimately reduces the number of women diagnosed only after a cancer becomes symptomatic between screenings, rather than just detecting it more often in hindsight, remains the open question the field will spend the next several years trying to answer.
Photo: Pexels / PIXABAY via Pixabay