Uncategorized

Most FDA-Cleared Health AI Was Never Tested to Work Equally Well on Women, New Analysis Shows

A September 2026 review found fewer than a third of 903 FDA-authorized AI medical devices were validated on sex-specific populations, and as of this year developers can still win clearance without proving their tools perform equally for women.

Most FDA-Cleared Health AI Was Never Tested to Work Equally Well on Women, New Analysis Shows

A wave of research examined in reporting published September 18, 2026 has documented a persistent blind spot in how the Food and Drug Administration clears artificial intelligence medical devices: the agency does not require developers to prove their tools work equally well for women before granting market authorization. A cross-sectional analysis of 903 FDA-authorized AI devices, published in JAMA Network Open, found that less than one-third had been validated on sex-specific populations at all, and separate figures show that of 168 machine-learning devices authorized in 2024 alone, only 15.5% provided any demographic performance breakdown to regulators.

How the gap slips through the FDA’s process

The core issue traces back to the regulatory pathway most AI devices use to reach the market. The FDA’s 510(k) fast-track clearance process, which covers the large majority of AI-enabled medical devices, does not require developers to conduct prospective clinical studies before authorization — instead allowing companies to demonstrate their device is “substantially equivalent” to an already-cleared product. Roughly a quarter of AI device submissions include no clinical performance study whatsoever. Even among devices that do report performance data, fewer than one in four address age-related subgroups, and overall sensitivity and specificity metrics — the basic statistics needed to judge how well a diagnostic tool performs — were reported for only 29.2% of 2024’s cleared machine-learning devices. The agency issued draft guidance in January 2025 encouraging developers to report demographic performance data, but as of September 2026 that guidance remains advisory rather than mandatory, meaning a device can still clear the FDA without ever demonstrating it performs equally well across sexes.

Where the bias has already shown up

The concern isn’t merely theoretical. Researchers at MIT’s Jameel Clinic, in a study led by Marzyeh Ghassemi and a colleague named Pan, found that large language models including OpenAI’s GPT-4, Meta’s Llama 3 and the healthcare-focused model Palmyra-Med recommended lower levels of care for women presenting with the same symptoms as men, with some models advising women to “self-treat at home rather than seek clinical care” in scenarios where equivalent male patients were directed toward in-person evaluation. A separate analysis published in BMC Medical Informatics by researcher Sam Rickman examined Google’s Gemma model — used by more than half of English local government authorities in social-care contexts — using 617 real care records, and found the model “downplayed women’s health needs” in its outputs. In imaging specifically, a study in the Proceedings of the National Academy of Sciences found that chest-imaging AI performance degraded measurably when the sex composition of the population being tested differed from the sex composition of the data the model was trained on. A widely cited example involves a DeepMind kidney function prediction model that was trained on VA hospital records that were 94% male.

Why the underlying data is skewed to begin with

Part of the problem predates AI entirely and reflects longstanding gaps in clinical research itself. Women made up only about 30% of cardiovascular research trial participants between 2017 and 2022, according to American Heart Association data, even though cardiovascular disease remains a leading cause of death for women. A 2020 analysis found women experience adverse drug reactions roughly twice as often as men, a disparity partly rooted in trial populations that historically skewed male. That imbalance traces back to a 1977 FDA guidance that excluded women of childbearing age from early-phase drug trials in the wake of the 1960s thalidomide crisis — a policy that remained technically in force until the NIH Revitalization Act of 1993. The consequences of that era persisted well beyond the policy change: the sleep drug zolpidem wasn’t found to require a lower dose for women, due to how differently their bodies metabolize it, until more than two decades after the drug reached the market, with the FDA halving the recommended female dose in 2013. AI models trained on decades of clinical data inherit those same historical imbalances.

Real-world stakes in emergency care

The disparities aren’t confined to research settings. An analysis of acute coronary syndrome cases found 74% of women reported chest pain compared with 79% of men, a gap that has long complicated diagnosis given chest pain’s central role in cardiac triage. In a review of 19,331 out-of-hospital cardiac arrest cases, only 39% of women received bystander CPR compared with 45% of men — a disparity researchers have partly linked to the fact that 75% of CPR training manikins depict male anatomy or provide no sex-specific anatomical detail at all, potentially shaping how bystanders are trained to recognize and respond to a woman in cardiac arrest.

Diverging paths on regulation

The United States and European Union are increasingly diverging on how they handle this problem. The EU’s AI Act classifies medical device AI as high-risk, which carries mandatory bias testing requirements, with full enforcement expected by 2027. In the U.S., by contrast, the FDA’s January 2025 draft guidance on demographic reporting remains non-binding, and a 2026 review of 574 published studies on health AI found that while 61% included both sexes in their study populations, only 44% actually analyzed their results broken down by sex — meaning even when women are included in a dataset, their outcomes often aren’t separately examined for signs the model performs worse for them.

What’s next

Patient advocacy groups and some researchers are pushing the FDA to convert its draft demographic-reporting guidance into a binding requirement, arguing that voluntary disclosure has produced years of data showing most developers simply don’t report the numbers unless required to. Absent a regulatory mandate, the pressure may instead come from purchasers: as more U.S. hospital systems examine AI vendor contracts more closely following consortium efforts around diagnostic AI governance, some health systems could begin demanding sex-disaggregated performance data as a condition of adoption, even where the FDA does not require it. Whether that market pressure moves faster than formal rulemaking is likely to determine how quickly this gap actually closes.

Photo: OsloMetX / PIXABAY via Pixabay