Uncategorized

New AI Model Reads Heart Disease Risk From an ECG Using 91% Less Training Data

Scripps Research published ECG-CLIP on September 1, 2026, an AI foundation model that detects heart disease from ECGs using 91% less labeled training data than existing tools, potentially speeding up detection of rarer cardiac conditions.

New AI Model Reads Heart Disease Risk From an ECG Using 91% Less Training Data

Scripps Research scientists published a new AI foundation model called ECG-CLIP in The Lancet Digital Health on September 1, 2026, trained on more than 1.7 million electrocardiograms collected from over 540,000 people and paired with clinicians’ written notes. The standout claim is efficiency: ECG-CLIP matches the performance of existing cardiac AI algorithms while requiring roughly 91% less labeled training data, and it can learn to detect a specific cardiovascular condition after seeing as few as a dozen confirmed examples.

Who Built It and How It Was Tested

The study, formally titled a contrastive learning foundation model for ECG-based prediction of cardiovascular diseases and outcomes, was led by senior author Giorgio Quer, an assistant professor of digital medicine at Scripps Research, together with co-first author Michael Ko, a graduate student and research assistant at the institute. After training on the 1.7 million-ECG corpus, the team validated ECG-CLIP against a separate holdout set of more than 800,000 additional ECGs, benchmarking it against standard deep learning models, simple linear models, general-purpose foundation models, and three other ECG-specific foundation models, all scored using area-under-the-curve (AUC) methodology. In head-to-head testing on three specific conditions, acute myocardial infarction, cardiac amyloidosis, and hypertrophic cardiomyopathy, ECG-CLIP consistently matched or outperformed the best conventional models while relying on a small fraction of the labeled examples those systems required.

Why Data Efficiency Matters More Than Raw Accuracy

Most AI diagnostic tools improve by training on enormous, carefully labeled datasets, a resource-intensive process that requires cardiologists to manually annotate thousands of ECGs for each new condition a model is meant to detect. That bottleneck has historically limited AI heart-disease tools to only a handful of conditions, common arrhythmias, for instance, where huge labeled datasets already existed. ECG-CLIP’s contrastive learning approach instead pairs raw ECG waveforms directly with the free-text clinical notes doctors already write, letting the model learn generalizable relationships between ECG patterns and disease language without needing every condition individually hand-labeled at scale. In this study, that approach was tested specifically against acute myocardial infarction, cardiac amyloidosis, and hypertrophic cardiomyopathy, three conditions where labeled ECG data is scarce or costly to compile, and the researchers say the same contrastive-learning approach could plausibly extend to atrial fibrillation, chronic kidney disease, and even type 2 diabetes risk prediction directly from ECG waveform signal alone.

Mimicking How Clinicians Actually Learn

Researchers describe the model’s few-example learning capability as mimicking the way physicians learn from general physiological principles rather than memorizing massive datasets. In practice, that means ECG-CLIP could potentially be adapted to flag rarer cardiac conditions, ones too uncommon to have generated the huge labeled datasets larger models require, simply by feeding it a small number of confirmed examples alongside its existing broad physiological understanding. Notably, the Scripps team also demonstrated that ECG-CLIP’s few-shot detection held up on single-lead ECG recordings, the simplified waveform format captured by consumer wearables rather than the full 12-lead setup used in hospitals, hinting at a more direct path toward wearable-based deployment than many earlier foundation models achieved.

Building on a Wave of Smartwatch Cardiac AI

ECG-CLIP arrives alongside a broader push to bring AI cardiac screening to wearable devices. Separate research has already demonstrated that AI algorithms paired with single-lead ECG sensors on consumer smartwatches can accurately flag structural heart disease, including weakened pumping function, damaged valves, and thickened heart muscle, conditions traditionally requiring an echocardiogram to detect. Other studies have shown AI-enabled smartwatch ECG monitoring can predict heart failure rehospitalization by identifying early physiological precursors, opening the door to earlier, less invasive intervention.

The Caveats Cardiologists Are Raising

Foundation models trained on hospital-derived ECG and clinical-note data inherit whatever biases exist in that underlying patient population and documentation style, and Scripps researchers acknowledge that external validation across more diverse patient populations, different hospital systems, different countries, different equipment, remains an essential next step before the model could be trusted in routine clinical use. There is also a familiar concern about explainability: even when the model performs well statistically, clinicians want to understand why it flagged a particular ECG as high-risk before acting on that flag, and foundation models trained via contrastive learning on unstructured text can be harder to interpret than simpler rule-based systems.

What Comes Next

The Scripps team’s next phase involves prospective validation studies testing ECG-CLIP against real-world patient outcomes rather than retrospective data, the standard bar any diagnostic AI tool must clear before regulatory clearance becomes realistic. If validated, the model’s low-data efficiency could meaningfully accelerate how quickly AI tools get built for rarer cardiac conditions, potentially compressing what has historically been a multi-year data-collection bottleneck into a matter of months for conditions affecting far smaller patient populations than common arrhythmias. Cardiologists outside the Scripps team have also flagged the model’s reliance on clinicians’ free-text notes as a double-edged sword: it lets the model learn richer, more nuanced associations than rigid structured labels would allow, but it also means the model absorbs whatever inconsistencies, abbreviations, and idiosyncratic language exist across different hospitals’ documentation habits, a variability that external validation studies will need to specifically test for before regulators consider broader clearance.

Photo: Los Muertos Crew / PEXELS via Pexels