Uncategorized

Insilico Medicine Opens Its AI Longevity Toolkit to the World in a Landmark Cell Study

Insilico Medicine published an open-source AI toolkit for aging research in Cell on September 17, 2026, with Liquid AI, the Buck Institute, and Harvard Medical School, nominating 328 candidate longevity gene targets.

Insilico Medicine Opens Its AI Longevity Toolkit to the World in a Landmark Cell Study

On September 17, 2026, Hong Kong-listed Insilico Medicine landed the cover of Cell with a paper that hands the rest of the aging-research world a set of tools it previously kept in-house. The release bundles three pieces: LongevityBench, an open benchmark for testing how well AI systems reason about aging biology; Longevity-LLMs, a family of five compact open-source language models tuned specifically on clinical and multi-omics aging data; and Longevity Claw, an agentic research platform that autonomously chains together specialized tools to hunt for therapeutic targets. The work was carried out with Liquid AI, the Buck Institute for Research on Aging, Harvard Medical School, and Brigham and Women’s Hospital, according to the company.

What the Toolkit Actually Does

LongevityBench evaluates AI reasoning across five domains that aging researchers juggle simultaneously: clinical records, genetics, epigenetics, transcriptomics, and proteomics. Longevity-LLMs range from 0.6 billion to 9 billion parameters, deliberately small so they can run on modest lab hardware rather than requiring frontier-scale compute. Longevity Claw then wraps a chosen model in an agentic loop that performs gene-set enrichment analysis, calculates aging clocks, and ranks candidate intervention targets without a human manually stitching each step together. Insilico founder and co-CEO Alex Zhavoronkov said the goal is to let the broader scientific community build on tools the company had been refining internally for years.

The Bottleneck This Is Meant to Fix

Aging research has a data problem before it has a drug problem. Genomic, epigenomic, proteomic, microbiome, and lifestyle datasets are collected in incompatible formats across dozens of labs, and there has been no agreed-upon way to test whether an AI model actually understands aging biology or has simply memorized correlations. Insilico built its reputation on generative models for drug discovery and has spent the past several years pushing candidates from its Pharma.AI pipeline into human trials. LongevityBench grew out of an internal need: before trusting a model to prioritize a therapeutic target, the company needed a rigorous way to score it. Making that scoring system public is a bet that a shared standard will pull the field together faster than proprietary silos would.

How the Models Actually Performed

The paper evaluated 26 AI systems, including 18 frontier models from OpenAI, Google, Anthropic, xAI, DeepSeek, and Moonshot AI. The best overall score came from L-Qwen3.5-9B, one of Insilico’s own fine-tuned models, but the more striking result was that its smallest variant, at just 0.6 billion parameters, outperformed most of the general-purpose frontier systems on aging-specific reasoning tasks. Running the pipeline through Longevity Claw, the researchers say the system nominated 328 genes as potential aging intervention targets, showing a 5.6-fold enrichment for genes already validated in prior aging research compared to random selection.

Not Everyone Is Convinced Benchmarks Translate to Biology

Longevity researchers outside Insilico have welcomed the transparency but caution against reading benchmark performance as a proxy for real-world drug success. The field has a history of enthusiasm outrunning evidence: epigenetic clocks and aging biomarkers have proliferated over the past decade, yet very few have been shown to predict which interventions actually extend healthy lifespan in humans. A model correctly flagging a gene that other researchers separately validated is encouraging, but it is a long way from a compound in a Phase 1 trial. Skeptics also point out that fine-tuning smaller models on curated aging datasets can inflate benchmark scores in ways that don’t generalize to messier, real-world clinical data, precisely the kind of gap that sank earlier waves of AI-in-biology hype. Insilico’s own answer is that LongevityBench was designed with input from wet-lab collaborators at the Buck Institute and Harvard specifically to guard against that kind of overfitting, though independent replication will take time.

What Happens From Here

Because the toolkit is open-source, smaller academic labs and biotechs that could never afford to build their own aging-specific foundation models now have a starting point, which could compress the time between identifying a target and getting it into a wet lab for validation. Insilico says it plans to use Longevity Claw’s target list to guide its own experimental partnerships, including further work with the Buck Institute and Harvard-affiliated Brigham and Women’s Hospital. The real test will be whether any of the 328 nominated genes survive contact with laboratory validation and, eventually, whether a therapy built around one of them reaches patients. That process typically takes years, not months, so the Cell paper marks the start of a much longer experiment rather than its conclusion.

Photo: PublicDomainPictures / PIXABAY via Pixabay