Insilico Medicine published a study in Cell on September 17, 2026 introducing LongevityBench, an open benchmark of 17 tasks testing whether today’s leading AI systems can actually make sense of the biological data scientists use to study aging — and the results were humbling. Eighteen frontier AI systems from six developer teams, including models from OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI, were put through the benchmark, and none of them dominated across all tasks.
What LongevityBench Actually Tests
The benchmark spans five categories of biological data: clinical measurements, genetics, epigenetics, transcriptomics and proteomics. Tasks range from comparing biological samples and classifying disease states to generating conditional molecular profiles — essentially asking a model to predict what a cell’s molecular signature would look like under a specified biological condition. The hardest task across the board, researchers found, was predicting biological age directly from omics data (the various layers of molecular information in a cell), which held up as difficult regardless of how large or capable the underlying model was.
The Surprise: Small Models Held Their Own
Perhaps the most striking finding wasn’t that large frontier models struggled — it’s that the researchers’ own compact models, dubbed Longevity-LLMs and ranging from just 0.6 billion to 9 billion parameters, matched or exceeded the performance of far larger general-purpose frontier systems on the aging-specific tasks. That suggests domain-specific fine-tuning on biological data can outperform sheer scale, a finding with implications well beyond longevity research, since it argues against the assumption that bigger general models will always win at specialized scientific tasks.
What Insilico Is Giving Away
In an unusual move for a commercial biotech company, Insilico released the full LongevityBench dataset, the trained Longevity-LLMs themselves, the evaluation code, and a tool called Longevity Claw — an agentic research interface that pairs the released models with domain-specific tools aimed at helping aging researchers run their own analyses. Opening up a benchmark and models this thoroughly is intended to let independent labs verify results and build on the work directly, rather than taking the company’s performance claims on faith.
Why This Matters for the Booming Biological-Age Testing Market
The timing lines up with rapid growth in consumer interest in biological age: roughly 4.2 million biological age tests have been administered globally as of 2026, according to market research, a roughly 31% jump over 2024 figures, with the broader biological-age and longevity diagnostics market projected to reach $5.58 billion by 2030. Most consumer-facing biological age tests today rely on epigenetic clocks measuring DNA methylation patterns rather than the kind of general-purpose AI models LongevityBench evaluates, but the benchmark’s findings raise a pointed question for that industry: if frontier AI systems still struggle with omics-based age prediction in a rigorous academic benchmark, how reliable are the proprietary algorithms behind consumer longevity tests marketed directly to the public?
Believers and Skeptics in the Longevity AI Race
Longevity researchers who welcome the benchmark argue that a shared, adversarial testing standard is exactly what the field needs after years of biotech and consumer-wellness companies making biological-age claims that were difficult for outsiders to independently verify. Skeptics counter that even a rigorous benchmark doesn’t resolve the deeper scientific uncertainty about whether biological age, as a construct, is a reliable predictor of actual disease risk or lifespan for any individual person, rather than just a population-level statistical pattern — a debate that predates AI entirely and that no benchmark alone can settle.
What Happens Next
Expect competing AI labs and biotech firms to run their own models against LongevityBench now that it’s public, both to validate Insilico’s findings and to see whether their own systems can beat the compact Longevity-LLMs on the hardest tasks. For the broader field, the benchmark’s central finding — that domain-specific small models can outperform frontier giants on aging biology — is likely to shape how future longevity AI products get built, pushing companies toward specialized, purpose-trained systems rather than simply wrapping consumer products around general-purpose large language models.
Photo: Tima Miroshnichenko / PEXELS via Pexels