The Single Aging Clock Is Over: How Foundation Models Are Changing Biological Age Measurement

Gaby Hayon, Co-Founder and CTO

Gaby Hayon

Co-Founder and CTO

June 22, 2026

Share:

The Single Aging Clock Is Over: How Foundation Models Are Changing Biological Age Measurement

For most of the last decade, a patient’s biological age arrived as a single number. You drew blood, sent it to a lab running an epigenetic clock, and weeks later a value came back. Forty-seven. Fifty-two. The number felt scientific because it was precise. Precision and accuracy are not the same thing, and the distance between them is exactly why many used these tools cautiously in the clinic. A confident-looking number built on one narrow slice of biology can still be wrong in ways that matter for a real patient.

That single-number era is now ending, and the reason is worth understanding for any physician building a preventive medicine practice.

The birthday number tells us almost nothing about how well a body is actually working. Two people sitting in the same clinic room at fifty-five can have hearts, immune systems, brains, and livers operating a decade apart in either direction. Anyone who has spent time in hospitals knows this intuitively. What is new — genuinely new, in a way that most physicians have not yet absorbed — is that we can now measure the gap, organ by organ, and increasingly cell by cell.

Why a single epigenetic clock was always a compression of biology

Aging shows up in many places at once. It appears in methylation patterns, in circulating proteins, in gene expression, and in the slow drift of dozens of clinical markers over years. A clock built to output one value from one data type was always a compression of something far richer. A model that reads only one of those streams is listening to a single instrument and trying to describe the whole orchestra. That is true even for the well-validated clocks clinicians know, such as GrimAge, PhenoAge, and DunedinPACE. Each is excellent at the narrow job it was trained for, and each sees only part of the picture.

The gap we can now measure organ by organ, cell by cell

The instruments doing the measuring are called biological aging clocks, and they have been quietly evolving for more than a decade. The first generation, developed largely by Steve Horvath and colleagues, used chemical marks on DNA — a process called methylation — to estimate how old a tissue looked at the molecular level. Those clocks were useful research tools. They opened a door. But what has happened in the last three years is a different order of change. Investigators can now draw a tube of blood, run thousands of proteins through a machine learning model, and generate an aging score for the heart, the brain, the immune system, the liver, the lungs, and the kidneys separately. In the most recent work, published in Nature Medicine just weeks ago, the same blood sample yielded aging signatures for more than forty distinct cell types across sixty thousand people.

Tony Wyss-Coray and Eric Topol laid out the whole landscape in an invited review in Nature Medicine this month. Their piece is dense and written for scientists, but the practical implications are worth understanding — carefully, because there is both a real opportunity here and a real risk of over-selling it.

What has changed is the runway. Organ-specific proteomic clocks were first demonstrated in a landmark paper in 2023. Cell-level clocks were published in 2026. The gap between those two milestones was under three years, and the volume of new work behind them is remarkable: dozens of validated cohorts, hundreds of thousands of individuals, and a rate of publication that has roughly doubled since the field moved from single-tissue to multi-organ measurement. We are not at the clinical inflection point yet. We are close enough that preparing for it is the right posture.

What foundation models for aging actually do

This is where the field is heading next, and the shift is captured well by a paper published this spring titled “The End of Aging Clocks: Training Foundation Models to Reason in Aging and Longevity”. The argument is straightforward: a narrow algorithm trained to predict one value is fundamentally limited, so the better path is a large model trained across many data types at once, then asked to reason about what it sees.

From one data type to multimodal reasoning

A foundation model does not simply output a verdict. It integrates genomic, proteomic, imaging, clinical, and longitudinal signals into one coherent read, and it can be interrogated about that read. The frame moves from measurement to reasoning. Instead of asking “what number does this assay produce,” the clinician can begin to ask “what in this patient’s biology is driving the result, and how confident is the model.”

Early results: Longevity-LLM and the first longevity foundation model

The early evidence is encouraging. Longevity-LLM, an initial version released this spring, was fine-tuned on DNA methylation, proteomics, clinical biomarkers, and RNA expression together. It predicted epigenetic age with a mean error of roughly 4.3 years, better than the original Horvath multi-tissue clock that helped define the field. The accuracy is not the real headline. The headline is that one model held four kinds of biology in view at the same time. The industry is moving the same way with serious capital behind it. In May, Insilico Medicine and Human Life Foundation Models announced a collaboration to build the first large-scale foundation model dedicated to human longevity science, trained on genomic, proteomic, imaging, clinical, and longitudinal health data. Insilico had already formed a dedicated longevity board in April to steer exactly this kind of work.

Why one clock will never be enough

The instinct people have when they hear about biological clocks is to ask which one is best. The right instinct is that no single clock is going to earn a place in clinical practice on its own. Recent clocks perform impressively — GrimAge v2, DunedinPACE, and the newer proteomic organ-specific clocks all predict disease and mortality with numbers most cardiologists would happily take. The reason to layer them is not that they are unreliable. The reason is that aging does not happen in one place at one rate.

The clearest demonstration comes from a British birth cohort study that has been running since 1946. Roughly eighteen hundred people, all born in the same week, were followed for six decades. When investigators measured organ-specific aging in this group at age sixty-three, they found a ten-year spread across individuals in how fast different organs had aged. Same chronological age. Very different biological trajectories. In a separate large analysis, about one in four people showed accelerated aging in at least one cell type — and the more organs or cell types running fast, the higher the risk of dying in the following fifteen years.

Two systems consistently stood out as gatekeepers: the brain and the immune system. When those two aged at a normal pace, fifteen-to-seventeen-year survival approached one hundred percent, even in individuals whose other organs looked considerably older. It is an early hint that the brain and the immune system may be doing something like systemic quality control — and that measuring them well may matter more than measuring anything else.

This is why the future is not going to look like “run the clock and act on the score.” It is going to look like layering. A proteomic aging signal means one thing on its own. It means considerably more when combined with a polygenic risk score built from the genome, the ten-year trajectory of blood pressure and lipids pulled from the electronic medical record, a coronary calcium score, wearable-derived data on sleep architecture and cardiorespiratory fitness, and — where clinically indicated — imaging-based aging signals from the retina, the heart, or the brain. Any single one of these can produce a false alarm. Concordance across several is a signal worth acting on. That principle — data layered against data, concordance and directionality before action — is exactly where the foundation-model approach and preventive medicine are both heading.

What reasoning models change at the exam table

A single clock hands you a verdict. A reasoning model opens a conversation. When a model integrates many signals, it can begin to surface which signals are driving an accelerated read, whether a patient’s proteomic profile and methylation profile agree, and where they part ways. That is the kind of output a physician can actually act on. A lone number tells you something is off somewhere. A model that reasons across biology starts to tell you where to look first.

A point I keep returning to with colleagues: these are probabilistic systems, not oracles. A model that reasons fluently can also reason confidently when the underlying data is thin or noisy. That is not a flaw to hide, it is a property to manage. The output is a well-informed estimate that a clinician weighs against the rest of the picture, not a decree that replaces judgment. Used that way, the probabilistic nature becomes a strength, because the model can express uncertainty and point to what would resolve it.

The immune system window

One implication of the new work is worth sitting with, because it is genuinely novel. In routine clinical practice today, there is no test for how well the immune system is aging. We can count white cells. We can measure inflammatory markers. We can react to specific infections. What we cannot do is answer the question a sixty-two-year-old should be able to ask: is my immune system running ahead of, at, or behind my calendar age?

That is becoming measurable. Immune-specific clocks — some proteomic, some built from single-cell profiling of blood — have now been developed and validated in tens of thousands of people. Accelerated immune aging has been linked with worse vaccine responses, higher infection severity, and, in early work, higher risk of neurodegenerative disease. Two recent analyses of the herpes zoster (shingles) vaccine — one from a natural experiment in Canada, one from a U.S. cohort — have separately reported an association between vaccination and lower incidence of dementia, and between vaccination and a slower measured pace of biological aging. Whether that link is causal, and whether it will survive randomized trials, is open. But the fact that we can now begin to ask the question rigorously — that we have an instrument for immune aging where six years ago we had none — is a meaningful shift.

Two stages of prevention

For diseases that develop over decades — Alzheimer’s, atherosclerotic cardiovascular disease, type 2 diabetes, several cancers — the most powerful use of biological clocks will not be in diagnosing sick patients. It will be in stratifying healthy ones early, and then verifying whether interventions are working.

Consider dementia. A person in their fifties with two copies of the APOE4 gene has been told for years that their genetic risk is high, without a good way to know how their brain is actually holding up. The recent data suggest that combining APOE status, blood-based markers like p-tau217, cell-specific clocks derived from brain cell types (astrocytes and microglia in particular), and cognitive testing produces a far more precise picture. In individuals with two APOE4 copies and older-than-expected astrocyte aging, the reported risk of developing Alzheimer’s was nearly forty-fold higher than in those with the same APOE4 status but younger-looking astrocytes. That difference is not marginal. It identifies a group who would benefit most from aggressive early intervention — lifestyle, cardiovascular optimization, sleep, and, as they mature, targeted disease-modifying therapies — and it identifies another group whose risk, despite the genotype, is much lower than we would have called it a decade ago. Just as importantly, the same clocks can be re-measured to check whether the intervention is actually shifting the trajectory. That closes a loop preventive medicine has never had.

A similar logic applies to cardiovascular disease. A middle-aged person with elevated apoB and a family history is one clinical profile. The same person with an accelerated artery clock and a blood-vessel-cell aging signature is a different clinical situation. The case for earlier and more aggressive lipid lowering, blood pressure optimization, and structured lifestyle intervention becomes materially stronger.

The catch: model output is only as good as the data underneath

All of which leads to the unglamorous truth underneath the breakthrough. The quality of what comes out depends entirely on the quality and completeness of what goes in. A patient with three lab panels scattered across two years and a wearable that synced intermittently does not give even a sophisticated model much to reason with. The most advanced longevity model in the world reads a fragmented record and produces a fragmented answer.

An honest word on where this stands today

Two things are true at once. First, we are not there yet. Organ and cell clocks do not belong on the annual physical. The tests are not standardized, not regulated, and not priced for routine use. No professional society has issued guidance on how to act on the results. And a marketplace of direct-to-consumer aging clocks — mostly first-generation methylation tests — is already being sold to the public at high prices and with inconsistent methodology. Topol and Wyss-Coray were direct in their recommendation: do not buy them at this stage. That advice deserves to be taken seriously.

Second, we are closer than most people — including most physicians — appreciate. Within a small number of years, well-validated proteomic organ and cell clocks will begin appearing in specialty preventive medicine practices, and shortly after that in larger health systems. The pace at which the underlying science produces new work is not slowing; it is accelerating.

What this means for longevity clinics

This is the part our team thinks about constantly. These models — and the organ and cell clocks that will feed them — are only as good as the longitudinal record they read from, and most clinics are still assembling that record by hand. Labs live in one system, wearable data in another, the clinical history in a third. Before any reasoning model or organ clock can do something useful for a specific patient, someone has to unify that patient’s data into a clean, continuous, structured picture. The model is the engine. The unified record is the fuel. We built our platform around that second half precisely because the first half is finally arriving, and clinics with a real data foundation will be ready to put these tools to work the moment they mature.

For anyone thinking about their own health in the interim, the useful move is not to chase every new aging test. It is to build the substrate on which those tests will eventually be interpreted: a longitudinal medical record, appropriate imaging, a lipid panel that includes apoB and Lp(a), fitness tracked over time, sleep data, and — when clinically appropriate — polygenic risk information. When organ and cell clocks arrive at scale, they will be additive to that substrate, not a replacement for it.

The shift from clocks to reasoning models is good news for preventive medicine. We are moving from a single annual verdict toward a living, multimodal read of a patient’s biology that a physician can interrogate and explain — organ by organ, and one cell type at a time. Medicine is not moving away from clinical judgment. It is giving physicians and patients something we have never had before: an objective, granular way to see how a body is actually aging, and to shape prevention accordingly. That is not marketing. That is a real shift, and the runway to the clinic is short. The number was never the point. Understanding what drives it always was.

Frequently asked questions

What is a foundation model for aging?

A foundation model for aging is a large artificial intelligence model trained across many types of biological data at once, including methylation, proteomics, gene expression, imaging, and clinical markers. Rather than producing a single biological age from one assay, it reasons across these inputs to estimate aging and to surface what is driving the result.

What are organ and cell aging clocks?

They are aging scores estimated for individual organs or cell types rather than for the body as a whole. From a single tube of blood, proteomic models can now generate separate aging signals for the heart, brain, immune system, liver, lungs, and kidneys, and the most recent work extends this to more than forty distinct cell types. Because organs age at different rates within the same person — a spread of up to a decade in cohort data — these clocks reveal where aging is accelerated, not just whether it is.

Are AI aging clocks accurate enough for clinical use?

Accuracy is improving quickly. Early multimodal models such as Longevity-LLM have already outperformed first-generation epigenetic clocks on age prediction, and organ-specific proteomic clocks predict disease and mortality well. For clinical use, the more important point is that these are probabilistic estimates meant to support a physician’s judgment, not to replace it, and their reliability depends heavily on the completeness of the patient data behind them. They are not yet standardized or regulated for routine practice.

Should I buy a direct-to-consumer aging clock test?

Not at this stage. Most consumer products on the market are first-generation methylation tests sold at high prices with inconsistent methodology, and the researchers leading this field — including Tony Wyss-Coray and Eric Topol — have advised against buying them for now. The validated organ and cell clocks are still emerging from research cohorts and are not yet available as reliable consumer tests.

Can we measure how well the immune system is aging?

Increasingly, yes. Immune-specific clocks, built from proteomics or single-cell profiling of blood, have been validated in tens of thousands of people. Accelerated immune aging has been associated with weaker vaccine responses, more severe infections, and, in early work, higher neurodegenerative risk. This is genuinely new — routine clinical practice has had no test for immune aging until now.

Do foundation models replace epigenetic clocks like GrimAge or DunedinPACE?

Not exactly. Established clocks remain useful and well-validated for the specific outcomes they were built around. Foundation models aim to integrate the signals those clocks capture, alongside many others, into a single reasoning system. The emerging consensus is not that one clock wins, but that clocks should be layered — proteomic, genomic, imaging, and clinical signals read together, with concordance across them as the thing worth acting on.

What does a clinic need to use these tools effectively?

The essential prerequisite is unified, longitudinal patient data. Labs, wearables, imaging, and clinical history need to live in one structured record so the model has something coherent to reason from. Clinics that solve data unification first will get far more value from longevity AI — and from organ and cell clocks as they arrive — than those running the most advanced model on fragmented inputs.

Read More

Start instantly. 10 minute onboarding. Risk-free.

And we'll get back to you asap with available slots

ask ai about us
chetgpt
ChatGPT
cloude
Claude

We use cookies to enhance your experience and analyze site usage. Your privacy matters to us.