← Longevity & Supplement Guides
Readiness scores: the most confident number on your wrist
There is a particular kind of morning that begins with a verdict. You wake, you reach for the wrist or the nightstand, and before your feet have found the floor a number has already decided what sort of day you are allowed to have. Sixty-one. Amber. Take it easier today.
It is a remarkable piece of theatre, and it is getting more crowded. Over the past weeks another major smartwatch platform added a morning readiness score to its lock screen, which means that essentially every mainstream wearable on sale now grades your recovery before breakfast. The scores arrive with two digits of precision and a colour. They arrive without a citation.
We build a wearable and an app, so this is our business, not a drive-by. Which is exactly why the honest version is worth saying out loud: the number on your wrist is assembled from decades of real physiology, and then tested almost not at all.

The ingredients are researched. The recipe is not.
In April 2025, a team led by Cailbhe Doherty at University College Dublin did something nobody had bothered to do: they went through the composite scores on the market and asked what each one is actually made of. They found fourteen of them across ten manufacturers — readiness, recovery, strain, resilience, energy, body resources, all the household words of the category.
The finding was not that the scores are wrong. It was that you cannot tell. The underlying algorithms are proprietary, the weightings are undisclosed, and the physiological meaning of the final figure is, in the authors' assessment, unclear. Two devices can look at the same night and disagree by twenty points, and there is no scale on which one of them is more right.
This is the part that gets skipped. Heart rate variability has a literature going back to the 1990s. Sleep architecture has a literature going back further. Training load has been argued over by sports scientists for forty years. But readiness — a specific secret function of those three things, expressed as a number out of one hundred — is a new object, and it does not inherit the evidence of its ingredients. You can build a bridge from certified steel and still not have a certified bridge.
Figure
Three measured inputs, one private number, one instruction — and where the testing stops
The inputs. Optical pulse variability, sleep timing and movement load all have published accuracy studies. They are imperfect and they are checkable.
The number. Fourteen composite scores across ten makers, all proprietary. No published validation of the final figure itself.
The advice. Almost no trials ask whether people who follow the daily instruction end up better off than people who ignore it.
What the sensor on your wrist is actually watching
Start one layer below the score, at the green light against your skin. A wrist optical sensor does not record the electrical signal of your heart. It records the timing of blood pulses, which is a close cousin of heart rate variability but not the same measurement — pulse rate variability, to give it its proper name.
How close a cousin? A 2026 meta-analysis in Sensors pooled 43 studies and ran the numbers on ten of them. Under controlled resting conditions, the agreement is respectable: a pooled standardised error of 0.188 for RMSSD and 0.134 for SDNN, the two variability measures most wearables lean on, with almost no heterogeneity between studies. In a quiet room, sitting still, the wrist does a decent job.
Then comes the sentence that should be printed on every readiness dashboard. The authors warn that their result should not be generalised to sleep, exercise, stress or free-living conditions, where agreement was markedly more variable. Which is to say: the measurement is best validated in precisely the circumstances in which nobody ever takes it, and least validated overnight, which is when your readiness score is computed.
That is not an accusation of fraud. It is a description of a gap, and gaps are worth naming before they are papered over with a colour.
The question nobody has answered yet
Suppose we grant the sensor its accuracy. There is still a harder question hiding underneath, and it is the one that actually matters to the person standing in the kitchen at seven in the morning: if I do what this number says, am I better off?
Writing in the Journal of Medical Internet Research in August 2026, Anna Zucker put the question to clinicians who work with active patients. The answers were more careful than the marketing. Readiness scores do not predict injury in any straightforward way. What is worth reading is the trend over weeks, not the digit on a given morning. Of the three ingredients, sleep turns out to be the strongest and most consistent signal. And smartwatch heart rate variability, taken on its own, should not be trusted as a verdict.
A narrative review published this year in the Journal of Orthopaedic Reports reached the same place from the clinical side. Across sixteen studies of physically active wearable users, the data reflected recovery capacity and load exposure — real things, worth knowing — rather than forecasting who was about to get hurt. Read longitudinally and combined with an actual assessment of the actual person, the trends were useful context. Read as a daily oracle, they were not.
The trial that would settle it is easy to describe and rarely run: take a few hundred people, randomise half to follow their score and half to ignore it, and see who is fitter, more consistent and less broken a year later. Until somebody runs it at scale, the instruction on your screen is a hypothesis wearing the clothes of a result.
Where the evidence does exist, it is a rule — not a number
Here is the part that makes this more interesting than a simple debunk. There is good evidence that letting day-to-day variability shape your training works. It just is not evidence for the score.
A 2020 meta-analysis by Granero-Gallegos and colleagues gathered six randomised controlled trials of heart-rate-variability-guided training in endurance athletes — 195 people in total, most of them running eight-week interventions after a few weeks of baseline. The athletes whose hard sessions were scheduled by their own morning variability improved their VO₂max more than the athletes following a fixed plan written in advance.
Look closely at what was tested. Not a proprietary index. A transparent, boring, inspectable rule: establish your personal rolling baseline over several weeks, train hard when today's reading sits inside your own normal range, and go easy when it falls below it. Anyone can state that rule, check it, argue with it, and reproduce it. That is what makes it evidence rather than a product feature. We wrote about the trials in more detail in what HRV-guided training actually shows.
The irony is neat enough to be worth dwelling on. The version with research behind it is the one you can understand. The version with no research behind it is the one that thinks for you.
How to use the number without being used by it
None of this is an argument for throwing the wearable in a drawer. Measurement is how you notice drift you would otherwise rationalise — the slow slide from well to fine that nobody feels happening. But a number earns its place by correcting your fantasies, not by replacing your experience.
Six ways to read a morning score honestly
Read the band, not the digit. One morning is noise. Your own rolling range over a week or two is the signal, and the only comparison worth making is with yourself a fortnight ago.
Ask what moved it. Open the breakdown before you accept the verdict. If the score dropped because you slept five hours, you did not need an algorithm to tell you that, and the fix is not a lighter workout.
Use it as a question, not an instruction. "Is there a reason I might be under-recovered?" is a useful prompt. Answer it from your body first — then look at the number and see whether the two of you agree.
Let it move the dose, not the decision. On most low mornings the sensible response is one notch down in volume or intensity, not a cancelled week. Genuinely low stretches that last are what a planned deload is for.
Be suspicious of a score that never disagrees with you — and of one that always does. A metric that only ever confirms the plan is decoration. A metric that vetoes every hard session is miscalibrated to your baseline, not to your life.
Never compare across brands. There is no shared scale. A 74 on one device and a 74 on another are two different opinions in the same costume.
This is the philosophy behind how we show recovery trends in the Agen app and on the Agen Band: trends you can interrogate, with the components visible, framed as wellness estimates rather than measurements of clinical grade. They are not diagnoses and they are not instructions. They are a way of seeing something that is otherwise invisible until it becomes a problem.
If you want the same scepticism applied elsewhere, we have pointed it at the sleep score and at biological age tests, both of which fail in the same family of ways. And the oldest instrument for all of this is still the one you were issued at birth: the ability to notice your own body, which turns out to be trainable.
The bottom line
Readiness scores are built from measurements that are real, combined by arithmetic nobody will show you, into advice that almost nobody has tested. That is three claims, and only the first one is on solid ground. The sensible response is neither to obey the number nor to sneer at it, but to demote it: from verdict to input, from instruction to question, from a single morning digit to a line you watch over weeks.
The transparent rule beats the confident number. It always has. The difference is that the confident number is easier to sell, and the wrist is a very good place to sell it from.


