← Longevity & Supplement Guides
There are fewer non-responders than you think
Somewhere this morning, a reasonable person is quietly deciding that they are broken.
They did the eight weeks. They logged the sessions, they held the intensity, they did not skip the boring ones. Then the test came back within a rounding error of where it started. And somewhere on the internet they found a word for it, which felt at first like relief: non-responder.
It sounds like a diagnosis. It is much closer to a filing error.

The word arrived before the evidence did
“Non-responder” has quietly become a personality type — a thing people say about themselves at dinner, in the same tone they might mention being bad at maths. Part of its appeal is that it explains the disappointment without asking anything further of you. It is the rare fitness idea that flatters fatalism and effort at the same time: you tried, and your genes declined.
The research that made the term famous never claimed that. What it claimed was narrower, stranger, and considerably more useful.
Where the idea actually came from
In the 1990s the HERITAGE Family Study put 481 sedentary adults from 98 two-generation families through an identical, supervised, 20-week cycling programme and measured what happened to their VO₂ max. The average gain was about 400 ml/min. But the spread was the finding: some participants improved by essentially nothing, and others gained more than a litre a minute — on the same programme, in the same lab.
The variance between families was about 2.5 times the variance within them, and the best-fitting models put the maximal heritability of the response at 47%. Relatives answered the same training more alike than strangers did.
That is a real and important result. It is also routinely asked to carry a conclusion it cannot hold. HERITAGE tested one dose. Everybody got the same programme — that was the entire point of the design, and it is precisely the boundary of what the study can tell you. It described how a population scattered around one stimulus. It never said that scatter was a permanent property of the people in it.
Then someone changed the dose
In 2017, David Montero and Carsten Lundby ran the experiment that HERITAGE structurally could not. Seventy-eight healthy young men did the same six-week endurance programme, split into five groups that differed only in how often they trained: one, two, three, four or five 60-minute sessions a week. Non-response was defined honestly, against the typical error of the measurement itself — a change in maximal power output falling within ±3.96% of your own baseline.
At one session a week, 11 of 16 men were non-responders. At two, 6 of 15. At three, 4 of 14. At four sessions a week: none of 17. At five: none of 16.
Then came the part that should have ended the conversation. Everyone who had been classified a non-responder did a second six weeks with two extra sessions per week. Afterwards, non-response was not observed in a single individual.
A label that dissolves when you add two hours a week was never describing the person. It was describing the prescription.
Dose, not destiny
Share of participants classified “non-responders,” by how much training they were given
Montero & Lundby 2017 — 6 weeks, 78 healthy young men
Ross et al. 2015 — 24 weeks, 121 sedentary middle-aged adults, 62% women
It held up in people who weren’t young men
The obvious objection to Montero and Lundby is the sample: 78 young men, six weeks, one outcome. So it matters that the same pattern had already turned up in a very different room.
Robert Ross and colleagues followed 121 sedentary, middle-aged adults with abdominal obesity — 75 of them women, average age 53 — through 24 weeks of five supervised sessions a week, with at least 90% attendance. Three arms: a low amount at low intensity, roughly double that amount at the same low intensity, and that higher amount at higher intensity (75% of VO₂ peak).
Non-responders at 24 weeks: 15 of 39 in the low-amount, low-intensity group. 9 of 51 when the amount roughly doubled. Zero of 31 when the amount was held and the intensity was raised instead. Doubling the amount at fixed intensity halved non-response; raising the intensity at fixed amount removed it altogether.
One detail in that paper deserves more attention than it gets. In the low-intensity groups, the number of non-responders fell early and then plateaued — people stopped converting. In the higher-intensity group, they kept converting across the whole 24 weeks. Patience is genuinely rewarded, but only by a stimulus capable of paying it back. Waiting longer at a dose that has stopped working is just waiting.
Sometimes it isn’t the dose, it’s the shape
There is a third possibility, and it is the one people almost never test on themselves: right amount, wrong kind.
Jacob Bonafiglia and colleagues had 21 recreationally active adults do both three weeks of steady endurance riding (30 minutes at about 65% of VO₂ peak work rate) and three weeks of sprint intervals (eight 20-second efforts at about 170%), in random order, separated by a three-month washout. Same people, two stimuli, everything else matched.
The correlation between an individual's endurance response and their interval response in VO₂ peak was r = 0.14, p = 0.57. Which is to say: none. Knowing how someone answered one protocol told you essentially nothing about how they would answer the other. Six participants failed to improve on some measure after one protocol, and every single one of them improved on at least one measure after the other.
It is a small study with short blocks, and “improved on something” is a generous bar. But the direction is what matters. The same person can be a responder and a non-responder in the same month, depending on which stimulus you hand them. That is not a trait. That is a mismatch — and mismatches are fixable in a way that traits are not.
The most uncomfortable finding is that you probably can’t tell
Here is where the story stops being reassuring.
To claim that people differ in trainability, you need to separate real individual differences from ordinary noise — the test's own error, the night you slept badly, the second coffee, the technician. Doing that requires a control group that did no training at all, so you can see how much a “change score” moves on its own. Very few exercise studies include one.
In 2024, a team led by James Renwick screened 32,968 studies and found 24 that met that bar. Their conclusions were blunt: most of the variation in observed change scores after an intervention is measurement error; a single study is far too small to estimate the true spread of individual response with any accuracy; and pooling across all 24 gave no strong evidence that the spread is greater than zero at all. They went further still, writing that it appears unlikely clinically relevant predictors of VO₂ max response will be discovered.
This is a live disagreement in the field, not a settled verdict — HERITAGE's familial signal is real and has not been explained away. But notice that both sides of the argument point the same direction for you, personally. If trainability differences are real, one eight-week block cannot measure yours. If they are mostly noise, one eight-week block certainly cannot. Either way, the number you are holding is not the verdict you think it is.
What to do instead of accepting the label
None of this is an argument for training harder as a matter of principle. It is an argument for exhausting the cheap explanations before reaching for the expensive one.
Change the dose before you change your story
Change the shape, not just the size
Give it longer than six weeks — but only at a dose that is still moving
Judge on more than one number, and measure it twice
The number is a proxy, not the point
There is a version of this post that ends with “so train harder,” and I don't want to write it. The deeper problem with the non-responder story is not that it is wrong about genetics. It is that it quietly assumes the number is the reason you were training.
VO₂ max is a good proxy — one of the best single numbers we have for how a body is holding up over decades, which is why we wrote a whole guide to it. But it is still a proxy. A test that fails to move has told you something about the test, the dose, and the interval you chose. It has told you nothing about whether the stairs feel easier, whether you sleep through the night, or whether you are the kind of person who still goes.
Use the numbers to correct the fantasy — including the fantasy that you are a special case. Don't let them replace the experience.
The bottom line
Genuine differences in trainability probably exist; the honest state of the evidence is that we cannot yet measure yours, and that most of what looks like non-response in the wild is under-dosing, a mismatched stimulus, or the noise in a single test. Before you accept a label that no study has been able to make stick to an individual, add the sessions, change the shape, extend the window, and measure more than one thing more than once. If it still hasn't moved, you will at least have learned something true — which is more than the label ever offered.
Related reading: how much exercise per week really, sprint intervals vs moderate exercise, HRV-guided training, and the evidence on deload weeks.


