← Longevity & Supplement Guides
How close to failure do you need to train?
There is a moment near the end of a hard set when the bar slows down and your face does something you would not want photographed. Gym culture has decided that moment is the point — that everything before it was rehearsal.
It is a seductive idea, because effort is almost impossible to measure from the outside and failure is not. Failure is binary. You either finished the rep or the rep finished you. So it became the proxy for sincerity, and everything short of it became suspect: you left something in the tank, said in the tone reserved for moral lapses.
The research of the last four years has quietly taken that idea apart — not by showing failure does nothing, but by pricing it. The last rep is real. It is also the most expensive thing you can buy in a gym, and for most of what most people want, it is optional.
One person, two legs, eight weeks
The cleanest experiment on this is also the strangest to picture. In 2024, Martin Refalo and colleagues at Deakin University took eighteen resistance-trained adults and randomised them against themselves. Each person's legs were assigned separately: one leg did the leg press and leg extension to momentary muscular failure, the other stopped with one to two repetitions in reserve. Twice a week, for eight weeks.
The design is the whole point. When you compare two groups of people, you are also comparing their genetics, their sleep, their protein intake and their enthusiasm. When you compare someone's left quadriceps to their right, all of that cancels.
Quadriceps thickness increased by about 0.18 cm on both sides. Not similar in the hand-waving sense — the same number to three decimal places, with the small remaining differences pointing in opposite directions depending on which muscle was measured.
One detail is easy to skim past and shouldn't be. The failure leg consistently lost more repetitions from the first set to the last. It spent harder early and then had less to give, so across eight weeks the two legs ended up doing roughly the same total work. Failure did not add effort. It moved it.
What happens when you pool the whole literature
One clever trial is an anecdote with error bars. So here is the pooled picture, which is less flattering to both sides of the argument than either would like.
Evidence
Training to failure vs stopping short — pooled effect sizes, with 95% intervals
Read the plot from the top and the argument takes care of itself. For strength, the estimate sits slightly on the wrong side of zero. For muscle size across all adults, the point estimate looks encouraging and the interval is wide enough to drive a bus through. Only two estimates clear or graze zero, and both are small.
The most telling row is the second from the bottom. When Refalo's team restricted the comparison to true momentary failure — not "a hard set", not "high velocity loss", but the rep that actually doesn't happen — the advantage vanished into the noise. The strictest test of the failure hypothesis is the one it passes least well.
Strength and size don't want the same thing
In 2024 a group led by Zac Robinson tried a better question. Instead of asking whether failure beats not-failure, they treated proximity to failure as a dial and asked what the dose–response curve looks like: 55 hypertrophy effects and 67 strength effects, regressed against estimated repetitions in reserve, adjusted for load, volume, study duration and training experience.
The two outcomes came apart. For strength, the slope's confidence interval contained zero in every best-fitting model — gains were similar across a wide span of repetitions in reserve. For muscle size, the slope was negative and clear of zero: the closer sets ended to failure, the more growth.
So the answer depends on what you are actually training for. If you lift to be strong — and if you are lifting for the next thirty years of your life, strength and the habit are what you are buying — the last rep appears to be nearly worthless. If you lift to add tissue, it buys something small and genuine.
The authors are unusually honest about the limits, and it matters: the repetitions in reserve were estimated from how each study described its training, not measured. That is the best available shape of the curve, not the curve.
What the last rep costs
The benefit side of this ledger is small and contested. The cost side is not.

In July 2026, Nathan Cowley and colleagues ran eighteen lifters through a crossover study: back squats and prone rows, at 65% and 85% of one-rep max, stopping at failure, one rep in reserve, or three. Then they brought everyone back a day later and measured what was left.
Neuromuscular fatigue, perceived fatigue and metabolic stress all rose as sets finished closer to failure, and the effect was still measurable twenty-four hours on. The surcharge also varied by lift: squats at the moderate load cost more than rows at the heavy one. Failure is not a flat fee. It is most expensive exactly where people most enjoy paying it — big compound lifts at weights that let you grind.
Then there is the part nobody models. In a companion analysis of that same two-legs study, the failure protocol was rated as more uncomfortable, more effortful, and left people feeling measurably worse afterwards. Adherence is the most powerful training variable in existence, and failure quietly taxes it. A programme you resent is a programme you will eventually negotiate with.
Nobody counts reps in reserve very well
All of this rests on an assumption that deserves more scrutiny than it gets: that you know when you are two reps out.
A 2025 trial from Brad Schoenfeld's lab — forty-two trained men and women doing a single set of nine exercises, twice a week for eight weeks, at failure or at two reps in reserve — tested that directly. Hypertrophy leaned slightly toward failure; strength and muscular endurance were the same either way. But the calibration result is the one worth keeping: people judged their reps in reserve far more accurately on the bench press than on the squat, and everyone got better at it over the eight weeks.
Two reps in reserve is a skill, not a setting. On a lift you know intimately, your estimate is probably close. On a squat, or on anything new, "two left" can quietly mean five. This is also why the meta-analytic curve is drawn in pencil — if lifters can't judge proximity to failure precisely, neither can the papers describing them.
Where to spend the failure budget
Here is the practical shape of it. Treat failure as a currency you have a limited weekly amount of, and spend it where it is cheap.
Applying it
Three decisions, in the order they matter
Most sets: stop with 1–3 reps in reserve
If you spend it, spend it last and spend it safe
Calibrate with bar speed, not with feeling
The cost also shows up in places you can watch. Soreness is a poor guide on its own, but a week where every session ends in failure tends to announce itself in your morning numbers — a resting heart rate that drifts up, a heart-rate variability trend that sags. If you track those trends with the Agen Band, you are not looking for a verdict on today's session; you are looking for whether the last two weeks were affordable. That is also what a deload is for.
Two footnotes for the over-forties, since that is where the stakes rise. The pooled subgroup that most clearly favoured failure was people who were already well trained — beginners had no such signal, which means the newest lifters have the least reason to grind. And if you are over 55 and lifting, one of the few supplement claims with authorised European standing is relevant here: creatine enhances the effect of resistance training on muscle strength in adults over 55, at 3 g per day.
The bottom line
Failure works. It just doesn't work much better than stopping a little short, it costs measurably more the next day, it makes training less pleasant to come back to, and the strictest version of the comparison shows the smallest benefit of all. If you lift for strength, the last rep is close to a rounding error. If you lift for size, it is worth a small amount and should be bought deliberately, at the end of an exercise, on a lift where failing is boring rather than dramatic.
Which leaves the thing the research keeps implying without saying. The grind was never the mechanism. It was the part you could see — and, as usual, the visible thing got promoted to the important thing. Most of your sets should end while you still look composed. That is not holding back. That is what the evidence looks like when you stop reading it as a moral test.
If you are assembling the rest of the picture, this sits inside the wider longevity protocol, alongside how you lower the weight and how to tell a hard week from a costly one.


