← Longevity & Supplement Guides
One Number for a Whole Night: What an AI Found in 10,000 Sleep Studies
A hospital sleep study watches thirteen channels of your body for eight hours and then reports, essentially, one number.
Brain waves from three scalp positions. Eye movement. Chin muscle tone. Airflow at the nose. Chest effort and abdominal effort, separately. Oxygen saturation. An ECG trace. The sound of your own snoring. Somewhere around a gigabyte of a human night, gathered by a technician who taped wires to your head — and the line the clinic circles on the report is events per hour.
On 3 August a paper in Nature Communications asked what else was in that recording. The answer is the most interesting thing published in sleep science this week, and it is not really about artificial intelligence at all. It is about what happens when you agree, for forty years, to describe a night with a single figure.
Ten thousand nights, thirteen channels, one number
A team led by Erhan Bilal and Jeffrey Rogers, with Reena Mehra as senior clinical author, trained a transformer model — 126 million parameters, the same broad architecture that underlies language models — on the raw signals from more than 10,000 clinical sleep recordings held in Cleveland Clinic's STARLIT registry. Not the summary reports. The waveforms.
The recordings ran from 2012 to 2022 and were linked to electronic medical records, so the researchers could follow what happened to those people afterwards: a mean of 14.5 years of it. Then they did something deliberately unglamorous. Instead of asking the model to predict a diagnosis, they let it learn the structure of sleep physiology and looked at how the nights sorted themselves.
They sorted into five groups. And the groups had very different futures.
The index that ate the night
To see why that matters, you need to know what the apnea–hypopnea index is. The AHI counts breathing interruptions and divides by hours asleep. Five to fifteen an hour is mild, fifteen to thirty moderate, above thirty severe. It has organised clinical sleep medicine, insurance reimbursement, and every conversation about disordered breathing since the 1970s.
It is also, by the field's own admission, a crude instrument. In 2021 a consensus review in SLEEP, led by Atul Malhotra and involving much of the discipline, laid out the case plainly: the AHI remains the best-studied metric available, and it is imperfect, and the two facts have to be held together. A ten-second pause and a ninety-second one count the same. A dip to 93% oxygen counts the same as a dip to 78%. Events in deep sleep count the same as events while dozing.
The AHI's problem was never accuracy. It was legibility. It survived because it fits on a form.
Five groups the index couldn't see
Here is the finding that should travel. After adjusting for age, sex, BMI, comorbidities and the AHI itself, the model's risk groups still separated on mortality: hazard ratios of 1.43, 1.54 and 1.75 for groups two through four, and 2.38 for the highest-risk group — more than double the lowest. In that same top group, atrial fibrillation ran at 2.23 (95% CI 1.47–3.38) and cognitive impairment at 1.93 (1.42–2.62).
Meanwhile, sorting the same patients by mild, moderate and severe AHI produced no such step-up at all.
Two details make this more than a modelling flourish. The pattern held in the Sleep Heart Health Study, an entirely separate national cohort — the groups still tracked with heart failure and mortality there. And it held equally in women and men, which the AHI has historically not: the index was built and validated in a population that skewed male, and it has been quietly worse at describing women ever since. A summary number encodes whoever it was calibrated on.
Figure — what the index threw away
One night, two readings of it
Five-year-plus mortality, by the model's risk group
The same patients, sorted by AHI severity
What this changes for someone who will never see a sleep lab
Three transfers, none of them medical advice:
1. A summary number is a choice about what to discard, made by whoever built it — not a property of your night.
2. When a number never moves, that can mean your nights are steady or that the number is too blunt to notice. Those feel identical from the inside.
3. Metrics built from patterns across many nights have held up better than metrics that average a single night down to a score.
Hazard ratios from Bilal et al., Nature Communications (3 Aug 2026), adjusted for age, sex, BMI, comorbidities and AHI; group one is the reference, and bar lengths are proportional to the reported hazard ratios. The AHI panel is drawn as open markers resting on the reference line because Cox regression in the same cohort found no increase in mortality or disease incidence across mild, moderate and severe — an absence of separation, not three equal measurements. Associations, not proof of cause.
Your sleep score has the same problem
This is where the story stops being about hospitals.
Every morning, tens of millions of people are handed a single number between 0 and 100 that claims to describe their night. It is computed from a handful of channels rather than thirteen, by a weighting somebody chose, and it is designed above all to be legible — a number you can read in the four seconds between waking and reaching for coffee.
That is the AHI's exact design brief. The wellness industry did not solve the compression problem; it re-solved it in a nicer typeface. Two genuinely different nights — one fragmented early, one with a shallow, restless second half — can both come back as an 82. And when the score sits at 82 for three weeks, you cannot tell whether your sleep is stable or whether the metric simply lacks the resolution to see what changed.
I want to be careful here, because the honest version of this argument cuts both ways. Consumer sleep tracking has become genuinely good at some things — timing, duration, night-to-night consistency, overnight heart rate — and those are the pieces most worth having. What it cannot do is compress a night without deciding what doesn't matter. Neither can a clinic. The difference is that the clinic has spent forty years arguing in public about its compression, and the score on your wrist has not. We wrote separately about what a wearable can and can't actually measure, and about what sleep stages do and don't tell you.
Depth and duration, not just how often
The AI didn't discover this problem. It priced it.
Sleep researchers have been building better summaries for years, and the clearest example is hypoxic burden. Instead of counting breathing events, Ali Azarbarzin and colleagues measured the area under each oxygen dip — depth multiplied by duration, summed across the night. In 2019, in the European Heart Journal, that measure tracked with cardiovascular-disease-related mortality across two independent cohorts, the Osteoporotic Fractures in Men Study and the Sleep Heart Health Study, where the event count did not.
Same recording. Same night. A different decision about what to keep, and a signal appears. That is the whole thesis of this week's paper, arrived at seven years earlier with arithmetic instead of a transformer.
The one summary number that survives compression
If single numbers are lossy, the useful question becomes: which ones lose the least?
On current evidence, the best answer for ordinary life is regularity. In 2024, Daniel Windred and colleagues calculated a sleep-regularity index from more than 10 million hours of accelerometry across 60,977 UK Biobank participants. Relative to the median, people at the 5th percentile of regularity carried a hazard ratio of 1.53 (95% CI 1.41–1.66) for all-cause mortality; those at the 95th percentile, 0.90 (0.81–1.00). Regularity outperformed sleep duration — the number everyone actually chases.
Why would regularity hold up where a nightly score doesn't? Because it isn't a within-night average. It is computed from the shape across many nights, so it keeps the pattern rather than flattening it. Compression that preserves structure survives; compression that preserves tidiness does not. We went into the mechanics in sleep regularity versus duration.
So, one thing to try this week. Stop reading the score for seven days and record two numbers only: the time you got into bed and the time you got up. Then look at your overnight resting heart rate and HRV as a trend against your own baseline, not against a population band. Consistency of timing, plus the direction those two are drifting, is a lower-resolution picture than a sleep lab produces and a higher-resolution one than a score — which is the trade most people actually want. This sits inside the wider measure, act, adjust loop.
What the study does not say
It does not say an algorithm can tell you how you slept. It ran on full clinical polysomnography, with electrodes on the scalp, in people referred to a sleep centre because something was already wrong. That is not a wrist, and it is not a general population.
It is retrospective, so it cannot separate two readings that both fit the data: that disordered sleep physiology contributes to later illness, or that early illness disrupts sleep years before it announces itself. The authors also note they lacked medication data and objective adherence data for airway pressure therapy — two large unmeasured influences on both the nights and the outcomes. The five groups came from a clustering step, which means their boundaries are a modelling decision rather than a natural kind; a different algorithm might have drawn four, or seven.
And nothing here has been tested prospectively as a clinical tool. The authors are explicit that validation in more diverse datasets comes next.
One thing that isn't a caveat: if you snore heavily, wake gasping, or feel wrecked after eight apparently normal hours, that is a conversation with a doctor, not with an app. No consumer device can diagnose anything, and none should try to.
The bottom line
An AI read ten thousand nights and found that the number medicine had agreed to use was hiding a doubling of risk. It's a good result. But the durable lesson isn't that better models are coming — it's that every summary number is an argument about what matters, and most of us accept those arguments without reading them.
Use numbers to correct fantasy, not to replace experience. A score is not your night. It is somebody's opinion about your night, rounded.


